changelog / testing / ai-agents

Tests and reusable scenarios

Twinbay team
Tests and reusable scenarios

The new Tests feature checks an agent's API calls against requirements you write in plain language. You can require an order lookup, for example, while forbidding a refund.

Describe the task and its outcomes, and Twinbay compiles them into rules that match captured requests. Results show a verdict and the calls behind it. When evidence is missing, an outcome can be inconclusive.

Use a checkpoint to check a run in progress, or a final evaluation to close it. Test and evaluator versions keep a record of the rules you used. To try another evaluator for the same test version, regrade the captured traffic without running the agent again.

Open Tests in the console, or use the same workflow through MCP.

Reusable scenarios

Save a set of twins and their starting data as a scenario to reuse across runs. You don't need to start an environment to create one. Scenarios save the starting definition; they don't capture changes from a previous run.

Describe the records you need, and Twinbay generates a seed before or during environment creation. Once it's ready, you can load it into compatible environments without generating the records again.

Open a scenario from the new card view to see its twins and seeds. You can also browse public scenarios from other organizations, which are read-only.

WebEDI twin

The new WebEDI twin gives browser agents a local version of the portal's Documents workspace. Agents can browse and filter documents, select a trading partner and save panel preferences.

Send confirmation and final submission aren't modelled in the browser workflow. The twin returns a not-modelled response if an agent tries either action.

Improvements

  • Shopify: Added manual collections and listings for draft orders and discounts.
  • Linear: Added reads for organization details, projects and cycles, including linked issues and progress. Project and cycle mutations remain unsupported.
  • Seeds: Generated records are now validated in full. Repairs can fix individual fields or records without changing the rest.
  • Tests: Outcomes, evaluators and runs now have their own tabs.
  • Website: Updated the homepage walkthrough and simplified the footer. Added privacy and terms pages and cookie preferences.
  • Docs: Added a HAR recording guide and a guide to testing agents that call third-party APIs.

Fixes

  • Fixed evaluator and run history when viewing older test versions.
  • Fixed homepage sign-in links and authentication callback recovery.
  • Fixed landing-page cookie consent dismissal and initial focus.
  • Fixed catalog rendering for twins without an SDK.

API

  • Added selected Shopify REST reads for shop details, products, customers, orders, draft orders and custom collections. These use the same records as GraphQL. REST writes remain unsupported.
  • Standardized paginated list responses and newest-first ordering.
  • Moved provisioned twin operations to /twins/{twin_id}, replacing nested environment routes.

Live · console.twinbay.ai

Start with one base URL.

Create an environment and point your existing SDK at the hostname it gives you. The environment logs each request your agent makes, and you can edit any record it holds.

Start building

An environment in under a minute, no card required. Twinbay is free for academic researchers.