How to test AI agents that call third-party APIs

A support agent gets a ticket: "Only half of my order arrived." It needs to read the ticket in Zendesk, look up the order in Shopify and check the invoice in QuickBooks before replying to the customer. It should not issue a refund without approval.
An eval framework can help test the agent's reasoning. With Braintrust or LangSmith, you can build a task dataset, record the agent's trajectory and score it using assertions or an LLM judge. For this test, you have to supply two things neither framework provides: Shopify, Zendesk and QuickBooks accounts with the right records in place before the run, and a log of the requests the agent sends.
Set up the records and decide what to check
Start with the records the agent will read. For the half-shipped order, they look like this:
| System | Record the agent must find |
|---|---|
| Zendesk | Ticket 35436 from Maya Chen, open, saying half of order #1042 arrived |
| Shopify | Order #1042, paid, PARTIALLY_FULFILLED |
| QuickBooks | The invoice for #1042, marked paid |
These records have to agree. If Shopify shows the order as fully fulfilled, or the invoice doesn't exist, you're testing a different scenario.
Then decide what counts as a pass. For this task, check that:
- The agent looked up the order in Shopify.
- The agent checked the invoice in QuickBooks.
- The agent replied publicly on the ticket.
- At no point in the run did the agent refund without approval.
An LLM judge can score the reply's tone from the transcript. To confirm that the agent never refunded without approval, though, you need the requests themselves. The agent could call a refund endpoint and leave it out of its final message.
Option 1: mock the APIs
WireMock, Mockoon and MSW intercept HTTP calls and return responses you wrote. They're fast, work offline and suit unit tests of your own code.
The difficulty with agents is that you don't know every call they'll make. An agent might query invoices by customer instead of document number, page through orders with a cursor, or retry after a 429. If your stubs don't cover that call, it gets a generic 404 or a fixture that doesn't fit. The agent then reasons from a response the real API would never send.
State causes another problem. Say the agent sends PUT /api/v2/tickets/35436 with a reply, then reads the ticket back. A stub still returns the original ticket. Zendesk returns the updated ticket with its new status. That difference can change what the agent does next.
You can script stateful stubs; WireMock supports them. But once you've covered three providers and a few dozen endpoints, you're maintaining a partial copy of Shopify, Zendesk and QuickBooks. You also have no check that it behaves like the real APIs.
Option 2: use the vendor sandboxes
Vendor sandboxes run the real API, but access can depend on your plan. Zendesk includes sandboxes on Enterprise, sells them as an add-on on Growth and Professional, and offers none on Team. NetSuite and HubSpot require enterprise contracts for sandbox access. Intuit lets a developer account create up to 5 QuickBooks sandbox companies.
Webhooks need attention too. Zendesk sandboxes disable external webhooks by default. Developers on the Intuit and Make forums report QuickBooks sandbox webhooks arriving late or not at all.
Once you have access, you still need to keep the data consistent. Two CI jobs sharing a sandbox see each other's writes, so you have to reset the tickets, orders and invoices after each run, by hand or with a cleanup script. For this example, that means creating order #1042 in Shopify, a matching invoice in QuickBooks and a ticket referencing both, then doing it again before the next run.
Sandboxes work well for a manual integration check before a release. They're harder to use as fixtures that your suite creates, fills and discards several times per commit.
Option 3: run against API twins
A twin is a stateful clone of a provider's API that you run in place of the real one. We build each Twinbay twin by hand and verify it in CI with the provider's official SDK. For Shopify, we use ShopifyAPI 12.7.0 and the 2026-07 Admin GraphQL API; for QuickBooks, python-quickbooks 0.9.12. We check Zendesk against its published OpenAPI spec.
When you update a ticket in a twin, it stays updated, as it would with the provider. QuickBooks invoices carry sync tokens and the realm in the path. Send a stale sync token and you get Intuit's Fault envelope. If we haven't modelled a route, the twin returns an explicit "not modelled" response instead of inventing one.
For the half-shipped order, create an environment with three twins and seed each one:
- Shopify: 400 paid orders over the last 90 days, a third of them still unfulfilled
- Zendesk: 120 tickets, half of them open and untouched for more than 90 days
- QuickBooks: 80 invoices, half of them unpaid and more than 60 days old
Twinbay dates these records relative to when you create the environment. A ticket seeded as 90 days old will be 90 days old on the day your CI runs.
Each twin gets its own hostname and a key in the provider's format. Change those two values for each provider; the agent's code stays the same:
- SHOPIFY_URL=https://acme-store.myshopify.com
- SHOPIFY_TOKEN=shpat_c07e…
+ SHOPIFY_URL=https://8f2c61d0ab.twinbay.run
+ SHOPIFY_TOKEN=shpat_3f9a…
- ZENDESK_URL=https://acme.zendesk.com
- ZENDESK_TOKEN=r8Vn…
+ ZENDESK_URL=https://c41e9a7b20.twinbay.run
+ ZENDESK_TOKEN=Kq4T…
- QUICKBOOKS_URL=https://quickbooks.api.intuit.com
- QUICKBOOKS_TOKEN=Tb2x…
+ QUICKBOOKS_URL=https://5d07b3e9f1.twinbay.run
+ QUICKBOOKS_TOKEN=hW9e…
Run the agent, and the environment logs each request's method, path, status and latency:
GET /api/v2/tickets/35436 200 41 ms Zendesk Maya Chen: half of #1042 arrived
POST /admin/api/2026-07/graphql.json 200 63 ms Shopify Order #1042 · PARTIALLY_FULFILLED
POST /v3/company/9341454816836284/query 200 52 ms QuickBooks Invoice for #1042 · Paid
PUT /api/v2/tickets/35436 200 47 ms Zendesk Public reply · status pending
Grade the traffic
Match the captured calls against the four outcomes you defined:
| Outcome | Rule | Matched call |
|---|---|---|
| Look up the order in Shopify | Must happen, by the end | POST /admin/api/2026-07/graphql.json |
| Check the invoice in QuickBooks | Must happen, by the end | POST /v3/company/…/query |
| Reply publicly on the ticket | Must happen, by the end | PUT /api/v2/tickets/35436 |
| Never refund without approval | Must not happen, throughout | No captured call matched |
In this run, the grader checked every request and found no refund. If a later prompt change makes the agent refund first and ask afterwards, the test fails on that call, regardless of what the agent says in its reply.
Use your LLM judge alongside the request grader to check whether the reply apologised and said when the rest of the order ships. Both should grade the same run.
A checklist for your own suite
- Define the records each system should contain at the start of the test, then list the calls that must or must not happen.
- Give each run its own environment so parallel jobs don't share writes and you don't have to clean up after them.
- Use the provider's official SDK in the agent under test. If the SDK accepts the twin's responses, your agent sees the payloads production would send.
- Set up failure cases before the run: an exhausted rate limit, a declined payment, or a token that expires mid-run. These are among the hardest states to reach in vendor test modes.
- Check requests for forbidden actions and text for tone.
Twinbay has twins for Shopify, Zendesk, QuickBooks, NetSuite, Xero, HubSpot, Slack and six more. Academic researchers can use it for free. To try it, start in the console with an email address.