Tutorial · Testing
How to debug an agent with Agent Lab
This tutorial shows you how to find out why an agent answered the way it did. You open Agent Lab from the playground, read the run graph for the turn that went wrong, fork the conversation from that message to try a change, and then use prompt knockout to see which sections of the system prompt are doing the work.
The run graph works on turns you have already sent. Knockout and consistency runs make extra model calls, and the lab shows the number before you start.
By the VegaDūta team · Last updated
Step by step
Send a message in the playground
Open an agent's playground and send the message you want to investigate. The lab reads turns from the conversation you have open, so there is nothing to look at until the agent has answered at least once.
Open the lab and read the run graph
Press the Lab button in the playground. The Run graph tab draws the selected turn as steps: user message, plan, tool calls, guardrail decisions, handoffs, and the reply. Use the Turn selector to pick an earlier turn. The totals show time, tool calls, tokens, and estimated cost.
Inspect a step
Select a step to see its details: the tool name, a preview of the arguments and the result, any error, and the duration. Previews are shortened to 400 characters and may be redacted, so treat them as a pointer to what happened, not a full payload.
Fork from a message
Choose Fork from here on a user message. Edit the message if you want to try different wording, then press Rewind and re-run. The conversation is rewound to just before that message and the message is sent again. Tool calls the original turn already made are not undone, and attachments are not sent again.
Compare the branch with the original
Open Compare. The original is shown as a snapshot next to the live branch, with the tools that only one version called listed for each side. Snapshots are kept in this browser only, for 24 hours.
Run a prompt knockout
Switch to the Prompt knockout tab. The system prompt is split at headings and blank lines; untick any section you want left out of the test. Choose an eval dataset or type probe questions, one per line, set the baseline repeats, check the stated number of model calls, and run. Each section comes back labelled load-bearing, harmful, or no measurable effect.
Good to know
A few limits decide what the lab can and cannot tell you.
- Knockout and consistency runs call the model directly: no tools, memory, knowledge, guardrails, or caching. They isolate the prompt and do not reproduce a full production turn.
- With too few questions or repeats the lab reports that it cannot tell a real change from noise, and labels nothing.
- If the agent has response caching on, re-sending an unchanged message can return the stored answer. Edit the message to get a fresh one.
- A turn that ran on the device, was answered from cache, failed, or was interrupted cannot be forked.
Frequently asked questions
How do I see which tools my agent called?
Open Agent Lab from the playground and look at the Run graph tab. Each tool call is a step, and selecting it shows the tool name, a preview of its arguments and result, and how long it took.
Does forking a conversation undo what the agent already did?
No. Fork rewinds the conversation and re-sends the message, but tool calls the original turn made, such as a sent message or a created record, are not undone.
How many model calls does a knockout run make?
The lab states the number before you start, based on the sections, questions, and baseline repeats you chose, and refuses a run over its per-run limit. Calls count toward your usual usage limits.
Where do knockout questions come from?
Either from an eval dataset for the agent, where cases with an expected answer are scored against it, or from probe questions you type in, where a judge scores how well each answer addresses the question.
See it working in two minutes
The sandbox provisions a real tenant — describe an agent in one sentence and test it, no account, no card. Or browse ~90 industry workflow recipes to see what teams build.