Platform · Agents
Agent Lab: see what a run did, then test why
Agent Lab is a panel in the agent playground for answering three questions about an agent: what did it do on that turn, which parts of its prompt are doing the work, and does it give the same answer twice. Open it from the Lab button in the playground and it works on the agent and conversation you already have open.
The lab has three tabs. Run graph draws a past turn as a graph of steps and lets you fork the conversation from any message. Prompt knockout removes sections of the system prompt one at a time and measures the effect. Consistency asks the same question several times and reports whether the answers mean the same thing.
By the VegaDūta team · Last updated
In short
- Agent Lab is a testing panel in the VegaDūta playground with three instruments: a run graph, prompt knockout, and a consistency check.
- The run graph shows each step of a past turn (plan, tool calls, guardrail decisions, handoffs, reply) and lets you fork the conversation from any message and compare the branch with the original.
- Prompt knockout removes one section of an agent's system prompt at a time and measures how the score changes on the same questions.
Run graph: every step of a turn
Each turn in the playground becomes a graph: the user message, the plan, each tool call, guardrail decisions, handoffs to other agents, and the reply. Selecting a step shows its tool name, a shortened preview of its arguments and result, any error, and how long it took. The panel also totals the turn's elapsed time, tool calls, tokens, and estimated cost.
Fork from any message
Fork rewinds the conversation to just before a message and sends that message again, edited if you want. Later turns are removed from the live conversation and the original is kept as a snapshot, so you can compare the branch with the original side by side, including which tools each version called. Two limits are stated in the dialog: tool calls the original turn already made are not undone, and snapshots are kept in your browser only, for 24 hours.
Prompt knockout: which instructions matter
Knockout splits the system prompt at headings and blank lines, then runs the same questions with one section removed at a time. Questions come from an eval dataset or from probe questions you type in. The unchanged prompt is run several times first to measure the noise floor, and each section is then labelled load-bearing, harmful, or no measurable effect. When there are too few questions to tell a real change from noise, the lab says so instead of labelling.
Consistency, and what lab runs leave out
The consistency check asks one question several times and groups the answers by meaning. Knockout and consistency runs ask the model directly with the agent's prompt, sampling settings, and usual model. They do not use tools, memory, knowledge lookups, guardrails, or caching, and they are not added to the conversation. The lab shows the number of model calls before a run starts, the calls count toward your usage limits, and Free workspaces get smaller run sizes.
Frequently asked questions
How do I see what my AI agent did during a conversation?
Open Agent Lab from the playground and pick the turn. The run graph shows the plan, each tool call with a preview of its arguments and result, guardrail decisions, handoffs, and the reply, with timings, token counts, and an estimated cost.
Can I re-run an agent from an earlier point in a conversation?
Yes. Fork rewinds the conversation to just before a chosen message and sends it again, optionally edited. The original is kept as a snapshot in your browser for 24 hours so you can compare. Tool calls the original turn already made are not undone.
How do I find out which parts of a system prompt matter?
Use prompt knockout. The lab removes one section of the prompt at a time, runs the same questions, and reports how the score changed against a measured noise floor. Sections come out labelled load-bearing, harmful, or no measurable effect.
Do Agent Lab runs behave exactly like production?
The run graph shows real playground turns. Knockout and consistency runs do not: they call the model directly without tools, memory, knowledge, guardrails, or caching, so they isolate the prompt. The lab states this next to the run button.
See it working in two minutes
The sandbox provisions a real tenant — describe an agent in one sentence and test it, no account, no card. Or browse ~90 industry workflow recipes to see what teams build.