VegaDūta

Platform · Quality

Versions and evals: change an agent without guessing

Every saved change to an agent or a workflow becomes a version. Versions can be labelled staging or production, and any earlier version can be restored. For agents, promotion to production can be tied to a quality gate, so a version that fails its eval dataset does not become the one your customers talk to.

The evaluation side lives on each agent's Quality page: datasets of test cases, LLM judges that score the answers, statistics that say how far to trust a score, and simulated conversations with difficult personas.

By the VegaDūta team · Last updated

In short

  • Every saved change to a VegaDūta agent or workflow is kept as a version that can be restored.
  • Versions carry staging and production labels, and promoting a version to production can be gated on a passing eval run.
  • Eval datasets are scored by six LLM judges (accuracy, faithfulness, helpfulness, relevance, safety and tone), with confidence intervals and paired comparisons between versions.

Versions, labels, and restore

The version list shows each saved configuration. Promote points the production label at a version, and runs follow the label from then on. Restore makes an older version the current configuration; the restore is itself saved as a new version, so it can be undone. Workflows work the same way: editors run the draft, and live triggers run the version labelled production.

Eval datasets and six judges

A dataset is a list of test cases for one agent, optionally with expected answers. You can save a real conversation turn as a case. A run sends every case to the agent and scores the answers with LLM judges for accuracy, faithfulness, helpfulness, relevance, safety, and tone. Runs have a budget cap and can repeat each case to measure how stable the result is.

A promotion gate that accounts for luck

An agent that passes a case once may fail it the next time. The gate can therefore require pass^k: a case counts only if it passes on every one of k repeats. Results are shown with confidence intervals instead of a bare percentage, and two runs can be compared case by case, so a two-point difference on twenty cases is not mistaken for an improvement.

Simulated conversations

Simulation runs a bounded multi-turn conversation between your agent and a persona, built-in or one you write, including adversarial ones. Each persona gets its own isolated run, and each conversation ends with a stated reason such as goal met, maximum turns reached, or no progress.

Frequently asked questions

Can I roll back an AI agent to an earlier configuration?

Yes. Every saved change is a version, and Restore makes any earlier version the current configuration. The restore is saved as a new version, so it can itself be undone.

Can I stop a bad agent change from reaching production?

Yes. Production is a label on a version, and a promotion gate can require a passing eval run before the label moves. The gate can use pass^k, which requires each test case to pass on every repeat.

What does pass^k mean?

A case passes only if it passes on all k repeated attempts. It measures reliability instead of a single lucky answer, which matters for agents because the same input can produce different outputs.

Are evals available on the Free plan?

Yes, with smaller limits on dataset size, repeats, and per-run budget than paid plans. The Quality page shows the limits that apply to your workspace.

See it working in two minutes

The sandbox provisions a real tenant — describe an agent in one sentence and test it, no account, no card. Or browse ~90 industry workflow recipes to see what teams build.

Related

Explore VegaDūta