VegaDūta

Integrations · Voice & language AI

Sarvam AI on VegaDūta: Indian-language agents

VegaDūta integrates Sarvam AI across the full surface of its API: speech-to-text, text-to-speech with the bulbul model, translation, transliteration, and language identification — all exposed to agents as callable tools — plus Sarvam's chat models (such as sarvam-105b), which you can select as an agent's LLM. If you are searching for how to build Sarvam-powered agents or Indian-language AI agents, this is that page: one platform where Sarvam handles both the language layer and, if you choose, the reasoning layer.

Every Sarvam call is logged with per-call records and governed by quotas, and the whole capability switches on per agent — so you decide which of your agents get Sarvam tools.

What agents get from Sarvam

The integration is not a single endpoint bolted on — it covers Sarvam's language capabilities and makes each one an agent tool the model can call mid-conversation.

  • Speech-to-text for Indian-language audio.
  • Text-to-speech via the bulbul model for spoken replies.
  • Translation across Indian languages.
  • Transliteration between scripts.
  • Language identification for routing and reply-language decisions.
  • Sarvam chat models (e.g. sarvam-105b) selectable as an agent's LLM.

Sarvam models as the agent's brain

Beyond tools, you can pick a Sarvam chat model as the LLM an agent runs on — useful when you want Indian-language reasoning end to end rather than a Western model with translation bolted on. Model selection lives in the agent builder alongside every other connected provider.

Governance: logging, quotas, per-agent control

Every Sarvam call is logged per invocation and counted against quotas, and the integration has a per-agent enable toggle. That means costs are visible, limits are enforceable, and an agent only carries Sarvam tools if you switched them on for it.

Why this matters for Indian-language agents

Agents on VegaDūta already answer in 22 Indian and global languages; Sarvam deepens that with speech and script handling built for India. A WhatsApp voice note in Kannada can be transcribed, understood, and answered — in text or speech — inside one conversation.

Frequently asked questions

Can I build an AI agent that speaks Indian languages?

Yes. VegaDūta agents answer in 22 Indian and global languages, and the Sarvam AI integration adds speech-to-text, text-to-speech (bulbul), translation, transliteration, and language identification as agent tools. You can also run the agent itself on a Sarvam chat model.

Can I use Sarvam's LLM as my agent's model?

Yes. Sarvam chat models such as sarvam-105b are selectable as an agent's LLM in the agent builder, the same way you would pick any other connected provider's model.

How is Sarvam usage controlled and billed for visibility?

Every Sarvam call is logged per invocation and governed by quotas, and the capability has a per-agent enable toggle. For plan-level details, see current pricing on the billing page.

What Sarvam capabilities does VegaDūta expose as tools?

Transcribe (speech-to-text), text-to-speech via the bulbul model, translate, transliterate, and identify-language — each callable by the agent mid-conversation when the toggle is on for that agent.

See it working in two minutes

The sandbox provisions a real tenant — describe an agent in one sentence and test it, no account, no card. Or browse ~90 industry workflow recipes to see what teams build.

Related

Explore VegaDūta