⚡ Product updates

What we shipped, and why it exists.

Every capability below is dated, with the problem it was built for. Where something is still landing, the entry says so.

Start free See pricing
Product updates

Sarvam’s Indian-language models are coming to Nutaan — 22 August 2026

We have partnered with Sarvam AI, who build speech and language models for Indian languages first rather than as an adaptation of English ones. It closes the gap our own callers hit most often: an agent that handles clean Hindi well, Marathi passably, and an accent from two districts over not at all.

Recognition trained on Indian speech instead of adapted to it. The difference shows up in names, numbers and places — which is precisely what a booking or a captured lead depends on getting right.

Code-switching treated as one sentence rather than two languages. “थोड़ा बताइए ना, कितने का padega” is how people actually talk, and it should not be the thing that breaks the call.

Voices that carry an Indian accent natively, rather than an English voice pronouncing Hindi words.

Where it stands: the agreement is signed and the integration is underway — this entry is the announcement, not the release note. Setu and every existing engine keep running unchanged, so nothing you have already built needs to move.

The same models that make an agent work for a caller in Kolkata make it work for an Indian business selling into Dubai or New Jersey. Bharat first, not Bharat only.

Setu — a bridge across Asia’s languages — 17 August 2026

Every speech-to-speech engine we could buy refuses most of the languages Asia actually speaks. Ask one for Marathi or Telugu and it answers with an error rather than an accent; the rest are English-first. Setu exists for that gap — 32 Asian languages spoken natively, 70 in all, from Hindi to Japanese and Bengali to Georgian.

South Asia — Hindi, Bengali, Marathi, Tamil, Telugu, Kannada, Malayalam, Gujarati, Punjabi, Urdu, Odia, Assamese, Nepali.

East and Southeast Asia — Chinese, Japanese, Korean, Indonesian, Malay, Filipino, Vietnamese, Thai, Khmer, Lao, Mongolian.

Central and West Asia — Arabic, Turkish, Hebrew, Kazakh, Uzbek, Azerbaijani, Armenian, Georgian.

It follows the caller. On one live call the caller moved between Bengali, Hindi, Gujarati, Odia, Marathi and English, and the agent changed with them each time without being told to.

Measured on real calls to a real phone: about 900ms to the first word of the greeting, 395ms to each reply — at roughly a sixth of the per-minute cost of the engines it replaces.

Stated rather than hidden: the engine understands these languages far better than it writes them down. For the eleven where we have a measured speech-to-text model, the caller is transcribed with that instead, so the saved transcript is accurate as well as the conversation.

See exactly how long every caller waited — 14 August 2026

Two silences decide whether a call feels human, and they break for different reasons. The wait before the agent says hello is the engine starting or the phone line; the wait after the caller stops talking is the model or the voice. Both are now measured on every call, separately.

Recorded automatically — nothing to switch on and nothing to send us.

The usual wait and the slow one side by side, because a one-second median with a three-second tail is a call that feels broken to one caller in ten.

Per-call breakdown with a bar for every answer, so a single stall is visible instead of averaged away.

Your knowledge base, fact-checked against real questions — 12 August 2026

Crawling a website gets you whatever the business chose to publish, which is rarely what callers ask. Nutaan now puts the real questions to your knowledge base, reads what comes back, and tells you only what is genuinely missing.

Judged by reading the retrieved passages, not by a similarity score — a menu page scores 0.44 against “products and services” and used to be reported missing.

Asks in your trade’s words: a cafe is asked about its menu, a clinic about treatments, a dealership about models.

A third verdict for things half-there — items listed with no prices, a booking form with no rules.

Shows the passages behind every verdict, so “we don’t have your prices” is checkable rather than asserted.

“Uh-huh” no longer interrupts the agent — 10 August 2026

People say “yeah”, “right”, “haan ji” while they are listening, not to take the floor. The agent now tells the difference by what was said rather than how many words it was — and a plain “yes” after a question still counts as the answer.

78 acknowledgements shipped across English, Hindi and Hinglish, applied by default.

Add your own per agent, for a language we don’t ship yet.

Judged by content, not word count: “haan ji haan ji” is four words and doesn’t interrupt; “no wait” is two and does.

Dial by hand, from your own number, without leaving — 8 August 2026

Not every call should be automated, and the ones that shouldn’t were forcing your team into a second tool. The dialer is now part of the same workspace, using the same number, writing to the same history.

Call straight from the browser on your configured Nutaan number.

Recording, transcript and outcome land in the same place as the AI calls.

No second subscription and no context lost between platforms.

Official benchmarks, published — 4 August 2026

Everyone in voice AI claims low latency and high accuracy. We published the measurements instead — per language, per engine, with the method written down so you can repeat them.

Speech recognition accuracy by language, including Indian English and regional languages.

Engine-by-engine response times, measured on real calls rather than a lab loop.

The method is stated, so the numbers can be checked rather than believed.

Agents that follow the caller’s language mid-sentence — 28 July 2026

Real Indian calls switch between Hindi and English inside one sentence, and a US caller expects plain English throughout. An agent now follows whoever it is talking to instead of dragging them back.

Answers in the language the caller used, switching whenever they do.

Technical and commercial terms stay in English, which is how people actually say them.

30+ languages, with recognition tuned per language rather than one global setting.