utter

Voice infrastructure for India

Built for how India actually speaks.

All 22 Indian languages plus English, Spanish and Arabic, with speakers labelled and code-switching followed, over one websocket, at the latency you pick.

Transcribing

hi-IN · 560 ms

मैंने रिपोर्ट भेज दी है, मीटिंग पाँच बजे शुरू होती है, बाकी काम कल सुबह करूँगा

Final

Partial

Coverage

The 22 Indian languages, plus English, Spanish and Arabic.

Available

  • Hindihi-IN
  • English (India)en-IN
  • Bengalibn-IN
  • Marathimr-IN
  • Telugute-IN
  • Tamilta-IN
  • Gujaratigu-IN
  • Kannadakn-IN
  • Malayalamml-IN
  • Odiaod-IN
  • Punjabipa-IN
  • Assameseas-IN
  • Urduur-PK
  • Maithilimai-IN
  • Nepaline-NP
  • Sanskritsa-IN
  • Sindhisd-IN
  • Konkanikok-IN
  • Dogridoi-IN
  • Kashmiriks-IN
  • Bodobrx-IN
  • Manipurimni-IN
  • Santalisat-IN
  • Spanishes-US
  • Arabicar-AR

Leave the language off and the stream works it out as you speak, follows a switch mid-sentence, and marks each word with the language it was spoken in.

Integration

First transcript in about ten lines.

One websocket: send raw PCM, read partial and final text as it arrives. There is no package to install, and the protocol is documented in full.

Pythonpip install websockets
import asyncio, json, os, websockets

URL = "wss://api.utter.cc/v1/stream?tier=standard&sample_rate=16000"
AUTH = {"Authorization": f"Bearer {os.environ['UTTER_API_KEY']}"}

async def main():
    async with websockets.connect(URL, additional_headers=AUTH) as ws:
        asyncio.create_task(send_mic(ws))     # float32 PCM at 16 kHz
        async for raw in ws:
            event = json.loads(raw)
            if event["type"] in ("partial", "final"):
                print(event["text"], event["type"])

asyncio.run(main())
JavaScriptno package to install
const url = "wss://api.utter.cc/v1/stream"
  + "?tier=standard&sample_rate=16000&token=" + token;

const ws = new WebSocket(url);

ws.onmessage = (message) => {
  const event = JSON.parse(message.data);
  if (event.type === "partial" || event.type === "final") {
    render(event.text, event.type === "final");
  }
};

mic((frame) => ws.send(frame));

Text to speech

Five voices in Hindi and English.

Post text, get a voice. Speech streams a sentence at a time, so playback starts while the rest is still being generated.

Languages
Hindi hi-INEnglish en-US
Voices
AriaJasonJohnLeoSofia
Audio
16-bit mono at 22,050 HzWAV or a stream
Per request
Up to 1,000 characters
Price
Free while in preview
Pythonpip install requests
import os, requests

response = requests.post(
    "https://api.utter.cc/v1/speech",
    headers={"Authorization": f"Bearer {os.environ['UTTER_API_KEY']}"},
    json={"text": "आपका ऑर्डर कल शाम तक पहुँच जाएगा।",
          "language": "hi-IN", "voice": "sofia"},
)
open("order.wav", "wb").write(response.content)

Voice infrastructure for India

Hear it as it is spoken.

मैंने रिपोर्ट भेज दी है, कृपया देख लीजिए।

Take a key, paste ten lines, hear your users as they actually talk.

utter

Voice infrastructure for India. Speech to text in all 22 Indian languages plus English, Spanish and Arabic, and text to speech in Hindi and English.