Voice infrastructure for India
Built for how India actually speaks.
All 22 Indian languages plus English, Spanish and Arabic, with speakers labelled and code-switching followed, over one websocket, at the latency you pick.
Transcribing
hi-IN · 560 ms
मैंने रिपोर्ट भेज दी है, मीटिंग पाँच बजे शुरू होती है, बाकी काम कल सुबह करूँगा
Final
Partial
Coverage
The 22 Indian languages, plus English, Spanish and Arabic.
Available
- Hindi
hi-IN - English (India)
en-IN - Bengali
bn-IN - Marathi
mr-IN - Telugu
te-IN - Tamil
ta-IN - Gujarati
gu-IN - Kannada
kn-IN - Malayalam
ml-IN - Odia
od-IN - Punjabi
pa-IN - Assamese
as-IN - Urdu
ur-PK - Maithili
mai-IN - Nepali
ne-NP - Sanskrit
sa-IN - Sindhi
sd-IN - Konkani
kok-IN - Dogri
doi-IN - Kashmiri
ks-IN - Bodo
brx-IN - Manipuri
mni-IN - Santali
sat-IN - Spanish
es-US - Arabic
ar-AR
Leave the language off and the stream works it out as you speak, follows a switch mid-sentence, and marks each word with the language it was spoken in.
Integration
First transcript in about ten lines.
One websocket: send raw PCM, read partial and final text as it arrives. There is no package to install, and the protocol is documented in full.
import asyncio, json, os, websockets
URL = "wss://api.utter.cc/v1/stream?tier=standard&sample_rate=16000"
AUTH = {"Authorization": f"Bearer {os.environ['UTTER_API_KEY']}"}
async def main():
async with websockets.connect(URL, additional_headers=AUTH) as ws:
asyncio.create_task(send_mic(ws)) # float32 PCM at 16 kHz
async for raw in ws:
event = json.loads(raw)
if event["type"] in ("partial", "final"):
print(event["text"], event["type"])
asyncio.run(main())
const url = "wss://api.utter.cc/v1/stream"
+ "?tier=standard&sample_rate=16000&token=" + token;
const ws = new WebSocket(url);
ws.onmessage = (message) => {
const event = JSON.parse(message.data);
if (event.type === "partial" || event.type === "final") {
render(event.text, event.type === "final");
}
};
mic((frame) => ws.send(frame));
Text to speech
Five voices in Hindi and English.
Post text, get a voice. Speech streams a sentence at a time, so playback starts while the rest is still being generated.
- Languages
- Hindi hi-INEnglish en-US
- Voices
- AriaJasonJohnLeoSofia
- Audio
- 16-bit mono at 22,050 HzWAV or a stream
- Per request
- Up to 1,000 characters
- Price
- Free while in preview
import os, requests
response = requests.post(
"https://api.utter.cc/v1/speech",
headers={"Authorization": f"Bearer {os.environ['UTTER_API_KEY']}"},
json={"text": "आपका ऑर्डर कल शाम तक पहुँच जाएगा।",
"language": "hi-IN", "voice": "sofia"},
)
open("order.wav", "wb").write(response.content)
Voice infrastructure for India
Hear it as it is spoken.
मैंने रिपोर्ट भेज दी है, कृपया देख लीजिए।
Take a key, paste ten lines, hear your users as they actually talk.
Voice infrastructure for India. Speech to text in all 22 Indian languages plus English, Spanish and Arabic, and text to speech in Hindi and English.