The voice layer for every engine.
Write a script, pick a real engine, and ship audio you own.
No credit card · 5 real engines · The audio is yours
[gently]The northern lights drifted across the sky, slow and silent, like breathing.
Live · Every click generates real audio
Quality Elo from the Artificial Analysis Speech Arena, June 10, 2026. Latency is our measured wall clock. * Notes on the full benchmark.
Same script, scored in the open
Press play on a use case.
Every disc is real output from the engine named on its use case, recorded unedited.
Every engine. One studio.
- Gemini Flash, Kokoro, Grok Voice, MAI Voice 2, and Zonos in one place
- Up to 30,000 characters per generation
- MP3 or WAV out, the audio is yours
{
"engine": "kokoro",
"voiceId": "af_heart",
"text": "Read this aloud."
}A developer API is on the roadmap. Today this runs in the studio.
An open benchmark, not vendor marketing.
- Real measured latency, dated and reproducible
- Quality Elo from a third-party arena, not self-scored
- Updated as engines change
Quality Elo from the Artificial Analysis Speech Arena, retrieved June 10, 2026. For context, the top model of all rated is Fun-Realtime-TTS at 1228.06. Latencies are our own measured wall-clock numbers.
Own your voice.
- Commercial rights, worldwide, no watermark
- Export every generation you make
- Flat pricing, not per-character credits
- Voice
- Kore
- Owner
- You
- Use
- Commercial, worldwide
- Export
- MP3 and WAV
- Watermark
- None
Plain English, not a credit balance. What you generate is yours to ship.
Dub episode 12 into Spanish. Keep the warm read.
EPISODE-12.MP3 · 14 MBTranscribed, translated, and re-voiced. Same pacing, same warmth, now in Spanish.
One take in. Every language out.
Dubbing
Your episode, re-voiced in another language, without losing the read.
- Carries the performance, not just the words
- 8 languages from one source recording
- Your file in, your file out, yours to keep
Text to Speech
Audiobook Studio
Sound & Music
From a blank script to audio you own.
Step 1: Write your script
Type or paste your text. Add bracketed [emotion] cues when you want a voice to act them.
Step 2: Pick an engine
Choose from the real engines, or follow the open benchmark to the one that fits your job.
Step 3: Generate
One request returns finished audio in seconds. Preview it right in the console.
Step 4: Own and export
Save it to your library and export MP3 or WAV. Commercial rights, no watermark.
The whole studio, at a glance.
The lighthouse had been dark for forty years.wistfulShe still set a cup out for him every night, the way her grandmother taught her.steadier nowTonight, the lamp would burn again.
Audiobook Studio
BetaChaptered narration, one voice held across the whole book.

Write the line. Hear it acted.
Bracketed cues are stage directions, and the voice performs them.
[softly]Some stories are meant to be heard.
This one starts with you.
Gemini Flash · acts the cue, not just reads it
- 5
- Real engines
- 30,000
- Characters per run
- ~1.0s
- Fastest measured
- $0
- To start
Latency measured 2026-06-10, wall clock to full audio
The honest answers.
Plain replies to the questions everyone asks before their first generation.
Still curious about how we build? Read the about page
How does pricing work?
Who owns the audio I generate?
Which engine should I pick?
Is the voice real or prerecorded?
What is the open benchmark?
What testers say.
I put [wearily] in front of one line in chapter nine and the read actually changed. Then I re-rendered that single chapter three times for a pronunciation fix and never watched a balance drain.

Start with the best voice for every job.
Free to start, no credit meter. Open the console and hear it for yourself.


