Skip to content
New · the open voice benchmark is liveRead it
cantari

The voice layer for every engine.

Write a script, pick a real engine, and ship audio you own.

No credit card · 5 real engines · The audio is yours

ScriptGemini Flash

[gently]The northern lights drifted across the sky, slow and silent, like breathing.

KoreGenerate

Live · Every click generates real audio

Voice console
Engine
RoutedGemini Flash·Expressive, acts your [cues]
Gemini Flash voices
696/30000
Live leaderboardFull benchmark →
01Gemini Flash1225
02Grok Voice1197
03Kokoro1060
04MAI Voice 21007*
05Zonos1000*

Quality Elo from the Artificial Analysis Speech Arena, June 10, 2026. Latency is our measured wall clock. * Notes on the full benchmark.

Same script, scored in the open

Routes across
Gemini FlashMP3 + WAV EXPORT30,000 CHARS PER RUNKokoroFLAT PRICINGOPEN BENCHMARKGrok VoiceCOMMERCIAL RIGHTS5 REAL ENGINESMAI Voice 2Zonos
01 / Hear it

Press play on a use case.

Every disc is real output from the engine named on its use case, recorded unedited.

Warm library, open book

Audiobooks & Publishing

One voice, every chapter

Game dev desk, teal glow

Game Dev

Barks and temp lines, in character

Recording corner, teal chair

Podcasts

Intros, ad reads, and pickups

Writer's desk at night, vintage microphone

YouTube & Video

A voiceover for every upload

02 / One studio

Every engine. One studio.

Five real engines today, in a single studio. Write once, switch engines with one tap, and hear the same script in any voice. New engines arrive in the same place, so your workflow never changes.
  • Gemini Flash, Kokoro, Grok Voice, MAI Voice 2, and Zonos in one place
  • Up to 30,000 characters per generation
  • MP3 or WAV out, the audio is yours
Open the console
One generationDeveloper API · soon
{
  "engine": "kokoro",
  "voiceId": "af_heart",
  "text": "Read this aloud."
}
Gemini FlashActs your [cues]KokoroFastest draftsGrok VoiceFive personasMAI Voice 2Style + speed controlsZonosAmerican + British

A developer API is on the roadmap. Today this runs in the studio.

03 / Measured, not marketed

An open benchmark, not vendor marketing.

We run the same script through every engine and publish what comes back: real wall-clock latency, languages, and a third-party quality score. No engine grades its own homework. You see the numbers and pick for yourself.
  • Real measured latency, dated and reproducible
  • Quality Elo from a third-party arena, not self-scored
  • Updated as engines change
Read the benchmark
Top enginesMeasured 2026-06-10
01
Gemini FlashActed scripts, audiobooks, dramatic reads
1225 Elo2770ms (measured 2026-06-10) · 24 langs
02
Grok VoiceCharacter and persona reads in English
1197 Elo2444ms (measured 2026-06-10) · 1 lang
03
KokoroDrafts, high volume, cost-sensitive jobs
1060 Elo973ms (measured 2026-06-10) · 8 langs

Quality Elo from the Artificial Analysis Speech Arena, retrieved June 10, 2026. For context, the top model of all rated is Fun-Realtime-TTS at 1228.06. Latencies are our own measured wall-clock numbers.

04 / No lock-in

Own your voice.

What you generate is yours. Export it as MP3 or WAV, ship it commercially, and read a plain-language license per generation instead of watching a credit meter. The opposite of credit lock-in.
  • Commercial rights, worldwide, no watermark
  • Export every generation you make
  • Flat pricing, not per-character credits
See pricing
LicensePer generation
Voice
Kore
Owner
You
Use
Commercial, worldwide
Export
MP3 and WAV
Watermark
None

Plain English, not a credit balance. What you generate is yours to ship.

You

Dub episode 12 into Spanish. Keep the warm read.

EPISODE-12.MP3 · 14 MB
Working...
Cantari

Transcribed, translated, and re-voiced. Same pacing, same warmth, now in Spanish.

Play
See dubbing →
05 / Dubbing

One take in. Every language out.

Dubbing

Your episode, re-voiced in another language, without losing the read.

  • Carries the performance, not just the words
  • 8 languages from one source recording
  • Your file in, your file out, yours to keep
Open dubbing →

Text to Speech

Audiobook Studio

Sound & Music

06 / How it works

From a blank script to audio you own.

Step 1: Write your script

Type or paste your text. Add bracketed [emotion] cues when you want a voice to act them.

Step 2: Pick an engine

Choose from the real engines, or follow the open benchmark to the one that fits your job.

Step 3: Generate

One request returns finished audio in seconds. Preview it right in the console.

Step 4: Own and export

Save it to your library and export MP3 or WAV. Commercial rights, no watermark.

07 / What you can make

The whole studio, at a glance.

A writer's desk with a vintage microphone, the right side falling into shadow
08 / The studio, mid-take

Write the line. Hear it acted.

Bracketed cues are stage directions, and the voice performs them.

Script

[softly]Some stories are meant to be heard.

This one starts with you.

Gemini Flash · acts the cue, not just reads it

09 / By the numbers
5
Real engines
30,000
Characters per run
~1.0s
Fastest measured
$0
To start

Latency measured 2026-06-10, wall clock to full audio

10 / Questions

The honest answers.

Plain replies to the questions everyone asks before their first generation.

Still curious about how we build? Read the about page

How does pricing work?
Flat, not per-credit. Free starts at $0 with the open-weight Kokoro engine. Creator is a flat $15 a month for all engines, and Studio is $49 for heavy production. You are never charged per character.
Who owns the audio I generate?
You do. Every generation is yours to export as MP3 or WAV and ship commercially, worldwide, with no watermark.
Which engine should I pick?
Use Gemini Flash when you want a voice to act your bracketed [emotion] cues, Kokoro when you want the fastest clean drafts, and Grok Voice for its five English personas. MAI Voice 2 adds real style and speed controls, and Zonos brings four American and British voices. The open benchmark shows the trade-offs side by side.
Is the voice real or prerecorded?
Real. The console above generates live audio through real engines on every click, and every sample on this page is real engine output, recorded unedited. No voice actors, no mockups.
What is the open benchmark?
We send the same script through every engine and publish what we measure: real wall-clock latency, languages, and a third-party quality score, each dated so you can reproduce it. No engine scores its own work.
11 / Early ears

What testers say.

cantari · creating what you love
I put [wearily] in front of one line in chapter nine and the read actually changed. Then I re-rendered that single chapter three times for a pronunciation fix and never watched a balance drain.
Maya R.Indie author, beta tester

Start with the best voice for every job.

Free to start, no credit meter. Open the console and hear it for yourself.