Pick the engine for the job, not the demo
A single engine is a demo. A routing table is how you ship without burning the wrong pass on every line.
Last updated July 29, 2026
The one-winner list is the wrong unit
Search "best AI voice" or "best TTS" and the page usually hands you a crown. One vendor. One model. One default voice that is supposed to cover the audiobook, the YouTube explainer, the training module, and the ad. That is marketing math, not production math.
Creators do not ship one job type. They draft, revise, cast, direct, and batch. The engine that wins a thirty-second demo is often the wrong tool for the fourth revision of a chapter, or for a plain accessibility read that should never sound theatrical. We run five engines in one studio for that reason. This post is the routing table: which job gets which engine, when the cheaper or faster one is the correct pick, and what numbers we actually publish so you can check us.
What we mean by routing
Routing is not a mystery backend. It is a decision you can make by hand: send this script to the engine whose traits match the job, not the engine whose homepage sounds the loudest.
On Cantari the traits live in the same registry the picker reads. Cue-following is a boolean. Personas are a cast list. Style and speed knobs only appear when the engine truly accepts them. Quality Elo is a third-party arena score. Latency is our own wall-clock measurement on the same path the studio uses. None of that is a vibe. It is a table.
- Ask what the listener needs: acted turn, plain clarity, a persona, a controlled style, or a specific accent.
- Ask how many times you will regenerate before the text stops moving.
- Ask whether the job is a keeper take or a draft loop.
- Then pick the engine. Switch mid-project when the job changes. That is normal.
The roster as instruments, not trophies
Here is the live roster with the numbers we already show on the open benchmark and the engines page. Quality Elo is third-party. Latency is ours. The "instrument" column is how we talk about the engine when a creator asks which one to open first.
* Quality Elo from the Artificial Analysis Speech Arena, retrieved 2026-06-10. User-vote arena ratings, not our scores.
* Latency: our own wall-clock measurement to full audio on the same routed path that serves the studio, median of 3 runs, measured 2026-06-10. Not a server SLA.
* MAI Voice 2: Score is for MAI-Voice-1; MAI-Voice-2 is not yet arena-rated.
* Zonos: Baseline rating with limited arena votes so far.
* Cue behavior from the engine registry (followsCues). Only Gemini Flash is registered to act bracketed cues today.
Job map: send the work where it belongs
The map below is the same grounding we use on use-case pages: recommended engines follow registry traits, not a sales ranking. If your project spans two rows, you are allowed to use two engines. That is the point.
* Job recommendations match the live use-case map in product copy. Always audition a short sample on your own script before locking a full project.
When the "worse" engine is the right engine
Arena Elo is useful and incomplete. Kokoro sits at 1060.25 while Gemini Flash sits at 1225.13, within a few Elo of Fun-Realtime-TTS at 1228.06. If you only ever pick the top Elo, you will punish every draft loop. On our wall-clock test, Kokoro returns the standard sample in 973 ms; Gemini Flash takes 2,770 ms. That gap is the entire drafting argument.
We have written the halves before. Why free drafting is unlimited is the allowance side. Cued TTS vs flat reads is the direction side. How to make an AI audiobook is the long-form production path. This post is the cross-job router: the same split applies to explainers, courses, and ad variants, not only novels.
Plain-read engines also win when drama would hurt the job. A dosage-style training line, a safety warning, or a screen-reader-style article pass should not perform. Route those to a clean plain instrument and stop trying to make them "more emotional."
A routing checklist you can run in ten minutes
If you want a short sequence instead of theory, use this.
- Write or paste the real script fragment, not a vendor sample line.
- Mark whether any line needs a performance turn. If yes, plan a keeper pass on Gemini Flash with [cues] (see what is a TTS cue).
- While the words still change, generate on Kokoro. Fix text at draft speed.
- If the job is a character host without stage directions, audition Grok's personas and lock one.
- If you need speed or intensity knobs more than brackets, try MAI Voice 2 controls.
- If the cast must be clearly American or British in plain English, audition Zonos.
- Render the keeper take only when the text is stable. Re-render the smallest unit that failed.
- Export the file you own. Compare engines on the benchmark when you want the receipts, not a demo reel.
What single-vendor guides usually skip
Vendor model matrices are real and useful inside one platform. What they rarely print is the awkward truth: the premium path is the wrong default for half the week. Draft on expensive expressive models and you train yourself to stop revising. Finish plain-read jobs on cue engines and you inject fake drama into material that should stay flat.
Multi-engine routing is also how you avoid tool-hopping. Exporting from one app to another just to get a second voice is how files and rights get messy. One studio, five instruments, commercial ownership on the export: that stack is the product promise, spelled out in you own what you make here.
Honest limits
Only Gemini Flash acts cues on this roster today. If that engine's safety classifiers refuse a lawful prompt, the English fallback engines still speak as plain reads. Multilingual work also rides the Gemini path, with the tradeoffs we already printed on regulated industries.
Arena Elo is a blind quality vote, not a use-case score. Our latency numbers are median wall-clock to full audio on a short fixed sample, not a promise about hour-long renders. MAI Voice 2 and Zonos carry the arena footnotes we show everywhere else; those scores can move.
Free tier still caps premium-class generation at 10,000 characters a month (about 10 minutes by house math). Kokoro drafting stays the unlimited loop so routing does not mean "burn the premium budget on every rehearse." Paid plans raise the flat allowance; they do not put you on a per-character meter.
When the roster gains another cue-following engine, or when a trait changes, the registry and this kind of page should move together. The live sources of truth remain engines, benchmark, and the engine cards inside Text to Speech.
Try two engines on the same line
The shortest proof is not another table. Open Text to Speech, paste one real line from your project, and generate it on Kokoro. Then switch to Gemini Flash with a single cue on the turn that matters, or to Grok if you are casting a persona. Listen for which instrument served the job.
When you want the full studio, create a free account. For chapter-scale work, Audiobook Studio keeps the same roster. For job-shaped landing pages, start from use cases. Pick the engine for the work in front of you. Leave the crown for the demo reel.
Quality Elo data: third-party, from the Artificial Analysis Speech Arena, retrieved 2026-06-10. Latency figures are our own measured wall-clock numbers, not a server SLA.
Check our work, then make your own.
The benchmark is live and the studio is free to start. Every claim above is one click from its source.
