Skip to content
New · the open voice benchmark is liveRead it
cantari
← All posts
ResearchJuly 21, 2026 · 6 min read

Your engine ignores stage directions (unless it doesn’t)

Emotional text to speech is not a filter you slap on at the end. It is direction in the script, and only some engines will follow it.

Last updated July 21, 2026

Flat reads taught us the wrong default

Shop for text to speech and you will hear the same promise: natural, lifelike, human. The demos are often good. What they rarely show is direction. A line read once, in one mood, is not the same job as a chapter that has to whisper, then pause, then break into excitement. Most engines are built to pronounce. Very few are built to take stage directions.

That gap is why "text to speech with emotion" became a search term, and why the results are still noisy. Vendors ship a voice library labeled warm or energetic and call it emotional TTS. That is casting, not directing. Casting picks who speaks. Directing tells them how this line should land.

We run five engines in one studio. Only one of them acts bracketed cues. The other four give a plain read on purpose. This post is the honest map of that split: what a cue is, what our roster actually does with one, when flat is the right tool, and when direction earns the slower, more expressive pass.

What a cue actually is

A cue is a performance direction written in square brackets inside the script. It is not meant to be spoken. A cue-following engine reads it the way an actor reads a stage direction: it shapes the line that follows. A plain-read engine either strips the brackets or, if the pipeline is careless, speaks the word out loud. Neither performs it.

On Cantari, cues look like this in the editor: the same shape as the studio palette and the glossary entry on TTS cues:

  • [whispering] The map was wrong. The door was never locked.
  • [pause] [excited] It opens from the inside!
  • [gently] With each breath out, let the tension go.

Anything descriptive can be a cue

Emotion is only one kind. Volume, pace, and physical state all work when the engine follows direction: [quietly], [nervously], [out of breath], [calm]. The studio ships a short palette (whispering, calm, quietly, nervously, trembling, breathing heavily, excited, pause), and you can type your own. Place the cue immediately before the line it directs. A single [calm] at the top of a page will not carry a whole scene.

Punctuation still matters. A question mark lifts pitch; a comma buys a beat. Cues cover what punctuation cannot say: read this like you are holding bad news, or like you just found the key.

Same roster, two behaviors

Cue-following is a property of the engine, not of the script. Write the same directed line for every engine on our roster and you get two honest outcomes. Gemini Flash is the only engine here with followsCues set true in the registry we ship: it is built to act bracketed directions. Kokoro, Grok Voice, MAI Voice 2, and Zonos are plain-read engines. The studio badges them that way so the behavior is never a surprise.

Here is the roster with the numbers we already publish elsewhere (third-party quality Elo and our own wall-clock latency), plus the cue behavior each engine is registered for.

EngineCue behaviorQuality EloOur measured latency
Gemini FlashActs [cues]1225.132,770 ms
Grok VoicePlain read: ignores cues1196.922,444 ms
KokoroPlain read: ignores cues1060.25973 ms
MAI Voice 2Plain read: style controls, not cues1006.962,426 ms
ZonosPlain read: ignores cues1000.004,523 ms

* Quality Elo from the Artificial Analysis Speech Arena, retrieved 2026-06-10. User-vote arena ratings, not our scores.

* Latency: our own wall-clock measurement to full audio on the same routed path that serves the studio, median of 3 runs, measured 2026-06-10. Not a server SLA.

* MAI Voice 2: Score is for MAI-Voice-1; MAI-Voice-2 is not yet arena-rated.

* Zonos: Baseline rating with limited arena votes so far.

* Cue behavior from the engine registry (followsCues). MAI Voice 2 exposes real style and speed controls; that is not the same as acting bracketed script cues.

Why the expressive engine is not always the right first pass

Gemini Flash rates 1225.13 on the arena, within a few Elo of Fun-Realtime-TTS at 1228.06, the top model of roughly 85 rated, and it is the engine we route to when a script needs to be acted. It is also slower on our wall-clock test: 2,770 ms to full audio for the standard sample, against Kokoro at 973 ms.

That spread is the whole drafting argument. If you are rewriting a paragraph five times, you want the sub-second plain read. If you are locking the keeper take of a reunion scene, you want the engine that will honor [warmly] and [pause]. Using the expressive engine for every micro-edit punishes the loop that makes the writing good. Using the plain engine for the final acted pass throws away the direction you wrote.

We have written the economics of that split before: why free drafting is unlimited is the pricing half. How we measure voice latency is the clock half. This post is the direction half.

When a flat read is the correct tool

Plain-read engines are not a downgrade. They are the right instrument for jobs that should not be performed.

  • Draft loops: change a word, regenerate, listen, again, without waiting on an acted pass.
  • High-volume batch work: product updates, internal explainers, bulk IVR prompts where consistency beats drama.
  • Neutral information: safety warnings, dosage-style education, compliance lines that should not sound theatrical.
  • Persona casting without cues: Grok Voice's five personas pick a character; the read stays plain unless you switch engines for direction.
  • Styled English with knobs: MAI Voice 2 exposes speed and intensity controls. That is production control at the engine UI, not bracketed stage direction in the script.

When cues earn their keep

Direction pays off when the listener is meant to feel a turn in the line, not just receive the words.

  • Audiobook dialogue and narration that shifts mood inside a chapter: the job Audiobook Studio is built around.
  • Ads and product demos that need emphasis on the offer line without recording a voice actor for every variant.
  • E-learning and meditation scripts where [gently] or [pause] is part of the pedagogy, not decoration.
  • Any script you already mark up like a play: if you think in stage directions, put them in the text and run them on the cue-following engine.

Enhance is direction help, not a rewrite

The studio's Enhance pass inserts [emotion] cues where they help (a [warmly] before a reunion, a [pause] before a reveal) without rewriting your words. One click to run it, one click to undo. On plain-read engines the control is disabled, because stuffing cues into a script an engine will ignore is just noise. Details live in the text-to-speech studio guide.

That is also why multi-engine routing matters more than a single "emotional voice" toggle. The same manuscript can draft on Kokoro, take an Enhance pass, and finish on Gemini Flash. One studio, two jobs, no export hop between tools.

Honest limits

We would rather you read the edges here than discover them mid-project.

Only Gemini Flash acts cues on this roster today. If that engine's safety classifiers refuse a lawful prompt, the English fallback engines will still speak as plain reads, without performing the brackets. Multilingual and dubbing work also routes through Gemini Flash as the only multilingual engine, so cue-following and multilingual capability sit on the same path. We said the same thing in plain language on regulated industries.

Cues are not a guarantee of a perfect performance. They are directions an engine may follow well, follow loosely, or undershoot on a given line. The fix is the same as with a human narrator: adjust the cue, split the line, regenerate. And arena Elo is a blind quality vote, not a cue-following score; we do not pretend the arena measured stage direction.

The live comparison surface is still the open benchmark and the engines page, with the method behind the quality numbers in the open benchmark post. When the roster gains another cue-following engine, this page should say so: the registry is the source of truth, not a slogan.

Try the same line both ways

The shortest proof is not another paragraph. Open Text to Speech, paste a line with a cue you care about, and generate it on Gemini Flash. Then switch to Kokoro with the same text and listen to the plain read. The brackets either shape the performance or they disappear into a clean, undirected take.

If you are building a long piece, draft unlimited on Kokoro, mark the turns with cues, and spend the expressive pass where the story actually turns. Create a free account when you want the full studio, including Audiobook Studio for chapter-scale work. You own the files either way: that stance is unchanged from you own what you make here.

Quality Elo data: third-party, from the Artificial Analysis Speech Arena, retrieved 2026-06-10. Latency figures are our own measured wall-clock numbers, not a server SLA.

Check our work, then make your own.

The benchmark is live and the studio is free to start. Every claim above is one click from its source.