Blaze Labs · AI Voice
Reading math out loud: Blaze 2.0 Pro vs ElevenLabs
Numbers and symbols are where most AI voices quietly fall apart. We put Blaze 2.0 Pro head-to-head with ElevenLabs on the unglamorous-but-brutal task of reading mathematical formulas out loud in Vietnamese — fractions, powers, integrals, Greek letters — and measured who actually says the right thing.
Why is reading a formula so hard?
Text-to-speech systems were trained mostly on prose. A formula is not prose — it is a tiny visual language that has to be linearised into the right spoken words, in the right order, in the right language. Three things go wrong constantly:
A fraction is two stacked things. Said aloud it needs "… over …" with the right grouping — not the numerator glued to the denominator.
Powers, subscripts, √, ∫, Σ, ±, ≥ all have spoken names. Skip one and the whole expression changes meaning.
For a Vietnamese audience, "x squared" must become "x bình phương", not English read with a Vietnamese accent.
Listen for yourself
Three real Vietnamese sentences, each with an embedded formula, increasing in difficulty. For each one we show the written sentence, how it should be read aloud, then both systems — with the audio underneath.
Blaze 2.0 Pro reads it loud, clear and accurate — it names “tang”, “alpha”, “beta” correctly and keeps the whole “tất cả chia cho 1 − tan α tan β” grouping intact.
Blaze 2.0 Pro reads it loud, clear and accurate — it articulates “a chỉ số n” and “b chỉ số n” distinctly and reads the full “tổng từ n bằng một đến vô cùng”. ElevenLabs, by contrast, reads “a không hai”, skipping the a0/2 fraction entirely.
Blaze 2.0 Pro reads it loud, clear and accurate — it says “M chỉ số y” and “tích phân từ a đến b”, without skipping the bounds or “dx”.
Where the errors happen
We transcribed every spoken formula back to text and compared it to the intended reading. Blaze 2.0 Pro leads in every category — 5.6% vs 8.4% overall — and the gap is widest exactly where structure is hardest: sums and integrals.
What human listeners accepted
Beyond raw accuracy, raters judged whether a clip was usable: read correctly, natural to listen to, and with nothing skipped. The two are close overall (90.1% vs 88.4%) — Blaze leads on correctness and on not skipping symbols, while ElevenLabs edges ahead on naturalness.
To be fair
These numbers don’t say ElevenLabs reads badly — they say this test deliberately aims at the hardest part of the job.
On short, symbol-light expressions — a lone number, “x plus one”, “two times three” — the gap all but disappears; both read them fluently. Most of ElevenLabs’ 8.4% comes from the multi-level cases: nested fractions, subscripts, and integral bounds. Take those away and the two are neck and neck.
And as the acceptance chart shows, ElevenLabs’ delivery is genuinely smooth — a hair ahead on naturalness. If your material is mostly narration with the occasional number, rather than dense notation, it stays a polished option.
Why this matters
If you build e-learning, exam prep, accessibility readers, or STEM tutoring in Vietnamese, the formula is the content. A voice that mangles “x squared” or skips an integral sign teaches the wrong thing. Blaze 2.0 Pro was tuned to read Vietnamese — including its mathematics — so the spoken answer matches the written one.
On methodology: 600 formulas (LaTeX source) rendered to a canonical Vietnamese reading, then synthesised by each system. “Formula WER” compares an automatic transcription of the audio against that canonical reading; “acceptance” is the share of clips three raters independently marked as correct-and-natural. The same yardstick applies to both sides. Models compared: Blaze v2.0_pro and ElevenLabs v3.
Illustrative audio clips are drawn from an internal evaluation set, used for research/illustration purposes.