← Blog

Blaze Labs · AI Voice

Reading math out loud: Blaze 2.0 Pro vs ElevenLabs

Numbers and symbols are where most AI voices quietly fall apart. We put Blaze 2.0 Pro head-to-head with ElevenLabs on the unglamorous-but-brutal task of reading mathematical formulas out loud in Vietnamese — fractions, powers, integrals, Greek letters — and measured who actually says the right thing.

Formula reading · Blaze 2.0 Pro vs ElevenLabs
Formula WER
5.6%Blaze 2.0 Pro
vs
8.4%ElevenLabs
word error rate when transcribing the spoken formula — lower is better
Human acceptance
90.1%Blaze 2.0 Pro
vs
88.4%ElevenLabs
share of clips raters marked "read correctly and natural" — higher is better
~1.5× fewer reading errors — and still ahead on human acceptance.

Why is reading a formula so hard?

Text-to-speech systems were trained mostly on prose. A formula is not prose — it is a tiny visual language that has to be linearised into the right spoken words, in the right order, in the right language. Three things go wrong constantly:

1 · Structure

A fraction is two stacked things. Said aloud it needs "… over …" with the right grouping — not the numerator glued to the denominator.

2 · Symbols

Powers, subscripts, √, ∫, Σ, ±, ≥ all have spoken names. Skip one and the whole expression changes meaning.

3 · Language

For a Vietnamese audience, "x squared" must become "x bình phương", not English read with a Vietnamese accent.

Listen for yourself

Three real Vietnamese sentences, each with an embedded formula, increasing in difficulty. For each one we show the written sentence, how it should be read aloud, then both systems — with the audio underneath.

Example 1 — Trigonometric addition formula
The sentence to read
Chúng ta biết rằng tan(α + β) = tan α + tan β1 − tan α · tan β là một công thức cộng quan trọng.

Blaze 2.0 Pro reads it loud, clear and accurate — it names “tang”, “alpha”, “beta” correctly and keeps the whole “tất cả chia cho 1 − tan α tan β” grouping intact.

Example 2 — Fourier series
The sentence to read
Khi nghiên cứu chuỗi Fourier, ta thường quan tâm đến việc chuỗi a02 + Σn=1 (an cos(nx) + bn sin(nx)) hội tụ tới hàm gốc hay không.

Blaze 2.0 Pro reads it loud, clear and accurate — it articulates “a chỉ số n” and “b chỉ số n” distinctly and reads the full “tổng từ n bằng một đến vô cùng”. ElevenLabs, by contrast, reads “a không hai”, skipping the a0/2 fraction entirely.

Example 3 — Centroid via an integral
The sentence to read
Một ứng dụng của tích phân là tìm khối tâm của một vật thể, sử dụng các công thức như My = ab x · f(x) dx.

Blaze 2.0 Pro reads it loud, clear and accurate — it says “M chỉ số y” and “tích phân từ a đến b”, without skipping the bounds or “dx”.

Where the errors happen

We transcribed every spoken formula back to text and compared it to the intended reading. Blaze 2.0 Pro leads in every category — 5.6% vs 8.4% overall — and the gap is widest exactly where structure is hardest: sums and integrals.

Fractions
4.8%
7.2%
Powers & subscripts
6.1%
9%
Sums & integrals
6.4%
9.8%
Greek & operators
5.1%
7.6%
Blaze 2.0 Pro ElevenLabsFormula WER by category — shorter is better

What human listeners accepted

Beyond raw accuracy, raters judged whether a clip was usable: read correctly, natural to listen to, and with nothing skipped. The two are close overall (90.1% vs 88.4%) — Blaze leads on correctness and on not skipping symbols, while ElevenLabs edges ahead on naturalness.

Read correctly
91%
88%
Natural prosody
89%
90%
No skipped symbols
90%
87%
In short: for reading math out loud in Vietnamese, Blaze 2.0 Pro reads it more accurately and with fewer errors — especially on fractions, subscripts and integrals — while keeping naturalness on par with ElevenLabs. It says the formula you actually wrote, in the language your audience speaks.

To be fair

These numbers don’t say ElevenLabs reads badly — they say this test deliberately aims at the hardest part of the job.

On short, symbol-light expressions — a lone number, “x plus one”, “two times three” — the gap all but disappears; both read them fluently. Most of ElevenLabs’ 8.4% comes from the multi-level cases: nested fractions, subscripts, and integral bounds. Take those away and the two are neck and neck.

And as the acceptance chart shows, ElevenLabs’ delivery is genuinely smooth — a hair ahead on naturalness. If your material is mostly narration with the occasional number, rather than dense notation, it stays a polished option.

Why this matters

If you build e-learning, exam prep, accessibility readers, or STEM tutoring in Vietnamese, the formula is the content. A voice that mangles “x squared” or skips an integral sign teaches the wrong thing. Blaze 2.0 Pro was tuned to read Vietnamese — including its mathematics — so the spoken answer matches the written one.

On methodology: 600 formulas (LaTeX source) rendered to a canonical Vietnamese reading, then synthesised by each system. “Formula WER” compares an automatic transcription of the audio against that canonical reading; “acceptance” is the share of clips three raters independently marked as correct-and-natural. The same yardstick applies to both sides. Models compared: Blaze v2.0_pro and ElevenLabs v3.

Illustrative audio clips are drawn from an internal evaluation set, used for research/illustration purposes.