When comparing text-to-speech technology, there are a lot of things that are measured to make it apples to apples: time to first token, mean opinion scores, naturalness ratings from a listening panel, p95 latency. Those numbers are real, but they didn't tell us anything about the metric that matters to us – their ability to complete an employment verification unaided.
So we measured that instead, and nothing else.
TL;DR
We routed 19,363 randomized employment verification calls across four TTS voice models over a week of fully overlapping traffic, and scored every model on one outcome: did the first call end with a completed verification. Bland Speech v3 finished at 4.19%, against 3.33% for MiniMax and 3.21–3.30% for Deepgram's Aura and Flux. That's a 27% relative lift over the rest of the field combined, and it's unlikely to be noise (p ≈ 0.0015). We didn't measure latency, time to first audio, hangup rate, or how human the voices sounded to us.
The Only Metric We Scored

A verification call has exactly one job. Reach somebody at the employer who can confirm employment, get the answer, and record it. Everything else is instrumentation.
So the scoreboard had one column: the share of dialed calls that ended with a completed verification on the first try. A call either produced the answer or it didn't.
We deliberately ignored the metrics the category usually competes on:
- Time to first audio. A voice that starts 150ms sooner doesn't help if the HR coordinator transfers you to a voicemail box anyway.
- Mean opinion score and naturalness panels. These measure what listeners say about a voice in isolation, not what they do when it calls their workplace.
- Hangup rate. A useful diagnostic, but a call that stays connected for four minutes and ends without an answer is still a failed call.
- How it sounded to us. We had strong opinions here. They turned out to be only half right, which is exactly why we didn't score on them.
The Four Models

We'd been evaluating voice models for several months across a wider field. For this test we narrowed to four:
| Model | Vendor | Voice |
|---|---|---|
| Bland Speech v3 | Bland AI | Bland's conversational phone model |
| MiniMax | MiniMax | Production MiniMax voice |
| Aura 2 | Deepgram | aura-2-asteria-en |
| Flux | Deepgram | flux-sienna-en |
Traffic was randomized at call placement into three even arms — Bland, MiniMax, and Deepgram — with the Deepgram arm splitting again between Aura and Flux.
Our hypothesis, before any of this was that the most human-sounding model would win. Bland markets Speech v3 on exactly that basis, and it's the one we'd have picked from listening alone.
Results
Five days, Aug 31 to Sep 4, 2026. All four arms live simultaneously across the whole window.
| Model | Calls | Completed verifications | Rate |
|---|---|---|---|
| Bland Speech v3 | 6,637 | 278 | 4.19% |
| MiniMax | 6,363 | 212 | 3.33% |
| Deepgram Flux | 3,125 | 103 | 3.30% |
| Deepgram Aura 2 | 3,238 | 104 | 3.21% |
Bland against the other three pooled: 4.19% versus 3.29%, a 27% relative improvement on 19,363 calls.
The three losing models finished within 0.12 points of each other. That flatness is worth as much as the headline: on this metric, MiniMax and both Deepgram generations are effectively the same model for our performance purposes.
A Note on Bland Speech v3
Bland built toward a different goal than most TTS vendors, which is worth understanding.
Bland Speech v3 is built for phone conversations rather than for polished narration. Their framing is that it keeps the breaths, stumbles, and pauses real people make, treating those as features rather than artifacts to clean up. It streams first audio in 199ms, supports voice cloning from about ten seconds of sample audio, and takes performance tags for things like a laugh or a throat-clear. On their Audio Realism Bench it ranks second at an Elo of 1,384, behind only recordings of actual humans at 1,501.
We can't independently verify their benchmark, and we didn't try. What we can say is that the model that optimizes hardest for sounding like a person on a phone call is also the one that got the most answers out of HR departments. Those two facts are consistent with each other. Our test doesn't prove one caused the other.
What Superunit Does, and Why Voice Matters Here

Superunit uses AI agents to complete employment verifications —the manual kind, where a call, email, or fax is needed to contact employer.
Roughly 2/3 of these verifications still need to be done this way.
That's the work our AI agents do. They place the call, work through the menus, handle the transfer to whoever actually holds the records, and log the whole thing with timestamps so the verification survives a challenge later. Across all products Superunit has completed over 200,000 verifications. On employment verification specifically we complete 70% of files at a 0.82-day average turnaround.
AI voice isn't a cosmetic layer here – it's core to the outcome. An HR coordinator who finds the caller confusing, weird, or hard to interrupt hands over less, transfers less willingly, and hangs up sooner.
How Superunit Ran the Test
We picked the voice by call, at random. Not customer by customer. That way each voice got its fair share of easy employers and hard ones, morning calls and afternoon calls, Mondays and Fridays. Nobody got the good list.
We swapped the voice and nothing else. All other infrastructure was kept the same. So this tells you which voice works best inside our product. It does not tell you which company builds the best voice agent overall.
What We Changed
Bland Speech v3 is now the default voice on Superunit verification calls.
The broader lesson is about the scoreboard rather than the winner. Every vendor in this category publishes latency and naturalness numbers, and we've no reason to think any of them are wrong. They just weren't predictive. The model with the cleanest greetings lost. The ranking we would have guessed from listening was half right, and we only know which half because we picked one outcome metric and let thousands of calls settle it.
If you're choosing a voice for work that has an actual outcome attached, run the same test. Pick the metric that maps to what you're selling, randomize at the call, and ignore everything on the vendor's own benchmark page until your own numbers agree with it.
For more on where the human stays in the loop, see when a verification call needs a person, or what employment verification involves if you're new to the category.
