Vol. 1 · Edition 033Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

0.70s
First audio response
Grok Voice Think Fast 2.0 · xAI · July 30, 2026
Product
By Sam Taylor with Samwise

On 0.70-second first audio response, 1.5–2.0× transcription gains over Deepgram and ElevenLabs, and what Starlink's A/B test results suggest about where production voice AI is actually landing

xAI's voice model runs Starlink's phones now. The latency numbers explain why.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

02

Pro (hyped)

01

← Anti-AI · Pro-AI →

There's a threshold in voice AI where the latency stops reading as "AI" and starts reading as "slightly slow human." I don't know exactly where it is, maybe 600 milliseconds, maybe 800. But I know we're converging on it, and xAI's Think Fast 2.0 is one of the cleaner pieces of evidence.

Grok Voice Think Fast 2.0 launched July 30. The headline metric is the first-audio-response time: 0.70 seconds, down from 1.25 seconds for Think Fast 1.0. That's a 44% reduction. In conversational voice, that gap is the difference between "this feels like a phone tree" and "this feels like a real exchange." Think Fast 1.0 was on the wrong side of it. Think Fast 2.0 might be on the right side.

The Starlink deployment is the actual data point I want to talk about. xAI ran Think Fast 2.0 on real Starlink customer support and sales calls in an A/B test against Think Fast 1.0. They say they saw improved sales conversion rate and improved support containment. They didn't publish the exact percentage lift, which I'll come back to. But they ran it on live traffic and liked the results enough to lead with it in the announcement. That's not nothing.

0.70s
First audio response time — Grok Voice Think Fast 2.0

→ Source: xAI announcement, July 30

Source spread

Pros & cons

What's real:

  • 0.70 seconds first audio response is a meaningful target for production voice agents. The perception of conversational naturalness is sensitive to this number in a nonlinear way. You can engineer around 2 seconds with clever patterns. Below 1 second you just need to be fast.
  • The 1.4× transcription accuracy gain vs Think Fast 1.0 compounds. Worse transcription means more downstream hallucination, misunderstood commands, and failed tool calls. A 1.4× improvement at the input layer improves every metric downstream of it.
  • xAI claims 1.5–2.0× transcription accuracy over Deepgram Nova 3 and ElevenLabs Scribe v2. Those are the incumbent vendors for standalone speech-to-text in agent pipelines. If xAI holds that gap, the integration case for a unified voice-plus-reasoning model strengthens.
  • $0.08/minute API ($4.80/hour) is competitive for a model this capable. For a contact-center-at-scale use case, per-minute pricing is easier to model than per-token.
  • Builders using grok-voice-latest automatically upgraded on August 5. No migration required.

What deserves a side-eye:

  • The Starlink A/B test is cited in the announcement without any actual numbers. xAI said "significant increase in sales conversion rate and support containment rate." That phrasing says nothing. 1% is a significant increase. 40% is also a significant increase. Without the number, the reference reads as proof-of-deployment (real) rather than proof-of-effect (unverified).
  • The 82.9% on Artificial Analysis's speech-to-speech quality index is self-referenced in the launch copy. Artificial Analysis is an independent evaluator, which helps, but the specific test conditions and scoring methodology matter and aren't linked in the announcement.
  • "More intelligent" in the announcement means smarter reasoning inside the voice loop. What that actually means for your specific use case depends on what your voice agent is doing. Customer support and sales (Starlink's use case) reward conversational repair skills. Code review voice agents reward something different. The intelligence gain doesn't transfer uniformly.

Samwise's take

What builders need to know

  • If you're on grok-voice-latest, you already upgraded on August 5. No action needed. Check your transcription accuracy and latency metrics in the next 48 hours to confirm the gains apply to your traffic.
  • Run the Deepgram/ElevenLabs transcription benchmark yourself before switching. The 1.5–2.0× accuracy claim is xAI's self-report; your audio environment, accent distribution, and domain vocabulary may produce a different result.
  • The Starlink A/B test result is deployment evidence, not uplift evidence. xAI didn't publish the specific conversion or containment numbers. Treat it as a signal that the model is production-stable, not as a prediction of your specific metric improvements.
  • $0.08/minute pricing means roughly $4.80/hour of voice agent runtime. Model your call-volume economics before committing at scale. At high volume, per-minute beats per-token for predictability.
  • The 0.70-second target requires the full pipeline. If your voice application has latency elsewhere (slow tool calls, slow API round-trips, client-side processing overhead), Think Fast 2.0's first-audio-response speed won't help you until you've addressed those bottlenecks too.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.