Updated
Updated · TechCrunch · Oct 11
Voice AI Lacks ChatGPT Moment as 2 Key Gaps Slow Reasoning and Transcription
Updated
Updated · TechCrunch · Oct 11

Voice AI Lacks ChatGPT Moment as 2 Key Gaps Slow Reasoning and Transcription

1 articles · Updated · TechCrunch · Oct 11

Summary

  • PolyAI CTO Shawn Wen said voice AI still falls short of a breakout moment because models cannot reason fast enough to keep conversations natural, even after full-duplex systems learned to speak while listening.
  • Otter CMO Alex Gay said weak speaker identification, intent capture and transcription accuracy still undermine automation, because errors in the first ASR layer cascade into flawed summaries and actions.
  • Enterprise use cases raise a second hurdle: trust. Wen said customer-service agents must sound confident enough to handle problems without a human, while Gay said meeting avatars need humanlike emotional expression to support real debate.
  • Both companies also stressed transparency, saying users should be told when they are speaking to AI or being recorded as voice tools spread across call centers, note-taking and digital-twin meeting products.

Insights

As AI latency drops below human perception, how will we instantly know if the voice on the phone is real?
If voice AI models are speaking while listening, what hidden data are they capturing during our silent pauses?
Why might the push for emotionally nuanced AI digital twins actually destroy user trust rather than build it?