ON ARTIFICIAL INTELLIGENCE -Voice assistants spent twenty years teaching us to talk like robots. What finally changed comes down to two-tenths of a second.
For twenty years I talked to computers the way you talk to a toddler holding a marker near a white wall: slowly, clearly, and with very low expectations.
"Call. Mom.” robotic voice: “Calling Tom.” You say “NO.” robotic voice: “Calling Tom on speaker.”
We all did it, and we all developed the Voice that flat, over-enunciated robot cadence. I call it my radio voice. But it still didn’t work well. Something changed recently, and it wasn’t the microphone.
THE 200-MILLISECOND PROBLEM
Human conversation runs on a clock most people never notice. Researchers recording ten languages across five continents found that the gap between one person finishing and the next starting averages roughly 200 milliseconds (Stivers et al., PNAS, 2009). Two-tenths of a second.
Here’s the part that should bother you, planning a single word takes the brain something like 600 milliseconds. Math only works if we’re all composing our replies while the other person is still talking. Rude? Yes. Universal? Also, yes. My wife has been telling me this for years.
Old voice assistants could not hit 200 milliseconds, because they weren’t one thing they were three things in a trench coat. Your speech got converted to text. The text got handed to a model. The answer got converted back into speech. Every hop added delay, and the total landed north of a second. Your brain reads a pause that long as a walkie-talkie, so you talk like it’s a walkie-talkie. Over.
WHAT ACTUALLY CHANGED
Newer systems cut out the middlemen. Audio goes in, audio comes out, no transcript in between. OpenAI’s GPT-4o, released in May 2024, answered speech in as little as 232 milliseconds and averaged about 320 squarely inside human range. Google’s Gemini Live and others are in the same neighborhood now.
Get under about half a second and something flips in your head. You stop announcing and start conversing. You can interrupt it mid-sentence the engineers call this “barge-in” and it stops, like a person would, instead of finishing its paragraph like a man who has waited all week to tell that story. I’m told I do this. Twice.
Because these models hear audio instead of reading a transcript, they also catch what the transcript throws away pace, emphasis, whether you’re annoyed. Fair warning, though the little “ums” and breaths you hear coming back are performance, not deliberation.
It isn’t thinking. It’s doing an impression of thinking the same way I nod through my boy’s explanation of a video game.
WHERE IT STILL FALLS OVER
Noisy rooms. Humans do the cocktail party trick, locking onto one voice in a crowd, and machines are still worse at it than you are on your worst day.
Accents and dialects. A study of five major commercial speech systems found word error rates averaging 35% for some accents versus 19% for others on identical speech (Koenecke et al., PNAS, 2020). Better since. Not solved.
Knowing when you’re done. Engineers call it endpointing. Pause to find the right word and it will confidently answer the question you hadn’t finished asking.
Confident wrongness. A warm voice makes a wrong answer sound friendlier. That’s a bug in us, not in it.
WHY THIS MATTERS AT WORK
You speak about 150 words a minute. You thumb-type about 40. Stanford clocked voice input at roughly three times faster than a phone keyboard, and that’s before you count the times autocorrect turns a client’s name into a snack food. Voice wins wherever hands are busy: the nurse mid-round, the tech up a ladder, the estimator walking a roof, the driver who should not be looking down.
Which is exactly the catch. Talking is the easiest way ever invented to hand your data to a stranger. If a call is recorded, transcribed, and stored, then patient details, CUI, and client files are now sitting somewhere with a retention policy nobody reads. Ask where the recordings live, who trains them, and how long they stay before they’re deployed. Not after.
So, can we talk to AI? Genuinely, finally, yes. Just talk to it the way you’d talk to a sharp new hire plainly and check their work (really check their work).
I asked one for a joke about a broken pencil. It said, “It’s pointless.” I laughed out loud. My boys did not. They just shook their heads in disappointment.
Progress.
SOURCES: Stivers, T., et al. “Universals and cultural variation in turn-taking in conversation.” PNAS 106(26), 2009. · Indefrey, P. & Levelt, W. J. M. “The spatial and temporal signatures of word production components.” Cognition 92, 2004. · OpenAI. “Hello GPT-4o.” May 2024 audio response in as little as 232 ms, averaging 320 ms. · Koenecke, A., et al. “Racial disparities in automated speech recognition.” PNAS 117(14), 2020. · Ruan, S., et al. “Speech is 3x faster than typing for English and Mandarin text entry on mobile devices.” Stanford University, 2016.
Originally published by Mark Putiyon on LinkedIn. Join the discussion there.
Read on LinkedInFounder of Technology Innovation Partners — 30+ years helping businesses secure and modernize their IT.




