Voice is giving AI the timing and social cues that make an exchange feel increasingly human.
You start asking a question, pause halfway through, and change direction. The AI waits. When it answers, you interrupt. It stops, follows the correction, and continues in a tone that responds to yours.
Together, those behaviors make the exchange feel like a conversation. We respond to the timing and social cues without knowing what the AI understands.
AI has produced human-sounding sentences for years. Voice adds another layer. Timing, pitch, emphasis, hesitation, interruption, and small acknowledgements all carry information that never appears in a transcript. They also supply many of the cues we use to decide whether someone is listening, confident, uncertain, impatient, or emotionally present.
The companies building these models expect those cues to change how we use them. OpenAI’s product lead for ChatGPT Voice, Atty Eleti, described voice as a possible replacement for much of the interface around computing:
“Over time, we think this will also unlock the ability to use voice as kind of the primary interface to computing.”
Google DeepMind says, “We believe conversation will be a key way we interact with AI.” A Jabra study led by a professor at the London School of Economics goes further: “Voice AI is set to become the default interface by 2028.”
These companies sell the technology they are predicting. Their forecasts tell us where they are investing, not what people will inevitably choose. The more interesting question is already here: as the interface gains a voice, do we experience the AI itself as more human?
The Cues We Recognize
Even simple computers can trigger social responses. Stanford researcher Clifford Nass demonstrated this in laboratory studies more than 30 years ago. Participants applied familiar social rules to machines even when they knew they were evaluating software. His 1992 description of the work noted that “computer-sophisticated individuals exhibit anthropomorphism”.
Anthropomorphism does not require a sincere belief that a machine is alive. It is the ordinary habit of applying human expectations to something that presents recognizable social behavior. We thank a voice assistant, become frustrated when it misunderstands us, and interpret a delayed answer differently from an immediate one.
A 2025 meta-analysis of human-like cues in text chatbots combined 800 effects from 199 datasets involving 41,642 participants. Human-like cues produced a small positive overall effect on responses including perception, emotion, rapport, trust, and behavior. The effect varied by context and sometimes became negative when the bot failed, the user was angry, or the interaction felt uncanny.
Voice brings several familiar parts of conversation into the exchange, and our social habits fill in the rest.
More Than Spoken Text
We have been speaking to computers for years. Siri, Alexa, automated phone menus, and dictation all accept spoken input. Much of that technology places audio around a text exchange. One component detects when a person has stopped speaking, another transcribes the recording, a language model generates text, and a speech synthesizer reads it aloud.
That sequence preserves the words while losing some of the interaction. Hesitation, emphasis, pacing, laughter, and uncertainty may disappear in transcription. The model waits for a finished message before it responds. The person waits for the model to finish before speaking again.
We call this half-duplex interaction. Human conversation is less orderly. People overlap, correct themselves, respond before a sentence is complete, and make small sounds to show they are listening. A response that arrives immediately can mean something different from the same words after a long pause.
Meta’s full-duplex dialogue research gives language models an internal sense of time so they can handle overlapping speech, backchannel responses, and dynamic turn-taking. Kyutai’s Moshi processes and produces audio directly on separate streams. It listens while speaking and reports practical latency of about 200 milliseconds.
Sam Altman described the intended experience in a 2023 interview with TIME:
“You’ll be able to do this with two-way voice, and it’ll feel real time.”
“Real time” changes the substance of the exchange along with its speed. A text prompt arrives as a completed object. Full-duplex voice has to interpret the exchange while it is still taking shape. The model responds to when we speak, how we speak, and whether we are still speaking at all.
A Human Voice Changes Trust
The same information can feel different when a voice delivers it. A 2024 experiment involving 2,165 people compared a text-only pseudo-LLM with one that presented the same kind of responses through text and synthesized speech. Participants rated the system with speech as more human-like. They also rated its information as more accurate.
The synthesized voice increased the credibility participants gave the information.
The Jabra study found a similar tension. Trust in AI rose by 33% when participants interacted through voice rather than text. Yet the people who preferred voice performed about 20% worse on certain tasks, often because they had difficulty articulating complex thoughts aloud or encountered accuracy problems.
The effect varies by context. Other experiments have found no meaningful difference, and the larger meta-analysis reached a modest overall result. Perceived humanness can still alter the credibility of an answer without improving the answer itself.
Engagement Feels Different
Voice can make AI available while we are walking, driving, cooking, or looking through a camera. It can also make the exchange feel more attentive. The model acknowledges an interruption, responds to frustration, remembers the earlier thread, and answers without requiring another carefully composed prompt.
Those behaviors can resemble listening and emotional presence. OpenAI and MIT examined that effect in a four-week randomized trial involving nearly 1,000 participants and more than 300,000 messages. The results were mixed. Brief use of voice was associated with better well-being in some conditions, while prolonged daily use was associated with worse outcomes. Heavy use across both voice and text correlated with emotional dependence and loneliness.
The study showed that emotional engagement was uncommon overall, concentrated among a smaller group of heavy users, and affected by how long and why people used the product.
The question of becoming more human matters because a model can affect our feelings without having feelings of its own. It can perform the recognizable behaviors of attention, patience, humor, reassurance, or concern. We respond to what the interaction gives us.
Someone in the Conversation
Voice can become a primary way people engage with AI while text remains essential. The more plausible interface moves between them. We speak when immediacy, mobility, or the flow of an unfinished thought matters. We return to text when we need precision, privacy, comparison, or a durable record.
Someone asks a question while cooking, interrupts when the answer heads in the wrong direction, and keeps talking while moving through the room. The voice waits, adjusts, and answers in a tone that fits the moment. The person knows it is software. The conversation gives them a familiar way to respond. Shall I put crackers on the list then?