Claude Voice Mode Adds Opus 5 and Sonnet 5 Support, Stays Turn-Based

Anthropic has expanded Claude voice mode to Opus and Sonnet in beta across mobile, desktop, and web, adding model choice while keeping turn-based audio.

TL;DR
  • Feature Expansion: Claude voice mode now supports Opus 5, Sonnet 5, and Haiku 4.5, with model switching during a conversation.
  • Plan Limits: Paying users gain stronger models and connected-app actions, while free users retain Haiku and one connector.
  • Voice Format: Claude still listens, pauses to process, and replies in turns rather than using continuous interruptible audio.
  • Product Limits: The beta spans mobile, desktop, and web, but excludes Anthropic’s Cowork workspace product and Claude Code coding agent.

Anthropic has expanded Claude voice mode to Opus and Sonnet. Plan users can now use more capable models in voice mode, while the speech layer remains turn-based. Users can choose a model and change it during a conversation.

Claude’s voice mode upgrade broadens model selection across a beta for mobile, desktop, and web. Paid plan accounts can combine Anthropic’s higher-capability Opus model family or balanced Sonnet model family with connected applications. Free accounts retain the speed-oriented Haiku model family and can authorize one connected app.

Anthropic’s model expansion makes Opus and Sonnet available in voice mode, but it does not introduce a new speech engine.

A listen-pause-reply sequence defines each spoken exchange with Claude. Continuous two-way audio instead lets a user interrupt an assistant while it is speaking. Permissioned links to services such as Gmail and Google Calendar let the stronger models handle more complex spoken requests and then act on connected accounts.

Stronger Models Reach Connected Accounts

Claude carries model selection across its text and voice interfaces. Voice mode starts with the model last used in text chat. Because voice and text share the same conversation, a user can change models or move between them without restarting, allowing a complex request to begin aloud and continue on screen.

Model switching gives users a practical speed-versus-capability choice without opening a new chat. Haiku 4.5 remains the quicker option for straightforward requests, while Opus 5 and Sonnet 5 can address work that requires more involved reasoning. Retaining the same conversation also keeps earlier instructions available when a user changes the model or moves from speech to text.

Connected services extend the feature beyond spoken question answering. Claude users can use connected tools, including Gmail, Google Calendar, and Slack, and can also use Canva. Those account connections build on Anthropic’s Google Workspace integration from 2025, with Claude requesting permission before it invokes a connected tool.

 

After a user grants tool access, Claude can for example move a calendar meeting by 30 minutes. Model reasoning interprets the spoken request, while the connector submits the change to Google Calendar. Connector permissions ultimately control what happens outside the chat, even when Opus or Sonnet processes a more involved instruction.

Voice mode can also turn a spoken request into a saved email. In that workflow, model capability affects how Claude interprets the request, but the connected service and its permissions determine whether the resulting action can proceed.

Claude saves voice conversations to chat history, giving users a text record for later review. Because the transcript stays in the chat, work that began hands-free can continue through the regular interface without forcing the user to reconstruct the request.

Claude Voice mode also includes hands-free and push-to-talk modes. It then listens for a turn without a held button, while push-to-talk receives audio only while the button is pressed. Users can choose between a more conversational control and tighter control over when Claude receives audio.

A Smarter Assistant, Not a New Speech Engine

Claude listens to one turn, processes the request, and answers. OpenAI’s GPT-Live uses continuous two-way audio designed to handle interruptions, while Google’s Gemini Live uses a turn-based approach on smartphones. Claude and Gemini preserve discrete turns, whereas GPT-Live permits interruption; no shared benchmark establishes a performance winner.

Claude’s original voice-mode rollout began in 2025. Immediately before Anthropic expanded the feature, it ran only on Haiku.

Anthropic chose the earlier voice model for speed. Adding more powerful models closes part of the capability gap between spoken and text requests. Despite the added reasoning capacity, Claude’s audio exchange is still neither continuous nor interruptible.

Claude voice mode will support 11 language variants: English, French, German, Hindi, Indonesian, Italian, Japanese, Korean, Brazilian Portuguese, Latin American Spanish, and Spain Spanish. Users switch languages manually, and each supported language retains the same model choices and listen-process-reply interaction.

Voice mode so far remains unavailable in Claude Cowork or Claude Code, Anthropic’s workspace product and coding agent.

Markus Kasanmascheff
Markus Kasanmascheff
Markus has been covering the tech industry for more than 15 years. He is holding a Master´s degree in International Economics and is the founder and managing editor of Winbuzzer.com.
Subscribe
Notify of
guest
0 Comments
Newest
Oldest Most Voted
Inline Feedbacks
View all comments