
Tech • IA • Crypto
OpenAI’s upgraded voice mode enables real-time, interruptible conversations that can control computers, run parallel tasks, and integrate across apps, signaling a shift toward fully agentic AI assistants.
The new Voice Mode replaces turn-based interaction with simultaneous listening and speaking, allowing users to interrupt mid-sentence without breaking context. This creates a more natural, human-like dialogue compared to earlier systems that required users to wait for responses to finish.
The feature is embedded in the ChatGPT desktop app, particularly within Work Mode and coding environments like Codex. Unlike the web or mobile-only versions, the desktop setup enables direct access to files, folders, and system-level actions, turning the assistant into an active operator rather than a passive responder.
Users can issue commands from a phone that execute on a connected computer in real time. Through a local connection setup, tasks such as file retrieval, renaming folders, or launching workflows can be triggered remotely, effectively turning the mobile app into a voice-based control hub.
A single voice command can initiate multiple processes simultaneously. The system can manage separate threads—such as searching files while recalling past activity—demonstrating early forms of multi-agent orchestration coordinated through natural language.
Voice mode can interact with connected services like Gmail, Google Calendar, and Slack, as well as third-party integrations. Using middleware such as Zapier MCP, the system can extend to over 9,000 applications, enabling workflows that span entire digital ecosystems.
The assistant can retrieve and act on information from previous conversations, including tasks completed days earlier. In one case, it revisited prior travel planning data to recommend updated flight options and initiate booking through a connected service.
With screen access enabled, the AI can interpret on-screen content and provide step-by-step assistance. This includes identifying software interfaces and guiding users through actions like applying video edits in real time.
Through integrations like Higgsfield MCP, users can generate multiple media assets simultaneously using voice alone. Examples include producing product videos and automatically assembling them into a functional animated website without manual prompting.
The system supports recurring tasks, such as generating daily summaries of emails and calendar events. These automations can be configured and executed entirely through voice, reducing routine administrative work.
By combining voice interaction, tool use, memory, and automation, the system behaves increasingly like an autonomous assistant. It can plan, execute, and adjust tasks dynamically, hinting at broader shifts in how users may interact with operating systems.
The enhanced voice mode transforms conversational AI into a hands-on digital operator, blending natural interaction with real task execution across devices and applications.