AI News (2026/5/8): OpenAI Launches Three Real-Time Voice Models

2026年5月8日 04:50

Executive Summary:

OpenAI has launched three new real-time voice models: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. These models are available through the Realtime API and offer GPT-5 level reasoning and tool calling capabilities, support for real-time translation in over 70 languages, low-latency transcription, and preservation of tone and emotion.

Details

On Friday, May 8, OpenAI announced the launch of three new real-time voice models: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. These models are available to developers through the Realtime API, aiming to enhance the efficiency and quality of voice processing. According to the official statement, these models have made significant improvements in reasoning capabilities, multilingual support, and low-latency transcription, meeting the needs of various scenarios.


Key Points

  • GPT-Realtime-2: This model possesses GPT-5 level reasoning and tool calling capabilities, enabling complex logical analysis and task execution in real-time conversations. For example, it can automatically summarize key points and generate action items during meetings, or provide more intelligent conversation support in customer service scenarios.

  • GPT-Realtime-Translate: It supports real-time translation in over 70 languages, with a cost of approximately 0.25 yuan per minute, which is about a hundred times cheaper than human simultaneous interpretation. This feature makes cross-language communication more convenient and cost-effective, suitable for international conferences, multilingual customer service, and online services supporting global users.

  • GPT-Realtime-Whisper: It achieves low-latency voice transcription, capable of converting speech to text within a few milliseconds. This characteristic is crucial for applications requiring immediate feedback, such as real-time subtitles and voice assistants. Additionally, the model can preserve the speaker's tone and emotion, making the transcribed content more natural and realistic.


AI-ALL In-Depth Analysis

The three real-time voice models launched by OpenAI not only reach new heights in technical specifications but also demonstrate significant potential in practical applications. The GPT-5 level reasoning capability of GPT-Realtime-2 makes it an ideal choice for enterprise-level applications, especially in scenarios requiring quick responses and complex task handling. The low cost and multilingual support of GPT-Realtime-Translate will greatly promote global communication, reducing the costs and time consumption associated with language barriers. The low-latency transcription feature of GPT-Realtime-Whisper provides more efficient technical support for applications like real-time subtitles and voice assistants.

The introduction of these models not only enriches OpenAI's product lineup but also injects new vitality into the AI ecosystem. By opening the Realtime API, developers can more easily integrate these features into their own applications, further advancing the adoption and application of AI technology across various industries.

About AI News

We use AI technology to automatically crawl and filter the latest AI news from around the world, providing you with the most valuable industry updates.