NEW
OpenAI releases new transcription models for audio and live speech
GPT Transcribe converts recorded speech to text, while GPT Live Transcribe produces text as speech arrives. Both models are released via the OpenAI API, but they are not new buttons for regular ChatGPT users.
Publicerad 29 July 2026, 08.09

OpenAI, the company behind ChatGPT, has released two models that turn speech into text. GPT Transcribe works with completed audio files and completed segments of a Realtime call. GPT Live Transcribe instead produces text continuously as the audio arrives.
The difference is noticeable in use. A newsroom can submit a completed interview to GPT Transcribe and receive a first text draft. A meeting or customer service app can use GPT Live Transcribe to display text while participants are still speaking.
Both models can get free background information, important keywords and several expected languages. It can help the system when the sound contains names, technical terms or language changes. On the other hand, OpenAI does not publish an independent Swedish quality test in the launch documentation.
GPT Transcribe costs $0.0045 per audio minute. GPT Live Transcribe costs $0.017 per minute. The prices apply to use via the OpenAI API, i.e. the technical connection that developers and companies use to build models into their own services.
This is not a new button in ChatGPT for all users. A business needs to build or use a service that is connected to the OpenAI API. It also needs to decide how the audio is sent, stored and removed.
A transcript is a working document, not a secure protocol. Names, numbers, quotes and words that are hard to hear can be wrong. Audio from meetings, interviews or care may also contain personal data and other sensitive information.
AI Kollen has not been able to run a test of its own because the project's publishing environment lacks API access to the new models. The article therefore describes the confirmed release, features and prices in OpenAI's documentation, not a separate result for Swedish.
Därför spelar det roll
Speech-to-text can reduce the time spent on a first draft of an interview, lecture or meeting. Live transcription can also make a conversation easier to follow. The usefulness depends on how well the model copes with the real sound and on someone checking the text before it is used as a quote, meeting record or basis for a decision.
Det här kan du göra
- Select complete transcription when the entire audio file can be processed afterwards. Choose continuous transcription when the text needs to be displayed during the call.
- Test with your own Swedish names, technical terms, dialects and background sounds before using the service in production.
- Decide who can record, how participants are informed and when audio and text are to be deleted.
- Check names, numbers and quotes against the original recording before publishing the text or guiding a decision.