
Google shipped Gemini 3.5 Transcribe on August 26, 2026, and the timing makes it a genuinely helpful comparability. OpenAI had launched its personal present flagship transcription mannequin, GPT-Transcribe, simply 4 weeks earlier, on July 28, 2026. Two labs, two new transcription fashions, launched shut sufficient collectively that evaluating them truly means one thing proper now as a substitute of stacking one mannequin era towards one other.
Each corporations cut up their providing the identical approach too — one mannequin constructed for real-time streaming, one constructed for pre-recorded audio — which makes the comparability unusually apples-to-apples. This is how every obtained to the place it’s, an actual use case and dealing code for each, and a side-by-side on the numbers that really matter.
Gemini 3.5 Transcribe
Gemini 3.5 Transcribe replaces Chirp 3, Google’s earlier transcription mannequin, and the advance Google is leaning on hardest is velocity: a 70% enchancment in time-to-final-transcription over Chirp 3, alongside higher accuracy. It ships as two distinct mannequin IDs moderately than one general-purpose endpoint: gemini-3.5-transcribe-live for steady, sub-second-latency streaming by means of the Reside API, and gemini-3.5-transcribe for pre-recorded audio, conferences, name logs, and related, by means of the Interactions API.
The actual numbers, as measured by Synthetic Evaluation and cited straight in Google’s announcement: a 4.0% phrase error charge (WER) for streaming use and a couple of.6% for non-streaming. On the FLEURS multilingual benchmark particularly, Google experiences 5.50% WER streaming and 5.04% non-streaming — price noting as a separate, more durable benchmark moderately than mixing the 2 numbers collectively.
Past uncooked accuracy, the pre-recorded mannequin contains built-in multi-speaker attribution (reliably as much as three audio system, with extra listed as experimental) and word-level timestamps out of the field, no separate mannequin wanted. It additionally helps over 85 languages, acknowledges customized vocabulary, and may delegate follow-up duties like picture era or file evaluation to different Gemini fashions by way of operate calling, presently reside within the Gemini app on macOS.
OpenAI’s GPT-Transcribe
Whisper was OpenAI’s unique open transcription mannequin, outdated by gpt-4o-transcribe in March 2025, OpenAI’s first transcription mannequin truly constructed on the GPT-4o structure moderately than Whisper’s older method. GPT-Transcribe, launched July 28, 2026, is the subsequent step in that very same line, and OpenAI now recommends it forward of whisper-1, gpt-4o-transcribe, and gpt-4o-mini-transcribe for transcribing recorded speech in its unique language. Like Gemini, it splits right into a streaming sibling, gpt-live-transcribe, for steady, low-latency classes.
The numbers: on OpenAI’s personal launch benchmark towards Widespread Voice throughout 22 languages, GPT-Transcribe roughly halves whisper-1’s phrase error charge, from 40.37% all the way down to 19.27%, whereas costing 25% much less per minute than its predecessor. Pricing lands at $0.0045 per minute for file transcription and $0.017 per minute of session audio for the streaming variant. It accepts key phrase hints and a number of language hints to assist with domain-specific phrases and code-switching, and experiences which languages it detected within the audio. The sincere hole price naming straight: plain GPT-Transcribe would not do speaker diarization or word-level timestamps — these nonetheless require the separate gpt-4o-transcribe-diarize mannequin or, for timestamps particularly, the older whisper-1.
Let’s take a fast take a look at some use instances.
Utilizing Gemini 3.5 Transcribe for a Multi-Speaker Assembly
Take into account an actual state of affairs the place the built-in diarization truly earns its hold: transcribing a recorded three-person assembly and getting again who mentioned what, not only a wall of undifferentiated textual content.
from google import genai
consumer = genai.Shopper(api_key="YOUR_GOOGLE_API_KEY")
with open("meeting_recording.mp3", "rb") as f:
audio_bytes = f.learn()
response = consumer.fashions.generate_content(
mannequin="gemini-3.5-transcribe",
contents=[
{"text": "Transcribe this meeting with speaker labels and timestamps."},
{"inline_data": {"mime_type": "audio/mp3", "data": audio_bytes}},
],
)
print(response.textual content)
The request sends the uncooked audio bytes alongside a plain-language instruction, since gemini-3.5-transcribe is constructed particularly to provide speaker-attributed, timestamped output without having a separate diarization step or mannequin. For an actual assembly, which means the returned transcript already distinguishes Speaker 1, Speaker 2, and Speaker 3 with timestamps connected — output a post-call analytics pipeline may eat straight.
Utilizing GPT-Transcribe for Reside Captioning
This is a state of affairs suited to streaming: real-time captions for a reside occasion, the place latency issues greater than diarization.
import asyncio
import websockets
import json
async def stream_captions(audio_chunks):
uri = "wss://api.openai.com/v1/realtime?intent=transcription"
headers = {"Authorization": "Bearer YOUR_OPENAI_API_KEY"}
async with websockets.join(uri, extra_headers=headers) as ws:
await ws.ship(json.dumps({
"kind": "transcription_session.replace",
"session": {"input_audio_transcription": {"mannequin": "gpt-live-transcribe"}},
}))
for chunk in audio_chunks:
await ws.ship(json.dumps({
"kind": "input_audio_buffer.append",
"audio": chunk,
}))
message = await ws.recv()
occasion = json.hundreds(message)
if occasion.get("kind") == "dialog.merchandise.input_audio_transcription.delta":
print(occasion["delta"], finish="", flush=True)
This opens a persistent WebSocket connection moderately than sending one request per audio clip, which is the entire level of a streaming mannequin. Partial transcription textual content arrives as delta occasions whereas the speaker continues to be speaking, not after the recording ends. Every audio chunk will get appended to an ongoing buffer, and gpt-live-transcribe returns incremental textual content because it turns into assured sufficient to commit — precisely the conduct a live-captioning show wants to remain in sync with the speaker.
Comparability Desk
| # | Gemini 3.5 Transcribe | OpenAI GPT-Transcribe |
|---|---|---|
| Launch date | August 26, 2026 | July 28, 2026 |
| Predecessor | Chirp 3 | gpt-4o-transcribe |
| Streaming mannequin | gemini-3.5-transcribe-live |
gpt-live-transcribe |
| File/pre-recorded mannequin | gemini-3.5-transcribe |
gpt-transcribe |
| Phrase error charge | 4.0% streaming / 2.6% non-streaming (Synthetic Evaluation) | ~19.27% on Widespread Voice, down from whisper-1’s 40.37% |
| Language assist | 85+ languages | Key phrase and language hints throughout 22+ benchmarked languages |
| Constructed-in speaker diarization | Sure, as much as 3 audio system reliably | No, requires separate gpt-4o-transcribe-diarize |
| Phrase-level timestamps | Sure, inbuilt | No, requires whisper-1 |
| Streaming pricing | Not printed per-minute as of this writing | $0.017 per minute of session audio |
| File pricing | Not printed per-minute as of this writing | $0.0045 per minute |
Wrapping Up
Gemini 3.5 Transcribe’s built-in diarization and timestamps make it the stronger choose the second your use case is a gathering, a name log, or something with a number of audio system you want instructed aside — that functionality alone saves a whole second mannequin name OpenAI’s stack nonetheless requires.
GPT-Transcribe earns its place on the opposite finish: a less expensive, faster-to-integrate choice when the job is easy single-speaker transcription or reside captioning, and you do not want attribution in any respect.
Shittu Olumide is a software program engineer and technical author obsessed with leveraging cutting-edge applied sciences to craft compelling narratives, with a eager eye for element and a knack for simplifying complicated ideas. You can too discover Shittu on Twitter.
















