GPT-5 Turbo is a distilled version of OpenAI's GPT-5 model optimized for speed and cost efficiency. It delivers comparable quality to the full GPT-5 while running three times faster and costing 50% less per token.

How does GPT-5 Turbo compare to GPT-5?

GPT-5 Turbo matches GPT-5 on most benchmarks, scoring within 2-3% on reasoning, coding, and general knowledge tasks. It is significantly faster with a time-to-first-token of under 200 milliseconds and costs half as much per API call.

What does GPT-5 Turbo cost?

GPT-5 Turbo is priced at $2.50 per million input tokens and $10 per million output tokens, making it one of the most cost-effective frontier-class models available.

OpenAI Launches GPT-5 Turbo — 3x Faster, Half the Cost

OpenAI has released GPT-5 Turbo, a leaner version of its flagship model that trades a small amount of peak capability for dramatically better speed and pricing. The model is aimed squarely at production developers who need frontier-quality responses without the latency and cost of the full GPT-5.

Speed and Pricing

The numbers are straightforward. GPT-5 Turbo delivers responses three times faster than GPT-5, with a time-to-first-token of under 200 milliseconds. API pricing drops to $2.50 per million input tokens and $10 per million output tokens — half the cost of standard GPT-5.

OpenAI achieved this through model distillation, a technique where a smaller model is trained to replicate the behavior of the larger one. The company says GPT-5 Turbo uses roughly 40% fewer parameters than GPT-5 while retaining most of its capabilities.

"The full GPT-5 is our research flagship. GPT-5 Turbo is what you ship to production," said Sam Altman during the announcement livestream.

Benchmark Performance

On standard benchmarks, GPT-5 Turbo scores within 2-3% of the full GPT-5 across reasoning, coding, math, and general knowledge tasks. On MMLU-Pro, it scores 89.1% compared to GPT-5's 91.4%. On HumanEval coding benchmarks, it achieves 93.2% versus 95.1%.

Where the gap widens is on the most demanding multi-step reasoning problems. On complex mathematical proofs and extended agentic coding tasks, the full GPT-5 maintains a more noticeable edge. For the vast majority of production use cases — customer support, content generation, data extraction, code completion — the difference is negligible.

Context Window

GPT-5 Turbo ships with a 256,000-token context window, matching the full GPT-5. OpenAI says there is no degradation in long-context retrieval accuracy, which was a common complaint with earlier Turbo variants.

Developer Reaction

The developer community has responded positively. Many teams had been using GPT-5 in development but switching to cheaper models for production due to cost constraints. GPT-5 Turbo eliminates that trade-off.

"We were spending $40K a month on GPT-5 API calls," said a startup CTO on X. "GPT-5 Turbo cuts that in half with no visible quality drop. This is what we were waiting for."

Competitive Pressure

The release puts pressure on Anthropic and Google, both of which charge premium rates for their flagship models. Anthropic's Claude Opus is priced at $15/$75 per million tokens, while Google's Gemini 3.1 Ultra sits at $12/$60. GPT-5 Turbo undercuts both significantly while claiming comparable performance.

OpenAI also announced that GPT-5 Turbo will replace GPT-4o as the default model in ChatGPT Free within the next two weeks, giving hundreds of millions of users access to near-frontier performance at no cost.

OpenAI Launches GPT-5 Turbo — 3x Faster, Half the Cost

Speed and Pricing

Benchmark Performance

Context Window

Developer Reaction

Competitive Pressure

More in Models

Moonshot Kimi K2.6 lands open-source, scales to 300 sub-agents and 4,000 coordinated steps

OpenAI's 'Spud' Caught Live in API Testing, Polymarket Jumps to 81% for April 23 Launch

OpenAI Launches GPT-Rosalind, Its First Domain-Specific Model Built for Life Sciences