OpenAI has reportedly cut ChatGPT inference costs for guest users by more than half, according to a report from The Information, marking one of the steepest efficiency gains the company has disclosed for running its models in production.
What OpenAI's engineers reportedly achieved
OpenAI engineers told colleagues earlier this month that they had managed to cut inference costs, the ongoing expense of running already-trained AI models, by more than half. That's according to a person familiar with the internal discussions, as first reported by The Information.
The optimizations were applied specifically to ChatGPT traffic from visitors who don't have an account. For that slice of usage, the number of Nvidia GPUs needed to keep the service running dropped to just a few hundred at times, a dramatic reduction from whatever baseline OpenAI was running before. It isn't clear exactly how many chips were required previously, and OpenAI hasn't disclosed the specific techniques its engineers used to reach the savings.
Guest users only get access to a limited slice of ChatGPT's features compared with signed-in accounts, so the open question is whether these efficiency gains would carry over if applied to the full, logged-in product, where usage patterns and model routing are more complex.
Why cutting ChatGPT inference costs matters
Inference is the single biggest recurring cost for any company running large language models at ChatGPT's scale, since every reply a user receives requires fresh computation on expensive AI accelerators. Even modest efficiency improvements translate into enormous savings when spread across hundreds of millions of weekly conversations, which is why AI labs treat inference optimization as seriously as they treat training breakthroughs.
It also matters competitively. OpenAI, Google, and Anthropic are all racing to make their assistants cheaper to run while keeping quality high enough to win over free-tier users who might eventually convert to paying subscribers. Lowering the cost of serving guest traffic, the users least likely to be paying anything at all, is a direct way to make that funnel more sustainable.
OpenAI isn't alone in chasing cheaper inference
The report lands alongside a similar move from Deepseek, which just released an open-source method that reportedly speeds up inference requests by 60 to 85 percent. Techniques like sparser model architectures, better caching of repeated computations, and smarter routing between smaller and larger models have all become common tools across the industry for squeezing more efficiency out of the same hardware. This broader trend in AI development shows labs increasingly optimizing for cost per query, not just raw capability.
Whatever specific method OpenAI used, the fact that a company running one of the world's most-used consumer AI products can more than halve costs for a segment of its traffic underscores how much headroom still exists in inference engineering, even for mature, heavily-optimized systems.
What OpenAI might do with the savings
The freed-up GPU capacity and budget could go toward several things: scaling ChatGPT to more users, training or serving better models, delivering faster responses, or simply improving OpenAI's margins as the company works toward profitability. Because data center buildouts move slowly and chip supply remains tight, analysts covering the AI business landscape suggest efficiency gains like this are more likely to give labs breathing room within their existing infrastructure than to meaningfully reduce overall demand for Nvidia chips in the near term.
For everyday ChatGPT users, the practical impact of this specific change is likely invisible, since the free, logged-out experience should look the same. But cost cuts like this are part of why free AI tools have remained free, or even improved, even as usage has exploded, and they help explain how OpenAI can keep offering broad public access without cutting off the guest tier entirely.
Frequently asked questions
Does this inference cost cut affect ChatGPT Plus subscribers?
The reported optimization was specifically applied to guest, logged-out ChatGPT traffic. It's unclear whether or how much the same techniques apply to paid or signed-in tiers of the product.
How did OpenAI reportedly cut inference costs so much?
The exact techniques haven't been disclosed. The report only confirms that the number of Nvidia GPUs needed to serve guest users dropped to a few hundred at times, without detailing the underlying engineering changes.










































