Last year, we were still in the ‘golden age’ of AI coding and productivity agents. I’m not talking about capability. Agents have massively improved in that area.
I’m referring to the freedom. In the beginning, even the cheapest LLM API subscription plans (from OpenAI and Anthropic) had lots of headroom for exploration, building and learning.
Oh how things have changed! Demand for AI inference has skyrocketed and major AI labs have been pulling back on subscription plans significantly.
Anthropic has been the poster child for this trend. When they re-launched their newest Fable model, they provided access to it via a Claude Code subscription. But only for a limited time. This led to some builders working day and night (or even missing the birth of their baby) to maximize their access to the model before it went behind the pay-as-you-go wall.
Although OpenAI seems to be taking a different approach with its latest Sol model, which appears to have much lower token consumption costs (and can support longer work sessions), they are also tweaking how they charge for LLM usage.
Overall, pretty much every major closed-source lab is moving toward more restrictive, rather than permissive, pricing strategies. Why? It’s all about supply, demand, and profitability.
- Demand for inference continues to skyrocket
- AI labs have been heavily subsidizing users’ costs.
- Anthropic and OpenAI are nearing IPOs. They’ll be under tremendous pressure to both continue to innovate and show a profit.
The result? Labs will pass on more costs to users, meter their best models and use every strategy they can to maximize revenue.
Bigger and Better? Yes!
Over, the last few years, we’ve gotten used to a steady tempo of newer and more powerful frontier model releases. And, as the Fable launch demonstrates, we’ve been trained to automatically reach for the most expensive models because they are better than the last one.
This makes a lot of sense. Consider Fable. Some users are reporting giving the model a complex task, putting it on a loop and returning the next morning with a fully finished (presumably relatively high-quality work product). Others praise Fable for being able to understand and work through very difficult problems with ease.
Each frontier model release has increased the wow factor. Highly advanced LLMs released by the major AI labs are becoming more capable of completing highly complex tasks, and many want to immediately upgrade to take advantage of that level of firepower.
But the downsides of this approach are real. Extensive AI use has real cognitive impacts. We don’t know whether they’re positive or negative. Another, less visible, impact is that the abundance of heavily subsidized inference has resulted in terrible LLM use habits. Many haven’t developed an intuition about how to get the most out of models, the best ways to communicate with LLMs and many other skills.
We shouldn’t expect many people know how to do this. AI is very new and a lot is still being figured out. However, in a world where access to the most powerful AI models is metered and expensive, we’re going to have to learn how to do more with less. This means using lower power models (yes, open source models are improving, but the frontier is still the frontier) and having less available time to work with the most advanced LLMs.
Uncovering the Secrets of the LLM Whisperer: Getting the Most Out of AI, At Lower Cost
In mid-June, Anthropic released interesting research revealing that the economic value of work conducted by users of Claude Code had increased between October 2025 and April 2026. The study stuck with me because, looking at the models that were available over the study period, the most significant value gains occurred after Opus 4.5’s release (Sonnet 4.6 was also launched during this period). This data demonstrates that today it’s possible to conduct high-value work using LLMs that are significantly less powerful than flagship models.
This has interesting implications for frontier model labs and LLM users, especially as open source LLMs (like the newly released Kimi K3) become more capable. Benchmarks indicate Kimi K3 outperforms Fable and GPT models on some frontend coding tasks.
But, there are still challenges to using open source models from a cost-efficiency perspective (as outlined in the table below).
| Trend | Cost and Efficiency Implications |
|---|---|
| Major AI Labs Meter Flagship Models | Users will have to carefully manage how and when they use the most advanced frontier models |
| Powerful Open Source Models Emerge | Open source models provide more options to users, enabling many tasks to be completed at cheaper rates. Kimi and models like it compete with closed source, less expensive models like Claude Sonnet (similar cost profile, different capabilities) |
| LLM User Skill Gap | Frontier models are more forgiving and easier to get high-quality results from. Operating open source and lower cost models optimally requires understanding concepts like prompt construction and retry reduction. If not run efficiently, models (regardless of capability) can consume time and human resources, which is another cost driver. |
So the question becomes: In a world where understanding model selection is only one part of the cost-efficiency maximization equation, what are the habits and behaviors that drive success, regardless of model quality?
To find out, I conducted a study of nearly 240,000 simulated LLM users. The goal was to find out which habits and behaviors cost (or save) people money on LLMs. Analysis of data from the study revealed a type of user I call the LLM Whisperer, representing people who get the most out of LLMs at the lowest cost.
Here are some of the key cost savings behaviors of LLM Whisperers:
- Use the right model for the job: This one isn’t very surprising. Being smart about picking the model you work with is a no-brainer way to save money. However, many people don’t switch models. First, it’s much easier to pick a primary model and stick with that. Second, there is a psychological barrier to working with less advanced models: trust. It can be hard to figure out if a model is reliable, or when it’s reached its limit. The other barrier is workflow. How to coax models to deliver the best output is challenging. It’s much easier to reach for Opus 4.8 for every task. You spend more money, but potentially save time. Interestingly, LLM Whisperer cost efficiency gains don’t come from always defaulting to the cheapest model.
- Slash retries as much as possible: The research highlighted the role of retries on token spend. We've gotten used to accepting retries as the cost of doing business with LLMs. After all, AI makes mistakes, we adjust (sometimes after a lot of effort) and then move on. But, retries add up, and sometimes make it less likely the AI will succeed on the next attempt. Managing retries and limiting them as much as possible is a critical LLM Whisperer strategy.
The other lesson from the research is that using AI cost-effectively will require more than software or other tools. For example, some have claimed that codegraph, a solution that pre-indexes a codebase so that agents spend less time (and tokens) finding the right content, can cut token use by 90%. However, these cost savings can be short-lived.
A long-term, sustainable LLM cost savings strategy requires:
- Knowledge: Knowing how LLMs work, especially how tokens are generated, the various token types and how they impact costs
- Understanding: Getting a grip on how to use LLMs and what habits might be costing money
- Education: Gaining skills that will, over time, will help deliver ongoing cost savings
To help, I have developed a free, three-part, research-backed system called the LLM Whisperer Method, which:
- Improves your knowledgeabout high-performing, cost-effective LLM use patterns
- Helps you identifyhabits and behaviors that may be costing money through a 5-minute assessment
- Delivers skillsvia 90-day courses that can be personalized and delivered by agents
The new reality: Using AI optimally takes a lot more than typing words into a text box. But having the right perspective, strategies and tactics can make the difference.