Start with the money, because it is the part every developer noticed. From 16 August, a million output tokens through DeepSeek-V4-Pro costs $3.96 at peak, up from $0.87.
V4-Flash goes from $0.28 to $1.32, Bloomberg reported. Off-peak rates are half the new peak, so $1.98 and $0.66.
Now hold those discount rates against last week’s full price. Off-peak V4-Pro costs $1.98 against an old peak of $0.87.
The new discount is more than double the old undiscounted price. A team that moves every workload to the quiet hours still pays over twice what it paid at lunchtime.
The company built a Claude Code rival first
The pricing makes more sense once you see what landed alongside it. On Thursday DeepSeek released a developer preview of DeepSeek Harness v0.1.
A harness is the scaffolding wrapped around a model that lets an agent do real work. It reads files, edits code, browses the web and keeps going until a task is finished.
That is the layer Anthropic sells as Claude Code, and it is where the money in agentic coding sits. The model is the engine; the harness is the car.
DeepSeek had been signalling this for days. It set up a “DeepSeek Harness Team” account on WeChat and posted job listings for the team, Bloomberg reported earlier in the week.
One of those postings described the goal plainly. The company wants to turn its models into cutting-edge agentic products.
The account sits under a Beijing entity that Chinese corporate records link to DeepSeek, and Tencent has verified it. For a company that communicates mainly through model releases, setting up a team account in public is itself the announcement.
The design choice is the interesting part
DeepSeek says its harness uses an open architecture. Users can plug in any component they choose, including models from other companies.
It framed that against American rivals, which it said hard-code their products. Whether that is a fair description of Claude Code is arguable, but the strategy behind it is not.
An open harness that runs anyone’s model is a bid to own the workspace rather than the engine. If developers work inside your tool, the model underneath becomes a component you can swap, including for your own.
The desk has watched this layer become the battleground. Cursor raised prices in a harness shift of its own, and Alibaba has banned Claude Code internally over tracking concerns.
The flagship left preview, quietly, then the notice vanished
Underneath both stories sits a third. DeepSeek-V4-Pro-0813 shipped as the general-availability build this week, ending a preview that ran nearly four months.
The company marked it with a brief statement on its website saying the model offered “significantly enhanced agent capabilities”.
By Thursday afternoon the statement had been removed, the South China Morning Post reported. DeepSeek has not explained why, and it did not respond to Bloomberg’s request for comment on the harness team.
A company does not usually delete a claim about its own flagship in the same week it raises prices on it.
Developers were not impressed
This is where the balance sits, and it cuts against the company. Early reaction to the 0813 build left developers underwhelmed on general capability and unhappy about the pricing, per the SCMP, though researchers were impressed in narrower areas such as cybersecurity.
The benchmark figures are vendor-reported and no independent evaluator has replicated them for this build. DeepSeek’s own model card puts V4-Pro at its maximum reasoning setting on 80.6% for SWE-bench Verified, level with Gemini 3.1 Pro and a fraction behind Claude Opus 4.6 at 80.8%.
On other tests it trails. The card shows 67.9% on Terminal Bench 2.0 against GPT-5.4 at 75.1%, and 37.7% on Humanity’s Last Exam against Gemini 3.1 Pro at 44.4%.
So the price rose more than fourfold on a model that its maker positions as roughly level with the frontier on coding and behind it elsewhere.
The underlying engineering is not in doubt. V4-Pro is a mixture-of-experts system with 1.6 trillion parameters, of which 49 billion are active per token, and DeepSeek says its attention design cuts the compute needed for a single token to 27% of what its previous generation used.
Efficiency of that kind is exactly what made the old prices possible. It is also why a fourfold rise reads as a decision about margin rather than a report about costs.
Cheap is relative, and it is still cheap
The headline multiple invites the wrong conclusion, so here is the ladder. At $3.96 per million output tokens, DeepSeek sits well below Moonshot’s Kimi K3 at $15 and Anthropic’s Fable 5 at $50.
Moonshot is the closer comparison, and the desk has covered its largest open model and its own funding run.
DeepSeek’s pricing has been so far outside the normal range that Bloomberg reports the industry now talks about a DeepSeek “death zone”, in which costlier or weaker models are simply obviated.
A fourfold rise does not end that. It narrows it, and it ends the claim that Chinese inference is effectively free.
The listing behind all of it
Bloomberg puts the increase directly against the IPO, reporting that it suggests a greater focus on profitability as DeepSeek eyes a stock-market debut.
The company is in the middle of a large fundraising and has begun preparations to list as soon as this year, most recently valued at around $71bn. Founder Liang Wenfeng now has to balance investors, expansion and the cost of compute at the same time.
Loss-leading is a private-company strategy. Owning the harness is how you stop needing it.
The desk covered the warning last week, when DeepSeek said a significant rise was coming without giving figures. These are the figures.
What would settle it
Two things are checkable rather than rhetorical. The first is whether the harness gets used outside China, because an open architecture is only a moat if developers actually build in it.
The second is the deleted statement. If the “significantly enhanced agent capabilities” claim returns alongside a benchmark somebody else can run, it was a publishing error. If it stays gone, it was a claim the company decided it could not stand behind in the week it started charging four times as much for it.
Get the TNW newsletter
Get the most important tech news in your inbox each week.