IDC ATLASINTELLIGENCE CONSOLE
SyncingColumns
Fact change
Research published
Brief edition
System scan
Verification
中文
IDC ATLAS
中文Back to intelligence
IDC ATLAS EXPLAINER · MODEL ECONOMICS · 40

Chips Cost More. DeepSeek Charges Less. Who Absorbs the Difference?

Count useful work, not cheap tokens.

Conceptual voxel illustration of cache compression and compute resources
IDC Atlas original editorial cover · MODEL ECONOMICS · 40

Higher equipment prices and lower model-service prices can coexist without proving that a provider is selling at a loss. A machine can become more expensive while the cost of completing useful work falls. The missing variable is how much reliable output that machine produces.

Two prices at opposite ends of the chain

DeepSeek released V4.1 Flash on September 10 and lowered API prices. Its announcement says the new model’s KV cache requires one-quarter of the previous generation’s HBM and one-eighth of its SSD storage. Those ratios concern cache requirements, not the purchase price, total memory or electricity consumption of an entire server.[1]

Reuters separately reported that some Chinese AI-chip suppliers had raised quotations as HBM costs increased. The reported quotations rely on unnamed sources and were not publicly confirmed by the manufacturers concerned. They are attributed market reporting, not substitute purchase contracts.[2]

There is no disclosed buyer–supplier link between these two items. We are not claiming that DeepSeek bought the products in the report. The research question is broader: can model and system improvements leave room for cheaper applications when hardware is under cost pressure?

Consider a delivery business. A truck might cost more, yet carrying more useful cargo and making fewer empty journeys could reduce the cost per parcel. That analogy explains the denominator. AI providers also have to preserve accuracy and response times; more incorrect answers cannot be counted as more useful cargo.

Between buying equipment and selling an API request sit depreciation, operating hours, concurrency, retries, networking, facilities and engineering. The two headline prices leave out most of that chain.

What is being saved?

When a model processes a long document or conversation, it can retain intermediate information for later generation. A KV cache is roughly like working notes: it helps avoid repeating some previous work. HBM provides fast memory close to the processor; SSD offers a different storage tier. Neither is interchangeable with the other.

The official model card describes architectural and cache changes and connects efficiency to input-heavy agent workloads. This establishes the developer’s proposed mechanism, not an independently reproduced improvement in every serving environment.[3]

Saving cache space can first relieve a capacity constraint. If long conversations previously filled available cache, compression may permit more concurrent users. If execution is limited instead by computation, networking or external tools, the benefit will be different.

Free space becomes an economic benefit only when it supports more paid work, avoids some equipment purchases or enables a better service. With insufficient demand, it may simply become spare capacity.

Long-running agents make this distinction important. Reusing earlier context may save work, but the full task may still wait for retrieval, an external application or human approval. A faster model stage need not shorten the whole job by the same proportion. Measure both model latency and end-to-end completion.

A small bill makes the arithmetic clearer

The following is a teaching example, not an estimate of any named company. Suppose an accepted task originally costs 1 unit. Resources affected by an optimization account for 0.30, while everything else accounts for 0.70. Halving the first component produces a new cost of 0.30 × 50% + 0.70 = 0.85.

A 50% improvement in one component therefore reduces total cost by 15% in this assumed case. This does not forecast the actual reduction. It shows why a cache ratio cannot be applied to the whole operating bill.

Swipe to read the full table →
Illustrative caseAccepted tasksTotal spendingCost per accepted task
Original system1001001.00
Lower spending, unchanged output100850.85
Higher spending, more useful output1501200.80
Cheaper calls, more failures7085About 1.21

All figures are hypothetical. Define acceptance criteria before testing. The denominator must contain work that meets those criteria, not merely attempts.

The final row explains a common procurement mistake. A cheaper model can require more retries or manual correction. A more expensive call can sometimes lower the total bill if it completes the task reliably. Compare a fixed set of real tasks, include tools and review effort, and avoid choosing only workloads favorable to the new system.

DeepSeek also retains peak and off-peak pricing. Scheduling flexible work outside peak periods can save a customer money, but that is a scheduling benefit, distinct from architecture.[1] Separate the effects so that later comparisons remain meaningful.

Who keeps the efficiency gain?

A model provider can retain savings, cut prices to attract usage, or spend more computation on improving task quality. A public price list reveals only part of those choices. It does not disclose the provider’s complete cost structure.

A cloud operator has a different exposure. If it rents machines, customers needing fewer machines may reduce rental demand. If it sells completed work, better utilization may improve its economics. The same technical change can have opposite effects under different business models.

Equipment demand depends on both resource intensity and volume. In a hypothetical case, if equipment time per task falls to 80% of its former level, accepted-task volume must grow by more than 25% for total equipment time to increase: 80% × 125% = 100%. This is a break-even condition, not a demand forecast.

Application providers may see lower API bills without retaining all the savings. Competitors can access the same models, and benefits may pass to customers through lower prices or additional service. Reliable workflows and valuable customer integration may help retain value, but these disclosures cannot establish which business has succeeded.

Facilities respond more slowly. Existing leases, equipment payments and power arrangements do not disappear with a model announcement. Efficiency can change scheduling first, then renewal and procurement decisions, and eventually construction. A benchmark observed over days and a capital decision spanning quarters require different evidence.

The strongest alternative explanations

Price competition may explain a reduction even when costs have not fallen enough to cover it. Architecture documentation does not prove the profitability of every price cut. Comparable service-cost or margin disclosure would be needed to resolve that question.

A second possibility is that benchmark improvements fail to translate into enterprise outcomes. Real workflows contain permissions, poor data, audit requirements and approvals. If model costs fall but delivery time does not, inspect the other stages rather than declaring the architecture irrelevant.

Hardware availability can also outweigh efficiency gains. Better use of existing equipment cannot automatically launch a project unable to obtain the necessary hardware. Delivery time, supported configurations, software maturity and replacement parts remain constraints.

Finally, customers may consume the improvement through higher standards. They may ask for evidence checking instead of a simple summary, or faster interaction instead of batch processing. Spending can stay constant while the service improves. An unchanged bill does not by itself prove that an optimization failed.

These explanations can coexist. Providers can improve efficiency, face competitive pressure and invest in better quality at the same time.

What would settle the question?

Application buyers should repeat a fixed task set and record acceptance, completion time, total cost and human intervention. Operators should measure whether spare resources become stable paid throughput without degrading service. Supply-chain observers need comparable purchase configurations and actual delivery evidence, not isolated quotations.

The base case is that efficiency affects task economics before it visibly changes aggregate facility requirements. Over one to four quarters, lower accepted-task costs combined with sustained usage growth and expansion would support the case for demand creation. Lower prices accompanied by margin pressure, weak usage or delayed projects would support a more cautious interpretation.

IDC ATLAS VIEW

The practical question is simple: how much accepted work did the total spending buy? That is more informative than either the latest chip price or the cheapest token headline.

Sources

  1. DeepSeek release, September 10, 2026. Company cache and pricing claims; not independently benchmarked here.
  2. Reuters via MarketScreener, September 10, 2026. Attributed, unconfirmed supplier quotations.
  3. Official DeepSeek model card, accessed September 15, 2026. Architecture context; disputed total-parameter counts are not used to infer equipment demand.

Evidence cutoff: September 15, 2026, Asia/Shanghai. Original industry analysis; no investment recommendation. Numerical examples are explicitly hypothetical. No company unit-cost estimate is supplied.

For information and research only. This is not investment advice.