IDC ATLASINTELLIGENCE CONSOLE
SyncingColumns
Fact change
Research published
Brief edition
System scan
Verification
中文
IDC ATLAS
中文Back to intelligence
IDC ATLAS COLUMN · MODEL ECONOMICS · 32

MiniMax and Z.ai H1: Token Scale Is Becoming Revenue, but the Compute Bill Is Still Open

MiniMax and Z.ai's 2026 interim results move Chinese foundation-model commercialization into a more verifiable phase. Open Platform and enterprise services reached 63.4% of MiniMax revenue, while Open Platform and API reached 86.5% at Z.ai. Z.ai also disclosed 100,000-plus domestic chips in inference, an 80% decline in unit-token inference cost since the start of the year and positive API gross margin. Revenue conversion is visible; training spend, borrowings and adjusted losses keep the compute bill open.

Comparison of MiniMax and Z.ai API revenue, token usage, inference clusters and loss structure
IDC Atlas original editorial cover · MODEL ECONOMICS · 32

MiniMax reported $116.6 million of first-half revenue, up 283.1%. Open Platform and other AI-based enterprise services generated $73.9 million, up 703.1%, and represented 63.4% of revenue. Total gross margin rose from 12.1% to 17.9%. Net loss narrowed 11.0% to $358.0 million, but adjusted net loss widened 111.2% to $293.0 million, leaving operating investment far above revenue.

Z.ai reported RMB953.9 million of first-half revenue, up 399.7%. Open Platform and API revenue was RMB825.2 million, or 86.5% of the total, up 2,735.7%. Segment gross margin moved from -0.4% to 24.6%, while group gross margin fell from 50.0% to 26.4% as cloud revenue diluted the legacy mix. Net loss narrowed 12.1% to RMB2.072 billion; adjusted net loss widened 12.1% to RMB1.964 billion.

The companies report in different currencies and use different revenue categories and non-IFRS adjustments, so this article does not impose an unreported exchange rate to rank their scale. The comparable direction is the shift toward API and subscription revenue, improving inference efficiency and continued reliance on R&D and financing.

The API is becoming the primary revenue-recognition channel

MiniMax's Open Platform and enterprise-services share rose from 30.3% to 63.4%, driven by paying users, enterprise customers, API calls and Token Plan adoption. AI-native product revenue also rose 100.9% to $42.6 million, preserving two revenue tracks across products and enterprise infrastructure. Revenue outside mainland China represented 60.8% of the total.

Z.ai's switch was larger. Open Platform and API moved from 15.2% of revenue a year earlier to 86.5%, while enterprise general-purpose large-model revenue fell 54.6% to RMB67.0 million. On-premise deployment fell from 73.7% of the prior full-year mix to 13.5%. Revenue is moving from project acceptance and one-time licenses toward metered calls and subscriptions, tying pricing, retention, concurrency and service reliability together.

Z.ai says MaaS token volume was more than 40 times its level at the start of 2026 by the announcement date, Coding Plan volume was more than 23 times higher and average API selling price rose about 101%. Volume and price rose together, but these remain company-defined statistics without independent cohort retention, net-price or promotion data.

MetricDisclosureBasis and boundary
MiniMax Open Platform63.4%Share of H1 revenue; up 703.1% year over year.
Z.ai Open Platform / API86.5%Share of H1 revenue; up 2,735.7% year over year.
Z.ai API gross margin24.6%Up from -0.4% a year earlier.
Z.ai token volume>40×Company measure from the start of 2026 to the announcement date.

Call disclosures show acceleration and increase the need for denominator discipline

Yicai reported MiniMax management's August ARR at more than $800 million, with roughly 80% from business customers. Caixin reported Z.ai management's August ARR at $1.6 billion using monthly revenue annualized. Neither number appears in the interim-results filings, and neither is accompanied by a public bridge across customers, contract duration, refunds, discounts or churn.

ARR annualizes a current month or contract state. It is not first-half recognized revenue, cash collection or guaranteed revenue for the next twelve months. The figures help measure August velocity; they should not be added to reported revenue or compared across the companies without aligned definitions.

The next filing must test three conversions: ARR into quarterly revenue, API margin durability and contract liabilities plus operating cash flow. If ARR rises without a matching movement in recognized revenue, margin or cash, promotions, customer concentration or the annualization method may explain part of the signal.

REPORTED

MiniMax ARR above $800M

Management disclosure reported by Yicai; roughly 80% business, absent from the interim filing.

REPORTED

Z.ai ARR at $1.6B

Management disclosure reported by Caixin; August monthly revenue multiplied by twelve, absent from the interim filing.

GATE

ARR is not recognized revenue

Quarterly revenue, margin, contract liabilities, retention and cash collection must complete the bridge.

Z.ai connects domestic-chip scale, cost and margin in one disclosure chain

Z.ai identifies domestic chips as a primary inference resource and reports a cluster of more than 100,000 chips plus an 80% decline in unit-token inference cost since the start of the year. The same filing puts Open Platform and API gross margin at 24.6%, attributing the improvement to scale, pricing and full-stack engineering. Physical infrastructure, volume, price and margin now sit in one operating chain.

Call reporting adds that owned clusters, leased compute and purchased services cover both training and inference. More than 100,000 chips therefore does not mean 100,000 company-owned assets or provide a direct capex estimate. Suppliers, effective online chips, power draw, throughput, SLA and the cost mix across sourcing modes remain undisclosed.

MiniMax discloses only that infrastructure efficiency lifted total gross margin to 17.9%. Yicai reported management saying M3 and H3 are being adapted to domestic chips, that a large domestic-compute cluster is expected to come online and that M3.1 targets inference cost at roughly one third of M3's launch level. Those are plans and targets, not yet an operating cluster-cost-margin bridge comparable with Z.ai's filing.

  1. 01
    Tokens enter billing

    API calls and subscriptions turn model use into recurring revenue.

  2. 02
    Clusters absorb concurrency

    Models, engines, networks and scheduling determine whether installed compute creates effective tokens.

  3. 03
    Unit cost falls

    Scale and engineering must cover cloud services, chips, networking and operations.

  4. 04
    Margin tests the economics

    Price and cost converge in segment gross margin; cash flow remains the next gate.

Revenue growth has not yet covered the training cycle or funding requirement

MiniMax R&D expense rose 138.8% to $296.9 million, primarily because of cloud-service expense for training. It ended June with a $1.323 billion cash balance and $133.6 million of bank borrowings. In July it raised approximately HK$9.44 billion net from a placement and HK$6.43 billion net from convertible bonds, extending the funding buffer for training and expansion.

Z.ai spent RMB2.131 billion on R&D, up 33.6%. It ended June with RMB3.994 billion of cash and cash equivalents and RMB2.225 billion of bank loans, while adjusted net loss was RMB1.964 billion. Positive API gross profit has arrived, but listing proceeds and debt still support a larger R&D and operating base.

Both companies reported no significant or outstanding period-end capital commitments. That does not imply low future compute demand. Cloud services, leases and purchased compute can be recognized as operating expense, short-duration contracts or other arrangements. Cash, debt, cloud-service expense, lease liabilities and later financing need to be read together.

The data-center demand signal is clearer; capacity conversion remains constrained

Both filings confirm expanding production inference, with coding and agent workloads adding long context, cache, tool calls and sustained execution. Z.ai's 100,000-plus-chip cluster is operating evidence for domestic accelerators, servers, interconnect and orchestration. MiniMax's international revenue mix implies inference capacity procured or deployed across regions.

Public disclosures omit compute per token, average sequence length, cache hit rate, batching efficiency, chip models and PUE. Forty-fold and twenty-fold token growth cannot be translated into forty-fold and twenty-fold GPUs or megawatts. Lower unit cost, model architecture and utilization can absorb a substantial share of the increase.

Over the next 12 to 36 months, the discriminating signals are sustained API-margin expansion, adjusted losses shrinking with revenue, effective online domestic-cluster scale and cost per completed task, and ARR converting into recognized revenue and operating cash flow. Improvement across all four would support financeable, durable data-center capacity.

DEMAND

Production use is in the revenue statement

API and subscription mix reduce reliance on leaderboards or free trials as demand proxies.

EFFICIENCY

Tokens are not fixed compute

Models, cache, concurrency and chip efficiency alter the GPU and megawatt conversion.

FINANCE

Cash follows gross margin

Training spend, debt and financing determine whether inference demand can keep expanding.

IDC ATLAS VIEW

MiniMax and Z.ai have shown that Chinese foundation models can convert large-scale use into recurring revenue. Z.ai also provides the first integrated operating snapshot across domestic inference scale, unit cost and API margin. The filings simultaneously show that the business still depends on heavy training investment and external capital. For IDC Atlas, the signal has advanced from model usage to billable usage with gross profit. Cash conversion and reproducible cluster efficiency now determine how much durable compute capacity it can support.

Cutoff: September 1, 2026, 8:45 AM Beijing time. Financial data, token volume, Z.ai's domestic-chip scale and cost change come from the companies' interim-results announcements. MiniMax's August ARR, business mix, domestic-chip adaptation and M3.1 cost target, together with Z.ai's August ARR and call Q&A, come from attributed media reports of the results calls. MiniMax's official event page verifies the call but does not provide a complete transcript licensed for republication. ARR is not recognized revenue, and tokens do not directly equal GPUs, megawatts or data-center revenue. This article contains no price target, rating or investment advice.

For information and research only. This is not investment advice.