IDC ATLASINTELLIGENCE CONSOLE
SyncingColumns
Fact change
Research published
Brief edition
System scan
Verification
中文
IDC ATLAS
中文Back to intelligence
IDC ATLAS COLUMN · AGENT ECONOMICS · 20

From GLM-5.3 to Claude: How Long-Running Agents Change Inference Demand

Stronger agent models, replaceable runtimes and Claude's commercial signals point in one direction: inference is moving from one-off answers toward repeatable long tasks. A single revenue curve does not prove the causal chain.

Long-running AI agent workstation, toolchain and inference-infrastructure cutaway
IDC Atlas original editorial cover · AGENT ECONOMICS · 20

Z.ai describes GLM-5.3 as using the same base model as GLM-5.2, with changes primarily attributed to post-training, and reports gains on coding, terminal and agent-oriented benchmarks. It is currently available through the GLM Coding Plan; the API is described as coming soon. These are company benchmarks, not independent replication.

DeepSeek Harness is an open-source developer-preview runtime. Its documentation emphasizes plugins, replaceable model adapters and tools, session logs, sandboxes, approvals and persistence. It is model-neutral agent infrastructure, not an official GLM-5.3 distribution or integration.

Anthropic said in 2026 announcements that its annualized run-rate rose from roughly $9B at the end of 2025 to more than $30B, and then beyond $47B in May. Anthropic is private; these are company-disclosed run-rates, not public quarterly P&L or profit.

GLM-5.3 signals post-training change, not independent performance proof

Z.ai says GLM-5.3 shares GLM-5.2's base model and credits the update primarily to post-training, while reporting improvements on Code Bench, Terminal-Bench, DeepSWE and Agents' Last Exam. The documentation also lists 1M context, 128K output, tool calls, structured output, context caching and MCP.

Those interfaces matter for long tasks: a task must read and write an environment, use tools, retain state across steps and recover from failure. Benchmark scores still do not establish production stability, cost or safety.

MetricDisclosureBasis and boundary
GLM-5.3 baseSame as GLM-5.2Z.ai attributes changes primarily to post-training.
Context / output1M / 128KOfficial documented interface limits.
AccessCoding PlanDocumentation says API is coming soon.

Harness places model capability inside a manageable execution loop

DeepSeek Harness composes a turn from model requests, tool calls and step records; models, tools, adapters and policies can be swapped through plugins. The model therefore need not own the entire runtime, while the execution layer can retain approvals, sandboxes and state.

That composability lowers friction in testing models and tools, but the developer preview explicitly warns of breaking compatibility changes. It discloses no usage, customers or revenue and is not evidence of realized compute demand.

MODEL

Inference and tool choice

At each step a model issues a request, uses tools and continues from the result.

RUNTIME

State and governance

Session logs, approvals, sandboxes and persistence determine how long tasks execute.

ECONOMICS

Billable work

Inference becomes recurring revenue only when tasks complete reliably inside customer workflows.

Claude run-rate is a signal, not a causal proof

Anthropic said run-rate exceeded $30B, with more than 1,000 customers spending more than $1M annualized, and its May financing announcement said run-rate had crossed $47B. It also announced multi-gigawatt compute partnerships.

That supports the direction that higher-value enterprise agent work is commercializing. It does not prove that GLM-5.3 capability or DeepSeek Harness architecture caused Claude growth, and run-rate cannot be directly translated into tokens, GPUs, MW, gross margin or free cash flow.

The countercase matters: better capability may reduce retries and tokens per task; runtime approvals may lower the autonomous-task share; private-company disclosure does not let outsiders validate customer, price or usage mix.

IDC ATLAS VIEW

The real demand variable is not a model leaderboard. It is whether long tasks complete in controlled environments, are adopted by enterprises and settle reliably. Models, runtimes and commercialization are converging, but public evidence does not yet lock them into a quantified causal chain.

Cutoff: August 16, 2026, Beijing time. GLM-5.3 benchmarks are Z.ai disclosures; DeepSeek Harness is an open-source developer preview; Anthropic is private, and run-rate is not quarterly revenue, profit, tokens, GPUs or MW.

For information and research only. This is not investment advice.