Kunlun Tech CEO: Piling Up Tokens Won't Build an AI-Native Organization, Models Are the Real Foundation
Key Highlights
Kunlun Tech CEO Fang Han poured cold water on the industry's current "token-burning" frenzy at a WAIC roundtable: simply stacking token consumption cannot measure the value AI creates. His core thesis is that model capability must rely on the engineering frameworks built by coding agents such as Claude Code to translate into real productivity. Put simply, no matter how many calls you buy, without a "skeleton" that embeds the model into the workflow, it is only a fleeting spectacle. The remark lands because many budgets still report AI spend as token volume, a number that grows with usage but says nothing about whether the output shipped; Fang is arguing for a different scoreboard, one tied to delivered work rather than consumed compute. The comment also implicitly critiques the habit of celebrating usage milestones, arguing that a rising token bill is a cost to manage, not a victory to publish, and that leadership should be judged on shipped capability instead.
Capabilities and What Happened
Fang Han's argument is direct: many enterprises equate AI value with how many tokens they spent, but that is only a cost-side metric and does not equal output. What truly makes models effective is the automated engineering system built by coding agents—code generation, testing, deployment, and monitoring chained into a pipeline—so the model moves from "can chat" to "can deliver." He also revealed that Kunlun Tech is still continuously training models and will release music models, embodied-world models, and game-world models, signaling the company treats "models and compute" as the long-term foundation of an AI company rather than a short-term wrapper. The emphasis on self-developed models is a strategic bet that the durable moat is owning the base layer, because wrappers can be cloned while a distinctive model trained on proprietary data and workloads is harder to replicate overnight. Fang's point about continuous training signals that Kunlun does not see model quality as a one-time achievement but as an ongoing operating cost, which reshapes how a company should budget for AI across the year rather than per launch.
Technical Details
Behind this view is a redefinition of the "AI-native organization": not everyone using a chat box, but distilling model capability into reusable, auditable engineering assets. The coding agent plays the role of connector here—it constrains the uncertain output of large models into the deterministic processes of version control, CI/CD, and accountability. Fang specifically warned that while AI coding brings efficiency, it also accumulates technical debt: auto-generated code without review can multiply production incidents severalfold, so code review and accountability must be strengthened in step. The warning is the part often missing from AI optimism: speed of generation does not equal safety of deployment, and unless the agent's output is gated by the same review rigor as human code, the organization is trading visible velocity for invisible risk that surfaces only during an outage. The warning on technical debt is also a management instruction: pair every agent-assisted commit with a human reviewer and a clear owner, so speed today does not become an outage next quarter that no one can trace.
Comparison With Competitors
Compared with players who only sell API calls or wrapped applications, Kunlun Tech emphasizes self-developed models and vertical world models, taking a "heavy model" route. Fang's judgment also echoes current industry reflection: as model capabilities converge, the real moat lies in who can engineer and productize the model. Downgrading the token economy to a cost item and upgrading model capability to an asset item is a more sober positioning. The contrast is between companies that rent intelligence and companies that own it; both can ship features today, but the latter retain optionality to differentiate on quality, cost, and latency as the market matures, while the former remain exposed to whoever controls the underlying API. Against pure application layers, Kunlun's bet on self-owned models and world models is a longer payoff curve, trading slower time-to-market for a foundation that competitors cannot simply rent away from the same supplier.
Industry Impact and Use Cases
To put it bluntly, this speech reminds teams blindly chasing volume: the KPI for AI investment should shift from "how much consumed" to "how much delivered." For enterprises preparing to build AI-native teams, the focus is not procurement quotas but building R&D and operations systems with agents as the skeleton, plus matching review and accountability. The layout of music, embodied, and game world models also hints that multimodality and simulation will become the main battleground of the next phase of model competition. The practical takeaway for leaders is to instrument AI spend against outcomes—tickets closed, features shipped, incidents avoided—so that the technology earns its budget on evidence rather than on enthusiasm. The speech is ultimately a call to measure AI like engineering, with reviews, ownership, and delivered outcomes, which if adopted widely would mature the whole sector past the demo phase and into accountable production.