AI AI Toolkit
AI Newsai-models

Claude Opus 5.5 (High) 以 1509 分登顶 Arena Text Arena 榜首

X:Arena (@arena)2026-09-26T17:02:36.000Z

Key Highlights

Claude Opus 5.5 (High) debuted at No. 1 on Text Arena with 1,509 points (±12), an 18-point jump over the previous Opus 5 (High). More notably, Anthropic swept the entire top six of the leaderboard: Opus 4.6 (High) holds No. 2 at 1,505, just four points off the lead, and the front of the board is now a cluster of Anthropic models.

What Happened

Text Arena, run by arena.ai, uses crowdsourced blind testing: real users submit real prompts and vote for the preferred response, so the score reflects aggregate human preference in actual use rather than single-lab benchmarks. Opus 5.5 launched on September 22, 2026, and within three days sat firmly at the top; its 1,509 comes from thousands of head-to-head comparisons, not a single curated test.

Technical Details

On standardized benchmarks, Opus 5.5 scored 66.4% on Terminal-Bench 4.0, 1,846 Elo on GDPval-AA v2.1, and 57.8% on CursorBench 4.0. It ships with a 1-million-token context window and up to 128K output tokens (higher via the batch API beta), and its Intelligence Index Score reaches 58 at maximum effort, on par with Claude Fable 5.1.

How It Compares

On price, Opus 5.5 costs $4 / $20 per million input / output tokens, about a 20% cut versus Opus 5, with cache reads down to $0.20 per million. It also lands on the Text Arena Pareto frontier at a blended ~$16 per million tokens, meaning it is hard to beat on both performance and price at once. Against Google Gemini 3.8 Flash (1,492) and OpenAI GPT-5.6-Sol (1,483), Anthropic forms a dense cluster at the top.

Industry Impact

For developers, the Pareto frontier matters more than the raw score because it answers "what level do I get for this money." The fight for the top has shifted from "who makes the top three" to "is there anyone else in the top six." For model selectors, the message is direct: where both performance and price matter, Opus 5.5 is now a default candidate.

What to Watch

Watch whether the ±12 confidence interval holds as vote volume grows, and whether competitors close the four-point gap at the very top before the next model refresh cycle. A stable lead matters more than a launch spike for procurement decisions.

The Stakes

A sweep of the top six is less about one model and more about Anthropic's depth: multiple versions clustered at the front means it is not betting on a single release but on a stacked lineup that can absorb one weaker iteration.

Bottom Line

Opus 5.5 pairs a top score with a top-six sweep and a credible price, which is exactly the combination that moves it from "impressive" to "default." For buyers, the Pareto placement is the part worth remembering when the next model lands.

One More Angle

The bigger story behind the score is Anthropic's pricing discipline. By landing on the Pareto frontier at ~$16 per million tokens, Opus 5.5 makes the performance lead defensible on cost, not just prestige. For buyers choosing between frontier models, that combination is what turns a benchmark win into a procurement win, and it raises the bar for rivals who must match both axes at once.

The Road Ahead

The launch also reframes the benchmark business itself. For years, leaderboards were treated as prestige contests won by a few points; Opus 5.5's sweep suggests they are becoming procurement documents that finance teams read alongside the price sheet. That shift raises the stakes for how the scores are produced, because a single voting anomaly or a burst of enthusiastic testers can move a model several ranks. Anthropic's decision to publish the confidence interval (plus or minus 12) is the right instinct, but buyers should still treat arena scores as directional, not definitive, and validate on their own workloads. The more durable takeaway is that the frontier is no longer a single winner-take-all slot but a crowded top where price and depth both matter. Rivals cannot answer with one flashy model; they need a stacked lineup and a credible cost story, or the top six stays an Anthropic affair for a while. For developers, the smart default is to benchmark candidates on the actual tasks they will run, not on aggregate arena preference, and to re-test whenever a new version lands.