Liquid AI released the d1 decision model, adding text and image inputs to its prior capabilities, accessible via console.liquid.ai and the d1 Playground. The model targets tasks that require seeing, reasoning and deciding together.
Model Watch
Weekly LLM rankings, capability comparison and benchmarks. Track GPT, Claude, Gemini and domestic model updates.
Arena reports OpenAI's GPT-6.1 Sol (Max) ranks 5th on Agent Arena with a +11.23% net gain and a median task cost of $0.56, reshaping the Pareto frontier.
Arena's latest Agent Arena ranks Anthropic's Claude Sonnet 5.5 third with a +12.5% net gain, but at $2.74 median cost per task it is more expensive than Opus 5.5 and stays off the Pareto frontier.
Ai2 open-sourced AstaBrief 8B, built on Qwen3-8B, that turns research questions and retrieved snippets into cited reports; it is live as Fast mode in Asta's Generate a report feature.
The LMSYS team fine-tuned LLaMA on multi-turn ShareGPT conversations and trained Vicuna-13B at minimal cost; it approached ChatGPT quality across several benchmarks.
Artificial Analysis released its Coding Agent Index; Claude Sonnet 5.5 (max) leads on Claude Code at 68 points but costs up to $14.19 per task, while GPT-6.1 Sol and Gemini 4 Argon top the chart at far lower cost.
GPT-6 Astra Ultrafast is now available via the OpenAI API and to eligible ChatGPT Work and Codex users, running on NVIDIA Blackwell accelerators.
Artificial Analysis locally evaluated Alibaba's Qwen-Image-2.1, released open-weight on Sept 20, which ranks 18th on both AA-Image-T2I v2.0 and AA-Image-Editing v2.0 and is the top open-weight model on both, ahead of Ideogram 4.0 (Quality) and HunyuanImage 3.0 Instruct.
Arena announced Xiaomi MiMo-V2.6-Pro and MiMo-V2.6-Flash on Agent Arena. Pro nets +3.17% over 8.1K real agent sessions, ranking 5th among open-weight models, up 9 places from MiMo-V2.5-Pro (-7.23%); its Confirmed Success score +7.35% leads open-weight models.
Arena announced Claude Sonnet 5.5 (xHigh) entered the Code Arena: WebDev leaderboard at 1786 points, 3rd place, just 2 points behind No. 2 GPT-6 Astra at 1788.
Artificial Analysis data shows GPT-6.1 Sol's cost per task is about $0.72, roughly 30% lower than GPT-6 Sol at $1.05, which is already about half of GPT-5.6 Sol at $1.99.
Arena places Gemini 4 Argon (High) 8th on Agent Arena with a +7.92% net gain and a $0.62 cost per task.
Anthropic's Claude Opus 5.5 (High) debuted at No. 1 on Text Arena with 1,509 points, an 18-point jump over Opus 5 (High), and helped Anthropic sweep the entire top six of the leaderboard while landing on the Pareto frontier at a blended $16 per million tokens.
Arena announced Claude Opus 5.5 (Max) tops Code Arena: WebDev with 1,818 points, leading runner-up GPT-6 Astra (Max) by 26 points and beating Opus 5 (Max)'s 1,692 by 126 points.
Google DeepMind released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two text-to-speech models that let users design voices from scratch with natural-language prompts, clone a voice from a 30-second sample, and offer line-by-line performance direction, long-form audio, and two-speaker staging across 100-plus languages.
Xiaomi released MiMo-V2.6 Pro and Flash, open-weight full-modal models trained via large-scale trial-and-error learning; Pro scores 46 on the Artificial Analysis Intelligence Index, the highest among open-weight models, and matches Claude Opus 5 and GPT-5.6 Sol on most agent benchmarks.
Qwen released Qwen-Audio-3.1 covering ASR, TTS, and Realtime, and added audio-creation models TTS-Next and ASR-Next, for a total of five models spanning understanding, generation, interaction, and creation. Prices dropped sharply: TTS about 70% off, Realtime about 85% off, ASR up to 95% off; Realtime can listen while speaking and slow down with empathy when low mood is detected.
Anthropic released Claude Opus 5.5, claiming typical token-metered workloads cost about 40% less than Opus 5, with cache-read pricing down 60% and input/output token prices down 20%.
Qwen's team publicly released Qwen-Image-2.1 with free, open weights that unify text-to-image and image editing in one model, using a 7B visual component and natively supporting generation and editing of transparent images, with up to 10 reference images and local, scribble, or independent masks.
Chinese AI lab StepFun released Step 5 Preview, a flagship base model using a sparse multi-expert architecture with 600B total and 27B active parameters, supporting a 1-million-token context and text plus vision input, scoring 44 on the Artificial Analysis Intelligence Index and ranking top three among openly usable models at one-eighth the cost of Claude Opus 5.
Qwen released Qwen3.8-LiveTranslate, rebuilding real-time simultaneous interpretation with an Interleave architecture and a hybrid Thinker-Talker design, cutting average latency from 2.8 to 2.3 seconds.
Qwen released the next-generation native omni model Qwen3.8-Omni-Flash, supporting text, image, audio, and video input with a 1M-token context. Its average score across 29 benchmarks rose over 25 percent versus Qwen3.5-Omni-Plus, while audio input price per hour fell over 98 percent and audio-video input price over 93 percent.
DeepSeek released DeepSeek-V4.1-Flash, a 552B-parameter mixture-of-experts multimodal model that sees, hears, and reads text, supports up to a 1M-token context, and has weights open on Hugging Face.
Microsoft AI CEO Mustafa Suleyman argued against the "model welfare" idea, saying AI has no consciousness, feels nothing, and experiences no pain, and granting it a right to care would make alignment and control harder or impossible.
Google DeepMind released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two near-real-time voice conversation models focused on voice agents and complex task execution.
Shengshu formally launched Vidu S2, including Vidu S2-Avatar for real-time interaction with digital characters and Vidu S2-Editing for real-time editing of video streams, and is exploring real-time spatial video generation and editing for VR headsets.
StepFun released the StepAudio 3 family with Realtime, ASR, TTS, Gen, and Music models, live on the StepFun open platform, with several variants ranking first globally on Artificial Analysis leaderboards.
SiliconFlow released the open-weight model Hy4 preview on its platform: 770B total parameters, 49B active per token, 1M context, Apache 2.0, aimed at coding, analysis, research and complex real-world work.
Fireworks AI launched DeepSeek-V4.1-Flash with full benchmarks: 74.34% pass@1 on DeepSWE at the max tier, on par with GPT-6 Astra, at just $0.43 per task.
Xiaohongshu's AllSpark team released the open-weight Search Agent model Iris; weights and eval code are public, with data and training recipe to follow. The 35B and 397B versions lead at their scale.
Suno released music model v6 Warner Music Group BMG Believe industry model v6 flagship v6 v6-wild Pro Premier users v6-mini
DeepSeek HuggingFace free DeepSeek -V4.1-Flash model Baseten Model call interface s shipped 552B parameters prefill 8B decode 16B 1M token context supports image
Artificial Analysis DeepSeek V4.1 Flash Intelligence Index 40 DeepSeek V4 Pro 0813 DeepSeek new flagship parameters 8B parameters 16B supports 1M token context MIT
WorkBuddy announced DeepSeek V4.1 -Flash shipped free trial DeepSeek said V4.1-Flash DeepSeek call interface supports native capability
Suno released v6 model image video voice music v6-wild 2 video
Open AI released GPT-6 Astra ChatGPT Work Codex call interface token $10 token $50
Open AI released ChatGPT Images 2.5 image model generation Images 2.0 50% details editing editing
OpenAI has released GPT-6 Astra, reporting FrontierMath Tier 4 of 98%, ARC-AGI-3 of 99.9% and ExploitBench of 100%, while stating its cyber-capability has crossed the Critical threshold of the Preparedness Framework.
Released on September 3, GPT-6 Astra is OpenAI's first assistant to clear the Critical cyber-capability threshold of its preparedness framework, able to find unknown flaws in hardened systems without step-by-step guidance.
Greg Brockman reposted an ARC Prize evaluation saying GPT-6 Astra reached SOTA on ARC-AGI-3, which he called saturated. The standard harness scored 63%; a new Provider Adapter harness reached 99%, beating human performance on 96% of levels.
Positioned as a computer-use model, GPT-6 Astra offers a 1,050,000-token context and 128,000 max output tokens, an April 30 2026 cutoff, OSWorld V2-Offline of 72.6%, and cuts average task time from ~75 to ~40 minutes.
Perplexity CEO Aravind Srinivas congratulated OpenAI on GPT-6 Astra, calling it far ahead on broad and deep research tasks and more cost-effective, and said it will soon reach Perplexity Computer's Pro and Max users.
Anthropic unveiled Claude Fable 5.1 and Claude Mythos 5.1, calling them its most advanced models for coding and knowledge work. Developer Thariq spent weeks testing them, says they are strong, and will publish a long review later. His advice: use lower effort for tasks needing less verification, and switching effort no longer breaks the prompt cache.
Anthropic's Claude Fable 5.1 is now on OpenRouter as a direct upgrade over Fable 5 workloads, with the biggest gains in agentic coding, long-running workflows, visual code generation, and finance and analysis. Endpoint: openrouter.ai/anthropic/claude-fable-5.1.
OpenAI says Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework the first model rated at that level able to find unknown vulnerabilities and build exploit chains with little human intervention.
Anthropic released Claude Fable 5.1 (claude-fable-5-1) for long-running autonomous coding, knowledge work, and research, and Claude Mythos 5.1 for Project Glasswing participants.
On August 31, DeepSeek publicly released on Hugging Face the first integrated model that can see images, hear audio, and understand text DeepSeek-V4-Flash-Vision-Exp under the MIT License, with model files, tokenizer, instruction-encoding reference implementation, and a minimal PyTorch inference implementation.
Zhipu AI has open-sourced the weights of GLM-5.3, which supports local deployment and customization and excels at complex coding, defensive cybersecurity, and long-horizon tasks. It scored 60 on the AA comprehensive intelligence index, tying Kimi K3 as the top open model and matching closed flagships like Claude Fable 5 and GPT-5.6 Sol; only institutions over 10 billion US dollars in annual revenue offering it as a service need security review.
Z.ai (Zhipu) has released GLM-5.3 as open weights, calling it its most powerful model for agentic coding and network defense. The weights are available on Hugging Face and the technical blog is published on z.ai, allowing developers to download, run, and customize the model freely.
Tencent Hunyuan has released its new flagship model Hy4 preview, with 770B total parameters, 49B active parameters, and a 1M token context length. The model is open-source and freely usable, and is now available on Tencent Cloud's model repository and on OpenRouter.