AI AI Toolkit
📊 Updated Weekly

Model Watch

Weekly LLM rankings, capability comparison and benchmarks. Track GPT, Claude, Gemini and domestic model updates.

Rankings
Benchmarks
📅 50 models
Updated 2026-10-06

Liquid AI released the d1 decision model, adding text and image inputs to its prior capabilities, accessible via console.liquid.ai and the d1 Playground. The model targets tasks that require seeing, reasoning and deciding together.

📰 Liquid AI 模型与工程博客(网页) · 10/5/2026

Artificial Analysis locally evaluated Alibaba's Qwen-Image-2.1, released open-weight on Sept 20, which ranks 18th on both AA-Image-T2I v2.0 and AA-Image-Editing v2.0 and is the top open-weight model on both, ahead of Ideogram 4.0 (Quality) and HunyuanImage 3.0 Instruct.

📰 X:Artificial Analysis (@ArtificialAnlys) · 10/2/2026

Arena announced Xiaomi MiMo-V2.6-Pro and MiMo-V2.6-Flash on Agent Arena. Pro nets +3.17% over 8.1K real agent sessions, ranking 5th among open-weight models, up 9 places from MiMo-V2.5-Pro (-7.23%); its Confirmed Success score +7.35% leads open-weight models.

📰 X:Arena (@arena) · 10/2/2026

Artificial Analysis data shows GPT-6.1 Sol's cost per task is about $0.72, roughly 30% lower than GPT-6 Sol at $1.05, which is already about half of GPT-5.6 Sol at $1.99.

📰 X:Artificial Analysis (@ArtificialAnlys) · 10/1/2026

Google DeepMind released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two text-to-speech models that let users design voices from scratch with natural-language prompts, clone a voice from a 30-second sample, and offer line-by-line performance direction, long-form audio, and two-speaker staging across 100-plus languages.

📰 Google DeepMind:Blog(RSS) · 9/23/2026

Xiaomi released MiMo-V2.6 Pro and Flash, open-weight full-modal models trained via large-scale trial-and-error learning; Pro scores 46 on the Artificial Analysis Intelligence Index, the highest among open-weight models, and matches Claude Opus 5 and GPT-5.6 Sol on most agent benchmarks.

📰 X:Aravind Srinivas(Perplexity CEO) (@AravSrinivas) · 9/23/2026

Qwen released Qwen-Audio-3.1 covering ASR, TTS, and Realtime, and added audio-creation models TTS-Next and ASR-Next, for a total of five models spanning understanding, generation, interaction, and creation. Prices dropped sharply: TTS about 70% off, Realtime about 85% off, ASR up to 95% off; Realtime can listen while speaking and slow down with empathy when low mood is detected.

📰 X:通义千问 / Qwen (@Alibaba_Qwen) · 9/23/2026

Qwen's team publicly released Qwen-Image-2.1 with free, open weights that unify text-to-image and image editing in one model, using a 7B visual component and natively supporting generation and editing of transparent images, with up to 10 reference images and local, scribble, or independent masks.

📰 Qwen:Blog Retrieval(API) · 9/20/2026

Chinese AI lab StepFun released Step 5 Preview, a flagship base model using a sparse multi-expert architecture with 600B total and 27B active parameters, supporting a 1-million-token context and text plus vision input, scoring 44 on the Artificial Analysis Intelligence Index and ranking top three among openly usable models at one-eighth the cost of Claude Opus 5.

📰 公众号:阶跃星辰(Step) · 9/20/2026

Qwen released the next-generation native omni model Qwen3.8-Omni-Flash, supporting text, image, audio, and video input with a 1M-token context. Its average score across 29 benchmarks rose over 25 percent versus Qwen3.5-Omni-Plus, while audio input price per hour fell over 98 percent and audio-video input price over 93 percent.

📰 Qwen:Blog Retrieval(API) · 9/18/2026

Microsoft AI CEO Mustafa Suleyman argued against the "model welfare" idea, saying AI has no consciousness, feels nothing, and experiences no pain, and granting it a right to care would make alignment and control harder or impossible.

📰 X:Mustafa Suleyman(Microsoft AI CEO) (@mustafasuleyman) · 9/16/2026

DeepSeek HuggingFace free DeepSeek -V4.1-Flash model Baseten Model call interface s shipped 552B parameters prefill 8B decode 16B 1M token context supports image

📰 Baseten 工程博客(网页) · 9/12/2026

WorkBuddy announced DeepSeek V4.1 -Flash shipped free trial DeepSeek said V4.1-Flash DeepSeek call interface supports native capability

📰 X:Tencent WorkBuddy (@WorkBuddy_AI) · 9/10/2026

Open AI released GPT-6 Astra ChatGPT Work Codex call interface token $10 token $50

📰 OpenAI:官网动态(RSS · 排除企业/客户案例) · 9/9/2026

Anthropic unveiled Claude Fable 5.1 and Claude Mythos 5.1, calling them its most advanced models for coding and knowledge work. Developer Thariq spent weeks testing them, says they are strong, and will publish a long review later. His advice: use lower effort for tasks needing less verification, and switching effort no longer breaks the prompt cache.

📰 X:Thariq (@trq212) · 9/2/2026

Anthropic's Claude Fable 5.1 is now on OpenRouter as a direct upgrade over Fable 5 workloads, with the biggest gains in agentic coding, long-running workflows, visual code generation, and finance and analysis. Endpoint: openrouter.ai/anthropic/claude-fable-5.1.

📰 X:OpenRouter (@OpenRouter) · 9/2/2026

Anthropic released Claude Fable 5.1 (claude-fable-5-1) for long-running autonomous coding, knowledge work, and research, and Claude Mythos 5.1 for Project Glasswing participants.

📰 Claude Platform:开发者版本说明(RSS) · 9/1/2026

Zhipu AI has open-sourced the weights of GLM-5.3, which supports local deployment and customization and excels at complex coding, defensive cybersecurity, and long-horizon tasks. It scored 60 on the AA comprehensive intelligence index, tying Kimi K3 as the top open model and matching closed flagships like Claude Fable 5 and GPT-5.6 Sol; only institutions over 10 billion US dollars in annual revenue offering it as a service need security review.

📰 IT之家(RSS) · 8/29/2026

Z.ai (Zhipu) has released GLM-5.3 as open weights, calling it its most powerful model for agentic coding and network defense. The weights are available on Hugging Face and the technical blog is published on z.ai, allowing developers to download, run, and customize the model freely.

📰 X:智谱 Z.ai (@Zai_org) · 8/28/2026