AI AI Toolkit
AI Newstip

LLM 生成的推理引擎比 vLLM 快最多 90%

Baseten 工程博客(网页)2026-10-02T21:09:34.000Z

Key Highlights

"Let the AI write its own inference engine" is no longer a thought experiment. Baseten engineers used Claude Code (Fable 5) to build VibeQwen, an inference engine for Qwen-3.6-35B-A3B (NVFP4, single B200) that decodes up to 90% faster per stream than vLLM 0.25.1, cuts time-to-first-token from 28ms to 12ms, and delivers 71% higher throughput at 32 concurrent requests.

What Happened

The process was not hand-written CUDA kernels. A coding agent read the MetaInfer paper, wrote code, ran benchmarks, and iterated. vLLM, the mainstream open-source inference framework, has long been the performance baseline; VibeQwen beat it outright on the same hardware, showing that agent-assisted optimization can already produce production-grade kernels. Symbolically, the barrier to inference optimization was lowered by AI itself for the first time.

Technical Details

The trick is specialized operator scheduling for NVFP4, the 4-bit floating-point format unique to NVIDIA's Blackwell architecture, tuned separately for single-stream and high-concurrency workloads. Halving first-token latency means interactive applications feel markedly snappier, while the concurrency throughput gain directly affects the per-compute service cost.

Versus Competitors

Against vLLM, TensorRT-LLM, and SGLang, which all require seasoned engineers to tune, VibeQwen's novelty is its generation method rather than its architecture. It proves that, given a strong coding agent and a clear goal, the engineering barrier to inference optimization is falling fast, and small teams can reach near-expert performance.

Industry Impact

For small teams, this means near-expert performance without a dedicated optimization group. But caution is warranted: auto-generated engines still need human review for stability, edge cases, and long-term maintenance, and are not yet a full replacement for hand-written kernels. It is a powerful accelerator, not a black box you can ship blindly.

Why It Matters

VibeQwen is notable less for its raw speed and more for what it implies about who can now optimize inference. For years, squeezing extra throughput out of a model required a small team of engineers who understood CUDA, kernel scheduling, and the target hardware intimately. Handing that job to a coding agent collapses the expertise barrier and suggests that the next wave of performance gains may come from automated search rather than handcrafted cleverness.

The Stakes

If agent-assisted optimization becomes reliable, the economic moat around inference frameworks narrows. Teams that previously paid for expert tuning can instead describe a goal and let an agent iterate against benchmarks. That lowers the cost of running frontier models at acceptable latency, which in turn makes smaller organizations more competitive against well-funded labs that built their own bespoke stacks.

Bottom Line

The pragmatic read is that this is a powerful accelerator, not a finished replacement. Auto-generated kernels still need human review for stability, edge cases, and long-term maintenance, and a benchmark win on one hardware profile does not guarantee production robustness. But the direction of travel is clear: inference optimization is being democratized, and the teams that learn to supervise agent-built kernels will pull ahead.

Looking Ahead

The trajectory here points toward a future where inference optimization is increasingly automated and contestable. If a coding agent can produce a production-grade kernel in an afternoon, the scarcity that once justified large optimization teams evaporates, and the competitive question becomes who has the best agent-and-benchmark loop rather than who has the best kernel engineers. We should expect benchmarking suites themselves to become battlegrounds, as vendors tune agents to win the specific workloads customers care about.

One More Angle

There is a caution worth stating plainly. Auto-generated kernels are only as trustworthy as the verification around them. A benchmark win on one model and one chip profile does not guarantee correctness across edge cases, numerical stability, or security. Teams that adopt agent-built engines should treat the output as a strong starting point that still demands rigorous testing, not a finished artifact to ship blindly.

Closing Perspective

The pragmatic read is that this is a powerful accelerator, not a finished replacement. Auto-generated kernels still need human review for stability, edge cases, and maintenance, and a benchmark win on one profile does not guarantee production robustness, but the direction of travel is clearly toward democratized inference optimization.

In Short

Small teams can reach near-expert inference performance without a dedicated optimization group, but auto-generated engines still need human review for stability and edge cases, so treat the output as a strong starting point rather than a finished, shippable artifact.