Liquid AI 发布 d1 决策模型并新增图像输入能力
Key Highlights
Liquid AI has released d1, positioned as a decision model — not another chatbot language model, but one that folds perception, reasoning and action suggestions into a single loop. The biggest addition in this version is finally supporting image input alongside text, which is what turns it from an interesting research object into something you can point at a live screen and expect a useful move.
The framing is a deliberate departure from the "bigger is smarter" race. Most labs sell raw capability; Liquid AI is selling a shape of capability — the ability to look at a state, reason about it, and propose the next action — that is purpose-built for agents rather than for conversation. That is a smaller market today, but it is the market that agent builders actually struggle with.
What Happened
Liquid AI's earlier models emphasized liquid neural networks and efficient inference; d1 shifts the focus to decision-making: given a situation, it does not merely describe, it proposes an executable next step. With image input added, the model can now read a text instruction and look at a picture or screenshot at the same time, making tasks like "glance at the interface, then decide where to click" far more natural. That pairing of words and vision is the unglamorous capability most real workflows actually need.
The experience is already open: via console.liquid.ai and the d1 Playground, developers can try it directly without waiting in a queue. Early access without a waitlist lowers the friction for exactly the builders who might embed it, which is smart distribution for a model whose value is in being embedded rather than being admired on a leaderboard.
Technical Detail
The difference between a decision model and a conventional large model lies in the output objective. A chat model optimizes "does the next word sound human"; a decision model optimizes "is this step correct". d1 leans harder in training on state-based value judgment — akin to the value-function idea in reinforcement learning, but wrapped in a developer-friendly calling interface. Adding image input means multimodal state can be fed straight into the decision loop, without first translating it into a text description, which is where most pipelines lose fidelity.
This architectural choice has a practical payoff. When you force a screenshot through a vision-to-text step before deciding, you throw away spatial and stylistic cues that matter for UI navigation. By keeping the image in the loop end to end, d1 can reason about layout directly, which is closer to how a human operator actually looks at a screen and acts.
Versus Competitors
Against Google's Gemini or OpenAI's multimodal models, d1 does not compete on breadth of knowledge but on density of action. It is more like a lightweight agent brain suited to products needing real-time reaction, rather than something you use to write long essays. Against similar agent frameworks, its difference is that decision capability is native to the model, not bolted on by external orchestration scripts that glue a planner to a generic LLM.
That native-versus-glued distinction is the whole ballgame for latency-sensitive agents. Every external planning layer adds a round trip, a failure mode and a prompt to maintain; a model that already outputs decisions collapses that stack. d1 will not beat a frontier model at trivia, but it may beat a frontier model plus a planner at "act now, correctly, on what you see".
Industry Impact and Use Cases
For developers: in scenarios like GUI automation, game AI and robot scheduling — see the state, make the call — d1 fits better than a general-purpose model that would need a separate controller on top. For enterprises: wiring a decision model into business flows can cut manual rules in customer triage and anomaly handling, where the logic is "look at the case, decide the route". For users: behind the agents you meet in future, there may be more of these small decision-specialized models rather than one universal giant model doing everything.
The quiet trend worth noting is specialization returning through the back door. For two years the assumption was "one model to rule them all"; d1 is part of a counter-movement where small, purpose-shaped models handle the repetitive, stateful decisions that a giant model is technically capable of but economically and latency-wise poor at. That is likely where a lot of real agent value gets built.
It is worth noting what d1 is not. It is not a general assistant you would ask to write a novel, and Liquid AI is not pretending otherwise. By refusing the temptation to be everything, it carves out a defensible niche in the messy middle between a dumb script and a brilliant but unreliable chatbot — the part of the stack where most real automation actually lives.