AI AI Toolkit
AI Newstip

AI 还要多久才能写得更好?

Nathan Lambert:Interconnects(RSS)2026-08-12T13:01:16.000Z

Core Highlights

After finishing a textbook about RLHF, which stands for reinforcement learning from human feedback, the author Nathan Lambert reflected on an awkward reality that many in the field would rather ignore. Progress on long-form nonfiction writing by models has clearly stalled, while the same models have raced ahead and nearly matched or surpassed humans on coding and mathematics. This gap between stagnant prose and soaring technical skill is the central puzzle he raises, and it has implications for anyone who hoped language models would soon write serious books on their own without heavy human oversight. The essay is a useful corrective to the hype that assumes every capability improves at the same pace and that writing is just around the corner from being solved.

Specific Capabilities and What Happened

The author observes that models once celebrated for writing, such as GPT 4.5 and Kimi K2, now look dated when it comes to producing long-form content that holds together for more than a few pages. Models can certainly help fix typos and perform local edits, but the moment you ask one to organize the logic of an entire chapter, its performance becomes messy and error-prone in ways that are easy to underestimate. This weakness stands in sharp contrast to the rapid advances in coding and math, and it leads the author to worry that models will struggle to autonomously push forward open-ended scientific problems that demand sustained, coherent argument across many pages of careful reasoning. The inability to plan at book length is, in his view, a real blocker for open science.

Technical Details

RLHF aligns model outputs using human feedback, and it works well on short texts and tightly scoped tasks, yet it struggles to cover the global consistency that long-form writing requires from beginning to end. A long document must keep its arguments, terminology, and style unified across dozens of pages, but current models still have weak context management and planning abilities, so chapter-level generation quality stays unstable no matter how polished each sentence feels. The training signal that makes a paragraph sound polite does not automatically produce a book that hangs together from start to finish in a way a reader can trust, because no amount of local correctness guarantees global structure.

Comparison with Competitors

On the writing dimension, the flagship models from various labs have not pulled clearly ahead of one another; instead they all stall at the same bottleneck that none has cracked. By contrast, coding and math advance far faster because their evaluations are clear and their feedback is immediate, and this imbalance is something the community should treat as a warning rather than a curiosity. Writing remains the hard case where benchmark scores do not capture the real difficulty of sustained composition, so leaderboards flatter models that still cannot outline a coherent monograph. The lesson is that progress is uneven, and prose is the laggard everyone hoped would lead.

Industry Impact and Use Cases

For users who depend on AI to assist research and writing, the conclusion is not cheerful: at least in the visible future, models are better suited to proofreading and polishing fragments than to independently authoring complete treatises. Put simply, handing an entire book to a model to write is still far from advisable, and human authors remain indispensable for structure, judgment, and the kind of coherence that only a mind trained over years can reliably supply. Teams should therefore deploy models as tireless editors, not as sole authors, and reserve the architecture of a long argument for people who can hold it in their heads. Until models can plan at the scale of a book, the realistic role for AI in serious nonfiction is the supporting one of editor and fact-checker rather than author of record. That framing protects readers from plausible but hollow prose and protects writers from overtrusting a tool that cannot see the whole arc of an argument they are building. The community should set expectations accordingly instead of promising autonomous manuscripts that do not yet exist and disappointing the people who trusted the marketing. Lambert's point is less a criticism than a clear-eyed map of where the frontier actually sits today, and it should guide how we fund and benchmark writing models going forward. It also suggests that evaluation methods need to measure coherence over length, not merely sentence-level polish, if we ever want to close the gap between fluent paragraphs and a finished book. Researchers who want help should treat models as tireless junior editors whose outlines still need a human hand on the structure.