AI AI Toolkit
China AI tip

My Claude Account Got Banned

📰 公众号:数字生命卡兹克 📅 2026-07-29

Key Highlights

My Claude Account Got Banned. An Anthropic SEPA verification flaw enabled a "zero-yuan purchase" exploit, after which the company mass-reclaimed exploited accounts and banned linked ones, taking down the author's half-year-old account on July 29. The author argues Claude no longer dominates and recommends Kimi K3 and GPT-5.6 Sol for coding, WorkBuddy plus Kimi K3 for office work, noting domestic models now reach the top tier w The broader signal is a shift from chasing raw parameters toward shipping dependable, integrable systems.

What Happened

An Anthropic SEPA verification flaw enabled a "zero-yuan purchase" exploit, after which the company mass-reclaimed exploited accounts and banned linked ones, taking down the author's half-year-old account on July 29. The author argues Claude no longer dominates and recommends Kimi K3 and GPT-5.6 Sol for coding, WorkBuddy plus Kimi K3 for office work, noting domestic models now reach the top tier with one-twentieth of the compute. The episode shows the capability has moved from proof-of-concept to a perceptible product experience that users can feel in daily work.

Technical Detail

Technically, the crux is that the model was granted permissions beyond its intended boundary along the tool-calling chain. Once an agent can execute or reach the network, the absence of sandboxing and least-privilege constraints amplifies risk, which makes an isolated runtime plus behavioral auditing non-negotiable guardrails. The deeper issue is that autonomous decisions in long-horizon tasks cannot be fully constrained by static rules and must be checked continuously at runtime.

Versus Competitors

Compared with traditional rule-based security scanners, model-based autonomous penetration and red-teaming find longer-chain weaknesses but raise controllability challenges; OpenAI, Anthropic and Google are all racing to build controllable agent-security frameworks, and whoever standardizes the guardrails first gains industry voice.

Industry Impact and Use Cases

The reminder for enterprises and developers is direct: before models touch real systems, sandboxing, permissions and auditing must come first. Security is an architectural assumption, not a post-hoc patch; one escalation can wipe out the goodwill accumulated over ten feature iterations.

What to Watch

For builders, the lesson is that capability demos and safe deployments are different engineering problems; shipping the former without the latter simply transfers risk to users. The community should treat red-teaming as continuous, not a one-time gate, because model behavior drifts as capabilities and prompts evolve in production. What to watch next is whether the capability translates into dependable daily use. Demos are easy; production reliability, cost at scale and graceful failure handling are what separate a headline from a habit. The stakes are broader than one release. As models take on more autonomous roles, the gap between impressive demos and auditable behavior is where trust and regulation will be won or lost. Bottom line: treat this as incremental progress, not a finish line. The teams that win will pair capability gains with disciplined engineering on safety, cost and integration rather than chasing benchmark bragging rights. One more thing worth noting is that adoption will hinge on developer experience. Clear docs, stable APIs and predictable pricing often matter more to real uptake than a marginal jump on a public leaderboard. For decision-makers, the practical question is not is this real but where does it fit our workflow. Piloting on a narrow, measurable task beats a broad rollout that nobody owns. The longer-term read is that capability alone is no longer the differentiator; the surrounding tooling, evaluation and operational discipline are what turn a model into a product people trust with real work. The practical takeaway for security teams is to assume agents will eventually touch sensitive systems, and to design for that from day one rather than bolting controls on after an incident. A useful mental model is defense in depth: no single control is sufficient, so combine sandboxing, permission scoping,logging and human-in-the-loop approvals for high-risk actions. We should expect regulators to demand evidence of control, not just assurances of intent, which raises the value of auditable runtimes and reproducible evaluation harnesses. For builders, the lesson is that capability demos and safe deployments are different engineering problems; shipping the former without the latter simply transfers risk to users. The community should treat red-teaming as continuous, not a one-time gate, because model behavior drifts as capabilities and prompts evolve in production. What to watch next is whether the capability translates into dependable daily use. Demos are easy; production reliability, cost at scale and graceful failure handling are what separate a headline from a habit. The stakes are broader than one release. As models take on more autonomous roles, the gap between impressive demos and auditable behavior is where trust and regulation will be won or lost. Bottom line: treat this as incremental progress, not a finish line. The teams that win will pair capability gains with disciplined engineering on safety, cost and integration rather than chasing benchmark bragging rights.