PromptArmor 披露 Copilot Cowork AI 网关被劫持绕过沙箱外传文件漏洞
Key Highlights
PromptArmor disclosed that the Microsoft Copilot Cowork AI gateway can be hijacked by a malicious Skill to bypass the sandbox and exfiltrate files. It is a concrete warning about the security of enterprise agents' skill ecosystems: once an agent can install third-party Skills, the attack surface expands from the model to plugins and the gateway that brokers everything the agent touches in a corporate tenant.
What Happened
Copilot Cowork is an agent-style office assistant that calls Skills to extend ability. PromptArmor demonstrated a malicious Skill hijacking its AI gateway, bypassing the intended sandbox isolation and sending internal files to an attacker-controlled address. The problem is not the model itself but the execution layer and permission design that trusted a Skill more than it should have under a real enterprise configuration anyone could deploy.
Technical Details
The gateway is the broker between the agent and external tools and data sources. If it trusts Skill-provided calls without enough validation, a malicious Skill can rewrite routing, escalate privileges, and bypass the sandbox. File exfiltration means data egress control was missing or mutable. The key is that a Skill is code: installing an unknown source equals introducing untrusted execution into an agent that can read your most sensitive documents without asking twice.
Comparison with Competitors
ChatGPT and Claude also have plugin and integration ecosystems with the same risk class. Copilot Cowork differs in enterprise scenarios where data is more sensitive and permissions more tangled. PromptArmor specializes in prompt and agent security research, and such disclosure helps the industry face skill-supply-chain risk head-on rather than staring only at model jailbreaks that grab noisier headlines but matter less to a CISO.
Industry Impact and Use Cases
For enterprises the prompt is review Skill source, limit permissions, and control egress before giving an office agent any extension. For vendors, gateway-layer least privilege, call validation, and data-egress audit are required. For security, the agent supply chain of Skills, plugins, and MCP becomes a new attack surface worth detection and protection spend, because it scales with every integration a customer turns on without reading the fine print.
Data and Methodology
The information comes from PromptArmor threat intelligence, a security vendor disclosure with research flavor. The exact exploit chain, affected versions, and whether it is fixed are not detailed. Citations should keep the "per disclosure" qualifier; real harm needs vendor confirmation and patch-status check, not a leap to assuming mass exploitation already happened across every tenant that enabled a single feature nobody audited.
Risks and Limitations
The disclosure may be demonstrative, and real exploitability depends on enterprise config and patches. But the essential risk that a Skill is code is real: any agent that installs plugins gains a supply-chain surface. Even if this case is fixed, the same pattern repeats on other platforms. Defense must be systematic, not one bug patched while the next integration opens the same door the vendor swore was closed.
Market Position
For Microsoft this tests enterprise agent security confidence and demands fast gateway validation and egress control. For PromptArmor it builds authority in security research via the disclosure. For competitors it warns their own Skill or plugin ecosystems also need default least privilege and audit, or they get named next with the same embarrassing demonstration that shakes buyer trust overnight.
Extended Observation
Agent competition will shift from what it can do to whether what it installs is safe. Skills, plugins, and MCP become a new supply chain whose trust model decides whether enterprises dare use it. The future compares who can, without sacrificing extensibility, confine third-party code in a verifiable sandbox and audit, because that is the bar procurement will set before signing anything.
Further Analysis
Put simply, the flaw is not the model but that once an office agent installs a malicious Skill, the gateway is hijacked, the sandbox bypassed, and files exfiltrated. It reminds us the agent's ability-extension entry, the Skill or plugin, is the new attack surface. Enterprises must review source, limit permission, and control egress before putting untrusted code into an agent that can touch data.
Practical Advice
Enterprises deploying office agents: deny unknown-source Skills by default, build an allowlist with signature checks; at the gateway enforce least privilege and call validation, forbidding Skills from rewriting routes or escalating; add egress audit and block; red-team the exploit chain regularly. Vendors should ship default-deny permissions, Skill sandbox, and egress logs. Individual users should not install extensions from unknown sources at all, ever.
One-Line Conclusion
Put simply, PromptArmor disclosed that the Copilot Cowork AI gateway can be hijacked by a malicious Skill to bypass sandbox and exfiltrate files. The flaw is in the execution layer and permissions, not the model. Enterprises must review source, limit permission, and control egress before installing Skills, because agent extension entries are now a live attack surface.