Tsinghua Open-Sources AgentWard to Secure Enterprise AI Agents

Tsinghua Open-Sources AgentWard to Secure Enterprise AI Agents

A Tsinghua University team has released AgentWard, an “end-to-end defense operating system” for autonomous AI agents, aiming to close a security gap that is increasingly blocking enterprises from granting models real execution rights in production systems.

The project—published in 2026 with code on GitHub—positions itself as infrastructure rather than a prompt-filter plug-in. It targets risks unique to agents that can plan, call tools, write to long-term memory and execute commands, where failures can translate into unauthorized data changes, privilege misuse and hard-to-trace automated attack chains.

Early deployments are being tested with OpenClaw-compatible agent stacks including Laikeclaw and “Lobster” agents, with field verification underway in Hainan province and Hangzhou’s Fuyang district. The team said the system has served more than 50,000 users and blocked over 95% of typical attack risks observed in tests.

Shifting Capabilities Forces Security From “Speech” to “Execution”

Large language models in China have been moving from dialog assistants to agents embedded in business workflows, expanding the threat surface from content compliance to full-stack operational risk. Once an agent can read files, load plug-ins, update memory and run shell commands, attacks shift toward indirect prompt injection, memory poisoning, intent drift and high-risk command execution—failure modes that traditional semantic filtering was not designed to audit or control.

That mismatch matters for buyers. Enterprises assessing agent rollouts increasingly price in not only model accuracy but also operational safety, traceability and the ability to enforce boundaries when agents interact with core systems such as data pipelines, internal tools and production environments.

Building Five Layers Reframes Agent Security as a “Closed Loop”

AgentWard’s architecture maps controls to an agent’s workflow: boot, perception, memory, decision and execution. It includes base scanning to verify dependencies, plug-ins and skills before loading; input sanitization to stop indirect prompt injection from files, logs or web content; cognition protection to monitor and block malicious writes into long-term memory stores such as MEMORY.md; decision alignment to catch actions that deviate from user-authorized intent; and execution control to deny high-risk commands such as destructive deletes or infinite loops at the final gate.

The emphasis on explainability—flagging file locations, rule hits and reasons—signals a push toward auditable controls, a key requirement for regulated sectors and internal risk committees that must sign off on delegating execution privileges to software agents.

OpenClaw Compatibility Positions It as a Supply-Chain Security Layer

By integrating with frameworks such as OpenClaw, AgentWard is effectively pitching itself as a “unified access and trusted runtime” layer for heterogeneous agents. That is a supply-chain proposition as much as a security one: scanning third-party skills and plug-ins addresses a growing weak point as agent ecosystems expand and teams reuse components that may be spoofed, repackaged or quietly granted excessive permissions.

If adoption scales, the likely impact will be upstream pressure on agent-framework providers and enterprise AI platforms to standardize policy enforcement, memory governance and execution-time guardrails—capabilities that customers may start to treat as table stakes rather than optional add-ons.

Field Tests Put Metrics on a Bottleneck for Commercialization

The team said real-world trials reduced unsafe or unstable incidents in Claw-based systems and intercepted more than 95% of typical attack risks, while serving over 50,000 users across pilot deployments. Those numbers, while not independently verified, reflect a market reality: agent pilots often stall when organizations cannot quantify residual risk or demonstrate controls across the full lifecycle from input to execution.

AgentWard’s release adds a concrete, developer-oriented option for teams trying to move agents from sandboxed demos into production—where the question is less whether the model can act, and more whether the system can veto.

Related Coverage:

Tsinghua Unveils Ultra-Flexible AI Chip That Bends 40,000 Times, Costs Just 2 Cents

Tsinghua-Backed eVTOL Maker Secures Series A Funding to Advance Hybrid Tilt-Rotor Fleet

Subscribe to ChinaBiz Insider

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe