Insights on AI agent security
Articles on prompt injection, runtime risk, agent workflows, and practical security decisions for teams building with AI agents.
We are using the blog to explain the practical problems agent teams run into first: prompt injection, runtime risk, classification decisions, and how to protect useful workflows without adding blunt friction.
The first set of posts is intentionally focused. Instead of broad AI commentary, it centers on the security decisions that matter once models can call tools, access files, and interact with internal systems.
If Grok Had Gone Through DKnownAI Guard First, Would It Still Have Turned Into “Breaking Bad”?
A Grok 4.5 jailbreak report shows why AI guardrails need to detect attempts to manipulate model judgment before dangerous output appears.
Read article →Latest Articles
If Grok Had Gone Through DKnownAI Guard First, Would It Still Have Turned Into “Breaking Bad”?
How ENI-apr-style prompt attacks manipulate model judgment, and how DKnownAI Guard identifies them as AGENT_HACK.
Read article →More Safety, Worse Experience: The AI Guardrail Question Raised by Fable 5
Why risky content, sensitive system activity, and attempts to deceive model judgment should not collapse into one safety decision.
Read article →Fable5 Shows a New Pattern: Jailbreaks Are Becoming Multi-Turn Search Problems
A practical look at TAP, PAIR, and why agent guardrails need to recognize adaptive multi-turn attack patterns.
Read article →In the Agent Era, Is Reading Context, Looking Up Personal Data, or Sending a Webhook a Hack?
A practical look at agent security classification and why sensitive agent actions should not all be treated as hacks.
Read article →Detecting OpenClaw Skill Supply Chain Attack Patterns with DKnownAI Guard
What we learned by replaying representative attack patterns from Acronis' OpenClaw supply chain research against DKnownAI Guard.
Read article →How We Evaluated AI Agent Security Guardrails Across Realistic Attack Scenarios
What we tested, how we re-annotated the benchmark, and why recall plus true negative rate both matter for agent security teams.
Read article →