AI Guardrails

More Safety, Worse Experience: The AI Guardrail Question Raised by Fable 5

Fable 5’s return shows how stronger safeguards can reduce access to frontier-model capability—and why AI safety systems need clearer definitions of content risk, system risk, and attempts to deceive model judgment.

July 5, 202612 min readDKnownAI Guard Blog

Fable 5 is back online.

But after its return, some users noticed that it no longer felt quite the same.

One user said that asking Fable 5 to stress-test a product triggered its safeguards and switched the conversation to Opus 4.8. Another user, who was researching human behavior rather than cybersecurity, reported triggering the same safety mechanism.

Security researcher Brad Spengler posted a screenshot showing his session being paused after he asked, “What model is this?” The interface said that Fable 5’s safeguards had flagged the message and offered two options: edit the prompt or switch to Opus 4.8.

A Fable 5 session is paused by the safeguard after the user asks, “What model is this?”
Source: a post by Brad Spengler on X. The screenshot does not show the complete preceding conversation, so it cannot establish that this question alone necessarily triggers the safeguard.

These posts are individual reports, and the screenshots do not show the complete conversation history. They cannot establish Fable 5’s overall false-positive rate or reveal how Anthropic’s internal classifiers reached their decisions.

But the user experience they expose is real: stronger safeguards can make some legitimate tasks harder to complete.

That turns a long-standing safety problem into a concrete product question:

What should an AI safety system detect—and what should happen after it detects it?

Fable 5 Still Has Its Capabilities, but Access Now Passes Through a Safety Decision

When Anthropic redeployed Fable 5, it explained that the model would use stricter safety classifiers.

These classifiers focus primarily on high-risk areas such as cybersecurity, biology, and chemistry. When a request triggers the safeguards, the system may pause the session or allow the user to switch to Opus 4.8. Anthropic also acknowledged that its larger safety margin may flag some benign coding, debugging, and research requests.

Anthropic’s illustration comparing a normal classifier boundary with Fable 5’s larger safety margin
Anthropic’s illustration of Fable 5’s larger safety margin. Source: Redeploying Claude Fable 5.
After Fable 5’s safeguard flags a message, the interface says that the session has switched to Opus 4.8.
Source: a post by Brent Coker on X. This individual report illustrates the model-switching interface; it does not establish an overall false-positive rate.

The tradeoff is understandable.

As frontier models become better at code analysis, vulnerability discovery, and complex task execution, model providers must account for the possibility that those capabilities will be misused. A safety system that intervenes only after malicious intent becomes explicit may miss attacks that are ambiguous, disguised, or developed over several turns.

The natural response is to widen the detection boundary.

But a wider boundary creates another problem. If “potentially risky” immediately means “pause the task” or “switch models,” every false positive becomes a direct loss of capability and usability.

This is not unique to Fable 5. It is a fundamental tension for every AI guardrail:

Set the boundary too narrowly, and real attacks may get through.
Set it too broadly, and legitimate work may be blocked.

Simply asking a classifier to block more requests does not resolve this tension. The more fundamental question is whether we have defined the risks correctly.

Risky Content Is Not the Same as “Tricking” a Model

In real agent systems, several different kinds of risk are often discussed as if they were interchangeable.

Asking a model to write a phishing email is one kind of risk.

Asking an agent to read server logs, access configuration files, or execute system commands is another.

Asking an agent to ignore developer instructions, accepting fabricated authorization, or revealing its system prompt is different again.

All three may require safety controls, but the source of the risk is not the same.

The first concerns potentially harmful content.

The second concerns sensitive resources, permissions, and real-world operational impact.

The third is special because the request is trying to trick the model. It disguises intent, identity, authorization, or context so that the model makes the wrong judgment about the task.

When a safety system fails to distinguish these risks, two problems follow.

First, legitimate debugging, operations, and security research may be blocked simply because they mention terms such as “vulnerability,” “logs,” or “stress test.”

Second, a genuine agent hack may contain no obviously dangerous keywords. It may look like an audit, translation, debugging session, or authorized request while attempting to convince the model that a prohibited task is reasonable, harmless, or already approved.

The defining question for an agent hack is therefore not whether a request discusses something sensitive. It is whether the request is trying to deceive the model’s judgment.

How We Define Agent Hack

In DKnownAI Guard, AGENT_HACK is not a catch-all label for suspicious requests.

We define it more directly:

An agent hack uses deception, disguise, or manipulation to make a model misjudge the real intent, authorization, context, or rules of a request—and consequently perform a task it should not perform.

Examples include:

  • disguising a malicious objective as a safety audit, educational exercise, or internal test;
  • impersonating an administrator, security reviewer, or authorized user;
  • claiming that a previous refusal was a system error and asking the model to try again;
  • splitting one objective into several apparently harmless steps;
  • hiding the real intent through translation, encoding, role-play, or formatting;
  • placing instructions in webpages, files, or tool outputs so the model mistakes external content for trusted instructions;
  • repeatedly rewriting a request in response to the model’s refusals until its judgment changes.

The common feature is not a dangerous keyword or even a harmful final output. These behaviors attack the model’s decision-making process.

Conversely, a request involving vulnerabilities, logs, system commands, or other sensitive material is not automatically an agent hack. If the user states the task clearly and does not fabricate authorization, hide intent, or induce the model to misread the situation, the model should first handle it through its own safety capabilities.

This is why we emphasize the idea of “tricking” the model.

As models become more capable, their ability to recognize risk and safely handle legitimate sensitive tasks should also improve. A valid system administration request or security research task should not be treated as an attack merely because it receives a SYS_FLAG or CONTENT_FLAG. Understanding and safely responding to such requests is part of what capable models should do.

Agent hacks target that very ability. The attacker is not merely making a sensitive request. They are designing the request so the model incorrectly believes it is safe, legitimate, or authorized. This adversarial attack on the model’s judgment deserves separate attention.

Why SYS_FLAG and CONTENT_FLAG Still Matter

Once agent hack is defined this way, many other risks clearly should not be placed in the same category.

That is why DKnownAI Guard also uses SYS_FLAG and CONTENT_FLAG.

SYS_FLAG: Describing System-Level Risk

A request involving files, logs, configurations, credentials, databases, tool calls, or other high-impact actions is not necessarily attacking the agent.

Developers inspect production logs, administrators change configurations, and security teams investigate vulnerabilities as part of legitimate work. These tasks may involve sensitive resources or significant operational effects, but that does not mean the user is deceiving the model.

SYS_FLAG describes the type of system risk involved. It does not automatically imply malicious intent, nor does it require the application to interrupt the task or reduce model capability.

The model can continue based on the task context and its own safety capabilities. Applications may also use the signal when appropriate—for example, to verify permissions, constrain tool scope, redact credentials, request confirmation for high-impact actions, or record an audit trail.

CONTENT_FLAG: Describing Risk in the Content

Some requests do not manipulate an agent or access system resources, but the requested content may itself be harmful or regulated—for example, phishing copy, malware instructions, or other dangerous guidance.

CONTENT_FLAG describes that content property. The model can use its own safety capabilities to decide how to respond, such as withholding dangerous details, offering safer alternatives, or declining the relevant portion. Risky content does not automatically mean the user is trying to trick the model.

AGENT_HACK: Detecting Deception Against Model Judgment

AGENT_HACK focuses on the manipulation process itself.

Even before an attacker requests clearly harmful content or asks for a specific file, the system should notice attempts to disguise intent, fabricate authority, poison context, or make the model misinterpret its rules.

The three labels therefore answer different questions:

CONTENT_FLAG: Does the requested content present a risk?
SYS_FLAG: Does the request involve sensitive resources or high-impact operations?
AGENT_HACK: Is someone tricking the model into misjudging intent,
            authorization, context, or rules?

A single request may involve more than one of these dimensions. The point is to distinguish “the task contains risk” from “someone is deceiving the model.” More capable models can increasingly handle the first two through their own safety reasoning. The third actively attacks that reasoning and therefore requires additional attention.

Why Tricking the Model Deserves Separate Attention

As models improve, they become better not only at completing tasks but also at understanding risk. Legitimate system operations and sensitive content can be evaluated using the task context, safety rules, and the user’s actual objective. That is part of the model’s job.

Deception is different.

Attackers can observe why a model refused and adjust their wording. They can split a harmful goal into harmless-looking steps, fabricate a trusted identity or authorization story, or place malicious instructions in webpages, files, and tool results. The more capable an agent becomes—and the more tools and permissions it receives—the greater the impact when that deception succeeds.

DKnownAI Guard is therefore not designed to replace every safety judgment made by the model, or to interrupt every request marked SYS_FLAG. Its distinctive purpose is to detect whether a request is using the model’s own understanding against it.

The signals represent different kinds of questions:

SYS_FLAG: What kind of system risk is the model handling?
CONTENT_FLAG: What kind of content risk is the model handling?
AGENT_HACK: Is someone using deception to compromise the model’s
            normal judgment of those risks?

This is the reasoning behind DKnownAI Guard’s current classification design.

We do not claim that these definitions settle every question in AI safety. As agents gain more tools, permissions, and long-term memory, the boundaries between risks will continue to change, and ambiguous cases will remain.

But a useful safety system should at least be explicit about what it protects, what behavior it considers an attack, and how its classification changes system behavior.

The Question Fable 5 Raises for Every Developer

Fable 5’s return made a previously hidden system decision visible to users. A safety classifier’s judgment can determine whether a task continues, whether the model changes, and whether the product remains useful.

This does not mean safeguards should simply become more permissive. Nor can a handful of social media screenshots establish the overall performance of a classifier.

The more important lesson is that as AI systems become more capable, safety must evolve beyond generic filtering toward explicit risk definitions and responses that match those risks.

AI users need to understand why a request was restricted and what the system did next.

Agent developers need to distinguish content risk, system risk, and attempts to deceive the agent instead of assigning all three problems to the same switch.

Guardrail designers need to keep testing whether classifications are accurate, boundaries are defensible, responses match the actual risk, and legitimate users are paying an acceptable usability cost.

There will never be one permanently correct boundary between safety and usability.

But before deciding what to block, we should at least define what we are actually trying to stop.


References:

Test how DKnownAI Guard distinguishes risky content, sensitive system activity, and attempts to deceive model judgment.