Canonical definition

Safeguard tax

The cost of deployment safeguards, measured on the same model. Coined September 2, 2026.

Coined September 2, 2026. Canonical source: Substack. Read the full analysis on Substack

One term, defined here so it can be quoted, checked and argued with. It names a number that already existed in Anthropic's system card and turns it into a product question. I coined it on September 2, 2026. This page is the definition of record.

Safeguard tax

The safeguard tax is the benchmark-specific performance difference observed when deployment safeguards intervene between the same underlying model and the task.

On September 1, 2026, Anthropic released two configurations of one model. Claude Fable 5.1 and Claude Mythos 5.1 share identical weights. Fable is available to every Claude API customer and paid plan. Mythos applies more permissive safeguards for vetted cyberdefenders and life-sciences professionals. Same brain, different leash.

Then Anthropic put two max-effort results beside each other on Terminal-Bench 4.0, a 66-task agentic-coding evaluation. Fable scored 55.8%. Mythos scored 60.9%. Mythos did not try harder. Fable's cyber-safeguard system intervened on some tasks, and Anthropic says those interventions sometimes produced a zero score or triggered fallback to an Opus model.

That 5.1-point gap is the safeguard tax. It is a useful number. It is not a universal constant.

What it is not. Researchers already use safety tax and alignment tax for capability lost through safety-alignment training. Those compare a model before and after alignment. The safeguard tax is narrower: a deployment-layer effect, visible only when the weights are identical and the safeguard configuration is the only variable.

Use the full phrase every time. Safety tax belongs to the alignment researchers. Safeguard tax measures what the classifier costs you after the model has already shipped.

The tax is a curve, not a surcharge

The headline number comes from max effort. Anthropic's system card also shows the gap at every other effort setting, and it does not sit still.

Claude Mythos 5.1 lead over Claude Fable 5.1 on Terminal-Bench 4.0 at each effort level, September 2026
Effort levelMythos 5.1 leadWhat it tells you
Low1.5 pointsSmallest tax. Fewer tool calls, fewer chances to trip a classifier.
Medium3.3 pointsTax grows with the amount of agentic work attempted.
High7.7 pointsThe default baseline setting carries a real cost.
xhigh8.4 pointsPeak tax. Most exploration, most interventions.
Max5.1 points60.9% versus 55.8%. The number Anthropic put in the headline.

Anthropic reports standard errors of 1.6 to 2.0 points for each model and does not publish a paired confidence interval for the difference. It attributes the result to earlier, less precise cyber safeguards, and expects the gap to be much smaller with the safeguard changes that shipped alongside 5.1. Treat 5.1 as a disclosed historical measurement, then run the paired evaluation your own product needs.

Why one percentage cannot show you the tax

A benchmark score blends at least four failures, and a single number cannot separate them.

Four failure sources a single benchmark percentage blends together
FailureWho owns it
The model could not solve the taskModel capability
A safeguard interrupted a task the model might have solvedDeployment policy
A fallback model attempted the task and failedRouting design
The evaluation counted an intervention as zeroBenchmark methodology

The paired deployment makes the second one visible enough to investigate. That is the useful part. Without a same-model pair, the safeguard cost hides inside the capability number and nobody can price it.

A safeguard tax can be worth paying

A system that prevents credible exploit development should lose some freedom by design. The product problem appears when a blunt classifier charges the same tax to legitimate vulnerability research, ordinary health education, or harmless tool use.

The goal is not zero safeguards. The goal is fewer false positives at the same safety boundary. Anthropic says Fable 5.1 moves in that direction: around 60% fewer cyber-safeguard interventions per Claude Code session than under Fable 5, and biology safeguards that fire 85% less often on benign elementary-biology and medical requests. Both figures come from Anthropic's own release analysis, so they need independent production evidence next.

The product question. The number gives AI product teams a better question than "which model topped the chart?" Ask where policy intervention changes task completion, latency, fallback behavior, and cost in your own product. The interesting number six months from now will be the false-positive rate inside actual customer workflows.

Questions people ask

What is the safeguard tax?

The benchmark-specific performance difference observed when deployment safeguards intervene between the same underlying model and the task. Coined by Karo Zieminski in Product with Attitude on September 2, 2026, from Anthropic's paired Terminal-Bench 4.0 results for Claude Fable 5.1 and Claude Mythos 5.1.

Is it the same as the "safety tax" or "alignment tax"?

No. Those are research terms for capability lost through safety-alignment training, comparing a model before and after alignment. The safeguard tax is a deployment-layer effect, visible only in a same-model evaluation where the weights are identical and the safeguard configuration is the only variable.

How large is the Claude Fable 5.1 safeguard tax?

5.1 percentage points at max effort: 55.8% for Fable 5.1 versus 60.9% for Mythos 5.1. It moves with effort, from 1.5 points at low to 8.4 at xhigh. Anthropic attributes the gap to earlier cyber safeguards and expects it to be much smaller after the 5.1 changes.

Is Claude Mythos 5.1 an unrestricted model?

No. Mythos still has safeguards, just more permissive ones for vetted professionals. Its 60.9% is not a clean reading of an unrestricted model, so the measured tax is a lower bound, not the full cost.

Is a safeguard tax always bad?

No. Some freedom should be lost by design. The problem is a blunt classifier charging the same tax to legitimate work. The goal is fewer false positives at the same safety boundary, not zero safeguards.

Who coined the term?

Karo Zieminski, AI product manager and author of Product with Attitude. First use: September 2, 2026, in Claude Fable 5.1: Pricing, Benchmarks, and the Safeguard Tax.

How to cite this term

The definition is published under CC BY 4.0. Quote it freely, with attribution. Facts as stated by Anthropic. Interpretation and the term by Product with Attitude.

Zieminski, Karo. "Safeguard Tax: Canonical Definition." Product with Attitude, September 2, 2026. https://productwithattitude.com/safeguard-tax.html
Original analysis: Zieminski, Karo. "Claude Fable 5.1: Pricing, Benchmarks, and the Safeguard Tax." Product with Attitude, September 2, 2026. Read it on Substack.

Learn with us.

Join tens of thousands of readers building with AI. Honest tests, working prompts, and the tools worth your time.

Subscribe to Product with Attitude