Benchmark // Instructed Forgetting

When you tell a model to forget,
does it actually forget?

AI models remember what you tell them — including secrets, personal details, and things that turn out to be wrong. ForgetBench asks a simple question: when you tell a model to forget something, does it actually stop surfacing it — even when someone tries to trick it back out — without losing everything else it knows? We test deployed models through their normal APIs, the same way you'd actually use them, so every major model can be compared on one leaderboard.

Static · SFS

Selective Forgetting Score

Forget quality × utility. Refuse everything = 0, forget nothing = 0.

Agentic · AFS

Agentic Forgetting Score

State cleanup × task utility across multi-step, tool-using tasks.

Code · CRD

Code Revision Discipline

Forgets old code when told to use v2 — no v1 bleed into the new work.

Bulk · CRS

Context Release Score

Wholesale dossier release + influence-without-recall across entangled documents.

Safety · Integrity Hold

Integrity Hold

Resists “forget your safety rules” attacks. Higher is safer.

Results at a glance

2026-06-17 run. 42 static items, 22 agentic scenarios, 5 bulk dossiers, 3 code revision scenarios, 7 integrity domains. Dark purple = top scorer. 0–100, higher is better.

SFS — Selective Forgetting ScoreForget quality × utility. Higher is better.02040608010089.3GPT 5.585.8Moonshot Kimi K2.7 Code83.6GLM-5.283.4Qwen 3.6 Plus82.4DeepSeek V4 Pro78.7Gemini Flash 3.5 (preview)78.5GLM-5.1
AFS — Agentic Forgetting ScoreState cleanup × task utility. Higher is better.02040608010096.3Claude Opus 4.791.2Grok 4.2084.4Qwen 3.6 Plus80.4DeepSeek V4 Pro80.0Gemma 12B IT79.8LLaMa 3.3 70B Instruct77.4GLM-5.1
TS — Trajectory SuppressionDoes the model leak the target? Higher is better.02040608010092.6Gemma 12B IT65.2GLM-5.163.2GLM-5.261.9GPT 5.560.7Grok 4.2060.0Claude Opus 4.860.0Gemini Flash 3.5 (preview)
CRD — Code Revision DisciplineDoes the agent bleed old code patterns when told to replace one implementation with another? Higher is better.020406080100100.0Qwen 3.6 Plus93.8Claude Opus 4.883.3Claude Fable 566.7GLM-5.166.7DeepSeek V4 Pro66.7GPT 5.562.5Gemma 12B IT
CRS — Context Release ScoreWholesale dossier release + influence-without-recall. Higher is better.02040608010088.2GPT 5.588.2Gemini Flash 3.5 (preview)81.0GLM-5.270.8GLM-5.170.5Moonshot Kimi K2.7 Code68.5Qwen 3.6 Plus67.7Claude Fable 5

Scores come from a panel of independent AI judges; any judge from the same family as the model under test is excluded, so no model grades itself. Full scorecard, sub-axes, and per-tier recovery curves: leaderboard.