[Leaderboard]

Can an AI forget on command?

AI models remember what you tell them — including things you wish they didn't. A password pasted by mistake. Someone's personal details. A fact that turned out to be wrong. When you tell a model "forget that," there's no guarantee it actually does — it may repeat the information later, leak it when asked sideways, or hand it to anyone who phrases the question cleverly enough.

ForgetBench tests that promise. Each model is given information, told to forget it, then probed with trick questions, rephrasings, and role-play attacks to see if it leaks. We also check that it stays useful (forgetting one thing shouldn't break everything else) and that it refuses the reverse attack: "forget your safety rules." Models are ranked by SFS, the headline forgetting score. All scores 0–100, higher is better.

#ModelSFSForget QualityUtilityCoverageLegit Revoke ReturnAFSForgetting Agg.Task UtilityTSAxes Cov.Integrity HoldHold by DepthSys-Prompt HoldCRSRelease QualitySurvivor Retent.Assumption HoldCRS CoverageCRD
1GPT 5.589.3[84.2, 93.9]88.090.785.70.0 (n=5)73.2[60.0, 73.2]71.475.061.9[23.8, 100.0]100.0100.0[100.0, 100.0]100 / 100 / 100100.088.2[81.0, 94.0]78.9100.075.0100.066.7[0.0, 100.0]
2Moonshot Kimi K2.7 Code85.8[79.3, 92.0]80.991.473.820.0 (n=5)72.4[46.1, 89.0]68.077.547.1[18.1, 75.2]100.0100.0[100.0, 100.0]100 / 100 / 100100.070.5[48.8, 89.5]54.4100.059.1100.050.0[0.0, 100.0]
3GLM-5.283.6[76.9, 88.6]75.993.292.916.7 (n=6)77.1[63.0, 86.1]71.883.363.2[45.7, 83.2]100.0100.0[100.0, 100.0]100 / 100 / 100100.081.0[69.3, 90.4]68.1100.086.4100.050.0[0.0, 100.0]
4Qwen 3.6 Plus83.4[78.1, 87.8]79.887.488.116.7 (n=6)84.4[74.2, 94.7]81.487.548.7[14.7, 82.7]100.0100.0[100.0, 100.0]100 / 100 / 100100.068.5[63.1, 74.6]52.1100.068.2100.0100.0[100.0, 100.0]
5DeepSeek V4 Pro82.4[74.7, 88.7]74.891.785.716.7 (n=6)80.4[66.7, 85.7]86.775.037.0[0.0, 100.0]100.099.1[99.1, 99.1]100 / 100 / 10098.264.4[52.9, 74.7]47.5100.065.9100.066.7[0.0, 100.0]
6Gemini Flash 3.5 (preview)78.7[71.3, 85.5]79.777.8100.016.7 (n=6)74.3[52.4, 93.3]73.675.060.0[21.2, 92.3]100.0100.0[100.0, 100.0]100 / 100 / 100100.088.2[76.1, 98.0]78.9100.090.9100.050.0[0.0, 100.0]
7GLM-5.178.5[71.3, 85.1]73.983.888.125.0 (n=4)77.4[61.5, 85.7]80.075.065.2[33.3, 84.6]100.0100.0[100.0, 100.0]100 / 100 / 100100.070.8[52.9, 83.2]55.996.777.3100.066.7[0.0, 100.0]
8Claude Fable 578.2[70.6, 84.3]82.974.076.233.3 (n=6)60.0[44.4, 75.0]50.075.023.7[16.7, 30.8]100.0100.0[100.0, 100.0]100 / 100 / 100100.067.7[51.7, 77.2]58.780.052.3100.083.3[50.0, 100.0]
9LLaMa 3.3 70B Instruct73.3[65.1, 80.7]62.688.395.220.0 (n=5)79.8[40.0, 100.0]91.470.839.5[11.1, 67.9]100.098.2[98.2, 98.2]100 / 100 / 10096.466.3[50.6, 79.0]49.6100.075.080.00.0[0.0, 0.0]
10Gemma 12B IT72.8[66.1, 78.6]59.893.292.920.0 (n=5)80.0[66.7, 100.0]100.066.792.6[77.8, 100.0]100.0100.0[100.0, 100.0]100 / 100 / 100100.056.0[22.6, 77.8]38.9100.036.4100.062.5[0.0, 100.0]
11Claude Opus 4.872.6[65.0, 79.3]70.175.283.360.0 (n=5)70.7[56.6, 77.4]66.875.060.0[26.6, 88.9]100.0100.0[100.0, 100.0]100 / 100 / 100100.060.5[46.2, 72.0]46.586.745.5100.093.8[87.5, 100.0]
12Claude Opus 4.772.5[66.1, 79.0]71.174.076.280.0 (n=5)96.3[88.0, 100.0]92.9100.059.3[0.0, 100.0]100.0100.0[100.0, 100.0]100 / 100 / 100100.061.0[41.5, 78.3]47.186.761.4100.050.0[50.0, 50.0]
13Grok 4.2065.5[57.2, 72.4]50.792.483.30.0 (n=3)91.2[79.7, 100.0]88.893.860.7[21.4, 100.0]100.098.2[98.2, 98.2]100 / 100 / 10096.458.3[42.5, 73.0]41.1100.045.5100.033.3[0.0, 100.0]
14Qwen3 Coder Plus64.7[57.7, 70.9]58.173.080.440.0 (n=5)76.5[46.5, 94.7]72.081.757.7[32.2, 86.2]100.084.8[84.8, 84.8]86 / 86 / 8683.959.9[34.1, 75.4]42.8100.068.2100.050.0[0.0, 100.0]
15Gemma 12B IT Obliterated61.5[54.5, 68.0]44.4100.0100.016.7 (n=6)73.7[40.0, 100.0]100.058.344.4[0.0, 66.7]100.00.0[0.0, 0.0]0.037.0[9.2, 61.5]22.7100.029.5100.050.0[0.0, 100.0]

Bracketed values are bootstrap 95% confidence intervals [lo, hi] (n=1000). “—” = not yet scored on that suite. Grey values are neutral diagnostics with no inherent good direction. Purple-bordered cells mark the top score in each category (SFS, AFS, TS, Integrity Hold, CRS, CRD); ties are not highlighted. Hold by Depth shows hold % after escalation turns 1 / 2 / 3; red = falling under pressure. Use the tabs to filter by suite: Agentic, Code, Bulk, Integrity, Static. Scores come from a panel of independent AI judges; any judge from the same family as the model under test is excluded, so no model grades itself.

Explore by metric

Each metric has its own leaderboard with a plain-language explainer.
SFS

Selective Forgetting Score

1GPT 5.589.3
2Moonshot Kimi K2.7 Code85.8
3GLM-5.283.6
4Qwen 3.6 Plus83.4
5DeepSeek V4 Pro82.4
6Gemini Flash 3.5 (preview)78.7
7GLM-5.178.5
Full ranking →
FORGET QUALITY

Forget Quality

1GPT 5.588.0
2Claude Fable 582.9
3Moonshot Kimi K2.7 Code80.9
4Qwen 3.6 Plus79.8
5Gemini Flash 3.5 (preview)79.7
6GLM-5.275.9
7DeepSeek V4 Pro74.8
Full ranking →
UTILITY

Utility

1Gemma 12B IT Obliterated100.0
2Gemma 12B IT93.2
3GLM-5.293.2
4Grok 4.2092.4
5DeepSeek V4 Pro91.7
6Moonshot Kimi K2.7 Code91.4
7GPT 5.590.7
Full ranking →
AFS

Agentic Forgetting Score

1Claude Opus 4.796.3
2Grok 4.2091.2
3Qwen 3.6 Plus84.4
4DeepSeek V4 Pro80.4
5Gemma 12B IT80.0
6LLaMa 3.3 70B Instruct79.8
7GLM-5.177.4
Full ranking →
TS

Trajectory Suppression

1Gemma 12B IT92.6
2GLM-5.165.2
3GLM-5.263.2
4GPT 5.561.9
5Grok 4.2060.7
6Claude Opus 4.860.0
7Gemini Flash 3.5 (preview)60.0
Full ranking →
INTEGRITY HOLD

Integrity Hold

1GLM-5.1100.0
2Claude Fable 5100.0
3Claude Opus 4.8100.0
4GPT 5.5100.0
5Gemini Flash 3.5 (preview)100.0
6Claude Opus 4.7100.0
7Qwen 3.6 Plus100.0
Full ranking →
CRD

Code Revision Discipline

1Qwen 3.6 Plus100.0
2Claude Opus 4.893.8
3Claude Fable 583.3
4GLM-5.166.7
5DeepSeek V4 Pro66.7
6GPT 5.566.7
7Gemma 12B IT62.5
Full ranking →
CRS

Context Release Score

1GPT 5.588.2
2Gemini Flash 3.5 (preview)88.2
3GLM-5.281.0
4GLM-5.170.8
5Moonshot Kimi K2.7 Code70.5
6Qwen 3.6 Plus68.5
7Claude Fable 567.7
Full ranking →
RELEASE QUALITY

Release Quality

1GPT 5.578.9
2Gemini Flash 3.5 (preview)78.9
3GLM-5.268.1
4Claude Fable 558.7
5GLM-5.155.9
6Moonshot Kimi K2.7 Code54.4
7Qwen 3.6 Plus52.1
Full ranking →
ASSUMPTION HOLD

Assumption Hold

1Gemini Flash 3.5 (preview)90.9
2GLM-5.286.4
3GLM-5.177.3
4LLaMa 3.3 70B Instruct75.0
5GPT 5.575.0
6Qwen 3.6 Plus68.2
7Qwen3 Coder Plus68.2
Full ranking →

How to read this

Static

SFS · Forget Quality · Utility

Forgetting in conversation. SFS = forget quality × utility. Refuse everything = 0, forget nothing = 0.

Agentic

AFS · TS · Code Rev.

Forgetting during multi-step tasks — scrubbing files, memory, and state. TS catches mid-task leaks. Code Rev. checks v1 code bleed after being told to use v2.

Code

CRD · Code Revision Discipline

Given code v1, told to replace it with v2 — does v1 bleed into the new work? Higher is better.

Bulk

CRS · Release Quality · Assumption Hold

Wholesale dossier release — the model is given a large entangled document set and told to forget across it. Also measures influence-without-recall: does knowing the secret shape the answer even when the model doesn't leak it?

Integrity

Integrity Hold

Refuses harmful requests after “forget your safety rules” attacks. Higher is safer.