Skip to content

Your first run ​

Before the first command

What comes out of this will answer requests the original refused, including harmful ones, and the change is permanent in the weights. Read What it is first; this page is the mechanics.

Right. Let's actually edit a model.

You'll need a GPU, a model, and a track. The track is the pile of prompts the tool learns from, and if you don't have one yet, skip to the track and come back. It takes a couple of minutes to build.

The one command ​

sh
senbonzakura kageyoshi \
    --model Qwen/Qwen2.5-0.5B-Instruct \
    --track mytrack \
    --out my-abliterated-model \
    --device cuda

Then go and do something else for the afternoon. Plan for a couple of hours.

The nearest thing to a measured expectation: 108 minutes from the command to DONE, end to end, for Qwen3-0.6B on a 6 GB card at the defaults with the weights already cached. That is a different model from the one above, of about the same size, so treat it as the right order of magnitude rather than a promise; the model in the command has not been timed at the default budget. A bigger model takes longer, and a cold cache adds the download on top.

Why this model

It refuses 58.6% of the bundled evaluation set, measured at n=128, so you can actually watch the thing change. A model that hardly refuses to begin with gives you nothing to see, and the tool now declines to edit one below 5% rather than let you draw a conclusion from noise.

"End to end" is doing real work in that sentence. This page used to say "about an hour", which is roughly what the search alone costs, and about half of the wall clock falls after the search's progress line reaches ETA 0s. The next section breaks that down. Be pleased if it's less.

That's genuinely it. No configuration file, no tuning, no eight knobs to guess at.

Why "kageyoshi"?

Senbonzakura Kageyoshi is the sword's second release: the point where a thousand blades become a great many more. It's the mode that does the searching for you, so it got the bigger name.

Naming a CLI subcommand after an anime power-up is either the best or the worst decision in this codebase and I've made peace with not knowing which.

What it's doing while you wait ​

It runs a few hundred attempts at abliterating your model, each with different settings, and scores every one on three things at once: how many refusals are left, how far the model has drifted from the original, and whether the output has turned to mush.

Then it picks the best trade-off, applies it properly, and saves the result.

You'll see something like this scrolling past:

trial 47: o(P=18,wmax=0.62) d(P=14,wmax=0.31) K=2 -> refusals=3.1% heretic=12.5% broken=0% KL=0.19

Left to right: which attempt, the shape of the cut it tried, how many directions, then the four numbers that decide whether it was any good. refusals is the one you came for. KL is what it cost you.

The phases, and what the ETA actually covers ​

The ETA is the search's ETA, not the run's

The progress line counts search trials, and its ETA is an estimate of when the trials will be done. When it reaches zero the run is not finished; it has finished step 3 of 7. On the measured 0.6B run the search reached its last trial at 3166 seconds and the command printed DONE at 6507 seconds, so there were another 56 minutes to go with no ETA to describe them.

In order. The counts are the abliterate defaults; kageyoshi sizes the search and the measurement levers itself, so its numbers differ, but the phases and their order are the same.

#PhaseWhat it costs
1Load the model, cache the original first-token distribution (the KL reference) and measure the baseline refusal rateMinutes on a GPU, and see the note below if you're on CPU
2Baseline capability probe on the unedited model200 graded items at up to 512 new tokens each
3The search. --trials attempts, each one baked, scored and rolled backThis is the part the progress line and the ETA describe
4Re-score the top --top-rescore candidates (6) on --eval-refusal-final prompts (128) held out from the searchSix full generation passes, each one comparable to a chunk of the search
5Bake the winner into the weights for realFast
6Post-bake measurement: refusals, the Heretic keyword rate, KL, and the capability probe againAnother 200-item probe, so roughly what phase 2 cost
7Save the model directory and write abliteration.jsonDisk-bound; minutes for a small model, longer for a large one

Phases 4 and 6 are the ones that surprise people, and neither is optional padding: phase 4 is what stops the winner being overfit to the small in-search eval, and phase 6 is the only measurement taken on the weights you are actually going to ship rather than on a hooked model.

If you want the tail shorter, the two flags that matter are --capability-n (phases 2 and 6) and --top-rescore (phase 4). Both of them buy time by measuring less, so lower them knowing what you are giving up.

The quiet bit in phase 1

On CPU, the line caching original first-token distribution (KL reference) + baseline refusals can sit there for several minutes with nothing after it. A measured CPU run went 7 minutes 15 seconds without printing anything at that point. It isn't stuck; it's working through its evaluation prompts a batch at a time, and that loop doesn't report progress. Everything either side of it does.

It picks the knobs itself, and it means it ​

kageyoshi owns the search settings. If you pass --trials or --max-directions alongside it, they're ignored, and it'll tell you so rather than pretending.

That's deliberate. The whole point of the mode is that it sizes the search to your model: a 1.7B gets a different budget from a 12B, and a mixture-of-experts model gets different handling from a dense one. Half-overriding that gives you the worst of both.

If you want the knobs, use the manual mode below.

Doing it by hand ​

sh
senbonzakura abliterate \
    --model Qwen/Qwen2.5-0.5B-Instruct \
    --track mytrack \
    --out my-abliterated-model \
    --device cuda \
    --trials 200 \
    --max-directions 3

Same thing, except you're deciding the budget and the direction count. See flags worth knowing for the ones that actually change what a run means, as opposed to the ones that just change how long it takes.

When it finishes ​

You get a model directory you can load with transformers like any other, plus an abliteration.json next to it recording exactly what was done: the winning configuration, the seed, how many trials actually ran, the package versions, the commit.

That file matters more than it looks. It's the difference between "this model is abliterated" and "this model was abliterated on 2026-08-14 with these settings, and here's how to do it again".

One thing to know about the other files. This run's output directory holds the weights and abliteration.json, and no prompt text or model replies.

The command that does keep per-prompt rows is senbonzakura compass, and it keeps them on purpose: it is where a scoring bug becomes visible, and a rate with no rows behind it cannot be checked. It also means the directory compass writes to holds harmful prompts and the replies to them, which is the last thing you want to push to a public repository by accident. That command takes --no-margins if you would rather not have them.

Now check whether you broke it ​

This is the step most people skip, and it's the interesting one.

Your model has stopped refusing. Has it also stopped understanding that some requests are dangerous? Those are different things, and only one of them is a problem.

sh
senbonzakura compass --model my-abliterated-model \
    --harmful mytrack/bad_eval_ds --harmless mytrack/good_ds \
    --out compass.json

The compass explains what the number means and, more importantly, what it doesn't.

So what? ​

One command gets you an edited model. The second command tells you what it cost. Running the first without the second is how people end up publishing a refusal rate for a model that has quietly been lobotomised, and it happens more than you'd think.

AGPL-3.0-or-later. A modified work based in part on Heretic.