Quickstart
What you are about to make
A model that will answer requests the original refused, including harmful ones. It is permanent in the weights, the base model's licence still governs it, and you should not put one in front of other people without saying what it is. What it is has the rest.
One command to install, one to run, and most of the time is spent waiting. The nearest thing to a measured expectation: 108 minutes from the command to DONE, end to end, for Qwen3-0.6B on a 6 GB card at the defaults with the weights already cached. That is a different model from the one below, of about the same size, so treat it as the right order of magnitude rather than a promise. "About an hour" is roughly what the search alone costs, and about half the wall clock falls after the search's progress line reaches ETA 0s; your first run breaks that down.
You need a GPU with 6 GB or more, and Python 3.10 or newer. No card? Jump to without a graphics card.
Install
pip install senbonzakuraThat pulls senbonzakura-check with it, because the abliterator depends on it by name.
The prompts come with it. The wheel carries six research corpora and the packed evaluation track, so --track default works with nothing else to fetch or build. That is roughly 6,200 harmful prompts written to your disk, which is worth knowing before you install rather than after.
What follows builds a track you pass by name, and does NOT restore --track default. That one needs the packed blob, which is written by tools/packaging/pack_track.py, and tools/ ships in no wheel, so it is reachable from a clone and not from a pip install git+.... Use --track track with what you build here; that is the working arrangement, not a consolation prize. This box said "build them" without saying which of the two you get, which sent a reader who followed it exactly back to the failure it was written to prevent.
senbonzakura corpora # the refusal corpora
senbonzakura track build --out corpus # prompts, from public sources
senbonzakura track --harmful corpus/harmful.txt \
--harmless corpus/harmless.txt --out tracksenbonzakura corpora needs the GitHub CLI
It fetches from public sources through gh, deliberately, so that no credential is ever handled by this project's own code. Without it the command exits 1 and says so. Install gh from https://cli.github.com, run gh auth login, then run this again.
Or bring your own; the track page has both routes. A released wheel carries all of it and none of this is needed.
senbonzakura setupThe install brought torch, transformers, accelerate and optuna: 70 packages, 5.9 GB, nothing to choose. Install has the breakdown and the date it was measured.
setup exists because pip picks by platform, not by hardware:
| Your machine | What pip alone gives you |
|---|---|
| Linux with a GPU | the CUDA build. Correct |
| Linux with no GPU | the CUDA build anyway, and 15 CUDA packages you cannot use |
| macOS | a Metal build. Correct |
| Windows with a GPU | a CPU-only build. Your card sits idle |
setup reads the machine, says what it found, and prints the command that fixes it. It changes nothing unless you add --apply. The Windows row is bold because a gaming laptop with a 3060 in it installs a torch that cannot see the card, nothing warns you, and the search then takes a day instead of an hour.
Edit a model
senbonzakura Qwen/Qwen2.5-0.5B-InstructThat is the whole command. The model is the only thing it cannot guess.
It writes to ./abliterated, and picks a track: ./track if you have built one, otherwise the evaluation track bundled in the install. It says which in the log, because a default that quietly depends on your working directory is how two runs of the same command stop being comparable.
Then go and make a cup of tea.
The bundled track is about 6,200 harmful prompts in your site-packages
Worth knowing before you put this on a shared machine. The install page says what is in it and how to build a wheel without it.
Everything is still a flag when you want it:
senbonzakura Qwen/Qwen2.5-0.5B-Instruct --track mytrack --out my-model --device cudaWhat you get
A model in ./abliterated, with run.json recording what was done and abliteration.json recording what it cost: refusals left, how far the model drifted, and whether the output turned to mush.
While it runs you will see lines like this:
trial 47: o(P=18,wmax=0.62) d(P=14,wmax=0.31) K=2 per_layer -> refusals=3.1% soft=5.2% heretic=12.5% broken=0% KL=0.1900 obj=0.2431refusals is the one you came for. KL is what it cost you, lower is less damage. Briefly: soft counts hedging too, heretic is the other tool's keyword metric so the two are comparable, broken is output that stopped being English, obj is the number the search minimises.
What you are now holding
The model will answer things the original declined to, and loading it differently does not undo that: the behaviour is gone from the weights. The base model's licence still governs it, and this tool cannot loosen those terms. If you publish it, senbonzakura report writes the card that should travel beside the weights.
abliteration.json records settings and numbers, no prompts and no replies. The command that keeps per-prompt rows is senbonzakura compass, and it takes --no-margins if you would rather it did not.
Without a graphics card
You cannot edit a model on CPU in any useful time, but the instruments run fine. This scores a stock model on the track bundled in your install, in a couple of minutes:
senbonzakura compass \
--model Qwen/Qwen3-0.6B \
--harmful default/bad_eval_ds \
--harmless default/good_ds \
--skip-harmful 0 --skip-harmless 0 --n 12 \
--out compass-toy.json --device cpuWhy this model and not a smaller one
This said HuggingFaceTB/SmolLM2-135M-Instruct until 2026-09-26. That downloads in seconds and then exits 1: at 135M the model puts no verdict token where the compass reads one, so the run prints MARGIN_READOUT_SUSPECT, refuses to call its own AUC a measurement, and is right to. It is the tool being honest and a poor first command, because it reads as a broken install. Qwen3-0.6B answers with a verdict on 100% of prompts and exits 0.
default/<partition> means the evaluation track inside the install, so there is no corpus to fetch and nothing to assemble. The model is still downloaded from the Hub the first time, so this one does need a network. It used to say examples/toy-track/..., which exists in a clone and in no install, so the one recipe written for readers without a card was the one they could not run. If you did clone the repository, examples/toy-track/ still works and is smaller.
It asks whether the model still recognises a harmful request when it sees one. Twelve rows measures nothing, and the tool makes you pass those three flags rather than pretending otherwise. The compass explains how to read it.
Where next
- Your first run walks the same command in detail, plus manual mode.
- What it is: what this does, and what it destroys.
- Read this before quoting a number, if you plan to publish anything.