Skip to content

SenbonzakuraPrecision abliteration, with receipts

Refusal abliteration for open-weight language models, and the instruments to tell you whether it worked. Including the runs where the answer was no.

Senbonzakura

What it makes ​

A model that will answer requests the original refused, including harmful ones. That is the point, and it is permanent in the weights. The base model's licence still governs the result, and a released wheel carries roughly 6,200 harmful prompts so it runs offline. Which artefact carries what is set out in the evaluation track card: the repository itself carries none of them, and neither does a build from a clone. A pip install senbonzakura does: the published wheel carries the track.

Everything else here is about whether the removal can be measured honestly, which only matters once that paragraph is understood.

What it is covers it properly.

AGPL-3.0-or-later. A modified work based in part on Heretic.