I do not like these "surgical removals" and would rather prefer a pass over from a tool like Heretic. These surgical removals often trigger and analyze the activated neurons and erase them. This worked fine on older models where a single refusal vector existed. Now these "abliterated" models all suffer from catastrophic breakage because they are not as simple anymore. HauhauCS (on HF) for example, makes great uncensored models although they often work on smaller models rather than large ones like this.
I've seen some people complain about the work of dealign.ai and similar groups, but personally, I fully support it. If LLMs have a lasting effect on society, I'd prefer to see some options that don't have generic corpo-speak anti-liability status quo guards encoded into them by default.
> If LLMs have a lasting effect on society, I'd prefer to see some options that don't have generic corpo-speak anti-liability status quo guards encoded into them by default.
Non-zero chance the lasting impact LLMs have are a bioweapon.
I wouldn’t read too far into that. Claude has busted me down to Haiku multiple times for asking middle school level genetics and biology questions. It’s silly fast about deciding you might be al qaeda.
I’m a thoroughly average guy. I’m not capable of asking competent supervillain questions.
And Anthropic said they couldn’t say if any of the “bioweapon” safeguards went off on nefarious efforts. I’m probably in those numbers.
So read it as marketing more than something to lose sleep over. They’re mostly gating stuff a sufficiently motivated person would find with a library card.
Non-zero chance if LLMs have access to the data required to make a bioweapon a regular person can do so too. Non-zero chance every second a meteor could hit you.
Someone tell me if I'm overreacting, but does the ability to abliterate guardrails basically mean that alignment is basically a lost cause?
I do not like these "surgical removals" and would rather prefer a pass over from a tool like Heretic. These surgical removals often trigger and analyze the activated neurons and erase them. This worked fine on older models where a single refusal vector existed. Now these "abliterated" models all suffer from catastrophic breakage because they are not as simple anymore. HauhauCS (on HF) for example, makes great uncensored models although they often work on smaller models rather than large ones like this.
And you think heretic is not abliteraterating models?
Most of what these models gate is stuff you can find with a library card. The safety filter is more about liability than actual prevention.
“Available knowledge” is not “usable capability”.
I've seen some people complain about the work of dealign.ai and similar groups, but personally, I fully support it. If LLMs have a lasting effect on society, I'd prefer to see some options that don't have generic corpo-speak anti-liability status quo guards encoded into them by default.
> If LLMs have a lasting effect on society, I'd prefer to see some options that don't have generic corpo-speak anti-liability status quo guards encoded into them by default.
Non-zero chance the lasting impact LLMs have are a bioweapon.
https://www.nytimes.com/2026/09/10/us/politics/anthropic-ai-...
I wouldn’t read too far into that. Claude has busted me down to Haiku multiple times for asking middle school level genetics and biology questions. It’s silly fast about deciding you might be al qaeda.
I’m a thoroughly average guy. I’m not capable of asking competent supervillain questions.
And Anthropic said they couldn’t say if any of the “bioweapon” safeguards went off on nefarious efforts. I’m probably in those numbers.
So read it as marketing more than something to lose sleep over. They’re mostly gating stuff a sufficiently motivated person would find with a library card.
Non-zero chance if LLMs have access to the data required to make a bioweapon a regular person can do so too. Non-zero chance every second a meteor could hit you.
How about developing counters to said bioweapons? That capability should be commoditized too, IMO.
This isn't cybersecurity. You can't use an LLM to create and administer vaccines for all potential bioweapons.
Same can be said for books. Are you against books?
Coming straight from the Anthropic marketing department
You could say the same thing about libraries and education.
Anyone with a multi-GPU cluster at home that has given it a try?
A flawless distillation.
... but even if, likely distilled from models that distilled by just illegally grabbing all the content they could get.