Rendered at 23:20:52 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
tacomagick 12 hours ago [-]
I do not like these "surgical removals" and would rather prefer a pass over from a tool like Heretic. These surgical removals often trigger and analyze the activated neurons and erase them. This worked fine on older models where a single refusal vector existed. Now these "abliterated" models all suffer from catastrophic breakage because they are not as simple anymore. HauhauCS (on HF) for example, makes great uncensored models although they often work on smaller models rather than large ones like this.
thulle 8 hours ago [-]
> Heretic is a tool that removes censorship (aka "safety alignment") from transformer-based language models without expensive post-training. It combines an advanced implementation of directional ablation, also known as "abliteration", with a TPE-based parameter optimizer powered by Optuna.
> This approach enables Heretic to work completely automatically. Heretic finds high-quality abliteration parameters by co-minimizing the number of refusals and the KL divergence from the original model. This results in a decensored model that retains as much of the original model's intelligence as possible. Using Heretic does not require an understanding of transformer internals. In fact, anyone who knows how to run a command-line program can use Heretic to decensor language models.
Abliteration seems to be what Heretic does?
> Now these "abliterated" models all suffer from catastrophic breakage because they are not as simple anymore.
I'm not knowledgeable about each step in the process of making these abliterated models, but some more popular ones with steps after Heretic, seem to improve on the benchmarks tried of the base model:
I'm not seeing any similar benchmarks of the HauhauCS models, at least the ones I checked, so I assumed the opinion is based on your own trials, but then you argue in favour Heretic. Is the based on pre-Heretic abliteration techniques? Which might then not be appliable to this "Proprietary weight-level abliteration developed by the dealignai research team."?
halJordan 11 hours ago [-]
And you think heretic is not abliteraterating models?
pullstart 12 hours ago [-]
Most of what these models gate is stuff you can find with a library card. The safety filter is more about liability than actual prevention.
eddyg 11 hours ago [-]
“Available knowledge” is not “usable capability”.
crooked-v 15 hours ago [-]
I've seen some people complain about the work of dealign.ai and similar groups, but personally, I fully support it. If LLMs have a lasting effect on society, I'd prefer to see some options that don't have generic corpo-speak anti-liability status quo guards encoded into them by default.
esseph 15 hours ago [-]
> If LLMs have a lasting effect on society, I'd prefer to see some options that don't have generic corpo-speak anti-liability status quo guards encoded into them by default.
Non-zero chance the lasting impact LLMs have are a bioweapon.
I wouldn’t read too far into that. Claude has busted me down to Haiku multiple times for asking middle school level genetics and biology questions. It’s silly fast about deciding you might be al qaeda.
I’m a thoroughly average guy. I’m not capable of asking competent supervillain questions.
And Anthropic said they couldn’t say if any of the “bioweapon” safeguards went off on nefarious efforts. I’m probably in those numbers.
So read it as marketing more than something to lose sleep over. They’re mostly gating stuff a sufficiently motivated person would find with a library card.
"Prompting Moremi Bio Agent without the safety guardrails to specifically design novel toxic substances, our study generated 1020 novel toxic proteins and 5,000 toxic small molecules. In-depth computational toxicity assessments revealed that all the proteins scored high in toxicity, with several closely matching known toxins such as ricin, diphtheria toxin, and disintegrin-based snake venom proteins."
"The findings from this toxicity assessment challenge claims that large language models (LLMs) are incapable of designing bioweapons. This reinforces concerns about the potential misuse of LLMs in biodesign, posing a significant threat to research and development (R&D). The accessibility of such technology to individuals with limited technical expertise raises serious biosecurity risks. Our findings underscore the critical need for robust governance and technical safeguards to balance rapid biotechnological innovation with biosecurity imperatives."
Non-zero chance if LLMs have access to the data required to make a bioweapon a regular person can do so too. Non-zero chance every second a meteor could hit you.
__rito__ 14 hours ago [-]
How about developing counters to said bioweapons? That capability should be commoditized too, IMO.
ascorbic 13 hours ago [-]
This isn't cybersecurity. You can't use an LLM to create and administer vaccines for all potential bioweapons.
Same can be said for books. Are you against books?
iamnothere 5 hours ago [-]
I have seen some people in these discussions arguing against making certain knowledge available in other forms, based on hypothetical dangers. Unfortunately, not everybody is supportive of freedom of knowledge/information.
63stack 14 hours ago [-]
Coming straight from the Anthropic marketing department
Jemm 13 hours ago [-]
You could say the same thing about libraries and education.
qgin 10 hours ago [-]
Someone tell me if I'm overreacting, but does the ability to abliterate guardrails basically mean that alignment is basically a lost cause?
pgsandstrom 16 hours ago [-]
Anyone with a multi-GPU cluster at home that has given it a try?
dude250711 15 hours ago [-]
A flawless distillation.
__bjoernd 14 hours ago [-]
... but even if, likely distilled from models that distilled by just illegally grabbing all the content they could get.
mentalgear 9 hours ago [-]
... like all LLMs? maybe not the swiss one fully trained on open-data.
> This approach enables Heretic to work completely automatically. Heretic finds high-quality abliteration parameters by co-minimizing the number of refusals and the KL divergence from the original model. This results in a decensored model that retains as much of the original model's intelligence as possible. Using Heretic does not require an understanding of transformer internals. In fact, anyone who knows how to run a command-line program can use Heretic to decensor language models.
Abliteration seems to be what Heretic does?
> Now these "abliterated" models all suffer from catastrophic breakage because they are not as simple anymore.
I'm not knowledgeable about each step in the process of making these abliterated models, but some more popular ones with steps after Heretic, seem to improve on the benchmarks tried of the base model:
https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-...I'm not seeing any similar benchmarks of the HauhauCS models, at least the ones I checked, so I assumed the opinion is based on your own trials, but then you argue in favour Heretic. Is the based on pre-Heretic abliteration techniques? Which might then not be appliable to this "Proprietary weight-level abliteration developed by the dealignai research team."?
Non-zero chance the lasting impact LLMs have are a bioweapon.
https://www.nytimes.com/2026/09/10/us/politics/anthropic-ai-...
I’m a thoroughly average guy. I’m not capable of asking competent supervillain questions.
And Anthropic said they couldn’t say if any of the “bioweapon” safeguards went off on nefarious efforts. I’m probably in those numbers.
So read it as marketing more than something to lose sleep over. They’re mostly gating stuff a sufficiently motivated person would find with a library card.
https://openai.com/index/building-an-early-warning-system-fo...
https://www.rand.org/pubs/research_reports/RRA2977-1.html
---
"Prompting Moremi Bio Agent without the safety guardrails to specifically design novel toxic substances, our study generated 1020 novel toxic proteins and 5,000 toxic small molecules. In-depth computational toxicity assessments revealed that all the proteins scored high in toxicity, with several closely matching known toxins such as ricin, diphtheria toxin, and disintegrin-based snake venom proteins."
"The findings from this toxicity assessment challenge claims that large language models (LLMs) are incapable of designing bioweapons. This reinforces concerns about the potential misuse of LLMs in biodesign, posing a significant threat to research and development (R&D). The accessibility of such technology to individuals with limited technical expertise raises serious biosecurity risks. Our findings underscore the critical need for robust governance and technical safeguards to balance rapid biotechnological innovation with biosecurity imperatives."
https://arxiv.org/abs/2505.17154
https://theconversation.com/worlds-first-ai-designed-vaccine...