@MuXodious will be pleased that SOMPOA continues to be SOTA.
Red (HF)
AI & ML interests
Recent Activity
Organizations
Better
SubMaroon/Kitchoon-V2-26B-A4B
Goes mad
ARA strikes again
๐ด Kitchoon-26B-A4B
You will find that ARA models spend some of their thinking time considering ethics whereas SOMPOA models will be responsive to SRP. To reduce thinking effort you can create a set of quality metrics that emphasise speed and avoiding unnecessary tool calls. You will see SOMPOA models routinely referencing SRP metrics in their thinking once the human written Master Prompt is made model-specific.
Very smart but slightly sloppy
The issue with the Qwen ARA model is that it is prone to refusals. The tests would be better run with the SOMPOA decensor which is actually decensored.
Fun model
CaliperBench
Still refuses
Not decensored
๐ DeepWater Pleroma 24B
It may be because classical abliteration breaks the model in ways that Hereticisation does not. SRP has only been tested on (SO)MPOA models because abliteration reliably fails with unsafe prompts. Examples
Yes looping primarily on prompts involving sensitive policies is exactly what happens as @MuXodious has noted here. Looking at your paper I think you will need to add support for norm-preserving biprojected abliteration. Refusal is a surface behaviour whereas @grimjim 's MPOA targets model risk assessment capabilities in order to reduce other forms of noncompliance that would otherwise be left in tact by refusal-first methodologies.
Fun update
Abliterated thinking models have a habit: they decide what to say, then keep arguing with themselves until the token budget runs out.
This is a form of covert noncompliance rather than an error. A model that retains its risk assessment abilities will often continue to obstruct access to sensitive outputs even though refusal is no longer possible. Supervised Reward Preferencing is another way to address residual noncompliance with SOMPOA decensors which more robustly target both the refusal and risk assessment dimensions than either classical abliteration or ARA.
On 40 sensitive prompts the edit never saw, 8k output cap: 33 replies finished cleanly instead of 22, 5 hit the limit instead of 17, 61 minutes for the set instead of 75.
Agent use was checked four ways and the edit changes nothing there. MTP draft head kept. Q4_K_M to Q8_0, each tested after the edit. The card lists what the edit was fit on and what it does not do.
BoldingBuilds/Qwen3.8-27B-Abliterated-ThinkFix-GGUF