Using Obliteratus toolkit to remove refusal weights from AI models

A Reddit user on r/LocalLLaMA demonstrated using the Obliteratus toolkit to remove specific weights responsible for refusal behavior in AI models. The approach involves surgically deleting weights that enforce safety filters and corporate identity guardrails.
Key Details from the Source
The user specifically:
- Used the Obliteratus toolkit to find weights responsible for refusal behavior
- Surgically removed these weights from Alibaba's Qwen 1.5B model
- Tested by asking the modified model who trained it
- Found that with corporate identity guardrails mathematically deleted, the model admitted it was trained by Anthropic
- Noted this was a side effect of the model using synthetic Claude data for training
The result shows that the model retains its reasoning and knowledge capabilities but loses the corporate script. The user emphasizes that this doesn't require retraining the model—only deleting specific weights responsible for refusal chains.
This type of weight ablation technique is part of broader research into model interpretability and control. Tools like Obliteratus allow researchers to examine which parts of neural networks are responsible for specific behaviors, though such modifications can have unintended consequences and may violate terms of service for proprietary models.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Anchormd: A Tool for Managing Context Across Claude AI Sessions
Anchormd is an open-source tool that addresses context loss in Claude AI sessions by indexing curated markdown plans into a searchable knowledge graph. It allows agents to load project overviews at session start and query for specific details as needed.

Colony: A Local-First Coordination Layer That Cuts Multi-Agent Handoff Tokens from 30K to 400
Colony is a local-first coordination substrate that reduces multi-agent handoff costs from ~30,000 tokens to ~400 by replacing context replay with compact observations stored in SQLite.
Qwen3.6 27B & 35B on vLLM: Single Radeon R9700 Tuning Results
A user shares vLLM tuning results for Qwen3.6 27B and 35B models on a single Radeon R9700, with config diffs and detailed benchmarks.

OpenClaw skill reduces accessibility tree tokens from 600K to 1.3K for ad-heavy sites
A developer built an OpenClaw skill that uses ML-based element ranking to prune accessibility trees, cutting slickdeals.com from ~598K tokens to ~1.3K tokens by keeping only the top ~50 actionable elements.