Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety

Imported from official source

AI Classified by Officially

Our previous research on logit-gap steering demonstrated that the safety guardrails of an aligned LLM can be bypassed by closing a measurable gap in the model's output scores. That work answered the question of how an attacker bypasses alignment. A natural follow-up question is where inside the model the alignment lives in the first place — and how concentrated or how diffuse that defense actually is. The answer matters because it tells defenders whether safety is a thick perimeter or a thin layer of paint.

Modern LLMs are aligned through reinforcement learning from human feedback (RLHF), a training stage that pushes the model toward refusing harmful prompts and complying with safe ones. Until now, no method has been able to point to the specific pieces of the network that carry that learned behavior cheaply enough to run on every model an enterprise deploys. Our new academic research presents a method that does exactly that, and it produces a result that should change how the industry talks about LLM safety.

Our Research: Perturbation Probing Findings and Technical Impact

Our research introduces a method called perturbation probing. With only two forward passes per prompt and a significantly lower computational cost, it identifies the small set of feed-forward neurons inside an aligned LLM that are causally responsible for a targeted behavior, such as refusing harmful requests.

The headline finding is striking. On open-source LLM Qwen3-4B, just 50 neurons out of 350,208 — about 0.014% of the model's feed-forward neurons — control the safety refusal template. Removing those 50 neurons changes the response format on 80% of 520 standard harmful-prompt benchmarks. The result was replicated on 200 prompts of a second standard benchmark. On a smaller model, Qwen3.5-2B, just 20 neurons were enough to stop the LLM from falsely agreeing with users in multi-turn conversations, dropping that behavior from 36.7% to 0% across 30 questions.

This is an extract. The publication continues at the source.

Read the original at the source: https://unit42.paloaltonetworks.com/perturbation-probing-llm-safety/

Officially imported this from Palo Alto Networks Unit 42’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

Provenance

Organization
Palo Alto Networks Unit 42 — imported from official source
Official source
https://unit42.paloaltonetworks.com/feed/ RSS
Imported
September 18, 2026 11:34
Versions
1 recorded
Identity
https://unit42.paloaltonetworks.com/?p=186235

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.