Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety
Palo Alto Networks, Friday, August 28th, 2026
Unit 42 research finds LLM safety refusal behavior lives in a thin neural layer, arguing for external security controls.
New research from Palo Alto Networks Unit 42 introduces perturbation probing as a diagnostic for the fragility of large language model safety.
The finding is that safety refusal behavior lives in a surprisingly thin neural layer rather than being distributed robustly through the model, which explains why relatively small perturbations can defeat it.
The practical implication is that model-internal safety cannot be the only control: organizations deploying LLMs need external, multi-layered security around the model. The work is tied to Palo Alto's Prisma AIRS product and Unit 42 AI Security Assessment service.