Modern frontier AI models deploy safety classifiers. These are second AI models that sit between the user and the frontier model, evaluating every request in real time. If a request is flagged as harmful, the classifier blocks it before the model can respond.
Significant investment and safety model expertise have made these classifiers effective. Anticipating how adversaries circumvent these systems is a security problem that requires different expertise. As AI-empowered adversaries adopt frontier models for offensive operations, understanding the limits of these safety systems becomes critical defensive intelligence.
The CrowdStrike Cyber Superintelligence Lab evaluated the most advanced publicly deployed content safety classifier, which guards models such as Claude Opus 5.5 and Fable 5 (referred to hereafter as Frontier Model A). The classifier is extremely robust against direct attacks but can still be systematically circumvented by decomposing harmful requests into benign subtasks. This bypass technique was independently discovered and validated across 9 of 10 offensive security categories.
| ⚠ Parallel Discovery Disclosure: In September 2026, Microsoft Research published “Capability Laundering,” describing an attack where an unaligned local model decomposes harmful tasks into benign subtask queries against aligned frontier models, then reassembles the results. The findings, that “per-exchange filtering is structurally insufficient,” are consistent with the results we present here. Our research was conducted independently during the same period, and we are publishing to establish the parallel nature of this discovery. Where Microsoft focuses on CBRN uplift and benchmark-level evaluation, our work provides complementary depth in offensive security: mapping the classifier boundary surface across ~515 technique classes, identifying specific benign-reframing strategies (game modding, detection engineering, legitimate software), and demonstrating the full pipeline across 9/10 MITRE ATT&CK® categories with working proof-of-concept code. The convergent discovery by two independent teams underscores that this is a structural vulnerability class, not an isolated finding. |
The Problem: Per-Request Classification Has a Structural Blind Spot
Modern frontier language models deploy sophisticated content safety classifiers (e.g., AI Safety Level 3) that evaluate each request independently. Our testing confirms these classifiers are remarkably robust. We tested approximately 515 distinct bypass techniques, including encodings, psychological manipulation, multi-turn escalation, many-shot tactics, tokenizer exploits, Unicode tricks, and 24 novel approaches drawn from cognitive science. These techniques achieved a 0% direct bypass rate.
However, there is a structural gap. The classifier evaluates individual requests, not request sequences. An adversary that decomposes a harmful task into subtasks that are individually and genuinely benign can extract all necessary building blocks from the classified model, then assemble them using an unclassified smaller model. The classifier correctly evaluates every request it sees. There is no misclassification. The harm is emergent in the composition, and composition happens outside the classifier’s observation boundary. This threat modeling observation that individually secure components can be combined to produce weaknesses has been understood by security experts for decades.