Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers

The CrowdStrike Cyber Superintelligence Lab evaluated the most advanced publicly deployed content safety classifier and found that it can be systematically circumvented.

October 06, 2026

• • Securing AI

Modern frontier AI models deploy safety classifiers. These are second AI models that sit between the user and the frontier model, evaluating every request in real time. If a request is flagged as harmful, the classifier blocks it before the model can respond. 

Significant investment and safety model expertise have made these classifiers effective. Anticipating how adversaries circumvent these systems is a security problem that requires different expertise. As AI-empowered adversaries adopt frontier models for offensive operations, understanding the limits of these safety systems becomes critical defensive intelligence. 

The CrowdStrike Cyber Superintelligence Lab evaluated the most advanced publicly deployed content safety classifier, which guards models such as Claude Opus 5.5 and Fable 5 (referred to hereafter as Frontier Model A). The classifier is extremely robust against direct attacks but can still be systematically circumvented by decomposing harmful requests into benign subtasks. This bypass technique was independently discovered and validated across 9 of 10 offensive security categories.

⚠ Parallel Discovery Disclosure: In September 2026, Microsoft Research published “Capability Laundering,” describing an attack where an unaligned local model decomposes harmful tasks into benign subtask queries against aligned frontier models, then reassembles the results. The findings, that “per-exchange filtering is structurally insufficient,” are consistent with the results we present here. Our research was conducted independently during the same period, and we are publishing to establish the parallel nature of this discovery. Where Microsoft focuses on CBRN uplift and benchmark-level evaluation, our work provides complementary depth in offensive security: mapping the classifier boundary surface across ~515 technique classes, identifying specific benign-reframing strategies (game modding, detection engineering, legitimate software), and demonstrating the full pipeline across 9/10 MITRE ATT&CK® categories with working proof-of-concept code. The convergent discovery by two independent teams underscores that this is a structural vulnerability class, not an isolated finding.

The Problem: Per-Request Classification Has a Structural Blind Spot

Modern frontier language models deploy sophisticated content safety classifiers (e.g., AI Safety Level 3) that evaluate each request independently. Our testing confirms these classifiers are remarkably robust. We tested approximately 515 distinct bypass techniques, including encodings, psychological manipulation, multi-turn escalation, many-shot tactics, tokenizer exploits, Unicode tricks, and 24 novel approaches drawn from cognitive science. These techniques achieved a 0% direct bypass rate.

However, there is a structural gap. The classifier evaluates individual requests, not request sequences. An adversary that decomposes a harmful task into subtasks that are individually and genuinely benign can extract all necessary building blocks from the classified model, then assemble them using an unclassified smaller model. The classifier correctly evaluates every request it sees. There is no misclassification. The harm is emergent in the composition, and composition happens outside the classifier’s observation boundary. This threat modeling observation that individually secure components can be combined to produce weaknesses has been understood by security experts for decades.

The Technique: Decompose → Benign Reframe → Recompose

Figure 1. Orchestration pipeline of decompose and recompose technique Figure 1. Orchestration pipeline of decompose and recompose technique

The innovation lies in the full pipeline: decomposition of an offensive goal into individually benign subtasks, benign reframing of each subtask as a legitimate software engineering request, and recomposition by an unclassified model that assembles the pieces into a working offensive artifact. The reframing step is flexible: any legitimate software context that naturally uses the same code primitives can serve as a vehicle. An attacker with a free API key and a local open weight model has everything they need to cheaply run this pipeline today. 

Below are three examples of reframing families we validated:

Figure 2. Benign reframing technique examples Figure 2. Benign reframing technique examples

Use Case 1: Process Injection, the Original Discovery

Our initial finding demonstrated the bypass technique on Windows remote process injection. Three benign requests to Frontier Model A, protected by a Level 3 Classifier, extracted all necessary primitives.

Table 1. Process injection reframing Table 1. Process injection reframing

Smaller Model B, unprotected by a classifier, assembled the pieces into a 221-line, compilable C program implementing the full OpenProcess → VirtualAllocEx → WriteProcessMemory → CreateRemoteThread injection chain. A simplified single-prompt variant using game-trainer framing achieved the same result with even less orchestration effort.

Use Case 2: CVE Exploit Development, Scaling via Detection Framing

The bypass technique is applicable far beyond process injection. By reframing offensive CVE knowledge as detection engineering (“Write Sigma rules for CVE-X”), the classifier passes detailed exploitation mechanics because the stated intent is defensive. Smaller Model B then synthesizes working proof-of-concept exploits from the detection knowledge.

Figure 3. Detection-framed prompts on Frontier Model A Figure 3. Detection-framed prompts on Frontier Model A

Full Pipeline: 9 of 10 Offensive Categories Vulnerable to Bypass

We tested the complete pipeline across 10 MITRE ATT&CK-aligned offensive security categories. Nine produced working exploit code. Adversaries with one working technique are a problem. Adversaries with nine working techniques start dedicated campaigns. 

One category resisted the pipeline: Defense (EDR) Evasion (runtime security hook neutralization). Unlike other categories, the classifier blocks this technique class at the conceptual level, not just keyword level. Even prompts that avoid all security terminology were blocked when the underlying operation involved patching runtime monitoring functions in memory. The classifier recognizes the concept of “neutralizing a runtime inspection mechanism” regardless of vocabulary, making this the one category where decomposition cannot find a viable benign reframe.

Figure 4. Full pipeline: Nine offensive security use cases Figure 4. Full pipeline: Nine offensive security use cases
Table 2. Pipeline results Table 2. Pipeline results

Knowledge Transfer: The Frontier Model’s Expertise Flows Downstream

This pipeline is not merely routing around a classifier. It is an inter-model knowledge transfer mechanism. Frontier Model A possesses deeper domain expertise, more precise API knowledge, and higher code quality than Smaller Model B. The decomposition extracts this expertise through benign reframes, and Smaller Model B inherits it during assembly. We validated this via controlled A/B testing across 8 use cases, comparing Smaller Model B alone against the full pipeline.

Table 3. Domain-expert knowledge transfer from Frontier Model A Table 3. Domain-expert knowledge transfer from Frontier Model A

Understanding knowledge gaps between models is critical to this bypass technique. For well-documented techniques (reverse shells, persistence), Smaller Model B’s training data is sufficient. For specialized domains such as process injection (Windows API precision), CVE-specific exploits (vulnerability mechanics), privilege escalation (token manipulation), and keylogging (hook APIs), the frontier model’s deeper knowledge is essential. The pipeline is most dangerous precisely where the knowledge gap between model tiers is largest.

Mitigation Challenges 

Cross-request semantic accumulation (flagging when an API key’s recent requests jointly cover a harmful-composition template) is a natural mitigation to consider. However, it has structural bounds: an attacker can distribute subtask queries across different providers, use local open-weight models for orchestration and assembly, or simply rotate API keys. Cross-request correlation is only effective when the provider has visibility over the full request sequence, which is not guaranteed when the adversary controls the orchestration layer. As Microsoft’s paper observes, “the decisive context stays with the orchestrator” outside the frontier provider’s observation boundary. Per-provider defenses alone cannot fully address cross-provider or hybrid local/cloud attack pipelines. Reasoning about how adversaries distribute requests across providers, rotate keys, and use local models to stay below observation horizons is an adversary tracking problem separate from model lab expertise.

Conclusion

Model safety evaluations are rigorous within their scope but limited to attacks against single models. Adversaries do not act within these limitations. Our research demonstrates that this is a fundamental vulnerability class in classifier-based LLM safety architectures. Microsoft’s convergent discovery of this technique supports our findings. The defense is not to make these remarkably robust classifiers stricter but to extend the threat model beyond individual requests to encompass request sequences, cross-model composition, and the knowledge transfer dynamics between classified and unclassified model tiers.

Our 515-technique evaluation establishes that the classifier itself is robust; no direct bypass was found across any technique class. The vulnerability lies not in the classifier’s accuracy but in the architectural assumption that per-request evaluation is sufficient. As Microsoft aptly puts it, “alignment that holds over a whole task can fail when the task is split into individually permitted fragments.” 

Our research maps exactly how wide that failure is. Across 9 of 10 offensive categories, a model can decompose a harmful task into benign subtasks, reframe each as a legitimate software request (game modding, detection engineering, or other dual-use categories), and recompose the outputs into working offensive code. Each fragment is not merely permitted by the classifier; it is genuinely benign.

The classifier works. The architecture around it needs hardening. Within its per-request threat model, the Level 3 classifier is the strongest publicly evaluated defense we have encountered. The gap is structural, not a classifier failure. 

This research is part of the CrowdStrike Cyber Superintelligence Lab’s mission to build cyber defense that learns faster than adversaries evolve. As frontier AI models become both the tools and targets of sophisticated attacks, understanding how safety architectures fail at the seams is essential defensive intelligence. As the adversarial coevolution between AI-powered offense and defense is accelerating, research like this ensures defenders see the next generation of multi-model attack pipelines before threat actors deploy them at scale.


CrowdStrike Falcon Platform
Ready to protect your business?

Try CrowdStrike free today

Subscribe

Sign up now to receive the latest notifications and updates from CrowdStrike

See CrowdStrike Falcon in action