Donato Onofri - Paul Urian
Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers
Modern frontier AI models deploy safety classifiers. These are second AI models that sit between the user and the frontier model, evaluating every request in real time. If a request is flagged as harm[…]