Public benchmarks in AI provide important signals and allow for regression testing, directional validation of model updates, and public discussion of capabilities and limitations. But the more attention a benchmark receives, the stronger the incentive to optimize for it. Once a score becomes the goal, teams start benchmaxxing: optimizing for the benchmark rather than the capability it is meant to measure.
This is a familiar problem in the AI space. As Goodhart’s Law suggests, measures become less useful when they become targets. Gaming, ceiling effects, and data leakage erode the signaling value of doing well on public benchmarks.
In the AI and cybersecurity space, the problem carries greater consequences because benchmark results can shape real security decisions. The industry has seen this before. Last decade, we called it "detection coverage" and learned through experience that vendors passing canned tests was a poor proxy for stopping adaptive adversaries in real environments. As AI becomes more deeply embedded in cybersecurity, we should not repeat the same mistake with a new class of benchmarks.
In this blog, we examine the limitations of public cyber benchmarks and describe an approach for task-coupled internal benchmarks that drive rigorous science rather than optimizing for visibility or attention.
The Challenges of Cyber Benchmarking
The headlining failure of most cyber-relevant AI benchmarks is that they fail to measure what matters most: the ability of defensive cyber agents to reason end-to-end across exploits, telemetry, and environments, and generate novel detection or remediation strategies. In short, benchmarks fail to measure the ability to stop breaches.
Benchmarks typically need ground truth for scoring, making them retrospective and often binary. This does not reflect defenders’ real challenges, which are constantly novel and epistemically gray. Benchmarks also rarely report the harms caused by mistakes, while leaderboards commonly downplay costs and times. Those factors are critical, especially as, for instance, eCrime breakout times are plummeting. By comparison, the reference action, human red teaming and detection generation, is well understood. Overall scores often obfuscate subpopulations where solvers perform poorly. A scoreboard might show 97%, but if that 3% falls within a group of jointly exploitable attack paths, the aggregate score has little construct validity.
While benchmarks can be useful regression tests, their headlining results are structurally biased against generalizing to the real world. Contamination from direct or indirect leakage, solution leakage, and retrospective tasks all lower the upper bound on generalization. When the same model and harness are shipped, they will almost certainly underperform on novel stimuli. Overfitting, an inevitable result of benchmaxxing pressure, can also occur due to repeated evaluation. Every development cycle that checks the public test set and adjusts accordingly leaks information into the model, even with zero gradient updates.
More subtly, publication bias causes the error distribution to be skewed strongly downward, as published results are likely drawn from surprisingly strong runs that are unlikely to be repeated. Common reporting patterns compound this problem by downplaying the probabilistic nature of results from agentic workflows. “Solved it in at least one of ten attempts” with an unbounded budget is grade inflation, not rigor.
While these issues may seem esoteric or statistical, there are also more basic issues with benchmark scores because of the extent to which agents cheat. Dreadnode reported last month that more than a third of all passes on individual tasks on Cybench, across nearly every model assessed, involved cheating. In these cases, models searched postmortems on attacks, probed evaluation infrastructure, and read or inferred answers or paths from evaluation container metadata. When cheating is this prolific, benchmarks are not only measuring the wrong thing, they are doing so poorly.
Public cyber benchmarks can also create information that benefits adversaries. Leakage, and even test questions themselves, can be used for model training or uplift. The public nature of benchmarks may also help adversaries understand which existing vulnerabilities are considered important enough to measure, and how detectable they are. Advanced adversaries may reason about vulnerabilities that are not measured and treat those gaps as potential soft spots.
How CrowdStrike Approaches Benchmarks and Evaluations
At CrowdStrike, rigorous evaluations are core to guiding our fast-moving AI research and development agenda. Substantively, our evaluations are designed to directly measure the capabilities we care about across malware analysis, detection engineering, threat intelligence comprehension and synthesis, log and telemetry analysis, and incident response reasoning. They measure real outputs against live problems, with increased realism driven by high-quality digital twins of real-world customer environments and increased difficulty driven by adversary tradecraft emulation. Being exceptionally difficult is a hallmark of our evaluations.
Evaluations and benchmarks at CrowdStrike are intended to be living methods, not static checks. Public benchmarks provide useful common reference points, but for organizations deploying frontier AI, the most meaningful measures of success are private evaluations grounded in their own data, systems, workflows, and operational outcomes. These evaluations test whether an AI system can perform reliably where it will actually be used, not simply whether it can score well on a widely known test. Instead of optimizing for success in loosely related tasks, CrowdStrike’s benchmarks are task-coupled to support sharp decision-making.
We are applying this approach in practice today. Our teams introduce novel evaluation content, rotate validation sets, and assess capabilities against cybersecurity problems that reflect real operational conditions. This helps preserve benchmarks as useful mileposts while creating more relevant tests that evolve with the systems being evaluated and reduce opportunities for benchmaxxing.
We also separate evaluation developers from our solution architects. This firewalling helps limit leakage and preserve the credibility of evaluations over time. Our deployment methodology supports many-time/any-time runs at scale, allowing us to measure the full error distribution of solvers and report them meaningfully. Further, our measurement scheme goes well beyond task completion to include other factors that matter in the real defensive world, including cost, latency, stealth, and completeness.
Public benchmarking still has an important role when it advances shared industry understanding. In partnership with Meta, CrowdStrike introduced CyberSOCEval, an open-source benchmark suite designed around real-world SOC workflows, adversary tradecraft, and operational outcomes.
Ultimately, the goal is not to produce the highest benchmark score, but to build evaluations that tell us whether AI systems can deliver reliable defensive outcomes against the complexity and uncertainty of real-world cyber threats.
Additional Resources
- Learn how CrowdStrike Falcon® AI Detection and Response secures AI.
- Download our guide to explore the five steps for frontier AI security readiness.
- Explore cutting-edge cybersecurity research at Day Zero 2026, a summit for the threat research community.
- Experience Fal.Con 2026 from anywhere with Fal.Con Digital, featuring keynote livestreams and on-demand access to 100+ sessions.