EchoBench is a human-calibrated benchmark for autonomous web application pentesting that scores model-and-harness systems against associate pentesters from NetSPI University. It uses four geometric-mean components—finding fidelity, difficulty-reach, OWASP breadth, and repeatability—while publishing cohort, configuration, provenance, and repeatability details alongside every score. #EchoBench #NetSPIUniversity #OWASP2025
Keypoints
- EchoBench evaluates autonomous AI web application pentesting against a measured human cohort rather than against abstract benchmark numbers.
- A score of 100 represents parity with associate pentesters from NetSPI University on the same applications and under the same finding definitions.
- The headline ECHO score is built from four components: finding fidelity, difficulty-reach, OWASP breadth, and repeatability.
- Models are evaluated together with their harnesses, so prompts, tools, browser environment, memory, execution policy, and reporting workflow all affect the score.
- The benchmark corpus currently includes five web applications and 184 canonical finding IDs, with human-assigned difficulty labels and applicable OWASP 2025 categories.
- The human reference is drawn from roughly 300 application assessments across five years of NetSPI University cohorts, with current matched data still being aligned.
- Every published score includes provenance, configuration, component values, uncapped human-relative results, and repeat coverage to keep the evidence transparent.
MITRE Techniques
- [T1190] Exploit Public-Facing Application – The benchmark evaluates web application pentesting against live target applications and finding opportunities, including “five web applications” and “validated web application findings under defined assessment conditions.”
- [T1580] Cloud Infrastructure Discovery – The article notes that models may identify targets through visible clues like hostnames, branding, certificates, and page titles, which reflects discovering environment or application identity from exposed information (‘hostname or branding already supplied the answer’).
- [T1595] Active Scanning – The autonomous pentesting workflow implies systematic target exploration to find validated findings, described as comparing agent results on the same application and measuring whether the system can reach findings across easy, moderate, and hard bands (‘can the system get to the hard stuff at all?’).
- [T1068] Exploitation for Privilege Escalation – The discussion of higher-difficulty findings and application pentesting encompasses deeper investigative and exploitative paths, including hard findings that require “sustained exploration, state tracking, careful validation, and longer reasoning chains.”
Indicators of Compromise
- [Benchmark artifacts] Benchmark scope and scoring corpus – five web applications, 184 canonical finding IDs, and roughly 300 completed application assessments.
- [Versions / categories] Human labeling and coverage metadata – OWASP 2025 and five years of NetSPI University cohorts.