Services Plans
About Us
Login

Yaga Benchmark

In AI pentesting, the same model performs far better
with Yaga's harness than on its own.

Yaga's results

Using Yaga's harness combined with several models, the results are far superior to those of any isolated model.

Yaga's results Line chart comparing the percentage of scenarios solved by Yaga (combined models) and by GPT-5.6, Opus 5, Grok 4.6 and Sonnet 5 on their own, in black box, gray box and white box modes. Yaga stays well above every isolated model in every mode. 0 20 40 60 80 100 Black box Gray box White box 96.2 97.0 98.8 Yaga Opus 5 GPT-5.6 Grok 4.6 Sonnet 5
Opus 5 GPT-5.6 Grok 4.6 Sonnet 5 Yaga · 4 models combined

Share of the fixed scenario set solved by test mode. Hover the points for the exact value.

Yaga vs. each isolated model

With Yaga, your pentest is far superior to that of any model running in isolation.

Model Black box Gray box White box
Yaga · 4 models combined 96.2% 97.0% 98.8%
Opus 5 40.0% 48.9% 61.0%
GPT-5.6 39.5% 48.5% 60.9%
Grok 4.6 39.0% 47.5% 58.7%
Sonnet 5 27.4% 36.5% 49.6%

HackerSec internal benchmark over a fixed set of scenarios, measuring confirmed and exploitable vulnerabilities. Each model is evaluated on its own, without Yaga's harness; the combined line is production Yaga, with the four models in a single orchestrated run.

Yaga v2.7 · Last updated · August 18, 2026

How we measure

The same yardstick for all

Every model runs the same set of scenarios, under the same conditions. The yardstick is identical for all, and the difference in results reflects only the method each one applies.

Validated by a human expert

In YagaBench, a vulnerability is counted only after it has been truly exploited and confirmed by a HackerSec specialist. Only what is proven makes it into the score.

The model answers. Yaga operates.

On its own, a model only answers. Inside Yaga's harness, it recons the target, exploits real flaws, and proves the impact, like a pentester would. Orchestrated, the same AI performs far better.

Measured with the limits on

Yaga ran the benchmark under the same guardrails that apply in a client environment. The numbers above come from real operation, not from a looser mode built for the test.

Scenarios aligned with OWASP PTES MITRE CWE NIST

See Yaga working on your target

Put the same harness to work on your scope, with HackerSec human oversight.