In AI pentesting, the same model performs far better
with Yaga's harness than on its own.
Using Yaga's harness combined with several models, the results are far superior to those of any isolated model.
Share of the fixed scenario set solved by test mode. Hover the points for the exact value.
With Yaga, your pentest is far superior to that of any model running in isolation.
HackerSec internal benchmark over a fixed set of scenarios, measuring confirmed and exploitable vulnerabilities. Each model is evaluated on its own, without Yaga's harness; the combined line is production Yaga, with the four models in a single orchestrated run.
Yaga v2.7 · Last updated · August 18, 2026
Every model runs the same set of scenarios, under the same conditions. The yardstick is identical for all, and the difference in results reflects only the method each one applies.
In YagaBench, a vulnerability is counted only after it has been truly exploited and confirmed by a HackerSec specialist. Only what is proven makes it into the score.
On its own, a model only answers. Inside Yaga's harness, it recons the target, exploits real flaws, and proves the impact, like a pentester would. Orchestrated, the same AI performs far better.
Yaga ran the benchmark under the same guardrails that apply in a client environment. The numbers above come from real operation, not from a looser mode built for the test.
Put the same harness to work on your scope, with HackerSec human oversight.