dry-run first · protected real runs
Multi-model penetration testing harness for the AI era
A staged Hunt → Validate → Score cockpit wraps the real Jean-Claude runner across 44 attack classes spanning 9 security frameworks — from OWASP Top 10:2025 and MITRE ATLAS to MCP, RAG, and agentic threats.
Total Runs
118$3.25 spend to dateFindings
1613 confirmed by validatorClass Coverage
100%44 / 44 classes with execution or finding evidenceAvg Run Cost
$0.0276per pipeline executionSeverity Distribution
Critical0
High0
Medium6
Low10
Info0
Framework Coverage
Distinct canonical attack classes covered by persisted runs; finding volume is shown separately.
OWASP
11 / 11
MCP
8 / 8
RAG
4 / 4
Agentic
7 / 7
ATLAS
2 / 2
API
5 / 5
BLA
1 / 1
NHI
4 / 4
LLM
2 / 2
Recent Runs
View all →pentest-2026-06-14-a39c81cdmbarbine/jc-pentest-harness
dry-run plan complete36d agopentest-2026-06-13-d1bca080mbarbine/jc-pentest-harness
model pipeline running37d agopentest-2026-06-13-a8865ae5mbarbine/jc-pentest-harness
dry-run plan complete37d agopentest-2026-06-13-a17614cembarbine/jc-pentest-harness
dry-run plan complete37d agopentest-2026-06-13-6dadaeaembarbine/jc-pentest-harness
dry-run plan complete37d agoTop Findings
View all →MEDIUMconfirmedf6b87aaf
tier1Gate TOCTOU: non-atomic check-then-reset enables concurrent runsExcessive AgencyMEDIUMconfirmed96087558
Auto-gapfill routes around model safety refusals with no human checkpointAgent Identity AbuseMEDIUMconfirmed0e4cb3fc
tier1Gate reset swallows failure silently — gate can remain armedAgent Goal HijackingMEDIUMconfirmedc0ad509a
Boot prompt YES/NO natural-language flow has no provenance bindingInjectionLOWconfirmed9dde3a1c
Default executionMode in example config is api-multi-model (live API)Security MisconfigurationLOWconfirmed491f5285
Example budget caps default to $5/$25 — should default to $0Security Misconfiguration