AI-Driven Penetration Testing Harness
- Sector
- Software products — with a module for AI applications
- My role
- Creator & lead
- Key decisions
- 6 covered below
In plain terms
A tool that uses AI to look for security holes the way a human tester would — thinking about each specific app — and proves every problem is real before reporting it.
The problem
Penetration testing doesn't scale well. Good testers are scarce and expensive, so most products are tested once a year. Scanners run all the time but only find what someone already wrote a signature for. Neither reasons about the particular application in front of it — and a growing number of products now ship AI features that classic tooling doesn't understand at all.
Architecture
Scroll sideways to see the whole diagram.
Key decisions
- 01
Use the model for judgement
Sending requests is cheap; deciding which ones are worth sending is the hard part. The harness studies how the application actually behaves, forms a view on where it's likely to be weak, and writes test cases for that system rather than replaying a generic wordlist.
- 02
A loop, not a one-off scan
Each response shapes the next round. A reply that hints at an underlying assumption becomes the basis for the following test — closer to how a human tester works than to how a scanner does.
- 03
Nothing reported until it's reproduced
Every candidate finding is re-run by the harness before anyone sees it. A report full of false positives is the quickest way to lose a team's trust, so verification is a gate, not a nice-to-have.
- 04
Coverage for AI products
A dedicated module for AI features: injection — including payloads that arrive through retrieved content rather than the user — misuse of tool and function calls, data leaking through model output, and how far an agent can reach when one step is manipulated.
- 05
Seeing it through to the fix
Findings are ranked by real impact, delivered with the specific control that closes each one, and re-tested afterwards so the fix is confirmed.
- 06
Authorised scope only
The harness runs only under engagement, against systems there is written permission to test. The scope boundary is built into the tool, not left as a policy beside it.
Built with
- Agentic test-generation loop
- Automated reproduction & verification
- AI-application coverage module
- Impact-ranked reporting
- Re-test workflow
- Scoped engagement controls