Skip to content
Ajith Thaduri

AI-Driven Penetration Testing Harness

Sector
Software products — with a module for AI applications
My role
Creator & lead
Key decisions
6 covered below

In plain terms

A tool that uses AI to look for security holes the way a human tester would — thinking about each specific app — and proves every problem is real before reporting it.

The problem

Penetration testing doesn't scale well. Good testers are scarce and expensive, so most products are tested once a year. Scanners run all the time but only find what someone already wrote a signature for. Neither reasons about the particular application in front of it — and a growing number of products now ship AI features that classic tooling doesn't understand at all.

Architecture

THE HARNESSAI-product moduleinjection · tool abuse · agent reachReconmap the surfaceHypothesisewhere it likely givesGenerate casestargeted, not a wordlistExecute & observewhat actually happenedwhat came back decides what to try nextVerify & reproducenothing reaches a human unprovenRank · report · re-testwith the fix, then proven closedruns only under engagement, against systems there is written authorisation to test

Scroll sideways to see the whole diagram.

Key decisions

  1. 01

    Use the model for judgement

    Sending requests is cheap; deciding which ones are worth sending is the hard part. The harness studies how the application actually behaves, forms a view on where it's likely to be weak, and writes test cases for that system rather than replaying a generic wordlist.

  2. 02

    A loop, not a one-off scan

    Each response shapes the next round. A reply that hints at an underlying assumption becomes the basis for the following test — closer to how a human tester works than to how a scanner does.

  3. 03

    Nothing reported until it's reproduced

    Every candidate finding is re-run by the harness before anyone sees it. A report full of false positives is the quickest way to lose a team's trust, so verification is a gate, not a nice-to-have.

  4. 04

    Coverage for AI products

    A dedicated module for AI features: injection — including payloads that arrive through retrieved content rather than the user — misuse of tool and function calls, data leaking through model output, and how far an agent can reach when one step is manipulated.

  5. 05

    Seeing it through to the fix

    Findings are ranked by real impact, delivered with the specific control that closes each one, and re-tested afterwards so the fix is confirmed.

  6. 06

    Authorised scope only

    The harness runs only under engagement, against systems there is written permission to test. The scope boundary is built into the tool, not left as a policy beside it.

Built with

  • Agentic test-generation loop
  • Automated reproduction & verification
  • AI-application coverage module
  • Impact-ranked reporting
  • Re-test workflow
  • Scoped engagement controls

Contact

Working on something
like this?

I'm open to AI engineering, architecture and training work. Tell me what you're building and what the constraints are — that's usually enough to start.

Prefer a short form? Send a project brief
  • Taking on new projects
  • Usually replies within a day