Skip to content
Ajith Thaduri

Evaluation

Measuring whether an AI system is actually getting better — with numbers, not impressions.

Every serious system I build carries an eval set. It turns 'this prompt feels better' into a number, and it's what lets a team change models, prompts or quantization without fear.

  • What this covers
  • Domain eval sets built from real cases
  • LLM-as-judge with rubrics, checked against human grading
  • Retrieval metrics alongside answer quality
  • Evals wired into CI so regressions can't ship

Work in this area

Product engineering teams

Evaluation Harness for AI Features

A domain eval set with automated scoring, wired into the delivery pipeline, so a prompt change, model swap or new quantization can't ship if it makes the task worse. It turned “it feels better” into a number the team could discuss.

Contact

Working on something
like this?

I'm open to AI engineering, architecture and training work. Tell me what you're building and what the constraints are — that's usually enough to start.

Prefer a short form? Send a project brief
  • Taking on new projects
  • Usually replies within a day