Skip to content
Ajith Thaduri

I'm Ajith — an AI engineer who likes the unglamorous parts.

Portrait of Ajith Thaduri
Based in
Hyderabad, India
Works with
Teams worldwide, remote
Focus
Secure, production AI
Sectors
Healthcare · Legal · Government

Most of my work is for organisations where the data is sensitive and the output gets scrutinised: health plans, legal practices, a government agency. In that setting, a clever prompt isn't enough. What matters is where the data goes, whether an answer can be reproduced, and whether someone can check how it was reached.

So I spend a lot of time below the prompt layer — isolating sensitive data, making pipelines deterministic, and fine-tuning, compressing and self-hosting models when an API can't meet the cost, latency or data-residency requirement. I also build AI tooling for security testing, which has taught me a lot about where systems break.

Alongside the engineering, I teach. Explaining this material to a room of engineers is the quickest way I know to find the gaps in my own understanding.

Model work

Fine-tuning, quantization and self-hosted serving — the layer below the prompt, where most of the cost sits.

Three things push you below the prompt layer: cost that stops being small once you're at volume, latency you can't fix from outside, and data that isn't allowed to leave the building. A hosted API doesn't solve any of them.

So you host the model yourself — and everything the API was quietly handling becomes your job: how much VRAM it needs, how fast it runs under concurrent load, and how much capability you give up to make it fit. It's less visible than prompt design, but it's where a lot of the real engineering happens.

THE STACK I OWNOpen-weights basechosen for the domainAdaptLoRA · QLoRA · SFTCompressAWQ 4-bit · GGUF K-quantvLLM · GPUPagedAttention · batchingllama.cpp · CPU / edgeGGUF · Metal · no GPURE-MEASURED AFTER EVERY CHANGEVRAM footprintdoes it still fitTokens/sec under batchat real concurrencyTask accuracy deltaon a held domain eval setresults decide the next quantizationperplexity is not the test — the domain eval set is
01

Adaptation

Teaching a model the domain without training from scratch.

  • LoRA and QLoRA adapters for domain behaviour on a single-GPU budget
  • Supervised fine-tuning and instruction tuning on curated task data
  • Dataset curation, deduplication and eval hygiene — the eval set never touches training
  • Knowing when not to fine-tune: retrieval solves more than people expect and is cheaper to maintain
02

Quantization

Making it fit without making it useless.

  • AWQ 4-bit for GPU serving — activation-aware, so the most important weights keep their precision
  • GPTQ compared against the same calibration set, rather than picked on reputation
  • GGUF K-quants (Q4_K_M, Q5_K_M) for llama.cpp on CPU, Metal and edge devices
  • bitsandbytes NF4 and int8 for quick experiments before choosing a serving format
  • Calibration sets drawn from the real domain data, not a generic web sample
03

Serving & runtime

Throughput is a design decision, not a number from a benchmark page.

  • vLLM — PagedAttention, continuous batching and tensor parallelism for concurrent GPU serving
  • llama.cpp — CPU and Metal inference where a GPU isn't available, affordable or permitted
  • KV-cache and prefix-cache management so repeated system context isn't recomputed
  • Balancing latency and throughput: batch size, context length, concurrency limits, queueing
  • Model routing — a small model answers first, a larger one only when the task needs it
04

Machine learning foundations

The groundwork that still shapes most of the decisions above.

  • Supervised and unsupervised learning, feature engineering, model selection
  • Neural network architectures and training loops
  • Custom NLP models for domain classification and structured extraction
  • Evaluation built around the cost of each kind of error — missing PHI is not the same as over-flagging it

Tools I use

Things I've shipped with in production, grouped by what they're for.

Languages

  • Python
  • TypeScript
  • Java
  • C

Agents & orchestration

  • LangGraph
  • LangChain
  • CrewAI
  • AutoGen
  • OpenAI Agents SDK
  • MCP
  • n8n

Retrieval

  • Embeddings
  • Hybrid search
  • Chunking & indexing
  • pgvector
  • FAISS
  • ChromaDB

Models & serving

  • Claude API
  • OpenAI API
  • Hugging Face
  • vLLM
  • llama.cpp
  • LoRA / QLoRA
  • AWQ · GPTQ · GGUF

Safety & evaluation

  • PII / PHI redaction
  • Prompt-injection defense
  • Structured outputs
  • LLM-as-judge
  • Domain eval sets
  • Audit logging

Voice & realtime

  • ElevenLabs
  • Speech-to-text
  • Streaming
  • Turn-taking & barge-in

Backend & data

  • FastAPI
  • PostgreSQL
  • Redis
  • SQLAlchemy
  • Alembic

Frontend

  • React
  • Next.js
  • Tailwind CSS

Cloud & identity

  • AWS
  • GCP
  • SSO / JWT
  • RBAC

In short

The quick version.

Common questions, answered plainly.

Ajith Thaduri is an AI engineer and technical consultant based in Hyderabad, India. He designs and builds production AI systems — agentic architectures, retrieval pipelines and self-hosted language models — for healthcare, legal and government organisations where the data is regulated and the output is closely checked. He also builds AI-driven security tooling and teaches engineering teams.

Contact

Working on something
like this?

I'm open to AI engineering, architecture and training work. Tell me what you're building and what the constraints are — that's usually enough to start.

Prefer a short form? Send a project brief
  • Taking on new projects
  • Usually replies within a day