Skip to content
Ajith Thaduri

Models & serving

Running AI models on your own hardware when sending data to an outside service isn't an option.

When cost, latency or data residency rule out a hosted API, I adapt, compress and serve open-weights models myself — and keep an eval set that proves the compressed model still does the job.

  • What this covers
  • LoRA / QLoRA fine-tuning on a single-GPU budget
  • AWQ, GPTQ and GGUF quantization, compared on the real task
  • vLLM and llama.cpp serving, batching and caching
  • Model routing — small model first, large model only when needed

Work in this area

Case study · Legal practices — plaintiff, defense & injuryMedical-Legal Intelligence PlatformHelps attorneys make sense of thousands of pages of medical records — without patient details ever reaching an outside AI service, and with the same answer every time the case is run.

Data-sensitive enterprise

Self-Hosted Model Deployment

Adapted an open-weights model to the client's domain, compressed it to fit their existing GPUs and served it inside their own network, so no request ever left. The project existed because a hosted API wasn't an option.

Contact

Working on something
like this?

I'm open to AI engineering, architecture and training work. Tell me what you're building and what the constraints are — that's usually enough to start.

Prefer a short form? Send a project brief
  • Taking on new projects
  • Usually replies within a day