Data-sensitive enterprise
Self-Hosted Model Deployment
Adapted an open-weights model to the client's domain, compressed it to fit their existing GPUs and served it inside their own network, so no request ever left. The project existed because a hosted API wasn't an option.
- QLoRA
- Quantization
- On-prem serving
