Local LLM Inference Platform
2026I run my own large language model on a single desktop machine — no cloud, no external API. I set up the model, the serving, and the tooling so it can be driven like any other API. On top of that I built an automated pipeline that turns 80 work records into a deduplicated, evidence-traced memory where every output carries its source. The hard part isn't the model — it's making a model that can be wrong produce output you can trust and trace. Under the hood: an NVFP4-quantized Qwen 27B model, a 262K-token context window, and an OpenAI-compatible API via SGLang (migrated from vLLM behind the same contract).
- A 27B model running locally on one 128 GB desktop box — no cloud
- 262K-token context, exposed as an OpenAI-compatible API
- Automated pipeline: 80 records, zero missed after fallback, every output traceable to its source
- Hybrid deduplication that only escalates ambiguous cases to the model