llm-serving — independent software & tools
-
Quartermaster
Run any model without tuning a single flag. Text, image and audio on one machine: Quartermaster works out what fits in your VRAM, launches each model with computed flags, and hot-swaps them on demand behind one OpenAI- and Anthropic-compatible API.
-
dgx-spark-llm-platform
Self-hosted multi-user LLM platform for the NVIDIA DGX Spark: OpenAI-compatible API (vLLM + LiteLLM), per-user keys & budgets, self-service portal with playground and an AI support assistant.
-
DIO
Drop-in OpenAI- and Ollama-compatible LLM gateway that learns each backend's latency online and routes vLLM / SGLang / TGI / Ollama with SLO-aware admission. No engine patches.