Apocrypha

llama-swap

Reliable LLM model swapping proxy for llama.cpp / vllm / etc.

llama-swap is an OpenAI/Anthropic-compatible HTTP proxy that starts and stops local LLM inference servers (llama.cpp, vllm, mlx-server, etc.) on demand based on the requested model. Lets a single API endpoint serve many models without keeping them all resident in GPU/NPU memory.

Available in

OverlayNewestEbuildsLast activity
stuff GitHub ↗ 242 2 2 d details ›

Versions & arches

VersionOverlay amd64arm64 Committed
242 stuff amd64 testing arm64 testing 3 d view · download · history ↗
241 stuff amd64 testing arm64 testing 6 d view · download · history ↗

Use flags of 242

  • openrc Install OpenRC init.d/conf.d files for a supervise-daemon-managed service
  • systemd Install a systemd template service unit ([email protected])
  • ui Build and embed the Svelte web UI (pulls net-libs/nodejs; runs npm at build time, requires network access during compile).