llama-swap
Reliable LLM model swapping proxy for llama.cpp / vllm / etc.
llama-swap is an OpenAI/Anthropic-compatible HTTP proxy that starts and stops local LLM inference servers (llama.cpp, vllm, mlx-server, etc.) on demand based on the requested model. Lets a single API endpoint serve many models without keeping them all resident in GPU/NPU memory.
homepage ↗ github: mostlygeek/llama-swap github: istitov/extra-stuff gitlab: istitov/extra-stuff codeberg: istitov/extra-stuff
Available in
| Overlay | Newest | Ebuilds | Last activity | |
|---|---|---|---|---|
| stuff GitHub ↗ | 242 | 2 | 2 d | details › |
Versions & arches
Use flags of 242
- openrc Install OpenRC init.d/conf.d files for a supervise-daemon-managed service
- systemd Install a systemd template service unit ([email protected])
- ui Build and embed the Svelte web UI (pulls net-libs/nodejs; runs npm at build time, requires network access during compile).