fastflowlm
NPU-first LLM runtime for AMD Ryzen AI (XDNA2) processors
FastFlowLM (FLM) is a lightweight LLM inference runtime purpose-built for AMD Ryzen AI NPUs (XDNA2 architecture). It provides an Ollama-style CLI and OpenAI-compatible server API for running language models entirely on the NPU with no GPU or CPU compute required. Supported hardware: Ryzen AI 300-series (Strix Point, Strix Halo), 400-series (Gorgon Point), and Z2 Extreme. XDNA1 (Ryzen AI 7000/8000) is NOT supported. The orchestration code and CLI are MIT-licensed. NPU compute kernels (xclbins) are proprietary binaries, free for commercial use under $10M annual company revenue.
homepage ↗homepage ↗ github: FastFlowLM/FastFlowLM
Available in
| Overlay | Newest | Ebuilds | Last activity | |
|---|---|---|---|---|
| stuff GitHub ↗ | 0.9.45 | 3 | 2 d | details › |
Versions & arches
Use flags of 0.9.45
- openrc Install OpenRC init.d/conf.d files for a supervise-daemon-managed service
- systemd Install a systemd template service unit ([email protected])