colibri
Run GLM-5.2 (744B MoE) on consumer hardware with expert streaming
Colibri runs GLM-5.2 (744B-parameter MoE) on consumer hardware with as little as 25GB RAM. Pure C engine with zero runtime dependencies, streaming experts from disk with an LRU cache. Supports MLA attention, MTP speculative decoding, and optional CUDA GPU offloading.
homepage ↗ github: JustVugg/colibri
Available in
| Overlay | Newest | Ebuilds | Last activity | |
|---|---|---|---|---|
| bennypowers GitHub ↗ | 1.1.1 | 4 | 5 d | details › |
Versions & arches
| Version | Overlay | amd64 | Committed | |
|---|---|---|---|---|
| 9999 LIVE ≈ | bennypowers | follows upstream — no keywords | 13 d | view · download · history ↗ |
| 1.1.1 ≈ | bennypowers | amd64 testing | 5 d | view · download · history ↗ |
| 1.1.0-r1 ≈ | bennypowers | amd64 testing | 6 d | view · download · history ↗ |
| 1.1.0 ≈ | bennypowers | amd64 testing | 6 d | view · download · history ↗ |
Use flags of 1.1.1
- bench Install dependencies for quality benchmarks (MMLU, HellaSwag, ARC)
- cuda Build with NVIDIA CUDA GPU offloading for resident expert tensors
- +python Install the coli Python CLI for model management, chat, serving, and weight conversion
- rocm Build with AMD ROCm/HIP GPU offloading for resident expert tensors