Apocrypha

colibri

Run GLM-5.2 (744B MoE) on consumer hardware with expert streaming

Colibri runs GLM-5.2 (744B-parameter MoE) on consumer hardware with as little as 25GB RAM. Pure C engine with zero runtime dependencies, streaming experts from disk with an LRU cache. Supports MLA attention, MTP speculative decoding, and optional CUDA GPU offloading.

Available in

OverlayNewestEbuildsLast activity
bennypowers GitHub ↗ 1.1.1 4 5 d details ›

Versions & arches

VersionOverlay amd64 Committed
9999 LIVE bennypowers follows upstream — no keywords 13 d view · download · history ↗
1.1.1 bennypowers amd64 testing 5 d view · download · history ↗
1.1.0-r1 bennypowers amd64 testing 6 d view · download · history ↗
1.1.0 bennypowers amd64 testing 6 d view · download · history ↗

Use flags of 1.1.1

  • bench Install dependencies for quality benchmarks (MMLU, HellaSwag, ARC)
  • cuda Build with NVIDIA CUDA GPU offloading for resident expert tensors
  • +python Install the coli Python CLI for model management, chat, serving, and weight conversion
  • rocm Build with AMD ROCm/HIP GPU offloading for resident expert tensors