← back to all tools
LOC · Local inference

Colibrì

A pure-C, zero-dependency inference engine that runs frontier Mixture-of-Experts models — like GLM-5.2's 744B — on hardware you already own. It keeps the dense core in RAM and streams experts from disk on demand, treating VRAM, RAM and storage as one tiered memory. It even serves an Anthropic-compatible API, so Claude Code talks to it directly.

view source ↗24k stars
token route ~25 GB RAM dense core LRU cache 744B experts · on disk streamed on demand

license

Apache-2.0

language

C

status

bookmarked

added

2026-08-11

categories

LOC · Local inference

tags

building blockinferencepure CMoE