LOC · Local inference
AirLLM
Run 70B-parameter LLMs on a single 4GB GPU — AirLLM does layer-by-layer inference so huge models fit on modest hardware, with no quantization, distillation, or pruning required.
view source ↗★ 20k stars
license
Apache-2.0
language
Python
status
bookmarked
added
2026-06-17
categories
tags
building blockpythonllminferencequantizationlow-vram