← back to all tools
LOC · Local inference

AirLLM

Run 70B-parameter LLMs on a single 4GB GPU — AirLLM does layer-by-layer inference so huge models fit on modest hardware, with no quantization, distillation, or pruning required.

view source ↗20k stars

license

Apache-2.0

language

Python

status

bookmarked

added

2026-06-17

categories

LOC · Local inference

tags

building blockpythonllminferencequantizationlow-vram