
open-source-ai• 15 min
ExLlamaV2: The fastest inference engine for quantized Llama/HF models (MIT, 5k stars)
Achieving 140+ tok/s on a single RTX 4090 with 4-bit quantization — the fastest inference engine for quantized Llama/HF models.
#exllamav2#inference