
ExLlamaV2: The fastest inference engine for quantized Llama/HF models (MIT, 5k stars)
Achieving 140+ tok/s on a single RTX 4090 with 4-bit quantization — the fastest inference engine for quantized Llama/HF models.
9 posts

Achieving 140+ tok/s on a single RTX 4090 with 4-bit quantization — the fastest inference engine for quantized Llama/HF models.

Nomic AI's local LLM ecosystem running models on consumer hardware with CPU-only support, a Python API, and local RAG capabilities.

A privacy-first desktop AI assistant running 100% offline with built-in model downloads, extensions, and a clean chat interface.

The definitive LLM inference engine — running quantized models on CPU and GPU with state-of-the-art performance via GGUF format.

Mozilla's single-file LLM runner distributing AI models as executable files that run on 6GB VRAM without installation.

A desktop application for running local LLMs — with a built-in model browser, chat interface, and OpenAI-compatible local API server.

A drop-in OpenAI API replacement running LLMs, image generation, audio transcription, and TTS entirely locally with no GPU required.

The easiest way to run LLMs locally — one-command setup for 100k+ models with OpenAI-compatible API and GPU acceleration.

Supporting 20+ model loaders, LoRA, RLHF, multimodal models, and custom chat templates — the most feature-rich local LLM interface.