
BentoML: The unified model serving framework (Apache 2.0, 7k stars)
Deploying ML models as production APIs with auto-scaling and GPU management — BentoML unifies packaging, serving, and observability in one framework.
8 posts

Deploying ML models as production APIs with auto-scaling and GPU management — BentoML unifies packaging, serving, and observability in one framework.

The memory-augmented LLM agent framework — giving LLMs persistent memory with virtual context management and tiered storage.

A memory layer for AI applications providing persistent, self-updating memory for LLMs with entity extraction and retrieval.

BentoML's LLM serving platform — one-command deployment of any open-source LLM with OpenAI-compatible API.

A scalable model serving library built on Ray for distributed inference with autoscaling, batching, and multi-model pipelines.

Deploying LLM inference at scale with continuous batching, tensor parallelism, and flash attention — Hugging Face's production-grade inference server.

Supporting multiple frameworks, GPU optimization, and dynamic batching — NVIDIA's production inference server.

Supporting 100+ models with built-in embedding, reranking, and multi-GPU inference — XInference unifies inference engines under one OpenAI-compatible API.