
Ray Serve: A scalable model serving library (Apache 2.0, 35k stars)
A scalable model serving library built on Ray for distributed inference with autoscaling, batching, and multi-model pipelines.
Explore our latest articles and technical deep dives.

A scalable model serving library built on Ray for distributed inference with autoscaling, batching, and multi-model pipelines.

Combining project management with AI agents for automated task assignment and workflow orchestration — an AI-native task management platform.

Deploying LLM inference at scale with continuous batching, tensor parallelism, and flash attention — Hugging Face's production-grade inference server.

Supporting multiple frameworks, GPU optimization, and dynamic batching — NVIDIA's production inference server.

Extracting and chunking text from PDFs, HTML, Word docs, and images for RAG pipelines — the enterprise document preprocessing library.

Supporting 100+ models with built-in embedding, reranking, and multi-GPU inference — XInference unifies inference engines under one OpenAI-compatible API.

An open-source automation platform with AI-powered workflow builder, pieces framework, and self-hosted deployment for enterprise automation.

Connecting LLMs to your company documents, wikis, and knowledge bases with permission-aware retrieval — Danswer is the open-source enterprise search platform.

A modern AI chat framework supporting multiple LLM providers, plugin system, TTS, vision, and a beautiful UI with multi-user management.