
Assembly AI: Speech-to-Text at Scale
Real-time transcription architecture, speaker diarization, and cost optimization strategies for processing 1M+ minutes of audio — cutting our bill by 45%.
17 posts

Real-time transcription architecture, speaker diarization, and cost optimization strategies for processing 1M+ minutes of audio — cutting our bill by 45%.

How we built a document processing pipeline handling 10K+ pages daily using ChatGPT's API — with structured outputs, function calling, and streaming patterns.

How we migrated from GPT-4 to Claude for code analysis — leveraging 200K token context windows, prompt caching, and multi-shot prompting to cut costs by 60%.

Running DeepSeek V3 and R1 locally on Kubernetes — how we cut API costs by 90%, the hardware requirements, and the quantization trade-offs we made.

Benchmarking Gemini 2.5 Pro against GPT-4o for code generation, image understanding, and audio processing — a 3-month study with production workloads.

Building a Grok-inspired real-time knowledge graph system that ingests live data streams and answers questions with up-to-the-minute accuracy.

Handling 200K+ token documents with Kimi — the architecture for processing entire codebases, legal contracts, and research papers in a single context window.

Benchmarking Kling against Sora 2, Veo 3, and other video generation models — a head-to-head comparison on quality, cost, latency, and API ergonomics.

Evaluating Manus as an AI agent framework for our automation pipeline — how it handles complex multi-step tasks and where it still needs human oversight.

From data curation and LoRA training to deployment and monitoring — our 6-month journey fine-tuning Llama 3 for production with 40% accuracy improvement.

Building a multi-agent system with NanoBanana — architecture patterns, tool-calling orchestration, and lessons from deploying agentic workflows in production.

OpenAI's open-source Symphony framework that monitors Linear boards, spawns autonomous coding agents, and delivers proof-of-work — without a single GPU.

Real-time web retrieval, citation-grounded generation, and the architecture behind sub-second answers — building a Perplexity-style search engine.

Benchmarking quality, cost per minute, and the 5 patterns that produced watchable results — experiments with OpenAI Sora 2 for technical explainer videos.

Prompt engineering patterns, temporal consistency techniques, and how we integrated Google Veo 3 into our documentation pipeline for product demo videos.

Terminal-native web agent achieving 86.7% on Online-Mind2Web — Microsoft Research's MIT-licensed harness beating every open-source alternative by 15+ points.

How we built a production RAG system that handles 10K+ queries daily — the architecture, the failures, and the hard-won optimizations that cut latency by 60%.