Webinar
Production-ready Distributed Inference with Ray Serve
Thursday, November 5 8:30 AM PST | 11:30 AM EST | 5:30 PM CETDesign scalable distributed inference systems for production AI applications, from classic ML models to large language models. Learn how Ray Serve delivers high-throughput, low-latency serving, and how to evaluate the performance tradeoffs in autoscaling, batching, routing, and inference-engine optimization.
Topics include:
Building Ray Serve applications with deployments, composition, and load-aware request routing
Deploying and scaling LLM endpoints with Ray Serve LLM on engines like vLLM and SGLang
Optimizing the inference engine with the KV cache, paged attention, prefix caching, and continuous batching
Scaling frontier models with KV-cache-aware routing, prefill/decode disaggregation, and wide expert parallelism for mixture-of-experts (MoE) models
Deploying Custom LLMs on Anyscale and Integrating Them with Agentic Platforms (Claude Code, Cursor, Codex)
Comparing costs across self-hosted GPUs, subscriptions, and token-based API billing.