HomeEventsProduction-ready Distributed Inference with Ray Serve

Webinar

Production-ready Distributed Inference with Ray Serve

Design scalable distributed inference systems for production AI applications, from classic ML models to large language models. Learn how Ray Serve delivers high-throughput, low-latency serving, and how to evaluate the performance tradeoffs in autoscaling, batching, routing, and inference-engine optimization.

Topics include:

  • Building Ray Serve applications with deployments, composition, and load-aware request routing

  • Deploying and scaling LLM endpoints with Ray Serve LLM on engines like vLLM and SGLang

  • Optimizing the inference engine with the KV cache, paged attention, prefix caching, and continuous batching

  • Scaling frontier models with KV-cache-aware routing, prefill/decode disaggregation, and wide expert parallelism for mixture-of-experts (MoE) models

  • Deploying Custom LLMs on Anyscale and Integrating Them with Agentic Platforms (Claude Code, Cursor, Codex)

  • Comparing costs across self-hosted GPUs, subscriptions, and token-based API billing.