
Marwan Sarieddine
AI Engineer

Wednesday, Oct 21 | 2:15 PM - 2:40 PM | Horizon Stage
Whether it comes from autonomy fleets recording sensor data, or simulators synthesizing data, physical AI now produces more data in a day than most teams can turn into training signal in a month. Scaling laws have crossed into the physical world. The bottleneck is quickly moving from collecting and generating data to converting it into intelligence.
This talk walks through the three workloads every physical AI team now runs at fleet scale, data pipelines, model training and simulation, and the three problem areas when running them on real infrastructure: utilization, when GPUs sit idle while the rest of the system catches up; fault tolerance, when a job cannot survive the failures a large cluster guarantees; and slow iteration speed, when teams spend the bulk of their time root-causing and wrangling with infrastructure failures instead of designing experiments and evaluating results.
The talk then showcases how Ray, the open-source library behind NVIDIA's Cosmos curation and GR00T training and used by Torc, Bedrock, Physical Intelligence, Waymo and many others, addresses these three problems at once.

AI Engineer
Wednesday, Oct 21 | 10:30 AM - 11:10 AM | Vista Room
Generating synthetic manipulation data looks embarrassingly parallel. Run the simulator 10000 times, vary the scene, write the episodes. A job queue is the obvious fit and the wrong one.
Every rollout is closed-loop: the simulator steps physics, asks a policy for an action, applies it, steps again. The policy is a multi-billion-parameter VLA and you cannot afford a copy per simulator. Physics and inference compete for the GPU. And when a run produces garbage, you need to know which of 10K rollouts, and why.
We run this on Ray: Isaac Lab simulators across a cluster, a VLA served as an autoscaling endpoint, episodes in LeRobot format. We will show where the naive architecture breaks, what a thousand episodes costs, and which parts do not need distributed infrastructure at all.

Member of Technical Staff, Field Engineer