HomeAssetGPU Health Observability Private Preview

GPU Health Observability Private Preview

Anyscale GPU Health Observability correlates GPU hardware signals, including XID errors, ECC memory error counts, SM Clock, and per GPU memory, directly to the exact Ray job and workspace running on each GPU. Every signal is enriched with that context automatically at the source, so a hardware fault never gets mistaken for a software bug again.

Built on DCGM, the industry standard GPU metric exporter, it works across KubeRay and VM deployments today, with support for the K8s Anyscale Operator coming soon. View GPU health across your entire fleet for a platform wide picture, or drill straight from a failed job to the exact GPU behind it, all in one correlated view instead of five disconnected surfaces to piece together by hand.

Fill out the form to request access to the private preview.