Case Study
Xoople Runs Physical-World Measurement at Planetary Scale on Anyscale
With Anyscale on Azure, Xoople runs an Earth data foundation model at planetary scale without the need to build a distributed AI infrastructure from the ground up.

500,000
km2 processed in <5 min on 12 A10 GPUs
15
Engineers across product/R&D on Anyscale
~1 petabyte
data generated per day
Xoople is a European data and AI company building Earth's System of Record™: a trusted, continuously updated representation of the physical world. By transforming vast volumes of geospatial, environmental and contextual data into an AI-ready intelligence layer, Xoople enables organisations to understand change, monitor assets and infrastructure, manage risk, and make better decisions with real-world context at global scale.
To make that possible at scale, Xoople runs on Anyscale on Azure, giving the team a managed platform for geospatial foundation model inference that scales from a handful of GPUs to planetary coverage and back down again, all without the need for a dedicated infrastructure team.
LinkChallenges
Xoople is building a scale of data and compute that demands serious infrastructure decisions from the start. As the product moves from a preview with a selected set of customers toward broader deployment, the volume of satellite imagery to ingest, process, and deliver will grow substantially with every new customer and coverage area added. To process that imagery at the pace and breadth the product requires, the team needed to address three interconnected challenges around data volume, pipeline efficiency, and engineering focus.
Satellite-scale inference demands infrastructure that can fan out across trillions of pixels on demand. At full operation, Xoople expects to generate close to a petabyte of satellite imagery per day. Even in the current preview stage, a single customer request can launch up to two trillion pixels for processing, and many such requests can arrive simultaneously. The team designed for that scale from the start which required a compute platform capable of spinning up large clusters on demand, fanning work out across many machines in parallel, and tearing everything back down cleanly once a job completes.
Processing demanding multimodal satellite imagery pipelines effectively to keep GPUs saturated. Unlike standard RGB images, satellite imagery from missions like Sentinel-2 carries roughly 12 spectral bands, spanning visible light, shortwave infrared, and aerosol channels. A single scene carries the data volume of roughly five standard RGB images. Loading, decoding, and chunking that imagery into model-ready tiles is heavily CPU-bound, leaving GPUs idle and slowing epochs by as much as five times relative to what the hardware should deliver. Compounding that, geospatial foundation models like TerraMind carry large, research-oriented Python dependency footprints that are far more fragile to install than typical production services, creating environment instability on worker nodes at scale.
Growing fast meant the team needed to stay focused on the product, not the infrastructure and dependency complexity. Xoople is deployed entirely on Azure, where its satellite data pipeline and enterprise customer integrations already live. Standing up and maintaining Kubernetes clusters on top of that, getting pods to communicate reliably, and keeping the whole thing running as the team grew would have pulled engineers away from the geospatial science and product work that actually differentiates the company.

Milos Colic | VP of Engineering, Xoople
LinkThe Solution
Xoople chose Anyscale to run Ray on Azure as a managed platform, so every engineering hour goes toward the product instead of infrastructure. With Anyscale, the team handles petabyte-scale inference throughput across their full Azure environment, keeps GPUs consistently fed across a complex pipeline, and operates without a dedicated infrastructure function.
With Anyscale, Xoople is able to:
Run planetary-scale geospatial inference with Anyscale Jobs on Azure. Xoople built extensions to Ray Data to read Sentinel-2 imagery directly from cloud-native Zarr stores on Azure Blob, tile it into 224x224 chunks, fan work out across many parallel actors, and write predictions back into output arrays. Anyscale Jobs orchestrate each run end to end within the same Azure environment, keeping the full pipeline within a single managed platform that scales from a handful of GPUs to a fleet covering hundreds of thousands of square kilometers.
Maximize GPUs utilization with fractional GPU allocation. Using fractional GPU Ray actors, Xoople runs multiple concurrent inference workers per A10, working with Anyscale's forward deployment engineers to tune the CPU-to-GPU handoff, so GPUs stay utilized while waiting on preprocessing to complete.
Focus a team of 15 engineers on the product, not the platform. Anyscale, a managed Ray offering, takes cluster and environment operation off the team's hands entirely, so roughly 15 engineers across product engineering and R&D ship product instead of running infrastructure. Anyscale's managed dependency environments handle Python package installation and versioning consistently across every worker node in the cluster, keeping geospatial foundation model environments stable and eliminating the crash-and-restart loops that occurred at scale when packages were installed per-worker.

Milos Colic | VP of Engineering, Xoople
LinkPhysical-world measurement at planetary scale
Xoople works with Sentinel-2 imagery stored in Zarr, a cloud-native chunked array format that standard data sources do not handle well. The team built Ray Data extensions to read it correctly, tile each scene into 224x224 chunks, run inference across the full tile set in parallel, and write predictions back into an output Zarr array.
Since the pipeline is model-agnostic, the team can swap or layer models, from lightweight cloud detection to land use classification to richer embedding generation with TerraMind, without reworking the surrounding infrastructure. On 12 A10 GPUs, the optimized pipeline classifies 500,000 km² of imagery in under five minutes, with throughput scaling near-linearly as additional GPUs are added.
Anyscale Jobs orchestrate each run end to end within the same Azure environment where Xoople's data already lives. The team submits a job, Anyscale spins up the cluster, runs the full pipeline including data loading, inference, observability tracing, and output writing, then tears everything back down. Because every run is version-controlled, the same input always produces the same output, and many jobs can execute in parallel as demand grows, with no coordination overhead between them and no one babysitting infrastructure while a large job runs.

Milos Colic | VP of Engineering, Xoople
LinkGPU utilization at scale
The CPU work of preparing multi-band satellite imagery, loading Zarr tiles, decoding multiple spectral bands, and assembling model inputs, posed a bottleneck in Xoople's early pipeline. GPUs sat idle while preprocessing caught up, resulting in epoch times running up to five times slower than the hardware was capable of delivering. Working closely with Anyscale's forward deployment engineers, the team tuned the CPU-to-GPU handoff, adjusting prefetch depth and concurrency, so inference actors stay occupied rather than waiting on the data pipeline.
In addition, by assigning 0.25 GPUs per Ray actor rather than one worker per device with fractional GPU allocation, Xoople runs multiple concurrent inference actors per A10, keeping each GPU consistently saturated across the full run. As a result, GPU utilization significantly increased, with throughput improving substantially relative to single-worker configurations.

Milos Colic | VP of Engineering, Xoople
LinkEngineering focus and developer velocity
Geospatial foundation models like TerraMind carry large, research-oriented dependency footprints that are far more fragile to install than typical production ML packages. Per-worker dependency installation with uv was exhausting disk space on worker nodes during multi-GPU runs, triggering crash-and-restart loops that stalled the pipeline. Moving dependency management into Anyscale's cluster configuration stopped the instability entirely, keeping all worker environments consistent without any ongoing maintenance overhead.
With infrastructure operation off the team's hands, roughly 15 engineers across product engineering and R&D ship products instead of running clusters. One engineer who came to the project from a purely geospatial background, with no prior Ray or Anyscale experience, picked up the platform, redesigned a core service, and delivered it to production within two months.

Milos Colic | VP of Engineering, Xoople
LinkWhat's Next
With a reliable batch inference foundation in place where models are used to process geospatial data, Xoople plans to expand to time-series change detection, enabling the platform to identify how a given area evolves across satellite passes rather than classifying a single moment in isolation.
Further out, the team intends to fine-tune its geospatial foundation model and serve it across multiple Azure regions. A separate R&D team already runs distributed model training, including an ImageNet workload on Ray Train, on the same Anyscale platform. As Xoople scales toward full planetary coverage, the infrastructure and the partnership are built to grow with it, positioning the company to set the standard for AI-ready Earth data at global scale.

Milos Colic | VP of Engineering, Xoople
