xoople-logo-green

Case Study

Xoople Runs Physical-World Measurement at Planetary Scale on Anyscale

With Anyscale on Azure, Xoople runs an Earth data foundation model at planetary scale without the need to build a distributed AI infrastructure from the ground up.

Digital globe visualizing Xoople’s planetary-scale Earth data system

500,000

km2 processed in <5 min on 12 A10 GPUs

15

Engineers across product/R&D on Anyscale

~1 petabyte

data generated per day

Xoople is a European data and AI company building Earth's System of Record™: a trusted, continuously updated representation of the physical world. By transforming vast volumes of geospatial, environmental and contextual data into an AI-ready intelligence layer, Xoople enables organisations to understand change, monitor assets and infrastructure, manage risk, and make better decisions with real-world context at global scale.

To make that possible at scale, Xoople runs on Anyscale on Azure, giving the team a managed platform for geospatial foundation model inference that scales from a handful of GPUs to planetary coverage and back down again, all without the need for a dedicated infrastructure team.

LinkChallenges

Xoople is building a scale of data and compute that demands serious infrastructure decisions from the start. As the product moves from a preview with a selected set of customers toward broader deployment, the volume of satellite imagery to ingest, process, and deliver will grow substantially with every new customer and coverage area added. To process that imagery at the pace and breadth the product requires, the team needed to address three interconnected challenges around data volume, pipeline efficiency, and engineering focus.

  • Satellite-scale inference demands infrastructure that can fan out across trillions of pixels on demand. At full operation, Xoople expects to generate close to a petabyte of satellite imagery per day. Even in the current preview stage, a single customer request can launch up to two trillion pixels for processing, and many such requests can arrive simultaneously. The team designed for that scale from the start which required a compute platform capable of spinning up large clusters on demand, fanning work out across many machines in parallel, and tearing everything back down cleanly once a job completes.

  • Processing demanding multimodal satellite imagery pipelines effectively to keep GPUs saturated. Unlike standard RGB images, satellite imagery from missions like Sentinel-2 carries roughly 12 spectral bands, spanning visible light, shortwave infrared, and aerosol channels. A single scene carries the data volume of roughly five standard RGB images. Loading, decoding, and chunking that imagery into model-ready tiles is heavily CPU-bound, leaving GPUs idle and slowing epochs by as much as five times relative to what the hardware should deliver. Compounding that, geospatial foundation models like TerraMind carry large, research-oriented Python dependency footprints that are far more fragile to install than typical production services, creating environment instability on worker nodes at scale.

  • Growing fast meant the team needed to stay focused on the product, not the infrastructure and dependency complexity. Xoople is deployed entirely on Azure, where its satellite data pipeline and enterprise customer integrations already live. Standing up and maintaining Kubernetes clusters on top of that, getting pods to communicate reliably, and keeping the whole thing running as the team grew would have pulled engineers away from the geospatial science and product work that actually differentiates the company.

"A single request can launch one and a half to two trillion pixels for processing. Now picture twenty of those arriving at once. That is the scale we have to design for before it arrives."
Milos Colic's profile

Milos Colic | VP of Engineering, Xoople

xoople-logo-green logo

LinkThe Solution

Xoople chose Anyscale to run Ray on Azure as a managed platform, so every engineering hour goes toward the product instead of infrastructure. With Anyscale, the team handles petabyte-scale inference throughput across their full Azure environment, keeps GPUs consistently fed across a complex pipeline, and operates without a dedicated infrastructure function.

With Anyscale, Xoople is able to: 

  • Run planetary-scale geospatial inference with Anyscale Jobs on Azure. Xoople built extensions to Ray Data to read Sentinel-2 imagery directly from cloud-native Zarr stores on Azure Blob, tile it into 224x224 chunks, fan work out across many parallel actors, and write predictions back into output arrays. Anyscale Jobs orchestrate each run end to end within the same Azure environment, keeping the full pipeline within a single managed platform that scales from a handful of GPUs to a fleet covering hundreds of thousands of square kilometers.

  • Maximize GPUs utilization with fractional GPU allocation. Using fractional GPU Ray actors, Xoople runs multiple concurrent inference workers per A10, working with Anyscale's forward deployment engineers to tune the CPU-to-GPU handoff, so GPUs stay utilized while waiting on preprocessing to complete.

  • Focus a team of 15 engineers on the product, not the platform. Anyscale, a managed Ray offering, takes cluster and environment operation off the team's hands entirely, so roughly 15 engineers across product engineering and R&D ship product instead of running infrastructure. Anyscale's managed dependency environments handle Python package installation and versioning consistently across every worker node in the cluster, keeping geospatial foundation model environments stable and eliminating the crash-and-restart loops that occurred at scale when packages were installed per-worker.

"Partnering with Anyscale means every engineering hour we have goes into the product, and that is exactly where those hours belong."
Milos Colic's profile

Milos Colic | VP of Engineering, Xoople

xoople-logo-green logo

LinkPhysical-world measurement at planetary scale

Xoople works with Sentinel-2 imagery stored in Zarr, a cloud-native chunked array format that standard data sources do not handle well. The team built Ray Data extensions to read it correctly, tile each scene into 224x224 chunks, run inference across the full tile set in parallel, and write predictions back into an output Zarr array.

Since the pipeline is model-agnostic, the team can swap or layer models, from lightweight cloud detection to land use classification to richer embedding generation with TerraMind, without reworking the surrounding infrastructure. On 12 A10 GPUs, the optimized pipeline classifies 500,000 km² of imagery in under five minutes, with throughput scaling near-linearly as additional GPUs are added.

Anyscale Jobs orchestrate each run end to end within the same Azure environment where Xoople's data already lives. The team submits a job, Anyscale spins up the cluster, runs the full pipeline including data loading, inference, observability tracing, and output writing, then tears everything back down. Because every run is version-controlled, the same input always produces the same output, and many jobs can execute in parallel as demand grows, with no coordination overhead between them and no one babysitting infrastructure while a large job runs.

"Anyscale Jobs packages our code and handles the whole run for us: spinning up a cluster, loading and running the data, writing the output, and tearing it down when the job is finished. With Anyscale, we can do this for multiple jobs, all running in parallel, without needing to directly manage compute or orchestration."
Milos Colic's profile

Milos Colic | VP of Engineering, Xoople

xoople-logo-green logo

LinkGPU utilization at scale

The CPU work of preparing multi-band satellite imagery, loading Zarr tiles, decoding multiple spectral bands, and assembling model inputs, posed a bottleneck in Xoople's early pipeline. GPUs sat idle while preprocessing caught up, resulting in epoch times running up to five times slower than the hardware was capable of delivering. Working closely with Anyscale's forward deployment engineers, the team tuned the CPU-to-GPU handoff, adjusting prefetch depth and concurrency, so inference actors stay occupied rather than waiting on the data pipeline.

In addition, by assigning 0.25 GPUs per Ray actor rather than one worker per device with fractional GPU allocation, Xoople runs multiple concurrent inference actors per A10, keeping each GPU consistently saturated across the full run. As a result, GPU utilization significantly increased, with throughput improving substantially relative to single-worker configurations.

"Anyscale isn't just a platform, it's also a team of experts. Their forward deployment engineers spotted that CPU-bound image loading was leaving our GPUs idle and running epochs up to five times slower than they should. They worked side by side with us to tune that handoff, so the GPUs stayed fed. That kind of hands-on partnership is what drives real cost-efficiency gains in our business."
Milos Colic's profile

Milos Colic | VP of Engineering, Xoople

xoople-logo-green logo

LinkEngineering focus and developer velocity

Geospatial foundation models like TerraMind carry large, research-oriented dependency footprints that are far more fragile to install than typical production ML packages. Per-worker dependency installation with uv was exhausting disk space on worker nodes during multi-GPU runs, triggering crash-and-restart loops that stalled the pipeline. Moving dependency management into Anyscale's cluster configuration stopped the instability entirely, keeping all worker environments consistent without any ongoing maintenance overhead.

With infrastructure operation off the team's hands, roughly 15 engineers across product engineering and R&D ship products instead of running clusters. One engineer who came to the project from a purely geospatial background, with no prior Ray or Anyscale experience, picked up the platform, redesigned a core service, and delivered it to production within two months.

"We are a zero-to-one team building our first product, and Anyscale has let us build a better service and a better architecture, with happier engineers who get to work with tools that excite them."
Milos Colic's profile

Milos Colic | VP of Engineering, Xoople

xoople-logo-green logo

LinkWhat's Next

With a reliable batch inference foundation in place where models are used to process geospatial data, Xoople plans to expand to time-series change detection, enabling the platform to identify how a given area evolves across satellite passes rather than classifying a single moment in isolation. 

Further out, the team intends to fine-tune its geospatial foundation model and serve it across multiple Azure regions. A separate R&D team already runs distributed model training, including an ImageNet workload on Ray Train, on the same Anyscale platform. As Xoople scales toward full planetary coverage, the infrastructure and the partnership are built to grow with it, positioning the company to set the standard for AI-ready Earth data at global scale.

"As AI moves into the physical world, it needs a trusted, continuously updated view of that world, delivered in a form it can use. That is what we are building, and Anyscale gives us the compute foundation to do it at scale."
Milos Colic's profile

Milos Colic | VP of Engineering, Xoople

xoople-logo-green logo

“With 15 engineers processing trillions of pixels per request, Xoople needs infrastructure it can trust at scale. Anyscale provides it, so every engineering hour goes into the product instead of managing infrastructure.”

Milos Colic

VP of Engineering, Xoople

Milos Colic
Want to give it a try?