Sovereign AI Cloud / Infrastructure
Platform Engineering · InfraOps · MLOps

Rebuilding a Sovereign Multi-Cloud AI Compute Platform That Actually Works

By Ehtisham Mubarik, Founder & Principal Engineer · Published November 15, 2025

15+
Cloud Providers Integrated
700
Data Centers Reachable
up to 85%
Cost Savings via GPU Pooling
2 months
Full Platform Rebuild
5 to 6
Eprecisio Engineers
5+ years
Client Relationship
Vanilla KubernetesMulti-Cloud GPU PoolNVIDIA GPU OperatorEncrypted Network MeshArgoCDTerraformKubeflowVMwareService MeshHelmCAST AIAgentic Architect Plugin

The client is a sovereign AI cloud platform: a single control plane that pools GPU compute across 15+ cloud providers and 700 data centers in 200 cities, so AI teams can mix and match capacity from hyperscalers (AWS, Azure, GCP, Oracle) alongside niche GPU providers (Civo, Runpod, Vultr, Scaleway, Hetzner, OVH, IONOS, Leafcloud, Equinix, DigitalOcean) without ever handing sovereignty over to any single one. A representative production cluster runs 10x H100 on GCP alongside 50x H100 on Civo, exposed as one unified Kubernetes cluster. The platform is backed by SFC Capital, Rule30.vc, and the NVIDIA Inception Program, and holds active membership in CNCF and the Linux Foundation. The vision was right. The infrastructure holding it together was not. Eprecisio joined as the founding engineering partner, rebuilt the entire platform from the infrastructure layer up, and has been the core delivery team through the product's growth to its current position as a recognised player at KubeCon.

Sovereign multi-cloud AI compute platform, UK - GPU orchestration across 15+ clouds with data sovereignty

How the relationship started

The engagement did not start with a Kubernetes platform. It started with a healthcare project.

Ehtisham joined the founder's team to work on a healthcare compliance project. The team was small, the stack was complex, and there were in-house challenges managing the infrastructure to the standard that healthcare compliance demands. Ehtisham stepped in individually to address those blockers.

That engagement built the trust that led to this partnership. When the founder started building the vision for a Kubernetes automation product, Eprecisio was the partner they turned to. The relationship that started with one engineer on a healthcare project is now a team of 5 to 6 engineers working full-time on a product that is being presented at KubeCon.

The product is commercially ambitious. Its differentiator is sovereign multi-cloud orchestration: teams pick the jurisdiction, provider mix, and price point that fits their AI or HPC workload, and the platform handles provisioning, networking, and billing end-to-end without the customer needing separate accounts with each provider. That means compute moves to where the data already lives, not the other way around, so egress overhead drops and data-residency requirements stay respected. The offering breaks down into three complementary products: an agentic solution architect that designs AI workflow deployments, a sovereign multi-cloud control plane that runs them, and a distributed edge orchestration layer that extends the fabric to the edge. The platform also ships a marketplace of plugins that teams deploy directly into their clusters: Kubeflow for ML pipelines, vLLM and KServe for inference serving, vector databases, Kafka streaming, Laravel stack integrations, and 100+ other open source Helm charts. All available through the platform interface without the customer needing deep Kubernetes expertise. Underneath sits a proprietary network fabric with sub-5ms latency, up to 100 Gbps bandwidth, and an encrypted mesh that knits 15+ clouds and 700 data centers together at L2/L3. To deliver that experience credibly, the platform itself has to be faultless.

The state of the platform when active development began

Seven months ago, when the current active engagement began in earnest, the platform was failing repeatedly. Not occasionally. Continuously.

The core problem was that the architecture had accumulated instability at every layer. Networking was unreliable when customers connected their own hardware from different environments. State management did not exist in any meaningful form, so the platform had no consistent picture of what was running, what had failed, or what needed attention.

AreaState at the startImpact on customers
Platform stabilityContinuously failing, no root cause trackingCustomers could not trust clusters they provisioned
State managementNo unified state layerNode status, provisioning state, and cluster health were inconsistent across views
NetworkingUnreliable when customers connected hardware from different environmentsWorkload connectivity failed silently when hardware was registered from mixed environments
GPU managementNo operator-level control over GPU allocationML teams could not rely on GPU provisioning
MarketplaceCharts deployed inconsistently, no deployment framework100+ open source charts had no reliable install path
Customer onboardingNode registration and NACL creation unreliableNew customer setup required manual intervention
AlertingNo structured alerting or status notificationsFailures went undetected until customers reported them
PricingConnectivity issues with external cloud provider billing APIsCost data was inaccurate or unavailable

When funding came in and the product needed to scale, the architecture underneath it was not ready. The decision was made to stop patching and do a full rebuild.

The team and how the engagement evolved

The engagement grew the way most of our strongest relationships do. It started with one person, proved its value, and expanded as the scope became clear.

RoleWhat Eprecisio owns
Platform engineering leadInfrastructure architecture, Kubernetes operator design, cross-cloud networking
DevOps engineers (x2)CI/CD, cluster lifecycle management, ArgoCD GitOps, Terraform modules
Full-stack engineerPlatform UI, customer-facing APIs, marketplace frontend
Product managerRoadmap, PRDs, delivery process, sprint management
AI-native developmentAI-assisted feature development and code quality processes

This is not a vendor relationship. Eprecisio owns the roadmap process, manages delivery, writes the PRDs, and makes architecture decisions. The founder focuses on business development, partnerships, and product vision. The engineering execution is ours.

The rebuild: what we actually did

The 2-month rebuild was not a rewrite of features. It was a reconstruction of the foundation the features run on.

Infrastructure and state management layer. The platform had no consistent state model. We designed and implemented a state management architecture that tracks every cluster, node, and workload across all three cloud environments in real time. Every provisioning operation now has defined state transitions with persistence and recovery paths.

Networking for bring-your-own-hardware. The platform does not provision managed Kubernetes services. It deploys vanilla Kubernetes clusters on hardware that customers register from wherever that hardware lives. That means the networking layer has to handle arbitrary hardware from arbitrary environments connecting into a single control plane. We rebuilt the networking layer to handle hardware registration from any environment, normalise the connectivity model across mixed infrastructure, and maintain stable cluster networking as customers add or remove nodes from different physical or virtual locations.

GPU operator and compute management. We integrated the NVIDIA GPU Operator with custom resource allocators that give the platform real control over GPU scheduling, allocation, and monitoring across customer clusters.

Service mesh. We designed and implemented the service mesh layer for inter-cluster communication, traffic management, and observability, resolving the connectivity issues that had made the platform unpredictable.

Marketplace and plugin framework. The platform ships a marketplace of plugins that customers deploy directly into their clusters from within the platform. This includes AI Architect for AI workflow orchestration, Kubeflow for ML pipelines, Laravel stack integrations, and 100+ other open source Helm charts. We rebuilt the framework that governs how plugins are packaged, versioned, deployed, and updated across customer clusters, so every chart in the marketplace installs reliably regardless of what hardware the cluster is running on.

Customer onboarding infrastructure. Node registration and NACL creation for new customers were manual and error-prone. We automated the full onboarding flow so new customer environments provision without manual intervention.

ComponentWhat we rebuiltTechnology
State managementUnified state layer across all cloud providersCustom Kubernetes operators, etcd
Networking for BYOHStable cluster networking across hardware registered from any environmentVanilla Kubernetes networking, custom node registration layer
GPU managementOperator-level GPU provisioning and allocationNVIDIA GPU Operator, custom allocators
Service meshFast, stable inter-cluster communicationCustom service mesh implementation
Plugin marketplaceDeployment framework for AI Architect, Kubeflow, Laravel stack, 100+ chartsHelm, ArgoCD, custom chart operator
Customer onboardingAutomated node registration and NACL creationTerraform, Kubernetes admission controllers
AlertingStructured cluster and node health alertingPrometheus, Alertmanager
GitOps pipelineAutomated cluster lifecycle managementArgoCD, GitHub Actions
Pricing integrationReliable connectivity to cloud provider billing APIsCAST AI integration, custom billing adapters

The hardest parts

Redesigning the infrastructure layer without taking the product offline. The platform had paying customers during the rebuild. The platform could not simply go dark for 2 months. The approach was to build the new infrastructure layer in parallel, migrate workloads incrementally, and cut over component by component.

Networking for arbitrary hardware configurations. Because the platform registers customer-owned hardware rather than provisioning managed cloud nodes, the networking layer has to handle a much wider range of physical and virtual configurations. Customers were registering nodes from bare metal, from private clouds, from VMware environments, and from various provider setups. Getting the control plane to maintain stable connectivity across all of these took significantly longer than a more constrained networking model would have.

State recovery for existing clusters. When we introduced the new state management layer, existing customer clusters had no state history. Building a reconciliation process that reconstructed accurate state for live clusters without disrupting them was the most technically delicate work of the rebuild. A single error would have made existing deployments unmanageable.

Dead code and architectural debt. The AI-assisted rebuild process surfaced a significant amount of duplicate and dead code. Removing it required understanding which code was genuinely unused versus which was reached through uncommon paths not obvious from static analysis. This took longer than a clean codebase would have, but it was the right call.

Results

MetricBeforeAfter
Platform stabilityContinuously failingStable. No recurring systemic failures since rebuild.
State managementNo consistent stateReal-time state tracking across all clusters and nodes
Customer onboardingManual intervention requiredFully automated node registration and environment setup
ML setup timeWeeks of manual GPU cluster configurationHours with automated GPU provisioning
Release velocityBlocked by instabilityRegular feature releases on structured sprint cadence
Chart deploymentInconsistent, manual troubleshootingReliable across all 100+ open source charts
Team model1 embedded engineer5 to 6 engineers, PM, roadmap ownership
Product positioningPre-funding, unstable productKubeCon presence, CAST AI partnership

"I've been working with Ehtisham personally since 2020, and when he brought the Eprecisio team in to rebuild and stabilize our platform, the quality of work matched everything I'd come to expect. The team hit the ground running, shipping features and cleaning up the platform without constant oversight. What I appreciate most is that the trust built over years of working together translated directly into how the team operates. Met Ehtisham a few times in person in London and the professionalism is consistent across the board. If you're a technical founder who needs a team that takes real ownership, Eprecisio delivers."

Founder, Kubernetes automation platform (UK)

Where the product is now

The product is no longer struggling to be stable. It is built on a clear and defensible position in the market: organisations that need production-grade AI compute across multiple sovereign jurisdictions, without surrendering control to any single hyperscaler. That means your jurisdiction, your provider mix, your rules on where workloads run and where data sits. The founder is now taking the product to KubeCon, presenting at Kubernetes automation working groups, building partnerships with infrastructure players like CAST AI around it, and holds active membership in both the Cloud Native Computing Foundation and the Linux Foundation.

The platform has attracted institutional backing from SFC Capital and Rule30.vc, and is a member of the NVIDIA Inception Program. Target sectors span government, defence, healthcare, research, finance, AI startups, and enterprise HPC teams across the UK, EU, and MENA regions. It is the same regulated-industry buyer profile that surfaces repeatedly in our EU AI Act and sovereign cloud work.

The Eprecisio team is not winding down. The engagement is actively growing. The founder has explicitly asked to scale the Pakistan-based engineering team further rather than continuing to hire in the UK, where previous direct hires did not work out.

For how we structure and manage production Kubernetes infrastructure at this scale, see our InfraOps service.

If you are building an infrastructure platform and need a team that can work at this level of technical depth and own the delivery process, book a free 30-minute call.

Keep going

Case study

On-Prem GPU Kubernetes: 5 to 100 Node Reference Architecture

As a research lab head or an AI-native founder, you need GPU compute your team can actually use in minutes, without shipping data to a public cloud. This is the exact reference architecture we deploy: 5 nodes on day one, scales to 100+, NVIDIA GPU Operator, KServe, Ray, and MIG partitioning included.

Read case study
Case study

Production Kubernetes for IoT: LoRaWAN Things Stack On-Prem

A healthcare analytics platform needed production-grade LoRaWAN infrastructure for IoT gateways and applications. Eprecisio deployed and configured a self-hosted Kubernetes-hosted Things Stack with plug-and-play gateway onboarding.

Read case study
MLOps

vLLM Autoscaling On Kubernetes: The Metric CPU-Based HPA Cannot See

On a Google Cloud G2 instance with an NVIDIA L4, the vLLM waiting queue went from 0 to 64 requests while host CPU stayed at 2.8% median and peaked at 22.2%. A CPU-based Kubernetes HPA would never have fired. The metric to alarm on is vllm:num_requests_waiting, plus a sharper by-reason variant most people miss. First-party evidence from the Eprecisio hardware baseline pilot.

Read post
InfraOps

Sovereign AI Cloud: Running GPU Workloads Under GDPR, PDPL, and the EU AI Act

Under oath before the French Senate in 2025, Microsoft France's Legal Director admitted he could not guarantee EU customer data would never be handed to US authorities. That single sentence is the sovereign-cloud thesis. Here is what actually meets EU AI Act, GDPR, and PDPL for GPU workloads in 2026, with real vendor pricing.

Read post

Want Similar Results for Your Business?

Let's discuss how Eprecisio can help you achieve your goals.

Book a Free 30-Min Call

Your infra shouldn't be the thing slowing you down.

Book a free 30-minute call. We'll look at your current setup and tell you exactly what's costing you money, what's a deployment risk, and what we'd fix first. No pitch, no fluff.

AWSAzureGCPKubernetesDockerTerraformPythonReactNext.jsArgoCDPrometheusGrafana