Capacity available now
We operate our own sites and partner with a trusted network. When you need GPUs, you get them — no waitlist, no lottery, no enterprise sales cycle.
AI Infrastructure
Infrastructure that adapts to your workload. B200, H200 and H100 — billed by the hour, or locked in when you're ready to commit.
NO HIDDEN FEES · NO SURPRISE EGRESS CHARGES · SCALE UP OR DOWN ON DEMAND
Workloads
Training, inference, agents, scientific simulation, geospatial analytics and more — if it needs GPUs, it runs on Kala.
Why Kala
We built the infrastructure ourselves — from the power source to the API. That means lower costs, more capacity, and no surprises on your bill.
We operate our own sites and partner with a trusted network. When you need GPUs, you get them — no waitlist, no lottery, no enterprise sales cycle.
Egress is free. The price you see on the pricing page is the price you pay — we don't recoup margin through data transfer fees.
Our sites run on stranded and curtailed energy that traditional data centres can't access. Lower power costs flow through to competitive GPU rates without sacrificing hardware quality.
We stock the current generation. No waiting for hardware allocation cycles or being pushed to older SKUs — the fleet you see in the pricing table is what's available.
Pricing
Transparent hourly rates for powerful GPUs. Scale your workloads without hidden costs or surprise charges.
Built for frontier-scale training, large-scale AI modelling and simulation.
| Memory | 96 GB HBM3e |
| Bandwidth | 8.0 TB/s |
| Tensor cores | 528 |
| Released | 2024 |
Optimised for high-throughput AI training, inference and data analytics.
| Memory | 141 GB HBM3e |
| Bandwidth | 4.8 TB/s |
| Tensor cores | 528 |
| Released | 2024 |
Reliable power for development, fine-tuning and scalable cloud workloads.
| Memory | 80 GB HBM2e |
| Bandwidth | 3.35 TB/s |
| Tensor cores | 456 |
| Released | 2023 |
NEED DEDICATED CLUSTERS, RESERVED CAPACITY OR CUSTOM CONFIGURATIONS? TALK TO OUR TEAM
Pricing models
No single model fits every team. Use on-demand for full flexibility, or commit to reserved capacity for a lower effective rate on sustained workloads.
Billed hourly with no commitment. Spin up GPUs in minutes and release them the moment your job finishes. Full published rates — exactly what you see in the pricing table above.
Lock in capacity for a set term and receive a lower effective hourly rate in exchange. Best for teams running predictable, sustained workloads — long training runs, batch pipelines, or production inference.
Isolated hardware reserved exclusively for your organisation. Uncontended performance for security-sensitive or large-scale production workloads, with full control of the stack.
Need a specific GPU, dedicated hardware, or a configuration we don't list? We'll procure and install exactly what your workload requires — talk to us about a custom deployment.
Need something we don't list?
If you have bespoke requirements — specific GPUs, dedicated hardware, or a particular configuration — we're happy to procure and install exactly what you need.
Kala Mesh distributes workloads across our energy-site and city-based infrastructure — training runs where power is cheapest, inference close to your users, with geo-aware placement. And if you need a GPU configuration we don’t have free, Kala Mesh sources it from our trusted partner network, so you’re never blocked on hardware. One platform, one API, one bill.
Spin up GPUs in minutes for development, testing and short-term jobs. Pay only for what you use.
Dedicated GPU resources with guaranteed availability for long-running training and production inference.
Expand from single nodes to multi-GPU clusters as demand grows, without re-architecting your systems.
Absorb traffic spikes and crunch periods with additional compute, available the moment you need it.
Platform
Standard tools and standard access methods — the interfaces your team already uses, without proprietary lock-in.
Provision and manage compute through a web console, a REST API, or a command-line interface.
Full CUDA access means existing training code runs without modification.
Containerised workloads via Docker with NVIDIA Container Toolkit, and cluster orchestration via Kubernetes.
Provision and manage GPU resources declaratively alongside the rest of your infrastructure.
Full control of the machine — install any driver, kernel module, or orchestration layer your stack requires.
Low-latency interconnect between nodes for distributed training at scale.
Secure, isolated compute environments with private networking between nodes.
Live workload and utilisation visibility through the Kala Mesh platform.
Getting started
Sign up on the Kala compute platform. It takes a couple of minutes.
Pick your GPU and configuration: on-demand instances, dedicated nodes or a cluster.
Run your workloads. Scale up, scale down or burst as needed, with full visibility into usage and spend.