Reach Data Center Decision MakersAdvertise

GPU Colocation: Where to Find High-Density Capacity for AI Compute

Quick Answer: GPU operators find colocation capacity in specialized high-density facilities concentrated in Dallas, Phoenix, Salt Lake City, Columbus, and Hillsboro. Pricing ranges $3,000-$6,000/mo per 20kW rack in primary markets, $2,000-$4,500 in Tier 2 markets. GPU-ready facilities offer 20-100kW per rack, liquid cooling, and InfiniBand networking. Not all colocation facilities can support GPUs. Most are designed for traditional compute with air cooling and Ethernet only.

Introduction: The GPU Infrastructure Challenge

The AI boom created an acute infrastructure crisis: billions in capital deployed toward model training and inference, but nowhere suitable to run it.

Cloud hyperscalers reached capacity limits in 2023-2024. GPU waiting lists stretched to months. Enterprise GPU operators (neocloud platforms, AI labs, ML-native companies) faced a critical decision: wait for cloud capacity or find alternative infrastructure.

Colocation emerged as the escape hatch. But not just any colocation facility. Traditional data centers built for web services (CPU-heavy, modest power, Ethernet networking) are physically unsuitable for GPU workloads. Moving 100 GPUs into a legacy air-cooled facility is a recipe for thermal failure.

This guide walks GPU operators (whether neocloud providers building platforms, enterprises scaling internal AI, or research labs) through finding and evaluating GPU-ready colocation capacity.

Understanding GPU Rack Requirements

Before searching for facilities, understand what GPUs actually demand from infrastructure.

Power Density: The Core Constraint

Modern GPU workloads require extraordinary power density:

NVIDIA H100: 700W under load (sustained training) + 150W for CPU/memory + 50W for networking = ~900W per GPU

  • 28 H100s per 20kW rack = typical baseline
  • 50+ H100s per 50kW rack = dense training clusters
  • 100+ H100s per 100kW rack = maximum density (rare, requires advanced cooling)

H200 (recent launch): 900W per GPU, similar constraints

NVIDIA L40S (inference): 350W per GPU, lower density requirements (50-60 GPUs per 20kW rack)

AMD MI300X: 750W per GPU, similar to H100

Compare to traditional server workloads: A 2-socket CPU server draws 500-800W. A 20kW rack holds 25-40 such servers. Adding GPUs cuts this to 15-20 servers per rack.

Critical implication: Most colocation facilities quoting “20kW available” expect you to deploy traditional compute. They have no cooling infrastructure to handle sustained GPU density. You need a facility that’s planned power distribution around GPU requirements.

Cooling: The Limiting Factor

Air cooling maxes out around 30-35kW per rack in typical row configurations. Above that, heat removal becomes the bottleneck.

Advanced air cooling (hot-aisle containment): Can push to 40-45kW/rack with precision cooling systems and in-row coolers

Liquid cooling (direct-to-chip or in-rack heat exchangers): Required for 50kW+ densities. Enables 100+ kW/rack in extreme cases.

Thermal consequences of undersized cooling:

  • Thermal throttling: GPUs reduce clock speed when temps exceed 80-85°C, slashing throughput by 20-40%
  • Hardware failure: Sustained temps above 85°C shorten GPU lifespan dramatically
  • Power instability: High load + poor cooling triggers voltage droop, causing hardware faults

A facility with 50kW/rack power capacity but only air cooling is a booby trap. You’ll deploy, hit thermal limits within hours, and be stuck renegotiating or relocating.

Networking: Cluster Communication

GPU-to-GPU communication during training is bandwidth-intensive:

  • Gradient synchronization: All-reduce operations across 100s of GPUs, terabytes of data per training step
  • Model distribution: Broadcasting model weights to compute nodes
  • Data loading: Distributed data pipelines shuffling terabytes across cluster

InfiniBand (industry standard): 200Gbps, 400Gbps. Ultra-low latency (<1 microsecond). Lossless (no dropped packets). Designed for HPC clusters.

Unified Ethernet Convergence (emerging): 400Gbps+ Ethernet with RDMA and advanced switching. Competitive with InfiniBand for newer deployments.

Traditional Ethernet (insufficient): 100Gbps max practical throughput after overhead. 5-10 microsecond latency. Packet loss requires retransmission, adding overhead.

For a 100-GPU cluster training a large language model, Ethernet bottlenecks can add 10-20% training time. Over a 60-day training run, that’s weeks of wasted compute.

Practical implication: If the facility doesn’t offer InfiniBand or UEC, evaluate whether your workload can tolerate Ethernet. Most serious training workloads cannot.

Physical Requirements

Beyond power, cooling, and networking:

Floor loading: GPUs and supporting infrastructure are dense. Verify the floor can handle 1.5+ tons per rack.

Rack footprint: Standard 19″ racks are the norm. Confirm spacing for cable management, liquid cooling loops, PSU overprovisioning.

Airflow management: Servers generate heat fronts and back. Facilities need hot-aisle containment or in-row cooling to manage it.

Training vs. Inference: Different Infrastructure Needs

GPU colocation isn’t one-size-fits-all. Training and inference have opposing infrastructure requirements.

Training Clusters

Characteristics:

  • High utilization (80-100% continuously)
  • Sustained power draw
  • Large clusters (50-1000 GPUs typical)
  • All-to-all network traffic (every GPU talks to every other GPU during synchronization)
  • 24/7 operation with rare scaling

Infrastructure requirements:

  • Maximum power density (50+ kW/rack)
  • Liquid cooling mandatory
  • InfiniBand or UEC networking
  • Massive east-west bandwidth (intra-cluster communication dominates)
  • Consistent power, stable cooling (any fluctuation disrupts training)

Facility type: Specialized neocloud-focused colocation or purpose-built training centers

Typical pricing: $4,000-$6,000/month per 20kW rack (premium for dedicated bandwidth and cooling)

Inference Deployments

Characteristics:

  • Variable utilization (20-80% depending on request volume)
  • Bursty network traffic (north-south: external requests to inference servers)
  • Distributed deployment (infer across 10s of edge locations, not 1 giant cluster)
  • Scaling: Add/remove racks based on traffic
  • Mixed hardware (H100, L40S, L4, depending on latency and cost requirements)

Infrastructure requirements:

  • Moderate power density (20-30 kW/rack typical)
  • Air cooling acceptable (if dense, add in-row coolers)
  • Standard Ethernet sufficient (external latency >> intra-facility latency)
  • North-south bandwidth critical (connections to users, not GPU-to-GPU)
  • Flexibility to scale up/down

Facility type: Traditional colo or Tier 2 neocloud (still fine for inference)

Typical pricing: $2,000-$4,000/month per 20kW rack (lower cost, standard colo features)

Decision framework:

  • Training? Find a specialized facility. Premium cost, but essential.
  • Inference only? Traditional high-density colo works. More options, lower cost.
  • Mixed? Look for facilities with hybrid capability: training zones (liquid-cooled, InfiniBand) and inference zones (air-cooled, Ethernet).

What Neocloud Operators Need From Colo Facilities

If you’re building a neocloud platform (renting GPU capacity to customers), infrastructure needs are even more specialized:

Powered Shell vs. Turnkey

Powered shell: You bring servers, NICs, switches, storage. Facility provides power, cooling, space, connectivity.

  • Pros: Full control, maximum flexibility, cost-effective at scale
  • Cons: You manage hardware procurement, lead times, installation
  • Typical timeline: 8-12 weeks from lease to full deployment

Turnkey: Facility provides pre-configured infrastructure, you add customers.

  • Pros: Fast to revenue (weeks), no hardware procurement headaches
  • Cons: Less flexibility, higher cost, lock-in to facility’s hardware choices
  • Typical timeline: 2-4 weeks to first customer

Most neocloud operators choose powered shell for scalability, but many use hybrid (turnkey for initial capacity, powered shell for scale-out).

Rapid Deployment

Neocloud growth is fast. You need:

  • Fast power availability (days, not months)
  • Quick cross-connect setup (connect to cloud providers or customer networks)
  • Flexible contract terms (3-month minimums, not 3-year locks)
  • Support for multi-tenant isolation (each customer’s racks isolated from others)

Not all facilities offer this. Traditional colocation expects multi-year commitments. Neocloud facilities (or colo operators serving neocloud) accept shorter terms and faster deployments.

Flexible Scaling

Demand forecasting is hard. You need:

  • Adjacent capacity available (when you need to add 50 racks, the facility has space)
  • No long procurement delays for additional power/cooling
  • Pricing that scales (not discounts only on long-term bulk deals)

Multi-Tenant Isolation

Your customers don’t want their models visible or accessible to other customers. You need:

  • Separate network VLANs or fabric isolation
  • Physical separation of customer racks (not just logical)
  • Access controls (your customers can access only their own racks)
  • Audit trails (who accessed what, when)

Most modern facilities support this; it’s table stakes for cloud-like services.

GPU Colocation Pricing

Pricing varies dramatically by market and facility quality. Here’s the realistic landscape:

Rack Pricing (20kW, Monthly)

Market TierPrimary Market (Dallas, Phoenix, Silicon Valley)Tier 2 (Columbus, Salt Lake, Hillsboro)Tier 3 (Secondary cities)
Standard colo$3,000-$4,000$2,000-$2,500$1,200-$1,800
GPU-optimized$4,000-$6,000$2,500-$4,000$1,800-$2,800

Note: GPU-optimized = liquid cooling, InfiniBand, higher power/rack available

Power (Overage Beyond Included)

Most leases include 2-4 kW power in the rack price. Beyond that:

  • Primary markets: $0.12-$0.18 per kWh
  • Tier 2 markets: $0.10-$0.14 per kWh

A 50kW GPU cluster running 24/7:

  • 50 kW × 24 hr × 30 days × $0.14 = $5,040/month in power alone
  • Plus $4,500 rack lease = $9,500/month per rack

Budget accordingly.

Cross-Connects (to Cloud Providers, Your Network)

Connecting your colo infrastructure to AWS, GCP, or your on-prem network:

  • 10Gbps cross-connect: $200-$500/month
  • 100Gbps cross-connect: $2,000-$5,000/month

Most GPU operators need at least one cross-connect for data/model transfer.

Setup Fees

One-time installation and configuration:

  • Powered shell: $1,000-$5,000 (depends on complexity)
  • Turnkey: Often included

NRE (Non-Recurring Engineering)

If the facility needs to upgrade cooling or power to support your deployment:

  • Modest upgrades: $5,000-$20,000
  • Major infrastructure work: $50,000+

Try to negotiate this into the lease or push back. Some facilities (especially those pursuing neocloud business) will absorb NRE to land a big customer.

Typical 100-GPU Deployment Cost (Monthly)

5 racks × 20 GPUs per rack:

  • Rack lease: 5 × $4,500 = $22,500
  • Power overage: 5 × $5,000 = $25,000
  • Cross-connects: 2 × $3,000 = $6,000
  • Total: $53,500/month

Over 36 months: $1.9M facility costs + hardware + operations = $3-4M all-in (discussed in the colocation vs. cloud blog).

Top Markets for GPU Colocation

Not all markets have GPU-ready infrastructure. Here are the leading markets (as of Q2 2026):

Dallas

Why: Central US location, abundant power (deregulated Texas grid), multiple facilities competing for GPU business

Facilities: DFT Dallas, QTS, Equinix DA (mixed), CyrusOne, others

Specs: 30-50 kW/rack typical, InfiniBand available in multiple locations, competitive pricing

Pros: Mature GPU ecosystem, fast deployments, price competition drives rates down

Cons: Limited liquid cooling options (in-row coolers more common than chilled water loops)

Phoenix

Why: Abundant renewable power (solar), newer facilities, aggressive capacity build-out

Facilities: Digital Bridge PHX, others in development

Specs: 40-100 kW/rack possible, liquid cooling increasingly available

Pros: Newest infrastructure, best power efficiency (solar baseload), rapid growth in GPU capacity

Cons: Some facilities still ramping up; not all fully operational yet

Salt Lake City

Why: Cheap power (hydroelectric), growing AI hub

Facilities: Digital Realty, XO (formerly Xyo), others

Specs: 30-40 kW/rack, InfiniBand in some locations

Pros: Lowest power costs, growing ecosystem

Cons: Smaller market, fewer facilities than Dallas/Phoenix, less competition

Columbus, Ohio

Why: Midwest location, reasonable power costs, major tech companies headquartered nearby

Facilities: QTS, DuPont Fabros (now Digital Realty)

Specs: 30-40 kW/rack

Pros: Central US location for distributed deployments, underrated market

Cons: Fewer pure-play GPU specialists than primary markets

Hillsboro, Oregon

Why: Proximity to Intel HQ, existing semiconductor/tech ecosystem, available capacity

Facilities: Multiple operators in development

Specs: Varies, newer builds targeting high density

Pros: Emerging market with growth potential

Cons: Smaller scale, fewer mature facilities

Also Worth Watching

  • Northern Virginia (Ashburn): AWS hub, extensive capacity, but expensive
  • Silicon Valley: Most expensive, but good for enterprises needing proximity to HQ or cloud provider offices
  • Singapore, Tokyo: If Asia-Pacific deployment needed, but emerging markets for GPU colo

How to Evaluate a GPU-Ready Facility

Use the AI-Readiness Checklist (covered in another blog post) to evaluate any facility. Key questions for GPU colocation specifically:

  1. Power density verified: Can they show 20kW+ per rack with actual deployments?
  2. Cooling deployed: Not “planned for,” but actually installed and operational?
  3. InfiniBand or UEC: Necessary for training clusters?
  4. GPU references: Current customers running GPU workloads?
  5. Expansion: Can they grow with you, or will you hit capacity walls?
  6. Pricing transparency: No hidden NRE or power overages?

If they fumble any of these, keep searching.

Frequently Asked Questions

Q: How much does GPU colocation cost compared to cloud?

A: Monthly: $2,000-$6,000 per 20kW rack depending on market and facilities. All-in over 3 years (hardware + facility + ops): $4-6M for a 100-GPU deployment vs. $13M+ on cloud. See the colocation vs. cloud blog for detailed breakeven analysis.

Q: Can I bring my own GPUs, or do I have to lease them from the facility?

A: Most GPU colocation operates on a bring-your-own-hardware model. You buy or lease GPUs separately, bring them to the facility, and pay for power/cooling/space. Some facilities offer turnkey (including hardware), but that’s less common and more expensive. GoDataCenters can help you evaluate both options.

Q: What power do I need for my workload?

A: Budget ~30W per GPU core. A single H100 (equivalent to 2-3 CPU nodes) needs 700W. For a cluster:

  • 10 H100s: ~7-8 kW
  • 50 H100s: ~35-40 kW
  • 100 H100s: ~70-80 kW

Add overhead for networking, storage, power distribution (~10%), and you get final rack densities. Most facilities quote conservatively (e.g., “We support 20 H100s per 20kW rack”).

Q: What’s the difference between training vs. inference facilities?

A: Training requires sustained high density, liquid cooling, InfiniBand, and stable power. Inference can use lower density, air cooling, Ethernet. Training facilities are specialized and pricier; inference can use traditional colo. See the detailed section above on this.

Q: How do I find a GPU-ready facility near my data or customers?

A: GoDataCenters.com filters facilities by market, AI-Readiness Score, and power specs. Start there. Then call references from current GPU customers at each facility you’re considering.

Q: Can I renegotiate my lease if my needs change?

A: Yes, but it’s easier with facilities pursuing neocloud business. Traditional colo operators are less flexible. Write provisions into your lease: upsize/downsize allowances, scaling clauses, early termination at a penalty. This is especially important if you’re uncertain about 3-year demand.

Q: How long does it take to go from lease signature to deployment?

A: Powered shell (you bring hardware): 6-12 weeks. Turnkey (facility provides everything): 2-4 weeks. This assumes you have equipment ready; procurement delays extend timelines significantly. Budget conservatively.

Conclusion

Finding GPU colocation capacity requires more than traditional colo search criteria. You need power density, advanced cooling, specialized networking, and proven GPU deployment experience.

The best markets today are Dallas, Phoenix, Salt Lake City, and Columbus. Pricing ranges $2,000-$6,000/month per 20kW rack depending on market and facility class. For serious training workloads, you need InfiniBand, liquid cooling, and references from current GPU customers.

Use the AI-Readiness Checklist to evaluate any facility. If they can’t demonstrate GPU deployments, power density, and cooling, they’re not ready for your workload.

Find GPU-Ready Colocation on GoDataCenters.com →

Sourcing capacity?

One requirement, matched against 4,562 facilities and 1,010 providers worldwide.

Get a quote