Reach Data Center Decision MakersAdvertise

The AI-Readiness Checklist: 10 Questions to Ask Before Signing a Colocation Lease

Quick answer: Most data centers claim AI-readiness but few deliver it. Separate real AI capability from marketing by asking: What is the maximum power density per rack (need 20kW+ minimum)? What cooling systems exist today (air-only disqualifies you)? What network fabric is available (InfiniBand required for training)? What is the actual PUE below 1.4 with 12-month data? Can they show GPU customer references? These 10 questions form a non-negotiable checklist.

Introduction: The AI-Readiness Myth

Visit the websites of 100 colocation facilities. You will find AI-readiness claims on roughly 80 of them. “Enterprise-grade power infrastructure.” “Designed for GPU workloads.” “Liquid cooling ready.”

Visit those same facilities and ask for references, current customers running GPU clusters. Most go silent.

The gap between claimed and delivered AI-readiness is enormous. Marketing departments promise; engineering teams struggle to deliver. Facilities built in 2015 for traditional compute do not have the power density, cooling, or network architecture that modern AI workloads demand.

This creates a critical risk: You sign a 3-5 year lease, move your GPU infrastructure into a facility, and discover in week three that rack temperatures exceed 45C, power fluctuates beyond your UPS tolerance, or the facility network cannot handle the all-to-all communication patterns your cluster requires.

By then, you are locked into a contract with relocation costs that dwarf your original lease terms.

This checklist separates facilities that are genuinely AI-ready from those using boilerplate marketing. It is grounded in what actual GPU operators need, not what sales teams claim.

The 10-Point AI-Readiness Checklist

1. What is the Maximum Power Density Per Rack?

Why this matters: Modern GPUs are power-hungry. An NVIDIA H100 consumes 700W under load. A single 20kW rack can hold just 28 H100s (accounting for CPUs, memory, networking, PSU overhead). Training clusters often run 40+ GPUs per rack to justify the facility footprint.

What to ask: “What is the maximum continuous power available per cabinet?” Not per row. Not per floor. Per cabinet.

Acceptable answers:

  • Minimum: 20kW/rack (inference workloads, sparse training)
  • Competitive: 30-40kW/rack (standard training workloads)
  • Exceptional: 50-100kW/rack (dense AI clusters, liquid-cooled)

Red flag: “We can support up to X kW, but you will need to work with our engineering team to deploy it.” Translation: We have never actually done it; you are the test case.

How to verify: Ask for the rack PDU specifications. Walk the floor. How many racks do they have deployed at their claimed maximum density? See them with your own eyes.

Current market reality:

  • Primary markets (Dallas, Phoenix, Silicon Valley): 30-40kW/rack available
  • Tier 2 markets (Salt Lake, Columbus): 20-30kW/rack
  • Legacy colo facilities: 10-15kW/rack (insufficient for AI)

2. What Cooling Systems Are Installed Today?

Why this matters: Air cooling maxes out around 30-35kW/rack. Above that, you need liquid cooling, either in-rack heat exchangers or hot-aisle containment with chilled water loops. A facility claiming 50kW/rack capability with only air conditioning is lying about utilization.

What to ask: “What cooling systems are deployed today? What percentage of racks have access to liquid cooling?”

Acceptable answers:

  • Full facility has hot-aisle containment plus chilled water distribution
  • Specific zones are liquid-cooled with In-Row Coolers (IRC)
  • Hybrid: Air-cooled primary, liquid-cooled secondary

Red flag: “We are air-cooled and can add liquid cooling for a premium.” This signals the facility was not designed for AI. You will be a pilot customer, with construction delays and upcharges.

How to verify:

  • Photograph the cooling infrastructure. Air ducts, chilled water lines, in-row coolers are visible.
  • Ask for PUE (Power Usage Effectiveness) trend data over 12 months. If they cannot produce it, they do not monitor it.
  • Request floor plan showing where liquid cooling is available. Point out your proposed rack locations. Confirm cooling can reach those racks.

Current state: Newer facilities (built post-2020) in competitive markets are liquid-cooled or hybrid. Legacy facilities offer liquid cooling as a retrofit or special service.

3. What Network Fabric Is Available?

Why this matters: GPU-to-GPU communication is bandwidth-intensive. Standard Ethernet (even 100Gbps) becomes a bottleneck for clusters training large models.

High-performance clusters require:

  • InfiniBand: 200Gbps+ with low-latency, lossless switching. Industry standard for training clusters.
  • Unified Ethernet Convergence (UEC): Advanced Ethernet with RDMA (Remote Direct Memory Access). Emerging alternative to InfiniBand.

Standard Ethernet forces workarounds: gradient compression, reduced batch sizes, longer training times. A 10% slowdown due to network architecture compounds to 2-3 month delays on a 12-month training project.

What to ask: “Do you offer InfiniBand fabric? If not, do you have UEC-capable switches and NICs?”

Acceptable answers:

  • “We have native 200Gbps InfiniBand across zones X, Y, Z”
  • “We offer 400Gbps Ethernet with RDMA-capable infrastructure”
  • “We have Ethernet, and we have successfully deployed GPU clusters here”

Red flag: “We support whatever networking your equipment requires.” Vague. Demand specifics.

How to verify:

  • Ask for a topology diagram showing InfiniBand/UEC segments
  • Request references from customers running GPU clusters (they will mention if networking was a bottleneck)
  • Ask about latency: “What is the typical latency within a cluster from our location?” Should be under 5 microseconds for InfiniBand.

Current reality:

  • Top-tier facilities (Dallas, Phoenix): 200Gbps+ InfiniBand standard
  • Tier 2 facilities: InfiniBand in specific zones; traditional Ethernet in others
  • Budget facilities: Ethernet only (not suitable for training clusters)

4. What Is the Facility Actual PUE? (Not Theoretical)

Why this matters: PUE (Power Usage Effectiveness) equals Total Facility Power divided by IT Equipment Power. A 1.2 PUE means the facility uses 0.2 watts of cooling and overhead for every 1 watt of IT equipment. Lower is better.

Facilities publish “design PUE” of 1.15-1.20. Actual trailing PUE is often 1.3-1.5 because design assumptions rarely hold (cooling is not perfectly distributed, ambient temps vary, equipment mix changes).

For GPU workloads, focus on your actual power costs: If you rent 1MW of capacity, is your bill for 1MW (PUE 1.0, unrealistic) or 1.3MW (PUE 1.3, realistic)?

What to ask: “What is your 12-month trailing PUE, published monthly? Can you provide data to a third party (auditor, me) for verification?”

Acceptable answers:

  • “Our trailing PUE is 1.25, audited by [third party]”
  • “We publish monthly PUE on our website”
  • “Our 12-month average is 1.32”

Red flag:

  • “Our design PUE is 1.18…” (not actual)
  • Refusal to share 12-month data
  • Vague answers about how they calculate PUE

How to verify:

  • Request audited PUE statements from recent years
  • Cross-reference with industry databases (Uptime Institute, TechUK data center rankings)
  • Ask about seasonal variance: “How does PUE change in summer vs. winter?” Dramatic swings signal cooling inefficiency.

Current market:

  • Efficient facilities (new builds, optimized operations): 1.20-1.30 PUE
  • Average facilities: 1.30-1.50 PUE
  • Inefficient facilities: 1.50+ PUE (watch out)

5. Can the Facility Demonstrate Current GPU Deployments?

Why this matters: Claiming AI-readiness and actually running AI workloads are fundamentally different. A facility that currently operates GPU clusters has:

  • Proved their power, cooling, and networking work at scale
  • Installed monitoring and support processes
  • Solved operational edge cases (thermal throttling, power spikes, network congestion)

A facility with zero GPU customers is experimental, you will be debugging infrastructure with your lease fees.

What to ask: “Which customers currently run GPU workloads at your facility? Can they be references?”

Acceptable answers:

  • “We have 3 customers running 200+ GPUs in our Phoenix facility. They are open to a reference call.”
  • “We deployed 50 H100s for a research lab in our Dallas facility in January; they are happy to discuss.”

Red flag:

  • “No current GPU customers, but we are designed for them”
  • Refusal to provide references
  • “Our other customers are under NDA” (possible, but often a deflection)

How to verify: Call the references. Ask:

  • How long have you been here?
  • Have you experienced thermal issues, power spikes, network congestion?
  • How responsive is the facility to GPU-specific requests (rapid equipment swaps, monitoring)?
  • Would you re-lease here?

What this reveals: Current GPU customers will tell you (honestly) about real operational challenges and how the facility handles them.

6. What Power Redundancy and Conditioning Exist?

Why this matters: Power disruptions kill GPU workloads. An unexpected outage during a 30-day training run costs weeks of wasted compute and data. But brief voltage fluctuations are just as damaging, they cause hardware failures and data corruption without triggering a full failover.

What to ask: “What is your power redundancy specification? What voltage tolerance does your infrastructure maintain?”

Acceptable answers:

  • “We have N+1 power redundancy with automated failover and plus or minus 3% voltage regulation”
  • “We maintain 2N power infrastructure with UPS to cover all equipment plus 15 minutes bridge to generators”
  • “Dual feeds from independent utility substations”

Red flag:

  • “We have backup generators” (not enough; you need conditioning)
  • “N redundancy is available at a premium” (it should be standard)
  • No specification for voltage tolerance

How to verify:

  • Request power SLAs in the contract. Should include voltage tolerance range (typically plus or minus 5% at worst, plus or minus 3% standard), availability commitment (99.99% uptime minimum), brownout procedures.
  • Ask about recent power incidents: “What happened in the last 24 months? How did it affect customers?”

Current standard:

  • N+1 or 2N redundancy
  • UPS capacity for 10-30 minutes bridge to generators
  • Plus or minus 3% to plus or minus 5% voltage regulation

7. What Is the Facility Expansion Runway?

Why this matters: You commit to a 3-5 year lease. Facility capacity constraints appear around year 2-3 (you have been successful, utilization grows). If the facility cannot expand, you are stuck: your options are renegotiate at inflated rates, relocate (expensive), or cap your growth.

What to ask: “How much additional capacity can you provision in the next 18, 36, and 60 months? What is the timeline?”

Acceptable answers:

  • “We have 50MW headroom on adjacent land. We are currently permitting 10MW for 2026 delivery, 15MW for 2027.”
  • “Our current utilization is 60%; we can absorb 40% growth without building new infrastructure.”

Red flag:

  • “We are at capacity and have no expansion plans”
  • “Expansion would require new buildings, timeline TBD”
  • Vague answers about future capacity

How to verify:

  • Check local permitting records for expansion plans
  • Ask about utility capacity (how much power/cooling can the local grid/utility support?)
  • Request a timeline: When does the facility hit 90% utilization? What happens then?

Current reality:

  • Expanding facilities in Tier 2 markets (Dallas, Phoenix, Columbus): 20-50MW planned over 2-5 years
  • Mature facilities in primary markets: May be capacity-constrained; ask about partner facilities
  • New builds: Often undersized within 3 years due to AI demand surge

8. What Security and Isolation Capabilities Exist?

Why this matters: Your AI models represent competitive IP. In a shared colocation facility, your code, models, and data live alongside competitors.

For regulated industries (healthcare, finance, government), isolation and compliance are non-negotiable.

What to ask: “Do you offer air-gapped zones? What audit trails exist for access to our infrastructure? What certifications do you hold?”

Acceptable answers:

  • “We offer dedicated cages with locked access, 24/7 monitoring, and audit logs of all facility access”
  • “We have SOC 2 Type II certification and HIPAA-compliant zones”
  • “We provide air-gapped racks for sensitive workloads”

Red flag:

  • “Your equipment is in a locked cage; that is the standard” (true for colo, but not sufficient for high-security workloads)
  • No mention of compliance certifications
  • Vague security descriptions

How to verify:

  • Request SOC 2 Type II report (should be available under NDA)
  • Ask about audit procedures: “Walk me through what happens if I request an audit of who accessed my equipment.”
  • Review cage access logs for a representative week

Current standard:

  • SOC 2 Type II certification (minimum)
  • Physical cages with biometric or badge access
  • 24/7 CCTV monitoring
  • Access audit logs (who, when, what was accessed)

9. What Certifications Does the Facility Hold?

Why this matters: Certifications are third-party proof of operational standards. They are boring but critical.

  • SOC 2 Type II: Verifies security, availability, and confidentiality controls
  • ISO 27001: Information security management
  • HIPAA/HITECH: Required for healthcare data (US)
  • PCI-DSS: Required for payment card data
  • FedRAMP: Required for US government contracts
  • Uptime Institute Tier: Availability/redundancy (Tier III minimum, Tier IV ideal)

What to ask: “What certifications do you hold? Can you provide audit reports?”

Acceptable answers:

  • “We are SOC 2 Type II certified, Uptime Institute Tier IV, and HIPAA-compliant”
  • “SOC 2 Type II audit is available under NDA”

Red flag:

  • No certifications listed
  • Tier I or II Uptime rating (suggests older facility)
  • Certifications that expired more than 1 year ago

How to verify:

  • Request current audit reports
  • Cross-reference on Uptime Institute database (public tier ratings)
  • Call the certification bodies if needed (rare, but possible)

Regulatory note: Different industries need different certs:

  • Healthcare: HIPAA plus SOC 2 Type II minimum
  • Finance: PCI-DSS plus SOC 2 Type II
  • Government: FedRAMP plus SOC 2 Type II

10. What Is the Grid Connection Timeline for New Capacity?

Why this matters: Power infrastructure is the bottleneck. A facility can promise 100MW of capacity, but if the local utility cannot deliver it for 3 years, you are stuck waiting.

Utility infrastructure upgrades take 18-36 months. Planning delays can extend this.

What to ask: “You mentioned 10MW of new capacity by 2026. What is the current status with the utility? Any delays anticipated?”

Acceptable answers:

  • “We have utility approval for 10MW delivery in Q2 2026. We have already ordered equipment.”
  • “We are in permitting for 15MW with expected approval by Q4 2025; utility delivery is scheduled for Q3 2026.”

Red flag:

  • “Timeline TBD”
  • “Waiting on utility approvals” (red flag if vague; yellow flag if they cannot show permitting progress)
  • Previously announced timelines that slipped

How to verify:

  • Ask for utility correspondence (formal capacity reservation letters)
  • Request permitting status from the facility
  • Check with the local utility directly (some publish expansion plans publicly)

Current reality:

  • Mature markets (Dallas, Phoenix): 12-24 month timelines for incremental capacity
  • Growing markets: 18-36 months
  • Constrained utilities: 36+ months (watch out)

Red Flags to Watch For

Claimed Density Without Matching Cooling

Facility says: “50kW/rack available” but has only air-cooled infrastructure. Not possible sustainably. Walk away or demand liquid cooling upgrade in writing before you lease.

“Designed For” Vs. “Deployed At”

Language matters. “Designed for GPU workloads” means nothing. “Currently operating GPU clusters” means they have solved it. Demand proof of deployments.

No Current GPU Customer References

If they claim AI-readiness but have zero GPU customers running production workloads, you are a beta customer with an inflated lease price. Negotiate hard or find another facility.

Vague Power Availability Timelines

“We will support your growth” sounds good. “We have utility approval for 5MW capacity by Q4 2025, with your rack slated for completion by Q2 2026” is specific and verifiable. Demand specifics and get them in writing.

Certification Requirements by Industry

IndustryRequired CertificationsNotes
HealthcareHIPAA, SOC 2 Type II, HITRUST optionalData residency often required (US only)
Financial ServicesPCI-DSS, SOC 2 Type II, ISO 27001Payment card data requires PCI; other data may need additional compliance
Government/DefenseFedRAMP, CMMC (if contractor), SOC 2 Type IIDoD contracts require CMMC Level 3+; intelligence work requires higher
Public CompanySOC 2 Type II, ISO 27001Audit requirements vary by industry; tech companies less stringent
Non-regulatedSOC 2 Type II minimumCovers security/availability; recommended even for non-regulated workloads

How GoDataCenters Maps to This Checklist

GoDataCenters AI-Readiness Score evaluates facilities across the exact dimensions this checklist covers:

  • Power Density (0-10 scale): Verified maximum rack density based on facility specs and current deployments
  • Cooling Capacity (weighted by type): Air-only equals lower score; liquid-cooled equals higher score
  • Network Fabric (presence and deployment): InfiniBand/UEC availability boosts scores; Ethernet-only penalizes
  • Power Conditioning: Redundancy specs and voltage tolerance verified
  • Security and Isolation: Certifications, audit trails, air-gapped zones
  • Expansion Capacity: Verified utility headroom and timeline

When you use GoDataCenters.com to search for facilities, you can filter by minimum AI-Readiness Score. Instead of calling 20 facilities and going through this checklist manually, you get pre-vetted options.

Frequently Asked Questions

Q: Can I ask a broker to do this vetting for me?

A: Some brokers help, but incentives matter. Brokers earn commission on lease value, not facility suitability. They are motivated to close deals, not rigorously evaluate your technical requirements. Do this vetting yourself or use GoDataCenters AI-Readiness Score as a shortcut.

Q: What if a facility will not answer these questions?

A: It signals they are either hiding something or do not take technical customers seriously. Find another facility. The best facilities welcome technical questions, it attracts serious customers and builds trust.

Q: Are there facilities that pass all 10 items?

A: Yes, but they are concentrated in primary markets (Dallas, Phoenix, Silicon Valley, Northern Virginia). Tier 2 markets have fewer fully equipped facilities. As a general rule, expect 7-8 of 10 in competitive markets, 5-6 in secondary markets.

Q: What if a facility is strong on power/cooling but weak on networking?

A: Networking can sometimes be upgraded or worked around. But it depends on your workload. Training clusters: Network bottlenecks are dealbreakers. Non-negotiable. Inference deployments: Standard Ethernet is often acceptable. Lower priority.

Q: How often should I re-validate a facility claims?

A: At minimum annually. PUE, cooling efficiency, and available capacity change. Request updated 12-month PUE data, capacity reports, and security audit refreshes during lease renewals.

Q: Should I insist on all 10 items before leasing?

A: Not necessarily. Weight by importance. Critical (must-haves): #1 (power density), #2 (cooling), #3 (networking), #6 (redundancy). Important (strong-haves): #4 (PUE), #5 (GPU references), #7 (expansion). Nice-to-haves: #8-10 (security/certifications, depending on industry). Regulated industries (healthcare, finance, government) elevate #8-10 to critical.

Conclusion

Finding an AI-ready colocation facility requires more than marketing review. It requires technical diligence, asking hard questions and demanding proof.

This 10-point checklist separates genuine AI-ready facilities from those using boilerplate promises. Current GPU deployments, verified power density, liquid cooling, InfiniBand fabric, actual PUE data, security certifications, these are non-marketing indicators of real capability.

Use this checklist before signing any colocation contract. If a facility cannot answer these questions clearly and with documentation, it is not ready for your AI workloads.

Find AI-Ready Facilities on GoDataCenters.com

Sourcing capacity?

One requirement, matched against 4,562 facilities and 1,010 providers worldwide.

Get a quote