The board approved the AI initiative. Legal said the training data cannot leave the enterprise. The data center was engineered for 8-kilowatt racks — an AI cluster pulls 120 kilowatts. That gap, between the mandate and the room it has to run in, is where most enterprise on-premises AI infrastructure deployments stall.
This is a build plan for closing that gap. It covers five essential layers, in the order you have to solve them: power, cooling, network, security, and operations. Skip the order and you will discover the power problem after the GPUs arrive, stalling the project. Maybe for good.
Why build AI on-premises at all?
Enterprises build AI on-premises ("on-prem") when the data cannot go to the cloud, the economics stop working at scale, or both. Regulated data in financial services, healthcare, and defense often cannot leave a controlled environment. Protected cloud enclaves like FedRAMP do exist for these workloads, but limited accelerator availability, capacity ceilings, and cost premiums often make them impractical for high-volume cloud inference. On-prem also changes the cost curve: at sustained training and inference volumes, owned infrastructure can undercut metered cloud while keeping data sovereignty intact. The trade is that you now own the physical plant, and the physical plant is the hard part.
The build plan at a glance
| Layer | What it demands in 2026 | Platform reference |
|---|---|---|
| Power | 100+ kW per rack; provisioned circuits and PDUs before hardware lands | Facility and UCS power planning |
| Cooling | Liquid cooling (direct-to-chip or immersion) above ~30 kW per rack | Direct-to-chip, rear-door heat exchangers |
| Network | Lossless, high-bandwidth east-west fabric for GPU-to-GPU traffic | Cisco Nexus, Cisco Hyperfabric AI |
| Compute | GPU clusters on validated reference architectures | Cisco UCS X-Series, NVIDIA GPUs, Cisco AI POD |
| Security and ops | Zero Trust, segmentation, AI governance, Day-2 handoff | QuickStrike, NIST AI RMF |
Step 1: Size the power before anything else
Power is the first constraint, and it is bigger than most facilities expect. A single modern AI rack draws enormous power. NVIDIA's GB200 NVL72 rack pulls roughly 120 to 130 kW. Most enterprise Data Centers were built for 8-to-12 kW air-cooled racks, as Network World documented in 2026. At national scale the trend is unmistakable: the U.S. Department of Energy's 2024 data center report, produced by Lawrence Berkeley National Laboratory, found data centers used about 4.4 percent of U.S. electricity in 2023 and projected 6.7 to 12 percent by 2028, driven largely by AI. You should provision circuits, distribution, and redundancy for the density you'll run in year three.
Step 2: Plan cooling for the density you'll actually run
Around 30 to 40 kW per rack, airflow struggles to carry the thermal output of dense GPU racks, and direct-to-chip or rear-door liquid cooling becomes necessary. Below that, air cooling still handles the majority of deployments. Uptime Institute's 2025 Cooling Systems Survey found only about 22% of operators using direct liquid cooling versus roughly 75% on perimeter air, and much of the current GPU market (including a large part of Cisco's UCS portfolio) ships air-cooled with liquid available as an option. Because retrofitting a facility for liquid is slow and costly, smaller teams should confirm their density target early rather than assume liquid is mandatory. By late 2025, only about 45 percent of data centers ran purely on air, with a majority planning liquid cooling within five years, per S&P Global 451 Research figures reported by Network World. Decide the cooling approach at design time. Retrofitting liquid into an air-cooled room is expensive and disruptive.
Step 3: Build the network fabric the GPUs need
AI clusters live or die on east-west bandwidth, so the fabric matters as much as the GPUs. Training and inference generate enormous GPU-to-GPU traffic that a traditional north-south enterprise network was never built to carry. You need a lossless, high-bandwidth fabric, the role Cisco Nexus and Cisco Hyperfabric AI play in reference designs. Get this wrong and expensive GPUs sit idle waiting on the network, which is the most common way enterprises waste an AI capital budget.
Step 4: Bake in security and AI governance from day one
Secure the AI environment with the same posture as the rest of your stack: Zero Trust, segmentation, and audit-grade evidence. Govern the models as well as the hardware. On-premises already keeps sensitive data behind your perimeter; the discipline is to segment the AI cluster, control egress, and prove it to an auditor. For the model side, NIST's AI Risk Management Framework gives you four functions to operationalize: Govern, Map, Measure, and Manage. Run AI governance as a defined process. AI inherits the security posture you already operate.
Step 5: Choose a deployment partner who hands the keys back
The right AI infrastructure deployment partner is certified on the hardware, builds with validated reference architectures, and trains your team to operate the cluster on Day-2. Evaluate AI hardware and systems integrators on three things: depth on the specific platforms (Cisco UCS X-Series, Cisco Hyperfabric AI, NVIDIA GPU integration, reference AI POD designs), a staged delivery method with defined exit criteria, and a real handoff. Two failure modes are worth naming: a partner that stands up the cluster and disappears, leaving you with hardware you cannot run, and a partner that never leaves and becomes a managed-service dependency.
How BTA delivers on-prem AI
BTA architects the AI data center, deploys it in production, and trains your team to run it — under one security and operating posture. As a Cisco MINT-certified partner across the AI Infrastructure and AI Solutions categories, BTA delivers across Cisco UCS X-Series, Cisco Hyperfabric AI, NVIDIA GPU integration, and reference AI POD architectures. Where speed matters, QuickStrike stands up a validated, secure on-premises AI environment on existing hardware in days, with zero cloud egress and controls aligned to HIPAA, CMMC, PCI DSS, GDPR, and SOC 2. The delivery follows BTA's SIMPLE methodology, with 1,000+ projects and zero project failures behind it, which is the same data-center discipline that helped Ballad Health consolidate five data centers into two with Cisco. Explore BTA's on-premise AI deployment practice, our AI deployment program, and BTA's Cisco partnership.
FAQ
Is it cheaper to run AI on-premises or in the cloud?
It depends on volume and data constraints, though at sustained scale on-premises often wins. Cloud costs less to start and grows harder to predict as inference volume rises. Owned infrastructure gives you a fixed, predictable cost base. For regulated data that cannot leave a controlled environment, on-prem is frequently the only compliant option regardless of cost.
How much power does an AI rack need?
A dense AI rack in 2026 can draw far more power than legacy infrastructure was built for. A single Cisco UCS C885A M8, with eight NVIDIA HGX GPUs, has a maximum system power of about 12.5 kW, so even a few per rack push well past the 8 to 12 kW most legacy racks were designed for. That is why power sizing comes first in the build plan. Provision for the density you will run at full scale.
Do we need liquid cooling for on-prem AI?
Above roughly 30 kW per rack, yes. Air cooling cannot carry the heat that dense GPU racks produce, which is why the industry is shifting to direct-to-chip and immersion cooling. Decide the cooling method at design time, because retrofitting liquid into an air-cooled room is expensive and disruptive.
How long does it take to stand up on-premises AI infrastructure?
A typical on-prem AI build runs around 90 days. BTA's QuickStrike stands up a secure AI environment on validated existing hardware in days. Full custom data-center builds take longer and depend on power and cooling lead times. Timeline is confirmed during scoping, before work begins.
How do we keep AI workloads compliant on-premises?
Segment the AI cluster, control egress, and govern the models with a defined framework. On-premises infrastructure keeps sensitive data behind your perimeter. Applying the NIST AI Risk Management Framework (Govern, Map, Measure, Manage) makes compliance a repeatable process. BTA applies the same Zero Trust and audit-grade evidence posture to AI that it applies to the rest of the stack.
What should we look for in an AI infrastructure deployment partner?
Look for certification on your hardware, validated reference architectures, and a partner that trains your team to own Day-2. The strongest AI hardware and systems integrators deliver the cluster in production and hand the operating skill back to your team.
Start the conversation
Talk to a BTA architect about your on-prem AI build, or start with a 30-minute scoping call. Let's get in touch.