In most AI data centers, the teams that run GPUs and the systems that cool them never talk to each other. The IT side schedules training and inference jobs on Kubernetes or Slurm. The OT side runs the liquid cooling plant: coolant distribution units (CDUs), pumps, and valves. Because they share no data, cooling reacts only after temperatures rise. Operators run pumps near maximum to stay safe, and GPUs still throttle when a large job ramps up faster than the cooling can follow.
The Unified IT and OT Closed-Loop System is part of Federator.ai Cortex, ProphetStor’s full-stack AI Ops platform for AI factories. It turns IT/OT convergence into a single control loop, bringing workload data and liquid cooling data together. Federator.ai Cortex knows which jobs are about to run and how much heat they will produce, so it can adjust cooling and workload placement together instead of treating them as separate problems.
How the Loop Closes
Sense
Collects IT telemetry, such as GPU power, job schedules, and Kubernetes metadata, alongside OT telemetry, such as CDU flow, supply and return temperatures, pump speed, and power feeds.
Correlate
Multi-Layer Correlation, ProphetStor’s patented technology, maps how a change in workload cascades into GPU power, heat, and cooling demand.
Predict
A heat forecast built from GPU power telemetry and job schedules predicts rack heat 30 to 60 seconds ahead. Model Predictive Control (MPC) uses that forecast to plan pump and valve moves before temperatures begin to climb.
Actuate
Sends pump speed and valve set-points to the cooling plant, and returns power-budget hints to the scheduler when thermal headroom runs short.
What the Closed Loop Delivers
Workload-Aware Cooling
Federator.ai Cortex reads Kubernetes/Slurm job schedules and live GPU power, so cooling ramps up before a training job starts. Each rack gets the coolant flow it will need next, not what it needed a minute ago.
Right-Sized Coolant Flow
Steady Clocks Under Load
AI workloads can swing rack power by up to 50 percent within seconds. Predictive control keeps GPUs below their throttle limits through these spikes, so clock speeds hold steady and throughput stays intact.
Thermal-Aware Scheduling
The loop runs in both directions. When a rack runs short of thermal headroom, Federator.ai Cortex sends power-budget hints back to the scheduler, steering new jobs toward racks with cooling capacity to spare.
Causal Root-Cause Analysis
Turn More Power into Compute
In many AI data centers, power is now a tighter constraint than GPUs. Every watt saved on cooling is a watt available for training and inference. By closing the loop between IT and OT, Federator.ai Cortex helps keep PUE at or below 1.15 and puts more of each facility’s power budget to productive work.
Frequently Asked Questions
What is a unified IT and OT closed-loop system?
It is a control system that combines workload data from the IT side, such as GPU utilization and job schedules, with facility data from the OT side, such as coolant flow and temperatures. Federator.ai Cortex uses this combined view to adjust cooling and workload placement together in real time.
How is this different from AI cooling optimization that only monitors the facility?
How far ahead does Federator.ai Cortex predict heat load?
Federator.ai Cortex forecasts rack heat 30 to 60 seconds ahead, using high-frequency GPU power telemetry, Kubernetes/Slurm job schedules, and ambient conditions. This lead time lets coolant pumps and valves ramp up before a workload spike arrives, so GPUs are already being cooled when the heat hits. Clock speeds stay clear of thermal throttling, and throughput remains steady even through sharp load changes.
Does Federator.ai Cortex replace the existing BMS or CDU controller?
No. Federator.ai Cortex publishes recommended set-points to existing controllers and building management systems. All commands respect vendor safety limits, and operators retain full control of the physical plant.