Federator.ai Cortex: Unified IT and OT Closed-Loop System

In most AI data centers, the teams that run GPUs and the systems that cool them never talk to each other. The IT side schedules training and inference jobs on Kubernetes or Slurm. The OT side runs the liquid cooling plant: coolant distribution units (CDUs), pumps, and valves. Because they share no data, cooling reacts only after temperatures rise. Operators run pumps near maximum to stay safe, and GPUs still throttle when a large job ramps up faster than the cooling can follow.

The Unified IT and OT Closed-Loop System is part of Federator.ai Cortex, ProphetStor’s full-stack AI Ops platform for AI factories. It turns IT/OT convergence into a single control loop, bringing workload data and liquid cooling data together. Federator.ai Cortex knows which jobs are about to run and how much heat they will produce, so it can adjust cooling and workload placement together instead of treating them as separate problems.

How the Loop Closes

The system runs continuously, on every control cycle, through four stages:

Sense

Collects IT telemetry, such as GPU power, job schedules, and Kubernetes metadata, alongside OT telemetry, such as CDU flow, supply and return temperatures, pump speed, and power feeds.

Correlate

Multi-Layer Correlation, ProphetStor’s patented technology, maps how a change in workload cascades into GPU power, heat, and cooling demand.

Predict

A heat forecast built from GPU power telemetry and job schedules predicts rack heat 30 to 60 seconds ahead. Model Predictive Control (MPC) uses that forecast to plan pump and valve moves before temperatures begin to climb.

Actuate

Sends pump speed and valve set-points to the cooling plant, and returns power-budget hints to the scheduler when thermal headroom runs short.

Every command stays within the equipment vendor’s limits for pump speed and valve travel, and is issued only when leak, flow, and pressure readings are within safe range. Operators keep full authority over the physical plant.

What the Closed Loop Delivers

Workload-Aware Cooling

Federator.ai Cortex reads Kubernetes/Slurm job schedules and live GPU power, so cooling ramps up before a training job starts. Each rack gets the coolant flow it will need next, not what it needed a minute ago.

Right-Sized Coolant Flow

Fixed-flow systems run pumps near maximum even when racks sit idle. Federator.ai Cortex holds coolant temperature rise within a target window and trims flow at low load, cutting pump energy by 25 to 30 percent.

Steady Clocks Under Load

AI workloads can swing rack power by up to 50 percent within seconds. Predictive control keeps GPUs below their throttle limits through these spikes, so clock speeds hold steady and throughput stays intact.

Thermal-Aware Scheduling

The loop runs in both directions. When a rack runs short of thermal headroom, Federator.ai Cortex sends power-budget hints back to the scheduler, steering new jobs toward racks with cooling capacity to spare.

Causal Root-Cause Analysis

With IT and OT data on one timeline, Federator.ai Cortex can tell whether a hot GPU comes from a heavy job or a weakening pump. Operators see the cause, not just an alarm, and can act before the problem spreads.

Turn More Power into Compute

In many AI data centers, power is now a tighter constraint than GPUs. Every watt saved on cooling is a watt available for training and inference. By closing the loop between IT and OT, Federator.ai Cortex helps keep PUE at or below 1.15 and puts more of each facility’s power budget to productive work.

Frequently Asked Questions

What is a unified IT and OT closed-loop system?

It is a control system that combines workload data from the IT side, such as GPU utilization and job schedules, with facility data from the OT side, such as coolant flow and temperatures. Federator.ai Cortex uses this combined view to adjust cooling and workload placement together in real time.

Facility-only systems adjust cooling to the heat load after it appears. Federator.ai Cortex knows which jobs are about to run and how much heat they will generate, so it can prepare cooling in advance and shift workloads when cooling capacity is limited.

Federator.ai Cortex forecasts rack heat 30 to 60 seconds ahead, using high-frequency GPU power telemetry, Kubernetes/Slurm job schedules, and ambient conditions. This lead time lets coolant pumps and valves ramp up before a workload spike arrives, so GPUs are already being cooled when the heat hits. Clock speeds stay clear of thermal throttling, and throughput remains steady even through sharp load changes.

No. Federator.ai Cortex publishes recommended set-points to existing controllers and building management systems. All commands respect vendor safety limits, and operators retain full control of the physical plant.

Multi-Layer Correlation (U.S. Patent No. 11,579,933 B2) maps how changes in application workload cascade into infrastructure resource demand. This gives Federator.ai Cortex its workload awareness. Combined with IT and OT convergence, it makes Federator.ai Cortex both workload-aware and thermal-aware, so compute and cooling can be tuned together to boost AI factory performance.

Please select the software/ platform you would like a demo of:

Federator.ai Cortex™

A Unified IT and OT Closed-Loop AIOps System for Modern AI Factories

Federator.ai GPU Booster™

GPU Performance Maximization with AI-Enhanced Dynamic Allocation for LLMs

Federator.ai Smart Liquid Cooling™

Predictive Workload-Aware Liquid Cooling for High-Density GPU Data Centers

Federator.ai GPU Booster Inference™

GPU Performance Maximization with AI-Enhanced Dynamic Allocation for LLM Inference

Federator.ai®

AI-Driven Compute Resource Optimization for Cloud and On-Premises Operations