Adaptive Distributed Parallelism (ADP): Right-Sized Training for Mixed GPU Fleets

GPU fleets now span several hardware generations, and older GPUs keep earning long after newer ones arrive. Training jobs have not caught up. Most are launched with a fixed mix of data, tensor, and pipeline parallelism, often paired with a sharded optimizer, chosen once by hand and tuned for a single GPU type. When the hardware, the network, or the facility changes, the job keeps running the same way and leaves performance on the table.

Adaptive Distributed Parallelism (ADP) is part of Federator.ai Cortex, ProphetStor’s full-stack AI Ops platform for AI factories. It works with NVIDIA’s open-source KAI Scheduler to choose the right parallelism strategy for each job, place it on the right mix of GPUs, and reconfigure it as conditions change, so every GPU generation in the fleet can contribute to the same training run.

How ADP Works with KAI Scheduler

KAI Scheduler, originally developed as part of NVIDIA Run:ai, is a Kubernetes-native scheduler that decides when and where AI jobs run. It manages queues, enforces fair sharing between teams, and gang-schedules the pods of a distributed job so they start together. ADP adds the layer above it: how each job should divide its work, and which GPUs should take each part. KAI Scheduler places the job, and ADP makes sure the job is shaped to fit the GPUs it lands on, even when they span more than one generation.

What ADP Delivers

Native KAI Scheduler Integration

ADP works alongside NVIDIA’s open-source KAI Scheduler instead of replacing it. KAI Scheduler still handles queues, fairness, and gang scheduling, while ADP adds GPU-aware placement and parallelism decisions.

Parallelism Matched to Each Job

ADP reads each job’s model topology and dataset size, then picks the parallelism mix that fits, combining data, tensor, and pipeline parallelism with sharded optimizers such as DeepSpeed ZeRO where needed. Launch scripts no longer need hand tuning.

Reconfiguration During Training

When a network link saturates, a GPU degrades, or new capacity frees up, ADP reconfigures the parallelism strategy, so the job keeps running near full speed instead of at a fraction of its potential.

Training Across GPU Generations

Jobs are no longer confined to a single GPU type. ADP gives each GPU generation a share of the model and data it can handle, so a mixed fleet of older and newer GPUs can train one model together.

Power- and Cooling-Aware Placement

Placement accounts for rack power budgets and cooling headroom, drawn from the IT and OT telemetry in Federator.ai Cortex. Jobs land where the facility can sustain them at full speed, not just where GPUs are free.

Put the Right GPU on the Right Workload

A GPU fleet is a long-term investment, and each generation should keep paying off as new hardware arrives. ADP builds on ProphetStor’s patented technologies, including Spatial-Temporal GPU Optimization (U.S. Patent US 12,596,580 B2), to keep every generation productive, whether the job is training a large language model or running an agentic AI pipeline.

Frequently Asked Questions

What is Adaptive Distributed Parallelism (ADP)?

ADP is a Federator.ai Cortex solution that selects and reconfigures the parallelism strategy for each AI training job. It considers model topology, dataset size, available GPU types, network conditions, GPU health, and power and cooling limits to decide how a job divides its work and where each part runs.

KAI Scheduler is NVIDIA’s open-source, Kubernetes-native scheduler for AI workloads. It manages queues, fair sharing, and gang scheduling. ADP integrates with KAI Scheduler and adds GPU-aware placement and parallelism decisions, so jobs are not limited to a single GPU type.

Yes. Adaptive Distributed Parallelism (ADP) in Federator.ai Cortex assigns each GPU generation a share of the model and data that matches its compute and memory capacity, so older and newer GPUs can train the same model without the slowest device setting the pace.
Federator.ai Cortex’s Adaptive Distributed Parallelism (ADP) works with widely used distributed training frameworks, including NVIDIA NeMo Megatron, DeepSpeed, and Ray.

Please select the software/ platform you would like a demo of:

Federator.ai Cortex™

A Unified IT and OT Closed-Loop AIOps System for Modern AI Factories

Federator.ai GPU Booster™

GPU Performance Maximization with AI-Enhanced Dynamic Allocation for LLMs

Federator.ai Smart Liquid Cooling™

Predictive Workload-Aware Liquid Cooling for High-Density GPU Data Centers

Federator.ai GPU Booster Inference™

GPU Performance Maximization with AI-Enhanced Dynamic Allocation for LLM Inference

Federator.ai®

AI-Driven Compute Resource Optimization for Cloud and On-Premises Operations