Meta and AMD have successfully ported PyTorch Monarch to ROCm, bringing fault-tolerant distributed training to AMD Instinct GPUs for the first time.

The collaboration addresses a critical challenge in large-scale AI training: hardware failures that can destroy weeks of progress. Traditional approaches rely on periodic checkpointing, which wastes computation and leaves clusters idle during recovery.

Monarch introduces a different paradigm. The system uses an actor-based runtime with hierarchical fault handling that isolates failures and enables rapid recovery without halting healthy nodes.

Engineering the ROCm Port

Porting Monarch from CUDA to ROCm required three main technical adaptations. The team used hipify_torch to convert collective communications code from CUDA to HIP, linking against RCCL instead of NCCL.

GPU memory management was extended through auto-detection that routes CUDA driver API calls through HIP equivalents. RDMA integration maintained the existing libibverbs path while swapping GPU-side bindings from CUDA to HIP for direct transfers.

Two platform differences shaped the implementation. ROCm lacks a static runtime library equivalent to NVIDIA's libcudart_static.a, forcing dynamic linking of libamdhip64. The team also created Rust compatibility shims that re-export HIP symbols under CUDA names, avoiding code forks.

All 1,171 tests now pass on ROCm 7.0+, with contributions upstreamed to the open-source community.

Production-Ready Architecture

The system integrates three layers for resilient training. Monarch orchestrates clusters and spawns ReplicaActors organized into Process Meshes. TorchFT handles step-level fault tolerance through quorum coordination and AllReduce operations that skip failed nodes. TorchTitan executes the actual training with FSDP, managing forward passes, backpropagation, and optimization.

This architecture enables checkpoint-less distributed training that dynamically recovers from node failures in seconds rather than minutes.

Monarch on ROCm supports full ecosystem integration across SLURM, Kubernetes, and SkyPilot environments. The system provides the foundation for production workloads requiring both scale and reliability on AMD hardware.

The port represents a significant step toward hardware-agnostic AI infrastructure, reducing dependence on single-vendor solutions for large-scale model training.