Distributed training ray train
Adapts existing PyTorch, TensorFlow, and HuggingFace training scripts to run distributed on Ray Train across multiple GPUs or nodes, including hyperparameter tuning, checkpointing, and cluster status checks. Use when the user wants to scale training to multi-G