Distributed Training 3 Megatron-LM Fine-Tuning Deep Dive: Qwen Recipes, Datasets, and the Training Loop Aug 11, 2026 NVSHMEM Deep Dive: One-Sided GPU Communication, and How It Differs from NCCL Jun 8, 2026 PyTorch DDP Deep Dive — Initialization, Buffer Broadcast, the Reducer, and the torch.compile Story May 11, 2026