pytorch 8
- Megatron-LM Fine-Tuning Deep Dive: Qwen Recipes, Datasets, and the Training Loop
- TVM FFI Deep Dive: Shipping Native Kernels Without Tying Them to One PyTorch Build
- NVSHMEM Deep Dive: One-Sided GPU Communication, and How It Differs from NCCL
- PyTorch DDP Deep Dive — Initialization, Buffer Broadcast, the Reducer, and the torch.compile Story
- Torch.compile Deep Dive II — Inductor Codegen and Buffer Lifetimes, End to End
- NCCL Deep Dive: From P2P/CUMEM to torchrun Init
- From FlashAttention to FlashAttention 2
- Torch.compile 101 - From Python Function to Triton Kernel