qwen3 2 Mixture-of-Experts Across the Stack: From a 20-Line Reference to Wide Expert Parallelism Jun 5, 2026 Qwen3 Architecture Layer-by-Layer — Reading the HuggingFace Implementation May 8, 2026