AMD FP8 training optimizations now mainstream in PyTorch's TorchTitan and TorchAO
AMD and Meta have upstreamed FP8 optimizations for Instinct GPUs into PyTorch: a 13.4% throughput gain on Llama3-8B, 89% recovery of quantization overhead on DeepSeek-V3 671B, and up to 6.2x faster MoE kernels. Teams can now get competitive FP8 train...