Enabling DeepSeek-V4-Flash Training on AMD Instinct MI355X GPUs With Primus
AMD, Thursday, September 3rd, 2026
AMD details how its Primus framework trains DeepSeek-V4-Flash, a 284B-parameter MoE model, on Instinct MI355X GPUs.
AMD's ROCm blog walks through enabling training of DeepSeek-V4-Flash on AMD Instinct MI355X GPUs using the Primus framework.
DeepSeek-AI released the MIT-licensed DeepSeek-V4 series on April 24, 2026, with the Flash variant carrying 284 billion total parameters, 13 billion activated, and a one-million-token context window.
The model pushes sparse attention further than any prior open-weight release, interleaving three different attention types across its 43 transformer layers. The post covers the engineering required to make that architecture train efficiently on AMD hardware.