Run Massive-Scale UMAP in Minutes Using Multiple GPUs Without Losing Accuracy
NVIDIA, Tuesday, August 18th, 2026
NVIDIA cuML and cuVS added multi-GPU support for UMAP, executing over 870 GB of vectors in eight minutes.
NVIDIA describes multi-GPU acceleration for UMAP, a dimensionality reduction technique widely used for visualization and feature extraction. NVIDIA cuML and cuVS 25.06 introduced multi-GPU support for the all-neighbors k-nearest-neighbor graph construction step, which is the main bottleneck.
That enables end-to-end scaling across GPUs rather than being limited by a single device's memory. NVIDIA reports executing UMAP over 870 GB of vectors in eight minutes. The post emphasizes that the speedup does not come at the cost of accuracy.