PaleBlueDot AI cluster clears NVIDIA training benchmark
A cluster in PaleBlueDot AI’s data center has spent a week running continuously at full load, the kind of sustained pressure that can expose weak links in large AI systems. The company says its NVIDIA HGX B300 cluster has now earned NVIDIA Exemplar Cloud status after exceeding NVIDIA’s benchmark threshold across every test recipe.
The result covers six large-model training workloads: DeepSeek-V3, GPT-OSS, Nemotron-H, Qwen3 and two configurations of Llama 3.1. PaleBlueDot AI says every run exceeded 98% of NVIDIA reference performance, across different model architectures, parameter scales and numerical precision formats. NVIDIA created Exemplar Cloud in 2025 as a standard against which buyers can compare infrastructure performance.
The hardware is built for the part of AI training that happens between the chips. Each compute node contains eight NVIDIA Blackwell Ultra GPUs connected through NVLink and NVLink Switch. An 800Gb/s non-blocking InfiniBand network — a high-speed system for linking computers — is designed to prevent communication bottlenecks between nodes, while each node can reach aggregate compute-network bandwidth of 6.4 Tb/s. A parallel storage system and 63.36TB of local NVMe cache are intended to speed data loading and checkpoint operations.
PaleBlueDot AI also says it tested compute, networking, storage and scheduling during a week-long, non-stop production simulation. Its quality process includes hardware burn-in, single-node acceptance and long-duration cluster testing, alongside automated monitoring and isolation of unhealthy nodes. The company presents these measures as a way to reduce interruptions and idle training costs.
So what changes in practice? For an AI laboratory or enterprise preparing a large training run, the NVIDIA benchmark offers a reference point for procurement and budget reviews, while the stability test addresses the risk of a job failing after days of operation. The figures are reported by NVIDIA’s assessment and PaleBlueDot AI’s own testing; no independent trial is cited, and the material does not establish a broad customer deployment.
Comentários
A carregar a conversa…
Inicie sessão para escrever um comentário. Iniciar sessão