Distributed inference / Personal lab
When a healthy cluster is still slow
EnvironmentTwo-box DGX Spark / GB10
FocusPrefill throughput
MeasurementPublished reference comparison
The problem
My two-box cluster was connected and its diagnostics looked healthy, but the interconnect was running at a fraction of its expected speed. I needed to check what the model workload was actually doing.
What I checked
I worked through the RDMA path, both ConnectX rails and NCCL configuration, then benchmarked distributed inference. The goal was a usable model service with performance measured on the workload, rather than a connection check alone.
Measured on my cluster
Qwen prefill reached 104–111% and GLM prefill reached 103–116% of the published reference figures for the same two-box setup. Prefill measures how quickly the model processes a prompt; it is separate from generation speed.
What this means for your system
I validate the communication path and benchmark your actual prompts before handover. Hardware, model, precision and request shape all affect the result.
Case study from my own lab. Results depend on the hardware, model and workload.