← News

GPU quality is something to plan for

2026-10-07 · 2 min read · Lucas Ewing


GPUs fail, and at scale they fail constantly. While training Llama 3 on 16,384 H100s, Meta logged 419 unexpected interruptions in 54 days, about one every three hours, and 78% of them traced back to hardware (Meta, 2024). Every neocloud inherits failure rates like that. What differs from one neocloud to the next is how quickly a failure is found and fixed, and whether the customer notices first.

What buyers now measure

SemiAnalysis's ClusterMAX has become the reference buyers use to judge GPU providers. Its latest edition, published in September, tested 77 providers in depth, and only 19 earned a medal (SemiAnalysis, 2026). The ratings weigh far more than hardware. Among the things the best-rated clouds do:

  • Burn-in that stresses GPUs and network together, before a customer sees the cluster.
  • Health checks that act on their own. A failing GPU is drained and replaced without waiting for someone to notice. SemiAnalysis injects real failures to test this and calls it its most critical test.
  • Current, patched images and a correctly configured fabric, checked rather than assumed.
  • A dashboard that shows the failed part, the affected workload, and when each check last ran. In SemiAnalysis's words, a stale green result is not evidence of health.

Quality shows up in the economics

SemiAnalysis's cost model assumes a Gold-rated cloud detects and replaces a failed GPU in 15 minutes, against an hour for Silver, and finds the better-run cluster 5 to 15% cheaper to operate at the same GPU price (SemiAnalysis, 2026). Buyers also pay more per GPU-hour for the best-run managed clusters (SemiAnalysis, 2025).

Most plans stop at the racks

Most companies planning a neocloud plan as far as ordering the hardware and getting it racked. Few plan for running it. Burn-in, health checks, repair, and the tooling behind them decide whether a cluster runs like a well-rated cloud, and they are much harder to add once customers are running on it.

That is the part of running a neocloud Lilac works on. If you are bringing a cluster online in the next few months, we would be glad to compare notes: contact@getlilac.com.


← All news