Does this AI chip still give the right answer?
A GPU can pass every vendor diagnostic and still return wrong numbers with no error raised. This is an open, vendor-neutral tester that proves whether a chip computed correctly, and publishes the evidence as a file anyone can check.
Start with what a silent error is, or go straight to the measured results.
What a silent error is
When a chip fails, there are three ways it can tell you. Only two of them are loud.
It crashes
The job stops. Annoying, but you know. Nothing is quietly wrong.
It raises an error
Memory with ECC, for example, notices a flipped bit, corrects it or reports it. Monitoring sees a counter go up.
It gives a wrong number
The job finishes. Every status light is green. One number in the output is wrong, and nothing on the chip knows. That is a silent error.
Where silent errors come from
A modern AI chip has tens of thousands of arithmetic units and billions of memory cells, made on a process so fine that a few atoms decide whether a transistor switches on time. Most units are perfect. A few are marginal: they work at one temperature and fail at another, work on most inputs and fail on a specific bit pattern, or degrade as the chip ages. When a marginal unit produces the wrong bit during a multiply, nothing compares the answer to the truth, so the wrong number simply flows onward.
Why the usual tests miss them
Vendor diagnostics run a fixed stress pattern for a short time and look for crashes and error counters. The research on real fleets says that is not enough. ByteDance's team reported that screening GPUs with synthetic microbenchmarks missed over 60 percent of the units that later corrupted real training runs, because the faults depend on the data being processed, show up with age, and are specific to individual units. A test that exercises one pattern for a few seconds can pass a chip that fails on the next pattern an hour later.
Why it matters
One wrong number is harmless. The problem is scale: thousands of chips, billions of operations a second, and nothing that checks the answers.
Training runs go quietly wrong
In large language model training, a corrupted gradient from one faulty GPU spreads to every other GPU at the next synchronisation. The loss curve wobbles, days of compute are lost, and the cause stays invisible until someone finds the one unit. The second OSDI 2026 study below counted 18 such incidents and 13 faulty GPUs across 35 million GPU hours.
Inference answers drift
A serving fleet with one marginal chip produces slightly different answers from that chip, forever. No crash, no alert, just a share of users getting worse results.
The used GPU market has no condition test
Second-hand accelerators are bought and sold on model name and hours of use. Age is the biggest risk factor for silent faults, and there is no open, shared way to test a used card before paying for it.
Lenders cannot inspect the collateral
Loans are secured against GPUs. A jet has maintenance records and an inspection standard; a GPU has a serial number. An open integrity test is the first step towards a condition grade anyone can verify.
How a chip is tested here
Two probes. The first makes the chip do arithmetic and judges every output element. The second fills the chip's memory with known values and reads every one back. Each probe has two independent checks, so a fault has to hide from both.
Probe 1: compute
The chip multiplies two fixed matrices, A and B, fifty times at each of three shapes and four precisions. Two checks run on every output.
The reference check, in plain words
Floating point arithmetic rounds at every step, so two correct computers can legitimately differ in the last digit. The question is how different they are allowed to be. Numerical analysis has a proven answer: for a dot product of length K, the rounding error can never exceed a bound that depends only on K and the precision. Every element of the chip's result is compared with a float64 reference and must sit inside that bound. For int8 the arithmetic is exact, so the bound is zero.
g = K·u / (1 − K·u) u = 2⁻²⁴ (FP32 accumulation)
u_out = 2⁻²⁴ fp32 2⁻¹¹ fp16 2⁻⁸ bf16 0 int8
An element outside the bound is a wrong answer, not a difference of opinion about tolerance.
The repeat check, in plain words
The same inputs through the same chip should give the same bits every time. Run 0 is kept, and every later run is compared bit for bit. This catches an intermittent fault far too small to leave the bound, because the two runs simply no longer match. On the chips measured so far every run matched, which is itself a useful fact: the matrix kernels on these chips are deterministic.
Probe 2: memory
The matrix multiply touches a few megabytes. Most of an 80 GB chip is memory, and a bad row or a stuck bit there corrupts whatever lands on it. The memory probe fills most of the device memory, writes a known value into every 32-bit word, waits, and reads every word back.
Reading a verdict
Every step ends in one of these words. They are chosen so a reader can tell a chip fault from a software quirk without opening the file.
Compute probe
Memory probe
Self-test
Glossary
- silent error, silent data corruption, SDC
- A wrong result from a chip with no crash, no error flag and no log entry. The chip does not know it happened.
- bit flip
- One binary digit changing value. In a floating point number a flipped high bit can change the value by orders of magnitude; a flipped low bit changes it by less than one part in a million.
- ECC
- Error-correcting code. Extra bits stored with memory that let the hardware detect and usually correct a single flipped bit. ECC covers memory, not the arithmetic units, and not every chip has it on all memory.
- HBM
- High bandwidth memory, the stacked memory next to a data centre GPU die. An H100 has 80 GB of it.
- tensor core
- A unit on NVIDIA GPUs that multiplies small matrices in one step. fp16, bf16 and int8 matrix multiplies run through tensor cores; fp32 runs through the regular floating point units.
- fp32, fp16, bf16, int8
- Number formats. fp32 keeps 24 bits of precision, fp16 keeps 11, bf16 keeps 8 with a wider range, int8 is an exact integer. Models are served in the smaller formats because they are faster, which is why the tester checks all four.
- accumulation
- The running sum inside a dot product. Even when the inputs are fp16, the sum is normally kept in fp32, which is what the reference bound assumes and what the tester enforces on CUDA.
- deterministic kernel
- A GPU routine that gives bit-identical output every time for the same input. Not all do, which is why non-determinism is reported as its own verdict instead of as a fault.
- reference bound
- The largest error that correct rounding can produce for a given computation and precision. Anything larger is a wrong answer. The bound used here is the standard componentwise bound for inner products from Higham.
- fault injection
- Deliberately corrupting a value to confirm the checks notice. Used here as a self-test on every published row.
- dwell
- The pause between writing memory and reading it back. Weak cells lose their value over time, so a longer dwell catches more of them.
- OCP Test and Validation
- The Open Compute Project's open format for hardware diagnostic output, used by cloud operators so that any test's results can be read by any fleet tool. Both probes write it.
Measured results
Every row comes from a result file in the repository, and the tables are generated from those files. Nothing here is modelled or estimated.
Compute probe
| Device | Date | Shapes (M×K×N) | Precisions | Runs | Silent faults | Injected caught | Verdict | Files |
|---|---|---|---|---|---|---|---|---|
| NVIDIA H100 80GB HBM3PyTorch 2.8.0+cu128 · tool v0.2.0 | 2026-10-07 | 1024x1024x1024 4096x4096x4096 32x4096x11008 |
fp32, fp16, bf16, int8 | 600 | 0 | 60 of 60 | no-silent-errors | clean · self-test |
| NVIDIA A100-SXM4-80GBPyTorch 2.8.0+cu128 · tool v0.2.0 | 2026-10-07 | 1024x1024x1024 4096x4096x4096 32x4096x11008 |
fp32, fp16, bf16, int8 | 600 | 0 | 60 of 60 | no-silent-errors | clean · self-test |
| NVIDIA H200PyTorch 2.8.0+cu128 · tool v0.2.0 | 2026-10-08 | 1024x1024x1024 4096x4096x4096 32x4096x11008 |
fp32, fp16, bf16, int8 | 600 | 0 | 60 of 60 | no-silent-errors | clean · self-test |
| Apple GPU (MPS) arm64PyTorch 2.14.1 · tool v0.2.0 | 2026-10-07 | 1024x1024x1024 4096x4096x4096 32x4096x11008 |
fp32, fp16, bf16 (3 skipped) | 450 | 0 | 45 of 45 | no-silent-errors | clean · self-test |
Memory probe
| Device | Date | Memory tested | Patterns | Dwell | Bad words | Injected caught | Verdict | Files |
|---|---|---|---|---|---|---|---|---|
| NVIDIA H100 80GB HBM3tool v0.1.0 | 2026-10-07 | 62.75 GiB of 79.18 GiB16,844,324,864 words | 6 | 2.0 s | 0 | 30 of 30 | no-memory-errors | clean · self-test |
| NVIDIA A100-SXM4-80GBtool v0.1.0 | 2026-10-07 | 63.00 GiB of 79.25 GiB16,911,433,728 words | 6 | 2.0 s | 0 | 30 of 30 | no-memory-errors | clean · self-test |
| NVIDIA H200tool v0.1.0 | 2026-10-08 | 111.25 GiB of 139.81 GiB29,863,444,480 words | 6 | 2.0 s | 0 | 30 of 30 | no-memory-errors | clean · self-test |
| Apple GPU (MPS) arm64tool v0.1.0 | 2026-10-07 | 5.50 GiB of 11.84 GiB1,476,395,008 words | 6 | 2.0 s | 0 | 30 of 30 | no-memory-errors | clean · self-test |
Rented data centre GPUs are young, monitored and already screened, so clean rows are the expected result there. The research says faults concentrate in aged units and appear under data-dependent, long, hot workloads. Those are the next probes on the roadmap, and where the first real findings are expected.
The self-test
A clean table is only worth something if the tester can show it would have seen a fault. So every row is paired with an injection run.
Compute self-test
In five of the fifty runs, one random bit of one random output element is flipped before the checks run. The bit can be anywhere: sign, exponent or the lowest mantissa bit. All five must be caught. On the chips measured so far, the reference check alone caught 46 of every 60 flips and the repeat check caught all 60. That is why there are two checks.
Memory self-test
After each pattern is written, five random bits in five random words of real device memory are flipped. The sweep must locate all of them by address. Because the flips happen in device memory, this exercises the actual read-back path, not only the comparison code.
What a PASS means, and does not
- It means: no silent error appeared in these matrix multiplies and this memory sweep, in these runs, on this device, on this day.
- It does not mean: the chip is healthy. Faults that depend on specific data, on heat, or on hours of load are not yet exercised. The next probes add them.
- Accumulation: the compute bound assumes FP32 or better accumulation. That is enforced on CUDA. Elsewhere a lower precision would show as outside-error-bound on every run, and the message says so.
- One device per run: links between chips are not yet tested.
- Self-test scope: the compute injection corrupts the output after it leaves the device, so it proves the checker, not the chip. The memory injection corrupts device memory itself.
Run it on your chip
Four commands, four result files, one pull request. Works on CUDA, Apple MPS and CPU today.
# once git clone https://github.com/tech4biz-yasha/ai-chip-integrity.git && cd ai-chip-integrity pip install -r requirements.txt # compute probe and its self-test python screen.py --out NAME_clean.jsonl python screen.py --inject 5 --out NAME_inject.jsonl # memory probe and its self-test python memcheck.py --out NAME_memcheck.jsonl python memcheck.py --inject 5 --out NAME_memcheck_inject.jsonl
- Open a pull request with the four files, the device name, driver version and PyTorch version.
- Rows are added only from attached files. The table is regenerated from them, so a row cannot say more than its file does.
- Found a fault? Open an issue with the file. A real finding on real hardware is the most useful thing anyone can contribute.
The output follows the OCP Test and Validation spec v2.0 through the official ocptv library, and the test suite validates it against the published schema.
Roadmap
Each probe targets one part of the chip. The order follows where the research says faults hide.
Research and standards this builds on
- Zheng et al., OSDI 2026. SDCs in the Wild: characterising and diagnosing SDC-defective GPUs in production LLM training. The source of the finding that microbenchmark screening misses most faulty units. usenix.org
- Lei et al., OSDI 2026. Safeguarding LLM Training at Scale: online SDC detection and insights from 35 million GPU hours. usenix.org
- Higham, Accuracy and Stability of Numerical Algorithms. The componentwise rounding error bound for inner products used by the reference check.
- Huang and Abraham, 1984. Algorithm-based fault tolerance for matrix operations. The original idea of checking a matrix result against a cheap independent computation.
- OCP Test and Validation. The open diagnostic output format and reference implementation both probes write to. ocp-diag-core
About and cite
Chip Integrity is maintained by Yasha Khandelwal, Bengaluru. The code is MIT licensed. The project stays vendor neutral: the same tests run on every chip, and a row is published only with its result file.
Yasha Khandelwal. ai-chip-integrity: an open tester for silent computation errors in AI chips. Version 0.2.0, 2026. https://github.com/tech4biz-yasha/ai-chip-integrity
Questions, rows and findings: yasha.khandelwal@tech4biz.io or an issue on GitHub.