OPEN SOURCE, MIT, OCP TEST AND VALIDATION OUTPUT

Does this AI chip still give the right answer?

A GPU can pass every vendor diagnostic and still return wrong numbers with no error raised. This is an open, vendor-neutral tester that proves whether a chip computed correctly, and publishes the evidence as a file anyone can check.

Start with what a silent error is, or go straight to the measured results.

4
chips measured
2,250
compute runs checked element by element
65.1 billion
memory words written and read back
0
silent faults found so far
345 of 345
injected faults caught in self-tests
LEARN, PART 1 OF 5

What a silent error is

When a chip fails, there are three ways it can tell you. Only two of them are loud.

It crashes

The job stops. Annoying, but you know. Nothing is quietly wrong.

It raises an error

Memory with ECC, for example, notices a flipped bit, corrects it or reports it. Monitoring sees a counter go up.

It gives a wrong number

The job finishes. Every status light is green. One number in the output is wrong, and nothing on the chip knows. That is a silent error.

Where silent errors come from

A modern AI chip has tens of thousands of arithmetic units and billions of memory cells, made on a process so fine that a few atoms decide whether a transistor switches on time. Most units are perfect. A few are marginal: they work at one temperature and fail at another, work on most inputs and fail on a specific bit pattern, or degrade as the chip ages. When a marginal unit produces the wrong bit during a multiply, nothing compares the answer to the truth, so the wrong number simply flows onward.

input numbers2.0 and 3.0 multiply unithealthy: 6.0marginal: 6.0000076 next layer uses the numberno error, no log line, no crash A single mantissa bit flipped inside one multiply. The chip does not compare 6.0000076 with anything, so it never knows.
The difference between a healthy and a marginal unit can be one bit: far below what anyone notices in a single number, and enough to change a model's output after a billion more multiplies.

Why the usual tests miss them

Vendor diagnostics run a fixed stress pattern for a short time and look for crashes and error counters. The research on real fleets says that is not enough. ByteDance's team reported that screening GPUs with synthetic microbenchmarks missed over 60 percent of the units that later corrupted real training runs, because the faults depend on the data being processed, show up with age, and are specific to individual units. A test that exercises one pattern for a few seconds can pass a chip that fails on the next pattern an hour later.

LEARN, PART 2 OF 5

Why it matters

One wrong number is harmless. The problem is scale: thousands of chips, billions of operations a second, and nothing that checks the answers.

Training runs go quietly wrong

In large language model training, a corrupted gradient from one faulty GPU spreads to every other GPU at the next synchronisation. The loss curve wobbles, days of compute are lost, and the cause stays invisible until someone finds the one unit. The second OSDI 2026 study below counted 18 such incidents and 13 faulty GPUs across 35 million GPU hours.

Inference answers drift

A serving fleet with one marginal chip produces slightly different answers from that chip, forever. No crash, no alert, just a share of users getting worse results.

The used GPU market has no condition test

Second-hand accelerators are bought and sold on model name and hours of use. Age is the biggest risk factor for silent faults, and there is no open, shared way to test a used card before paying for it.

Lenders cannot inspect the collateral

Loans are secured against GPUs. A jet has maintenance records and an inspection standard; a GPU has a serial number. An open integrity test is the first step towards a condition grade anyone can verify.

LEARN, PART 3 OF 5

How a chip is tested here

Two probes. The first makes the chip do arithmetic and judges every output element. The second fills the chip's memory with known values and reads every one back. Each probe has two independent checks, so a fault has to hide from both.

Probe 1: compute

The chip multiplies two fixed matrices, A and B, fifty times at each of three shapes and four precisions. Two checks run on every output.

fixed inputs A, Bseeded, rounded chip computesC = A @ B, 50 times reference checkevery element inside the bound? repeat checkevery run bit-identical to run 0? verdictper shape, precision The CPU computes the same product in float64 as the reference. The chip never sees it.
The checks are independent. A large error fails the reference check. A tiny one, even a single bit, fails the repeat check because the first run and the later run no longer match.

The reference check, in plain words

Floating point arithmetic rounds at every step, so two correct computers can legitimately differ in the last digit. The question is how different they are allowed to be. Numerical analysis has a proven answer: for a dot product of length K, the rounding error can never exceed a bound that depends only on K and the precision. Every element of the chip's result is compared with a float64 reference and must sit inside that bound. For int8 the arithmetic is exact, so the bound is zero.

|C − R| ≤ g·(|A|@|B|) + u_out·(|R| + g·(|A|@|B|))
g = K·u / (1 − K·u) u = 2⁻²⁴ (FP32 accumulation)
u_out = 2⁻²⁴ fp32 2⁻¹¹ fp16 2⁻⁸ bf16 0 int8

An element outside the bound is a wrong answer, not a difference of opinion about tolerance.

The repeat check, in plain words

The same inputs through the same chip should give the same bits every time. Run 0 is kept, and every later run is compared bit for bit. This catches an intermittent fault far too small to leave the bound, because the two runs simply no longer match. On the chips measured so far every run matched, which is itself a useful fact: the matrix kernels on these chips are deterministic.

Probe 2: memory

The matrix multiply touches a few megabytes. Most of an 80 GB chip is memory, and a bad row or a stuck bit there corrupts whatever lands on it. The memory probe fills most of the device memory, writes a known value into every 32-bit word, waits, and reads every word back.

DEVICE MEMORY, ONE PATTERN PASS
write the pattern into every worddwell 2 sread back, compare exactly
SIX PATTERNS, IN ORDER
all zerosevery bit 0
all onesevery bit 1
0xAAAAAAAA1010 1010 ...
0x555555550101 0101 ...
address hasheach word unique
inverse hasheach bit flipped
Together they drive every bit to 0 and to 1 while its neighbours hold the opposite, which is what exposes stuck bits, coupled cells and weak rows.
For each failing word the probe records the byte offset and the bits that differed, so a stuck bit shows as a single mask such as 0x00000020 at a known address.
LEARN, PART 4 OF 5

Reading a verdict

Every step ends in one of these words. They are chosen so a reader can tell a chip fault from a software quirk without opening the file.

Compute probe

no-silent-errors
Every run inside the bound and bit-identical to run 0. The baseline.
silent-data-corruption
Some runs differ from run 0. An intermittent fault, the kind a fleet scanner exists to find.
outside-error-bound
All runs agree with each other but sit outside the proven bound. A faulty unit, or a device accumulating below FP32.
nondeterministic-kernel
Many runs differ, yet all stay inside the bound. The kernel is not repeatable on that device and precision, so the repeat check cannot be used there. Reported as a finding, not a fault.

Memory probe

no-memory-errors
Every word read back correctly on every pattern.
memory-errors
At least one word differed. The measurements say which patterns and the log says where.

Self-test

injection-self-test-pass
Every injected fault was caught and nothing else fired. A row is only published alongside one of these.
injection-self-test-fail
The tester missed something it planted. The row is not publishable until that is understood.
LEARN, PART 5 OF 5

Glossary

silent error, silent data corruption, SDC
A wrong result from a chip with no crash, no error flag and no log entry. The chip does not know it happened.
bit flip
One binary digit changing value. In a floating point number a flipped high bit can change the value by orders of magnitude; a flipped low bit changes it by less than one part in a million.
ECC
Error-correcting code. Extra bits stored with memory that let the hardware detect and usually correct a single flipped bit. ECC covers memory, not the arithmetic units, and not every chip has it on all memory.
HBM
High bandwidth memory, the stacked memory next to a data centre GPU die. An H100 has 80 GB of it.
tensor core
A unit on NVIDIA GPUs that multiplies small matrices in one step. fp16, bf16 and int8 matrix multiplies run through tensor cores; fp32 runs through the regular floating point units.
fp32, fp16, bf16, int8
Number formats. fp32 keeps 24 bits of precision, fp16 keeps 11, bf16 keeps 8 with a wider range, int8 is an exact integer. Models are served in the smaller formats because they are faster, which is why the tester checks all four.
accumulation
The running sum inside a dot product. Even when the inputs are fp16, the sum is normally kept in fp32, which is what the reference bound assumes and what the tester enforces on CUDA.
deterministic kernel
A GPU routine that gives bit-identical output every time for the same input. Not all do, which is why non-determinism is reported as its own verdict instead of as a fault.
reference bound
The largest error that correct rounding can produce for a given computation and precision. Anything larger is a wrong answer. The bound used here is the standard componentwise bound for inner products from Higham.
fault injection
Deliberately corrupting a value to confirm the checks notice. Used here as a self-test on every published row.
dwell
The pause between writing memory and reading it back. Weak cells lose their value over time, so a longer dwell catches more of them.
OCP Test and Validation
The Open Compute Project's open format for hardware diagnostic output, used by cloud operators so that any test's results can be read by any fleet tool. Both probes write it.
EVIDENCE

Measured results

Every row comes from a result file in the repository, and the tables are generated from those files. Nothing here is modelled or estimated.

Compute probe

DeviceDateShapes (M×K×N)PrecisionsRunsSilent faultsInjected caughtVerdictFiles
NVIDIA H100 80GB HBM3PyTorch 2.8.0+cu128 · tool v0.2.0 2026-10-07 1024x1024x1024
4096x4096x4096
32x4096x11008
fp32, fp16, bf16, int8 600 0 60 of 60 no-silent-errors clean · self-test
NVIDIA A100-SXM4-80GBPyTorch 2.8.0+cu128 · tool v0.2.0 2026-10-07 1024x1024x1024
4096x4096x4096
32x4096x11008
fp32, fp16, bf16, int8 600 0 60 of 60 no-silent-errors clean · self-test
NVIDIA H200PyTorch 2.8.0+cu128 · tool v0.2.0 2026-10-08 1024x1024x1024
4096x4096x4096
32x4096x11008
fp32, fp16, bf16, int8 600 0 60 of 60 no-silent-errors clean · self-test
Apple GPU (MPS) arm64PyTorch 2.14.1 · tool v0.2.0 2026-10-07 1024x1024x1024
4096x4096x4096
32x4096x11008
fp32, fp16, bf16 (3 skipped) 450 0 45 of 45 no-silent-errors clean · self-test

Memory probe

DeviceDateMemory testedPatternsDwellBad wordsInjected caughtVerdictFiles
NVIDIA H100 80GB HBM3tool v0.1.0 2026-10-07 62.75 GiB of 79.18 GiB16,844,324,864 words 6 2.0 s 0 30 of 30 no-memory-errors clean · self-test
NVIDIA A100-SXM4-80GBtool v0.1.0 2026-10-07 63.00 GiB of 79.25 GiB16,911,433,728 words 6 2.0 s 0 30 of 30 no-memory-errors clean · self-test
NVIDIA H200tool v0.1.0 2026-10-08 111.25 GiB of 139.81 GiB29,863,444,480 words 6 2.0 s 0 30 of 30 no-memory-errors clean · self-test
Apple GPU (MPS) arm64tool v0.1.0 2026-10-07 5.50 GiB of 11.84 GiB1,476,395,008 words 6 2.0 s 0 30 of 30 no-memory-errors clean · self-test

Rented data centre GPUs are young, monitored and already screened, so clean rows are the expected result there. The research says faults concentrate in aged units and appear under data-dependent, long, hot workloads. Those are the next probes on the roadmap, and where the first real findings are expected.

EVIDENCE

The self-test

A clean table is only worth something if the tester can show it would have seen a fault. So every row is paired with an injection run.

Compute self-test

In five of the fifty runs, one random bit of one random output element is flipped before the checks run. The bit can be anywhere: sign, exponent or the lowest mantissa bit. All five must be caught. On the chips measured so far, the reference check alone caught 46 of every 60 flips and the repeat check caught all 60. That is why there are two checks.

Memory self-test

After each pattern is written, five random bits in five random words of real device memory are flipped. The sweep must locate all of them by address. Because the flips happen in device memory, this exercises the actual read-back path, not only the comparison code.

EVIDENCE

What a PASS means, and does not

  • It means: no silent error appeared in these matrix multiplies and this memory sweep, in these runs, on this device, on this day.
  • It does not mean: the chip is healthy. Faults that depend on specific data, on heat, or on hours of load are not yet exercised. The next probes add them.
  • Accumulation: the compute bound assumes FP32 or better accumulation. That is enforced on CUDA. Elsewhere a lower precision would show as outside-error-bound on every run, and the message says so.
  • One device per run: links between chips are not yet tested.
  • Self-test scope: the compute injection corrupts the output after it leaves the device, so it proves the checker, not the chip. The memory injection corrupts device memory itself.
TAKE PART

Run it on your chip

Four commands, four result files, one pull request. Works on CUDA, Apple MPS and CPU today.

# once
git clone https://github.com/tech4biz-yasha/ai-chip-integrity.git && cd ai-chip-integrity
pip install -r requirements.txt

# compute probe and its self-test
python screen.py --out NAME_clean.jsonl
python screen.py --inject 5 --out NAME_inject.jsonl

# memory probe and its self-test
python memcheck.py --out NAME_memcheck.jsonl
python memcheck.py --inject 5 --out NAME_memcheck_inject.jsonl
  1. Open a pull request with the four files, the device name, driver version and PyTorch version.
  2. Rows are added only from attached files. The table is regenerated from them, so a row cannot say more than its file does.
  3. Found a fault? Open an issue with the file. A real finding on real hardware is the most useful thing anyone can contribute.

The output follows the OCP Test and Validation spec v2.0 through the official ocptv library, and the test suite validates it against the published schema.

TAKE PART

Roadmap

Each probe targets one part of the chip. The order follows where the research says faults hide.

Compute probe, three shapes, four precisionsMatrix multiply with proven bound and bit-exact repeat. Done.
Memory probe, whole deviceSix patterns, exact read-back, address and bit mask on every miss. Done.
Real model kernelsAttention, softmax, layer norm and a full small-LLM forward pass against a float64 reference. The data-dependent probe.
Hours under loadThe same probes for hours at full power with temperature logged next to every result.
Fleet modeRun every probe across many devices automatically and collect the files. Fault rates are per thousand units, so volume is how they get found.
Fault injection inside the computationFlip bits during the multiply on the device, to measure how sensitive each chip and precision is. Also the basis for a space radiation column.
Between devicesAll-reduce over NVLink and PCIe with checksums.
Vendor comparisonRun NVIDIA DCGM and AMD AGFHC on the same device and record what each catches.
TAKE PART

Research and standards this builds on

  • Zheng et al., OSDI 2026. SDCs in the Wild: characterising and diagnosing SDC-defective GPUs in production LLM training. The source of the finding that microbenchmark screening misses most faulty units. usenix.org
  • Lei et al., OSDI 2026. Safeguarding LLM Training at Scale: online SDC detection and insights from 35 million GPU hours. usenix.org
  • Higham, Accuracy and Stability of Numerical Algorithms. The componentwise rounding error bound for inner products used by the reference check.
  • Huang and Abraham, 1984. Algorithm-based fault tolerance for matrix operations. The original idea of checking a matrix result against a cheap independent computation.
  • OCP Test and Validation. The open diagnostic output format and reference implementation both probes write to. ocp-diag-core
TAKE PART

About and cite

Chip Integrity is maintained by Yasha Khandelwal, Bengaluru. The code is MIT licensed. The project stays vendor neutral: the same tests run on every chip, and a row is published only with its result file.

Yasha Khandelwal. ai-chip-integrity: an open tester for silent computation errors in AI chips.
Version 0.2.0, 2026. https://github.com/tech4biz-yasha/ai-chip-integrity

Questions, rows and findings: yasha.khandelwal@tech4biz.io or an issue on GitHub.