For each token, a linear layer computes y = Wx + b. Huang and Abraham’s 1984 insight is that sums pass straight through a matrix multiply.24 Split the outputs into blocks B of a few hundred. The sum of a block’s outputs is then fixed by the input alone:
cB is the sum of the block’s weight rows, a single vector. Flipless works it out once, when the model loads. At run time it compares the actual sum with that prediction:
With no fault, rB is only rounding noise. If a bit flips in any output of the block, rB jumps by the size of the error.
How big can honest rounding get?
In floating point, a dot product of length n can be off by at most25
where u is the unit roundoff of the arithmetic. A GPU matrix multiply adds up its products in fp32, where u = 2−24, and rounds each output to fp16 once, which moves it by at most 2−11 of its size. Put those together over a block, counting the check’s own fp32 sums too, and to first order the worst honest error, in the units of the threshold below, is
That’s 1.2 for most of Qwen2.5-0.5B’s layers and 2.2 for the down projection, where n = 4,864. In practice the roundings point in different directions and mostly cancel, so the error grows like √n·u rather than n·u.26 Flipless doesn’t use the worst case. It sets each block’s threshold from the inputs it actually sees, with a scale κ measured per layer on clean prompts and multiplied by eight for safety:
On Qwen2.5-0.5B the calibrated κ is 0.11 for a typical layer and 1.62 at most, with the eight-times margin included. So a typical threshold is about ten times tighter than the worst case. That works because real rounding noise is far smaller still: in the NumPy twin of the check, the largest clean error came to 0.005 in these units. The thresholds raised no false alarms in 15.2 million checks.
Why a big error can’t hide
Suppose one output picks up an error e. The residual is that error plus the usual noise η, and calibration keeps η under the threshold. So:
Any single error bigger than twice the threshold fails the check. We tested exactly that on a layer shaped like Qwen’s MLP, with 960 single-bit flips in emulated fp16: every flip that moved a value by more than 2τ was caught, no flip ever failed a block other than its own, and every flip that slipped through moved its value by at most 1.5τ.
Finding the exact value
A second checksum weights each output by its position i in the block. One error e at position p moves the plain residual by e and the weighted residual by p·e, so their ratio points at the culprit:
The prototype measures this during fault-injection runs. In the 1,500-trial run it pinpointed the flipped value in 725 of the 927 caught flips in layer outputs. In normal running Flipless simply replays the token, which fixes the error wherever it is, so finding the exact spot is a bonus, not a requirement.