flipless v0.4 · 47 of 47 tests pass

Space flips bits Flipless flips them back

AI is moving into orbit on ordinary GPUs. Up there, radiation flips bits in memory and in the middle of calculations, and one flipped bit can change what a model says. Flipless checks every matrix multiply and norm layer while the model runs, catches the flip and repairs it before the answer goes out. It’s software, built to run on the commercial GPUs now being launched.

Simulation of 12 GPU satellites in low orbit. Click the orbit to fire a particle.

Fleet logFlips 5Caught 4Wrong answers 0

    Results so far

    268 of 268
    bit flips that changed a model’s answer were stopped
    0
    right answers turned wrong by Flipless, in 2,100 trials
    15.2 million
    checks run, with no false alarms
    98.6%
    lower bound on the share of answer-changing flips stopped, at 95% confidence

    Single-bit flips injected in software into Qwen2.5-0.5B-Instruct on an RTX 4060 Laptop GPU, 7 and 8 October 2026, sign and exponent bits only. Every run is on the Evidence page.

    One bit is enough

    Most numbers inside a model are stored in 16 bits. In fp16, a common format, that’s one bit for the sign, five for the exponent and ten for the digits. The exponent bits set the scale, so they’re the dangerous ones. Flip the top one and 0.0123 becomes 806. The next layer takes 806 at face value, and the error spreads through everything after it.

    Click any bit to flip it.

    In our runs, 112 of 900 exponent flips in a layer’s output changed the model’s answer. None of 194 sign flips did. NVIDIA and UBC researchers found the same pattern in 2017: in neural networks, it’s the high exponent bits that do the damage.21

    Compute is leaving the planet

    AI runs on electricity, and data centres are on course to use around 945 TWh a year by 2030, more than double 2024.49 In the right orbit a solar panel collects up to eight times more energy a year than one on the ground, and the sunlight barely stops.4, 5 So the biggest names in AI are building for space.

    1. Starcloud-1 launches with an NVIDIA H100, which NVIDIA said would bring “100x more powerful GPU compute than any previous space-based operation”. In orbit it later trains a small model on Shakespeare and runs Google’s Gemma.35, 36

    2. Google announces Project Suncatcher: solar-powered satellites carrying its TPUs.4

    3. Axiom Space launches its first two orbital data centre nodes.46

    4. SpaceX merges with xAI. Elon Musk: “within two to three years, the lowest cost way to generate AI compute will be in space.”42

    5. NVIDIA launches space computing at GTC, with a Space-1 Vera Rubin Module to follow. Jensen Huang: “Space computing, the final frontier, has arrived.”39

    6. SpaceX’s S-1 filing says it expects to start deploying orbital AI compute satellites “as early as 2028”.41

    7. SpaceX and NVIDIA design the Starmind AI1 payload around Rubin GPUs. Starcloud is valued at $2.3 billion, with NVIDIA among its new investors.43, 44, 38

    8. Google’s first Suncatcher prototype reaches orbit and starts operating.45

    Starcloud and Google are already flying chips built for data centres on the ground, an H100 and Google’s TPUs, and SpaceX plans to fly Rubin GPUs.

    It already happens on the ground

    Silent data corruption isn’t only a space problem. In every large fleet, a small share of chips compute wrong answers without raising an error, through manufacturing defects, wear or stray neutrons. Fleet scanners test machines between jobs. Flipless checks the real workload while it runs.

    • 1 in 1,000server devices in Meta’s fleet silently corrupt data.11
    • Every week or twois how often Google expected silent corruption to hit Gemini’s training.15
    • 15 of 18silent corruption incidents in 35 million GPU-hours of LLM training caused no visible failure.16
    • Over 60%of defective GPUs in another 2026 study were missed by synthetic tests.17

    How Flipless works

    1. Check every multiply

      Each matrix multiply carries a checksum built from its weights when the model loads. If a block of outputs doesn’t add up to what the inputs predict, a flag goes up.

    2. Double-check the rest

      Norm layers run twice and must match bit for bit. The KV cache keeps a spare copy and is compared with it before every read.

    3. Roll back and replay

      Flags stay on the GPU, so the model never pauses. Once per token Flipless reads them, and if one is up it rolls the token back and runs it again.

    4. Repair

      During the replay, corrupted weights come back from a clean copy, and in version 0.4 corrupted cache rows come back from their spare. Then the token goes out.

    Read the maths and try the checksum yourself

    Who it’s for

    NVIDIA

    Its GPUs are the ones going up. A reliability layer in software makes every one of them easier to sell for orbit.

    SpaceX and xAI

    Orbital AI compute as early as 2028, at Starlink scale, where bit flips stop being rare events and become a rate.

    Google

    Its own beam test saw silent data corruption in a TPU, with no error flag raised.

    Starcloud and the start-ups

    They fly commercial GPUs and need reliability they can buy rather than build.

    AI labs and clouds

    Meta finds about one server device in a thousand corrupting results quietly. Flipless checks the real workload while it runs.

    Space agencies

    Perseverance’s hardened RAD750 runs at up to 200 MHz. Flipless lets missions use modern GPUs with checks they can audit.

    Why each of them needs it

    One line to harden a model

    Flipless wraps a Hugging Face model in place. There’s no retraining and no change to the weights, and when nothing goes wrong the answers are identical, token for token.

    python
    from flipless import Runtime, calibrate, harden
    
    model = harden(model, Runtime(mode="deferred"))
    calibrate(model, lambda: [model.generate(**b, max_new_tokens=32)
                             for b in calibration_batches])
    
    out = model.generate(**inputs)   # every layer checked
    print(model.flipless.stats)      # checks, replays, repairs

    What happens next

    1. DoneVersions 0 to 0.3: checksums, replay and norm checks, tested in 1,800 trials.
    2. DoneVersion 0.4: batched checks, CUDA graphs and a guarded KV cache, at 21% overhead.
    3. 15 OctoberAn RTX PRO 6000 Blackwell with 96 GB arrives, for 7B-class models.
    4. ThenFused kernels, attention coverage, and a radiation beam test.

    The full roadmap