Invalid float64 PTX in tinygrad

2026-09-05

At the commit I branched from, line 240 of test/backend/test_dtype.py skips test_float64_increased_precision for any PTXRenderer, with this comment:

"conversion not supported on CI CUDA, PTX, and NIR")  # TODO: why not?

That TODO has been sitting there since at least c5467e5bd, March 2024. For PTX, I have the answer. Every float64 reciprocal, square root and transcendental that tinygrad's PTX renderer emitted was invalid PTX. The numbers weren't a little off. The assembler refused to take the code at all.

How I got to it

I wasn't looking. On September 5 I was checking two bug bounty drafts that an earlier agent session on my box had written, one for a QCOM A630 emulator and one for GPUOcelot on CUDA 12. Both turned out dead. The ocelot draft had exactly one real line in it, rcp.approx.f64 changed to rcp.rn.f64.

I wanted to know if that one line was right. It was, and it was one of five.

The answer, in ptxas's words

asm_for_op in tinygrad/renderer/ptx.py, lines 19 to 21, renders rcp.approx, sqrt.approx, and ex2.approx, lg2.approx and sin.approx for every float dtype without checking which one. At f64, PTX allows none of it. rcp.approx.f64 needs .ftz. sqrt has no .approx form at f64. ex2, lg2 and sin have no f64 form at all.

I picked ten ops (reciprocal, sqrt, rsqrt, exp, log, sin, cos, tanh, a/(a+1) and a**2.5), ran each on Tensor([2.0,4.0]) in f64 and f32 with DEV=MOCK+CUDA:PTX MOCKGPU=1, patched the header to .version 8.0 and .target sm_86, and fed the output to ptxas 12.9.86 with -arch=sm_86. All ten f32 kernels assembled. All ten f64 kernels didn't:

cos_f64.ptx            error   : Unexpected instruction types specified for 'sin'
div_f64.ptx            error   : '.ftz' modifier required for instruction 'rcp.ap
exp_f64.ptx            error   : Unexpected instruction types specified for 'ex2'
log_f64.ptx            error   : Unexpected instruction types specified for 'lg2'
pow_f64.ptx            error   : Illegal modifier '.approx' for instruction 'sqrt
reciprocal_f64.ptx     error   : '.ftz' modifier required for instruction 'rcp.ap
rsqrt_f64.ptx          error   : Illegal modifier '.approx' for instruction 'sqrt
sin_f64.ptx            error   : Unexpected instruction types specified for 'sin'
sqrt_f64.ptx           error   : Illegal modifier '.approx' for instruction 'sqrt
tanh_f64.ptx           error   : Unexpected instruction types specified for 'ex2'

Division is on the list because f64 division lowers to RECIPROCAL. The scope matters here: a recheck on the 11th found f64 add, max, trunc, comparisons and every cast assembling fine, and every other dtype from half through uint64 and bool clean. It's exactly the f64 ops that land on those five instructions. Under gpuocelot, running one of them doesn't raise an exception. It takes Python down with Fatal Python error: Aborted.

Why nobody saw it

Four things stood between that bug and a red CI run.

GuardWhereWhat it did
skipIf(DEV.interface.startswith("MOCK") and Device.DEFAULT in {"NV", "CUDA"}, "crashed")test_transcendental.py lines 18, 29skipped f64 tests by DEV string, whatever the renderer
early return marked # crashes in CI CUDAtest_transcendental.py lines 70, 92same, inline
skip any PTXRenderer, # TODO: why not?test_dtype.py line 240skipped the one test that checks f64 precision
no test_float64_unarytest_dtype_alu.pytest_float32_unary is at 195, test_float16_unary at 199, nothing for f64

The first two are the interesting ones, because the crash they describe was real. The DEV string doesn't tell you the renderer. select_first_inited tries CUDARenderer first, so with nvrtc installed DEV=MOCK+NV gets cstyle, which is what CI does, and with nvrtc missing it quietly falls back to PTX. On the cstyle path gpuocelot really does fail on the bfe.u64 that nvrtc emits for f64. So "crashed" was a true statement about one renderer, written as a skip that covered both. The CI job that does run PTX never rendered an f64 rcp, sqrt or transcendental, because every test that would have was skipped.

The fix

The change in ptx.py is five lines. At double, rcp and sqrt use .rn, and EXP2, LOG2 and SIN get rewritten in ptx_matcher to xexp2, xlog2 and xsin. I didn't invent the second part. AMDLLVMRenderer already decomposes f64 LOG2 and EXP2 because the LLVM intrinsics don't support double, and transcendental.py already had full f64 paths. PTX just never asked for them.

The legal alternative for reciprocal, rcp.approx.ftz.f64, is about 2^-23 relative error and flushes subnormals. That's f32 accuracy on a double. The precision test asserts rtol 1e-12, and nvcc already gives the cstyle renderer an exact 1/x, so .rn makes the two backends agree.

The test side is where I'd have gotten it wrong alone. I ran six critic rounds against the patch. Round 2 caught that deleting the skip would have turned the tests on for the nv job, which runs cstyle and would crash on bfe.u64. So the skip got narrowed instead:

mock_cuda_f64_crash = DEV.interface.startswith("MOCK") and Device.DEFAULT in {"NV", "CUDA"} and not isinstance(Device[Device.DEFAULT].renderer, PTXRenderer)

Round 3 caught something worse. The re-enabled tests forced TRANSCENDENTAL=2, which decomposes everything before ptx_matcher runs, so four of my five changes had zero coverage. I added test_no_approx_float64 in test_renderer_failures.py, a string match that checks .approx.f64 is absent and rcp.rn.f64 and sqrt.rn.f64 are present. The whole thing is branch ptx-f64-rcp, commit 68df7a76a, four files, +26/-11.

Checking it with no GPU

I don't have NVIDIA hardware. pip install nvidia-cuda-nvcc-cu12 gives you a real ptxas, which is a host compiler and runs on NixOS without patchelf. For execution I used gpuocelot. Its shipped libgpuocelot.so needed the PF_X bit cleared on PT_GNU_STACK by hand, since the nixpkgs patchelf was too old for --clear-execstack, plus an LD_LIBRARY_PATH pointing at a gcc lib dir for libstdc++.so.6. To get CI's renderer choice you need nvidia-cuda-nvrtc-cu11, not cu12, because 12.x dropped sm_35 and the mock reports sm_35.

With the fix, all 20 kernels assemble. Under ocelot, f64 reciprocal, sqrt, exp, log and sin come back with 0 to 1.7e-16 relative error. Four tests that crashed pytest on master pass. Reverting any one of the five changes makes a specific subtest fail by name. The ptx job went from 1289 to 1292 passing at the same runtime, and a junit-xml diff of the nv job showed zero changed outcomes. ruff and mypy are clean.

What I can't tell you is how rcp.rn.f64 and the decompositions behave on a real card. That part is reasoned, not measured.

Why it's still on my disk

As of September 22, upstream master still has rcp with .approx and sqrt.approx on lines 19 to 21. I haven't opened a PR.

Here's what's on record. On the 5th the session stopped at the branch because opening a PR is my call, not its. When I asked whether it would merge, the answer was no: role-played critic odds of 12% in round 4 and 40% in round 5, which are guesses, not data. An earlier small PR of mine to tinygrad, #16876, was approved in a comment and closed anyway. The fix is worth $0 and satisfies neither bounty it came out of. On the 11th a scout re-verified it on 4e9a08437, found the bug still live and a minimal +10/-4 version passing, and ranked it below other candidates "because nobody runs f64 on NVIDIA in practice."

That same week I sent #18150, fixing var on int input, which merged, and #18170, the gradient of an assign through a strided slice, which is still open. Both are bugs people run into.