It Wasn't the Hardware
2026-09-02
On the night of September 2 I tried to serve a 195 GB checkpoint, dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4, on two H200s on NCShare. It never ran. I decided early that the reason was the GPUs: NVFP4 is a Blackwell format, the H200 is sm90, so the weights had nowhere to go. I wrote that into a memory note, a README and a preflight check that hard-fails on it. All three are wrong. sglang 0.5.18 has a Marlin fallback for NVFP4 on sm80 through sm90, and the thing that actually killed the run was that no version of sglang, release or main, knew what a glm5_next model was. The checkpoint died at config parsing, and the check I built to catch it looks at the wrong thing.
What I wrote down
The download was running by 21:31. It was 120 shards, and by 21:49 all of them were on /work. While it ran I was building rsi, a self-improvement loop meant to sit on top of this model, and at 21:42 the assistant in that session read sinfo, saw gpu:h200:8 on every node, and said the NVFP4 weights were "a Blackwell-native format. That's the key risk."
That became the working theory. At 22:17 rsi went public as a single commit, 0fbfee3. It includes python/rsi_tools/preflight.py, whose first line is "Say why the model will not serve here, before Slurm spends 40 minutes proving it." check_capability, at preflight.py:146-160, fails hard when the GPU is below what the quant format needs, and serve.py:69-71 says NVFP4 needs (10, 0). The other checks cover the model dir, shards, shard sizes, whether sglang imports, and VRAM. None of them asks whether sglang supports the checkpoint's model_type.
At 22:18 I saved the memory note. It says NVFP4 tensor cores are sm100, "so either sglang refuses the checkpoint at load, or it dequantizes to bf16, which needs all 8 GPUs." The README says the same thing under "On NVFP4 and Hopper" and adds that preflight reports the gap as a hard failure.
What the live job showed
At 22:40 job 720737 started on two H200s. nvidia-smi inside it reported NVIDIA H200, 9.0 on both cards, 143771 MiB each. The venv had sglang 0.5.18 and torch 2.13.0+cu130. So this was exactly the setup my note said couldn't work.
Then I read the installed sglang instead of reasoning about it. ModelOptNvFp4FusedMoEMethod in modelopt_quant.py sets use_marlin_fallback = (8, 0) <= capability < (10, 0). The dense path's error message tells you what to do on Hopper: "Use --fp4-gemm-backend marlin on SM80-SM90." And in this checkpoint only the routed experts are 4-bit. The quant config's ignore list has 32 patterns covering lm_head, the embeddings, the attention projections and the dense and shared MLPs. At 22:54 the session said it plainly: NVFP4 is servable on Hopper here, and the memory note predates this.
At 22:57 I launched sglang inside the allocation. At 23:00 sglang-720737.log ended with this:
ValueError: The checkpoint you are trying to load has model type `glm5_next` but Transformers does not recognize this architecture.
It never reached the weights. The traceback goes through sglang/srt/configs/model_config.py into transformers' configuration_auto.py. The config's architectures field is Glm5NextForConditionalGeneration and auto_map is None, so --trust-remote-code had nothing to load. A grep for glm5_next or Glm5Next over the installed sglang came back empty. Installed transformers was 5.12.1, with a dozen glm* model dirs and no glm5_next. pip index versions sglang said 0.5.18 was also the latest. The sglang repo on main had glm4, glm4_moe, glm4_moe_lite and friends in its models dir, and nothing for GLM-5. transformers main did have glm5_next, which didn't help, because sglang would still need its own model implementation.
The earlier optimism had its own mistake in it. At 21:29 the assistant found glm5 mentioned in sglang's deepseek_v2.py and called the architecture supported. That was GLM-5, not glm5_next. The two claims I leaned on that night were both backwards: the one I trusted about the architecture was false, and the one I feared about the hardware was false too.
The switch
At 23:04 GLM-4.5-Air-FP8 was downloading at 330 MB/s. At 23:06 I committed 8502aa2, "Serve GLM-4.5-Air instead of GLM-5.3-Flash, and fix the CUDA_HOME probe." The message says sglang has no glm5_next model, on 0.5.18 or on main, and the checkpoint ships no remote modeling code, so the NVFP4 weights can't be served at all. The comment in cluster/sglang.sh is sharper: "The NVFP4 quant was never the blocker." Job 720737 expired the same minute with nothing loaded.
That commit also added cluster/model-info.py, which reads the shard count from the index and turns on --quantization modelopt_fp4 and --fp4-gemm-backend marlin only for NVFP4 below (10, 0), and --attention-backend dsa only when the config has deepseek_sparse_attention or index_topk. So the launcher got fixed to read the checkpoint. The preflight didn't.
Why the preflight is the part that matters
The note is one file. The README is prose. The preflight is code that runs in every rsi_tools slurm job before the server starts, and on failure it prints "preflight failed, releasing the allocation." I ran its check against the real config at sm90 on September 22 and it still says NVFP4 "needs compute capability 10.0; this GPU is 9.0," and that sglang "either refuses the checkpoint or dequantizes it to bf16." The suite passes, 49 passed in 0.08s. Green tests around a wrong claim.
The check gets two cases wrong, in opposite directions. Hand it an NVFP4 checkpoint of an architecture sglang does implement and it will refuse to start on the H200s, even though 8502aa2 says that exact case "would serve fine here." Hand it an FP8 checkpoint of an architecture sglang has never heard of and it passes every check, and the job dies at config parsing the same way 720737 did.
The fix is small. Read model_type from config.json, look for it in the installed sglang, and fail with that reason if it's missing, which is the grep I ran by hand at 23:00. Then take check_capability down from a hard failure to a note about which GEMM backend will be used.
The best case for leaving it alone
The strongest objection is that none of this changes the outcome. The preflight would have failed this checkpoint on sm90 anyway. A capability check that says no to a model that can't run is doing its job even if the message is wrong, and a wrong reason attached to a correct verdict costs nothing on this checkpoint. You could go further. A native NVFP4 GEMM really does need SM100, sglang's own error says so. Marlin never got to run on this model, because the load died before the weights. Being conservative about a format the hardware doesn't natively support is a defensible default, and the assistant said three different things about Hopper within about 90 minutes, so maybe the cautious one was the one to write down.
I don't think that holds up. The verdict was right by accident, and a preflight's whole value is its reason. Its docstring promises to say why the model won't serve here. The next time I pull a checkpoint, the reason decides what I do: an sm90 capability failure sends me looking for Blackwell time or a different quant, and a missing architecture sends me looking for a different model. On September 2 the true reason sent me to GLM-4.5-Air within six minutes of reading the log. The false one had me sizing a bf16 fallback across all eight GPUs of a node. And the conservative default isn't free. It blocks the NVFP4-on-Hopper case the installed sglang explicitly handles, and lets through the unknown-architecture case that actually burned an allocation. Being cautious about the wrong axis just makes the check wrong in a different way.
The memory note still says sm100. The README still says the preflight reports the capability gap as a hard failure, and rsi still has one commit.