flashinfer.fused_moe.prepare_nvfp4_w2_data_k256

flashinfer.fused_moe.prepare_nvfp4_w2_data_k256(w2)

Prepare immutable 128-byte-row W2 panels for the large-M token-tile route.

The contiguous uint8 input is [E,K,N/4] (K output rows of N/2 packed FP4 intermediate values per expert). The result is [E*(K/128)*(N/512),128,128], an exact byte permutation with no dequantization: panel (e*(K/128)+ob)*(N/512)+j holds rows ob*128..+128 and bytes j*128..+128 (256 intermediate values, one K-chunk-256 MMA stage) of expert e. Keep the raw weights and this tensor alive and immutable while using the routed operator. Preparation is outside routed execution.