flashinfer.fused_moe.prepare_nvfp4_w2_data

flashinfer.fused_moe.prepare_nvfp4_w2_data(w2)

Prepare immutable packed W2 panels once during model loading.

The contiguous uint8 input is [E,K,N/4] (K output rows of N/2 packed FP4 intermediate values per expert). The result is [E*(K/128)*(N/256),128,64], an exact byte permutation with no dequantization: panel (e*(K/128)+ob)*(N/256)+j holds rows ob*128..+128 and bytes j*64..+64 of expert e. Keep the raw weights and this tensor alive and immutable while using the routed operator. Preparation is outside routed execution.