flashinfer.fused_moe.prepare_nvfp4_w2_data¶
- flashinfer.fused_moe.prepare_nvfp4_w2_data(w2)¶
Prepare immutable packed W2 panels once during model loading.
The contiguous uint8 input is [E,K,N/4] (K output rows of N/2 packed FP4 intermediate values per expert). The result is [E*(K/128)*(N/256),128,64], an exact byte permutation with no dequantization: panel
(e*(K/128)+ob)*(N/256)+jholds rowsob*128..+128and bytesj*64..+64of experte. Keep the raw weights and this tensor alive and immutable while using the routed operator. Preparation is outside routed execution.