flashinfer.fused_moe.prepare_nvfp4_w1_data¶
- flashinfer.fused_moe.prepare_nvfp4_w1_data(w1)¶
Prepare immutable packed W1 panels once during model loading.
The contiguous uint8 input is [E,N,K/2] in the original gate/up row order. The result is [E*(N/128)*(K/256),128,128], an exact byte permutation with no dequantization. Keep the raw weights and this tensor alive and immutable while using the routed operator. Preparation is outside routed execution.