flashinfer.fused_moe.prepare_nvfp4_w2_scales_k256¶
- flashinfer.fused_moe.prepare_nvfp4_w2_scales_k256(w2_scale)¶
Return 16-scale CP-layout U8 panels of the W2 scales as
[E*(K/128)*(N/512),16,128].w2_scalekeeps the original contiguous[E,K,N/32]E4M3 ABI. Each panel carries the scale factors of one 128-output-row x 256-intermediate K-chunk-256 MMA stage in the layoutprepare_nvfp4_w1_scales()uses for W1 (same byte permutation applied to the W2 rows). The result is a separately owned tensor on the same device; preparation is outside routed execution.