flashinfer.fused_moe.prepare_nvfp4_w2_scales_k256

flashinfer.fused_moe.prepare_nvfp4_w2_scales_k256(w2_scale)

Return 16-scale CP-layout U8 panels of the W2 scales as [E*(K/128)*(N/512),16,128].

w2_scale keeps the original contiguous [E,K,N/32] E4M3 ABI. Each panel carries the scale factors of one 128-output-row x 256-intermediate K-chunk-256 MMA stage in the layout prepare_nvfp4_w1_scales() uses for W1 (same byte permutation applied to the W2 rows). The result is a separately owned tensor on the same device; preparation is outside routed execution.