flashinfer.fused_moe.prepare_nvfp4_w1_scales

flashinfer.fused_moe.prepare_nvfp4_w1_scales(w1_scale)

Return native CP-layout U8 panels as [E*(N/128)*(K/256),16,128].

w1_scale keeps the original contiguous [E,N,K/16] E4M3 ABI and gate/up order. The result is a separately owned tensor on the same device. For row r=32*g+8*c+d and four-byte word k, destination byte offset inside its 2048-byte panel is ((4*c+k)*8+d)*16+4*g.