flashinfer.fused_moe.prepare_nvfp4_w1_gate_up_scales¶
- flashinfer.fused_moe.prepare_nvfp4_w1_gate_up_scales(prepared_scales, original_weight_shape)¶
Pair existing prepared W1 scale bytes with adjacent gate/up data.
prepared_scalesis the uint8 output ofprepare_nvfp4_w1_scales();original_weight_shapeis the raw W1 shape[E,N,K/2]. The result is[E*(N/256)*(K/256),32,128]and preserves every E4M3 byte. Reuse it with data prepared from the same device-local weight load.