flashinfer.diffusion_ops.quantize_minimax_h3_qkv_weight_nvfp4¶
- flashinfer.diffusion_ops.quantize_minimax_h3_qkv_weight_nvfp4(qkv_weight: Tensor) Tuple[Tensor, Tensor, Tensor]¶
FlashInfer NVFP4 quantization of the BF16
[21504, 5376]fused QKV weight.Returns
(qkv_weight_q uint8 [21504, 2688], qkv_weight_sf uint8 (128x4 swizzled layout), qkv_weight_global_scale float32 [1])exactly asflashinfer.fp4_quantize()withsf_vec_size=16andis_sf_swizzled_layout=Trueproduces them.