flashinfer.diffusion_ops.quantize_minimax_h3_o_weight_nvfp4¶
- flashinfer.diffusion_ops.quantize_minimax_h3_o_weight_nvfp4(o_weight: Tensor) Tuple[Tensor, Tensor, Tensor]¶
FlashInfer NVFP4 quantization of the BF16
[5376, 7168]attention output weight.- Parameters:
o_weight (torch.Tensor) – BF16 (or any float)
[5376, 7168]attention output projection weight on a CUDA device.- Returns:
(o_weight_q, o_weight_sf, o_weight_global_scale): uint8[5376, 3584]E2M1x2 codes, uint8 UE4M3 block-16 scales in the 128x4 swizzled layout and the FP32[1]global scale448 * 6 / amax(W), exactly asflashinfer.fp4_quantize()withsf_vec_size=16andis_sf_swizzled_layout=Trueproduces them.- Return type:
Tuple[torch.Tensor, torch.Tensor, torch.Tensor]