flashinfer.diffusion_ops.quantize_minimax_h3_o_weight_nvfp4

flashinfer.diffusion_ops.quantize_minimax_h3_o_weight_nvfp4(o_weight: Tensor) → Tuple[Tensor, Tensor, Tensor]

FlashInfer NVFP4 quantization of the BF16 [5376, 7168] attention output weight.

Parameters:

o_weight (torch.Tensor) – BF16 (or any float) [5376, 7168] attention output projection weight on a CUDA device.

Returns:

(o_weight_q, o_weight_sf, o_weight_global_scale): uint8 [5376, 3584] E2M1x2 codes, uint8 UE4M3 block-16 scales in the 128x4 swizzled layout and the FP32 [1] global scale 448 * 6 / amax(W), exactly as flashinfer.fp4_quantize() with sf_vec_size=16 and is_sf_swizzled_layout=True produces them.

Return type:

Tuple[torch.Tensor, torch.Tensor, torch.Tensor]