flashinfer.diffusion_ops.quantize_minimax_h3_o_weight_fp8¶
- flashinfer.diffusion_ops.quantize_minimax_h3_o_weight_fp8(o_weight: Tensor, chunk_rows: int = 1792) Tuple[Tensor, Tensor]¶
Per-output-channel E4M3 quantization of the BF16
[5376, 7168]attention output weight.- Parameters:
o_weight (torch.Tensor) – BF16 (or any float)
[5376, 7168]attention output projection weight on a CUDA device.chunk_rows (int) – Rows quantized per chunk (bounds the FP32 temporary; the result does not depend on it).
- Returns:
(o_weight_q, o_weight_scale): float8_e4m3fn[5376, 7168]codes and FP32[5376]dequant multipliers witho_weight_scale[n] = RN(amax(row n) / 448)ando_weight_q = RN(row / scale).- Return type:
Tuple[torch.Tensor, torch.Tensor]