flashinfer.diffusion_ops.quantize_minimax_h3_o_weight_fp8

flashinfer.diffusion_ops.quantize_minimax_h3_o_weight_fp8(o_weight: Tensor, chunk_rows: int = 1792) → Tuple[Tensor, Tensor]

Per-output-channel E4M3 quantization of the BF16 [5376, 7168] attention output weight.

Parameters:
  • o_weight (torch.Tensor) – BF16 (or any float) [5376, 7168] attention output projection weight on a CUDA device.

  • chunk_rows (int) – Rows quantized per chunk (bounds the FP32 temporary; the result does not depend on it).

Returns:

(o_weight_q, o_weight_scale): float8_e4m3fn [5376, 7168] codes and FP32 [5376] dequant multipliers with o_weight_scale[n] = RN(amax(row n) / 448) and o_weight_q = RN(row / scale).

Return type:

Tuple[torch.Tensor, torch.Tensor]