flashinfer.diffusion_ops.quantize_minimax_h3_qkv_weight_fp8¶
- flashinfer.diffusion_ops.quantize_minimax_h3_qkv_weight_fp8(qkv_weight: Tensor, chunk_rows: int = 2048) Tuple[Tensor, Tensor]¶
Per-output-channel E4M3 quantization of the BF16
[21504, 5376]fused QKV weight.Returns
(qkv_weight_q float8_e4m3fn [21504, 5376], qkv_weight_scale float32 [21504])withqkv_weight_scale[n] = RN(amax(row n) / 448)andqkv_weight_q = RN(row / scale).