flashinfer.diffusion_ops.quantize_minimax_h3_qkv_weight_fp8

flashinfer.diffusion_ops.quantize_minimax_h3_qkv_weight_fp8(qkv_weight: Tensor, chunk_rows: int = 2048) → Tuple[Tensor, Tensor]

Per-output-channel E4M3 quantization of the BF16 [21504, 5376] fused QKV weight.

Returns (qkv_weight_q float8_e4m3fn [21504, 5376], qkv_weight_scale float32 [21504]) with qkv_weight_scale[n] = RN(amax(row n) / 448) and qkv_weight_q = RN(row / scale).