flashinfer.diffusion_ops.quantize_minimax_h3_qkv_weight_nvfp4

flashinfer.diffusion_ops.quantize_minimax_h3_qkv_weight_nvfp4(qkv_weight: Tensor) → Tuple[Tensor, Tensor, Tensor]

FlashInfer NVFP4 quantization of the BF16 [21504, 5376] fused QKV weight.

Returns (qkv_weight_q uint8 [21504, 2688], qkv_weight_sf uint8 (128x4 swizzled layout), qkv_weight_global_scale float32 [1]) exactly as flashinfer.fp4_quantize() with sf_vec_size=16 and is_sf_swizzled_layout=True produces them.