flashinfer.diffusion_ops¶
Fused operators for diffusion-transformer inference.
|
Run the fused BF16 pre-attention projection for MiniMax-H3. |
|
Fused FP8 (W8A8) MiniMax-H3 pre-attention for SM120 (RTX 5090 / RTX PRO 6000 Blackwell). |
|
Fused NVFP4 MiniMax-H3 pre-attention for SM120 (RTX 5090 / RTX PRO 6000 Blackwell). |
|
Per-output-channel E4M3 quantization of the BF16 |
|
FlashInfer NVFP4 quantization of the BF16 |
|
MiniMax-H3 FP8 (W8A8) attention output projection with the fused gated residual for SM120 (GB202: RTX 5090 / RTX PRO 6000 Blackwell). |
|
MiniMax-H3 NVFP4 (W4A4) attention output projection with the fused gated residual for SM120 (GB202: RTX 5090 / RTX PRO 6000 Blackwell). |
|
Per-output-channel E4M3 quantization of the BF16 |
|
FlashInfer NVFP4 quantization of the BF16 |
|
FP8 (E4M3) non-causal packed-varlen self-attention for MiniMax-H3 on SM120 (GB202). |
Experimental NVFP4 non-causal packed-varlen self-attention for MiniMax-H3 on SM120 (GB202), following the SageAttention3 FP4 recipe. |
|
|
Fused FP8 (W8A8) RMSNorm + indexed AdaLN + FC1 GEMM + SwiGLU of the MiniMax-H3 video DiT block for SM120 (RTX 5090 / RTX PRO 6000 Blackwell). |
|
Quantize the fused FC1 weight for |
|
Fused QK RMSNorm + 3D RoPE + V copy for video generation DIT self-attention. |
Fused residual + LayerNorm + scale/shift for DIT self-attention. |
|
Fused gate + residual + LayerNorm + scale/shift for DIT self-attention. |
|
Fused gate + residual + LayerNorm(gamma, beta) for DIT self-attention. |