flashinfer.diffusion_ops

Fused operators for diffusion-transformer inference.

minimax_h3_bf16_pre_attention(x, ...[, eps])

Run the fused BF16 pre-attention projection for MiniMax-H3.

minimax_h3_fp8_pre_attention(x, ...[, eps, ...])

Fused FP8 (W8A8) MiniMax-H3 pre-attention for SM120 (RTX 5090 / RTX PRO 6000 Blackwell).

minimax_h3_nvfp4_pre_attention(x, ...[, ...])

Fused NVFP4 MiniMax-H3 pre-attention for SM120 (RTX 5090 / RTX PRO 6000 Blackwell).

quantize_minimax_h3_qkv_weight_fp8(qkv_weight)

Per-output-channel E4M3 quantization of the BF16 [21504, 5376] fused QKV weight.

quantize_minimax_h3_qkv_weight_nvfp4(qkv_weight)

FlashInfer NVFP4 quantization of the BF16 [21504, 5376] fused QKV weight.

minimax_h3_fp8_out_proj(attn_out, ...[, ...])

MiniMax-H3 FP8 (W8A8) attention output projection with the fused gated residual for SM120 (GB202: RTX 5090 / RTX PRO 6000 Blackwell).

minimax_h3_nvfp4_out_proj(attn_out, ...[, ...])

MiniMax-H3 NVFP4 (W4A4) attention output projection with the fused gated residual for SM120 (GB202: RTX 5090 / RTX PRO 6000 Blackwell).

quantize_minimax_h3_o_weight_fp8(o_weight[, ...])

Per-output-channel E4M3 quantization of the BF16 [5376, 7168] attention output weight.

quantize_minimax_h3_o_weight_nvfp4(o_weight)

FlashInfer NVFP4 quantization of the BF16 [5376, 7168] attention output weight.

minimax_h3_sm120_varlen_attention_fp8(q, k, ...)

FP8 (E4M3) non-causal packed-varlen self-attention for MiniMax-H3 on SM120 (GB202).

minimax_h3_sm120_varlen_attention_nvfp4(q, ...)

Experimental NVFP4 non-causal packed-varlen self-attention for MiniMax-H3 on SM120 (GB202), following the SageAttention3 FP4 recipe.

minimax_h3_fc1_swiglu_fp8(x, x_norm_weight, ...)

Fused FP8 (W8A8) RMSNorm + indexed AdaLN + FC1 GEMM + SwiGLU of the MiniMax-H3 video DiT block for SM120 (RTX 5090 / RTX PRO 6000 Blackwell).

prepare_minimax_h3_fc1_weight_fp8(fc1_weight)

Quantize the fused FC1 weight for minimax_h3_fc1_swiglu_fp8() (SM120).

fused_qk_rmsnorm_rope(qkv, q_weight, ...[, ...])

Fused QK RMSNorm + 3D RoPE + V copy for video generation DIT self-attention.

fused_dit_residual_layernorm_scale_shift(...)

Fused residual + LayerNorm + scale/shift for DIT self-attention.

fused_dit_gate_residual_layernorm_scale_shift(...)

Fused gate + residual + LayerNorm + scale/shift for DIT self-attention.

fused_dit_gate_residual_layernorm_gamma_beta(...)

Fused gate + residual + LayerNorm(gamma, beta) for DIT self-attention.