Welcome to FlashInfer’s documentation!¶
Blog | Discussion Forum | GitHub
FlashInfer is a library and kernel generator for Large Language Models that provides high-performance implementation of LLM GPU kernels such as FlashAttention, PageAttention and LoRA. FlashInfer focus on LLM serving and inference, and delivers state-of-the-art performance across diverse scenarios.
Get Started
Tutorials
PyTorch API Reference
- FlashInfer Attention Kernels
- flashinfer.attn_scores
- flashinfer.cake_fmha
- flashinfer.cake_sampling
- flashinfer.gemm
- BF16 GEMM
- FP4 GEMM
- SVDQuant NVFP4 GEMM (SM100)
- BF16 x FP4 GEMM (W4A16)
- MXFP8 GEMM
- FP8 GEMM
- Mixed Precision GEMM (fp8 x fp4)
- Router GEMM (DeepSeek-V3 / Mistral / GLM / Kimi-K2 / Kimi-K3)
- Blackwell SM100 GEMM
- cuTile GEMM
- Grouped GEMM (CuTe-DSL, Blackwell)
- Grouped GEMM (Ampere/Hopper)
- PrimsTS dense FP8 and NVFP4 GEMM (SM100, SM103, SM107)
- flashinfer.grouped_mm
- flashinfer.fused_moe
- Types and Enums
- Unified MoE API
- Utility Functions
- AlphaMoE Router (SM100/SM103)
- Multi-LoRA MoE (BGMV)
- CUTLASS Fused MoE
- cuTile Fused MoE
- TensorRT-LLM Fused MoE
- AlphaMoE FP8 Block-Scaled MoE (SM100/SM103)
- AlphaMoE NVFP4 (SM100/SM103)
- AlphaMoE NVFP4 prepared weight scales
- AlphaMoE NVFP4 prepared gate/up data
- AlphaMoE NVFP4 adjacent gate/up panels
- Cake NVFP4 Warp Decode (SM100/SM103)
- Prims-TS Unified MoE (SM100/SM103)
- Prims-TS Fused MoE
- Standalone TRT-LLM Gen Routing
- CuteDSL Fused MoE
- MonoMoE (Single-Kernel Block-FP8, SM90a)
- flashinfer.cascade
- flashinfer.comm
- CUDA IPC Utilities
- DLPack Utilities
- Mapping Utilities
- All-Gather Matmul
- TensorRT-LLM AllReduce
- Unified AllReduce Fusion API
- FP8 Quantized AllReduce
- vLLM AllReduce
- PCIe IPC AllReduce
- PCIe IPC AllGather and ReduceScatter
- Ulysses Context-Parallel All-to-All
- MNNVL (Multi-Node NVLink)
- TensorRT-LLM MNNVL AllReduce
- MNNVL A2A (Throughput Backend)
- DCP All-to-All (Context-Parallel Attention Reduction)
- NCCL LSA DCP All-to-All + LSE Reduce
- Mixed Communication
- flashinfer.sparse
- flashinfer.msa_ops
- flashinfer.msa_ops.msa_proxy_score
- flashinfer.msa_ops.msa_proxy_score_fp4
- flashinfer.msa_ops.MSASparseAttentionWorkspace
- flashinfer.msa_ops.supports_packed_kv
- flashinfer.msa_ops.msa_sparse_attention
- flashinfer.msa_ops.msa_sparse_decode_attention
- flashinfer.msa_ops.msa_topk_select
- flashinfer.msa_ops.prepare_msa_nvfp4_sparse_decode
- flashinfer.pod
- flashinfer.cudnn
- flashinfer.cute_dsl
- flashinfer.concat_ops
- flashinfer.diffusion_ops
- flashinfer.diffusion_ops.minimax_h3_bf16_pre_attention
- flashinfer.diffusion_ops.minimax_h3_fp8_pre_attention
- flashinfer.diffusion_ops.minimax_h3_nvfp4_pre_attention
- flashinfer.diffusion_ops.quantize_minimax_h3_qkv_weight_fp8
- flashinfer.diffusion_ops.quantize_minimax_h3_qkv_weight_nvfp4
- flashinfer.diffusion_ops.minimax_h3_fp8_out_proj
- flashinfer.diffusion_ops.minimax_h3_nvfp4_out_proj
- flashinfer.diffusion_ops.quantize_minimax_h3_o_weight_fp8
- flashinfer.diffusion_ops.quantize_minimax_h3_o_weight_nvfp4
- flashinfer.diffusion_ops.minimax_h3_sm120_varlen_attention_fp8
- flashinfer.diffusion_ops.minimax_h3_sm120_varlen_attention_nvfp4
- flashinfer.diffusion_ops.minimax_h3_fc1_swiglu_fp8
- flashinfer.diffusion_ops.prepare_minimax_h3_fc1_weight_fp8
- flashinfer.diffusion_ops.fused_qk_rmsnorm_rope
- flashinfer.diffusion_ops.fused_dit_residual_layernorm_scale_shift
- flashinfer.diffusion_ops.fused_dit_gate_residual_layernorm_scale_shift
- flashinfer.diffusion_ops.fused_dit_gate_residual_layernorm_gamma_beta
- flashinfer.page
- flashinfer.sampling
- flashinfer.sampling.sampling_from_probs
- flashinfer.sampling.sampling_from_logits
- flashinfer.sampling.softmax
- flashinfer.sampling.top_p_sampling_from_probs
- flashinfer.sampling.top_k_sampling_from_probs
- flashinfer.sampling.min_p_sampling_from_probs
- flashinfer.sampling.top_k_top_p_sampling_from_logits
- flashinfer.sampling.top_k_top_p_sampling_from_probs
- flashinfer.sampling.top_p_renorm_probs
- flashinfer.sampling.top_k_renorm_probs
- flashinfer.sampling.top_k_mask_logits
- flashinfer.sampling.chain_speculative_sampling
- flashinfer.topk
- flashinfer.logits_processor
- flashinfer.norm
- flashinfer.norm.rmsnorm
- flashinfer.norm.rmsnorm_quant
- flashinfer.norm.fused_add_rmsnorm
- flashinfer.norm.fused_add_rmsnorm_quant
- flashinfer.norm.fused_add_rmsnorm_fp8_block_quant
- flashinfer.norm.gemma_rmsnorm
- flashinfer.norm.gemma_fused_add_rmsnorm
- flashinfer.norm.layernorm
- flashinfer.norm.layernorm_quant
- flashinfer.norm.fused_rmsnorm_silu
- flashinfer.norm.fused_qk_rmsnorm_rope
- flashinfer.norm.fused_dit_residual_layernorm_scale_shift
- flashinfer.norm.fused_dit_gate_residual_layernorm_scale_shift
- flashinfer.norm.fused_dit_gate_residual_layernorm_gamma_beta
- flashinfer.rope
- flashinfer.rope.apply_rope_inplace
- flashinfer.rope.apply_llama31_rope_inplace
- flashinfer.rope.apply_rope
- flashinfer.rope.apply_llama31_rope
- flashinfer.rope.apply_rope_pos_ids
- flashinfer.rope.apply_rope_pos_ids_inplace
- flashinfer.rope.apply_llama31_rope_pos_ids
- flashinfer.rope.apply_llama31_rope_pos_ids_inplace
- flashinfer.rope.apply_rope_with_cos_sin_cache
- flashinfer.rope.apply_rope_with_cos_sin_cache_inplace
- flashinfer.rope.rope_quantize_fp8
- flashinfer.rope.rope_quantize_fp8_append_paged_kv_cache
- flashinfer.rope.mla_rope_quantize_fp8
- flashinfer.activation
- flashinfer.gdn_decode
- flashinfer.gdn_fused_decode_step
- flashinfer.gdn_prefill
- flashinfer.gdn2_prefill
- flashinfer.gdp_prefill
- Kimi Delta Attention (KDA)
- flashinfer.mamba
- flashinfer.mhc
- flashinfer.quantization
- flashinfer.green_ctx
- flashinfer.fp4_quantization
- flashinfer.testing