flashinfer.fused_moe.CuTileNvfp4Bf16Config

class flashinfer.fused_moe.CuTileNvfp4Bf16Config

cuTile NVFP4-weight x BF16-activation backend.

Uses the same prepared weight view as CuTileNvfp4Config. Expert parallelism and fused shared experts are not supported.

__init__() → None

Methods

__init__()

prepare_weights(w1_fp4, w1_block_scale, ...)

Build the cutile_nvfp4 view from checkpoint NVFP4 weights.

supported(arch)