flashinfer.fused_moe.CuTileMxfp4Mxfp8Config

class flashinfer.fused_moe.CuTileMxfp4Mxfp8Config

cuTile MXFP4 weights with MXFP8 inputs to both GEMMs.

BF16 API inputs are dynamically quantized. Shares packed weights and scales with CuTileMxfp4Config and CuTileMxfp4Bf16Config on SM12x. Expert parallelism and fused shared experts are not supported.

__init__() → None

Methods

__init__()

prepare_weights(w1_fp4, w1_block_scale, ...)

Build the shared cutile_mxfp4 weight view.

supported(arch)