flashinfer.fused_moe.CuTileMxfp4Mxfp8Config¶
- class flashinfer.fused_moe.CuTileMxfp4Mxfp8Config¶
cuTile MXFP4 weights with MXFP8 inputs to both GEMMs.
BF16 API inputs are dynamically quantized. Shares packed weights and scales with
CuTileMxfp4ConfigandCuTileMxfp4Bf16Configon SM12x. Expert parallelism and fused shared experts are not supported.- __init__() None¶
Methods
__init__()prepare_weights(w1_fp4, w1_block_scale, ...)Build the shared
cutile_mxfp4weight view.supported(arch)