flashinfer.fused_moe.CuTileMxfp8Config¶
- class flashinfer.fused_moe.CuTileMxfp8Config¶
cuTile MXFP8 weights and activations, with E8M0 scales per 32 values.
Inputs are BF16; both GEMM inputs are dynamically quantized to MXFP8. Expert parallelism and fused shared experts are not supported.
- __init__() None¶
Methods
__init__()prepare_weights(w1_fp8, w1_scale, w2_fp8, ...)Prepare E4M3 weights with logical
[E, N, K/32]E8M0 scales.supported(arch)