flashinfer.fused_moe.CuTileFp8PerTensorBf16Config¶
- class flashinfer.fused_moe.CuTileFp8PerTensorBf16Config¶
cuTile per-tensor E4M3 weights with BF16 inputs to both GEMMs.
Shares prepared weights with
CuTileFp8PerTensorConfig. Expert parallelism and fused shared experts are not supported.- __init__() None¶
Methods
__init__()prepare_weights(w1_fp8, w1_scale, w2_fp8, ...)Prepare E4M3 weights with one FP32 dequantization scale per expert.
supported(arch)