flashinfer.fused_moe.CuTileFp8PerTensorBf16Config

class flashinfer.fused_moe.CuTileFp8PerTensorBf16Config

cuTile per-tensor E4M3 weights with BF16 inputs to both GEMMs.

Shares prepared weights with CuTileFp8PerTensorConfig. Expert parallelism and fused shared experts are not supported.

__init__() → None

Methods

__init__()

prepare_weights(w1_fp8, w1_scale, w2_fp8, ...)

Prepare E4M3 weights with one FP32 dequantization scale per expert.

supported(arch)