flashinfer.fused_moe.CuTileFp8PerTensorBf16Runner¶
- class flashinfer.fused_moe.CuTileFp8PerTensorBf16Runner(config: MoEConfig, device: device)¶
-
Methods
__init__(config, device)build()check_support()forward(inputs[, tactic, do_preparation])Forward pass for tunable runners.
get_cache_key_extras(inputs)Return extra values to include in the autotune cache key.
get_valid_tactics(inputs, _profile)One tactic corresponding to one cuda kernel normally, but how to interpret the meaning of tactic is pure internal details of the runner.
launch_kwargs_for(inputs)Return only the launch kwargs required by this packed call.
launch_state_for(inputs)Return immutable per-call launch metadata paired with
inputs.pack_inputs(act, weights)precompile_tactics(inputs, tactics, profile, ...)Optionally perform compile-only work for all tactics before profiling.
supports_quant(quant)Return whether this runner can execute
quant's three format axes.tuning_config_for(inputs)Return the tuning config paired with
inputswhen one is present.Attributes
backend_keysupported_activation_classessupported_activation_classes_by_quantsupported_output_formatssupported_quant_variantssupported_routing_modessupports_expert_parallelismsupports_fused_shared_expertssupports_nvfp4_4over6config