flashinfer.fused_moe.PrimsTsRunner¶
- class flashinfer.fused_moe.PrimsTsRunner(config: MoEConfig, device: device)¶
Unified adapter over the inner
PrimsTs*MoERunnerclasses.Routing and finalize stay on the TRT-LLM Gen module loaded by
_TrtllmRunnerBase._build. The middle GEMM is the Prims-TS CuTe-DSL batched path. Weight/activation layouts match the corresponding TRT-LLM prepare helpers so oneMoEWeightPackview can be registered under both"prims_ts"and"trtllm_fp4_routed"/"trtllm_bf16_routed".Methods
__init__(config, device)build()check_support()forward(inputs[, tactic, do_preparation])Forward pass for tunable runners.
get_cache_key_extras(inputs)Return extra values to include in the autotune cache key.
get_valid_tactics(inputs, profile)One tactic corresponding to one cuda kernel normally, but how to interpret the meaning of tactic is pure internal details of the runner.
launch_kwargs_for(inputs)Return only the launch kwargs required by this packed call.
launch_state_for(inputs)Return immutable per-call launch metadata paired with
inputs.pack_inputs(act, weights)precompile_tactics(inputs, tactics, profile, ...)Optionally perform compile-only work for all tactics before profiling.
supports_quant(quant)Return whether this runner can execute
quant's three format axes.tuning_config_for(inputs)Return the tuning config paired with
inputswhen one is present.Attributes
backend_keysupported_activation_classessupported_activation_classes_by_quantsupported_output_formatssupported_quant_variantssupported_routing_modessupports_expert_parallelismsupports_fused_shared_expertssupports_nvfp4_4over6config