flashinfer.fused_moe.allocate_alphamoe_route_plan¶
- flashinfer.fused_moe.allocate_alphamoe_route_plan(logits: Tensor, *, top_k: int, block_m: int = 8, has_shared_expert: bool = False) AlphaMoERoutePlan¶
Allocate a reusable
AlphaMoERoutePlanforlogits.This allocation helper does not launch the router. Pass the result back as
plan=toalphamoe_fused_router(); that form is CUDA-graph-safe after the JIT module has been loaded once outside capture.