flashinfer.fused_moe.allocate_alphamoe_route_plan

flashinfer.fused_moe.allocate_alphamoe_route_plan(logits: Tensor, *, top_k: int, block_m: int = 8, has_shared_expert: bool = False) → AlphaMoERoutePlan

Allocate a reusable AlphaMoERoutePlan for logits.

This allocation helper does not launch the router. Pass the result back as plan= to alphamoe_fused_router(); that form is CUDA-graph-safe after the JIT module has been loaded once outside capture.