flashinfer.fused_moe.AlphaMoERoutePlan¶
- class flashinfer.fused_moe.AlphaMoERoutePlan(topk_weights: torch.Tensor, topk_ids: torch.Tensor, sorted_token_ids: torch.Tensor, expert_ids: torch.Tensor, num_tokens_post_padded: torch.Tensor, expert_counts: torch.Tensor, expert_offsets: torch.Tensor, expert_scatter_offsets: torch.Tensor)¶
Device-resident routing outputs and reusable private workspace.
The first five fields are the route-plan ABI consumed by AlphaMoE compute kernels.
expert_counts,expert_offsets, andexpert_scatter_offsetsare implementation-owned workspace retained in the tuple so repeated calls can avoid allocation. Consumers must use only the prefix ofsorted_token_idsnamed bynum_tokens_post_padded[0]and the corresponding prefix ofexpert_ids; reading that device scalar on the host is not required to launch a compatible compute kernel.- __init__()¶
Methods
__init__()count(value, /)Return number of occurrences of value.
index(value[, start, stop])Return first index of value.
Attributes
expert_countsAlias for field number 5
expert_idsAlias for field number 3
expert_offsetsAlias for field number 6
expert_scatter_offsetsAlias for field number 7
max_padded_pairsAllocated number of padded token/route entries.
max_route_blocksAllocated number of route blocks.
num_tokens_post_paddedAlias for field number 4
sorted_token_idsAlias for field number 2
topk_idsAlias for field number 1
topk_weightsAlias for field number 0