flashinfer.min_block_table_width¶
- flashinfer.min_block_table_width(seq_len: int, block_size: int) int¶
Return the minimum
block_tablescolumn count for a sequence length.This is the natural width
ceil(seq_len / block_size)– one entry per KV block the sequence occupies, exactly the table a paged-KV serving stack already keeps per request. The kernels predicate every block-table read on the row’s own block count, so no padding columns are needed and entries past a sequence’s blocks are never read (they may hold anything).Kept as a helper so callers name the contract in one place rather than hard-coding the rule; a future kernel that needs a wider table can change it here without touching call sites.
Allocate
block_tablesas:[batch_size, min_block_table_width(int(seq_lens.max()), block_size)]
(using
max_seq_lenas the bound also works; see the Example infp8_paged_mqa_logits()).