flashinfer.plan_cake_fmha_request_ordered_paged_decode

flashinfer.plan_cake_fmha_request_ordered_paged_decode(kv_lens: Sequence[int], q_len: int, *, request_order_case: Literal['identity', 'length_desc'] = 'length_desc', real_batch_size: int | None = None, write_lse: bool = False) → CakeFmhaRequestOrderedDecodePlan

Select an exported schedule from host-visible immutable metadata.

The returned plan contains no device data. Call this before CUDA Graph capture, then update the contents of the device request_order tensor in place between replays.