flashinfer.cake_sampling.cake_sampling_route

flashinfer.cake_sampling.cake_sampling_route(probs: Tensor, top_k: int | Tensor | None, top_k_max: int | None = None) → str

"pipeline" when the frozen kernels serve this request, else "fallback:<reason>".

top_k_max avoids a device sync when top_k is a tensor.