flashinfer.mla.dsv41_fp4_quantize_append_sparse_mla_cache

flashinfer.mla.dsv41_fp4_quantize_append_sparse_mla_cache(latent_kv: Tensor, slot_mapping: Tensor, cache: Tensor) → None

Quantize and append DeepSeek-V4.1 latent KV by physical cache slot.

Parameters:
  • latent_kv (torch.Tensor) – Contiguous CUDA BF16/FP16 tensor with one 512-element latent-KV row per entry in slot_mapping.

  • slot_mapping (torch.Tensor) – Contiguous 1D CUDA int32 or int64 tensor. slot_mapping[i] is page_id * page_size + entry_id. Negative and out-of-range slots are padding and are ignored. If a valid slot occurs more than once, the lowest-index input row is written deterministically.

  • cache (torch.Tensor) – Destination opaque uint8 paged cache (2D [num_pages, page_bytes], 3D [num_pages, page_size, 288], or the 4D HND/NHD forms). Page-strided views may place the cache inside vLLM’s packed physical block allocation. Only addressed data rows and scale slots are written, so prefill history is reused directly by decode.