flashinfer.mla.dsv41_fp4_quantize_pack_sparse_mla_cache

flashinfer.mla.dsv41_fp4_quantize_pack_sparse_mla_cache(latent_kv: Tensor, *, kv_layout: str = 'HND') → Tensor

Quantize complete DeepSeek-V4.1 latent-KV pages to the V41_FP4 ABI.

Parameters:
  • latent_kv (torch.Tensor) – Contiguous CUDA BF16/FP16 tensor with shape [num_pages, page_size, 512]. A singleton latent-head axis is also accepted in HND or NHD position. All 512 values per token (RoPE dims included) are quantized in groups of 16 to packed E2M1 with E4M3 scales (scale = amax/6), mirroring the FlashMLA V41_FP4 trajectory.

  • kv_layout (str) – Output layout, either "HND" or "NHD".

Returns:

Opaque uint8 paged cache with logical shape [num_pages, 1, page_size, 288] for HND or [num_pages, page_size, 1, 288] for NHD. Within each physical page it stores page_size * 256 data bytes followed by page_size * 32 scale bytes. Consumers must not interpret the last dimension as a contiguous per-token record.

Return type:

torch.Tensor