flashinfer.mla.dsv41_fp4_quantize_pack_sparse_mla_cache¶
- flashinfer.mla.dsv41_fp4_quantize_pack_sparse_mla_cache(latent_kv: Tensor, *, kv_layout: str = 'HND') Tensor¶
Quantize complete DeepSeek-V4.1 latent-KV pages to the V41_FP4 ABI.
- Parameters:
latent_kv (torch.Tensor) – Contiguous CUDA BF16/FP16 tensor with shape
[num_pages, page_size, 512]. A singleton latent-head axis is also accepted in HND or NHD position. All 512 values per token (RoPE dims included) are quantized in groups of 16 to packed E2M1 with E4M3 scales (scale = amax/6), mirroring the FlashMLA V41_FP4 trajectory.kv_layout (str) – Output layout, either
"HND"or"NHD".
- Returns:
Opaque uint8 paged cache with logical shape
[num_pages, 1, page_size, 288]for HND or[num_pages, page_size, 1, 288]for NHD. Within each physical page it storespage_size * 256data bytes followed bypage_size * 32scale bytes. Consumers must not interpret the last dimension as a contiguous per-token record.- Return type:
torch.Tensor