flashinfer.quantization.make_nvfp4_global_scale¶
- flashinfer.quantization.make_nvfp4_global_scale(input_tensor: Tensor, per_token_activation: bool, global_scale: float | None = None, nvfp4_4over6_config: NVFP44Over6Config | None = None) Tensor¶
Build the NVFP4 global scale implied by a resolved 4over6 recipe.
The recipe must match the quantize call’s: the E4M3 clamp appears both in this scale and in the kernel’s candidate search, and a mismatch silently rescales the whole tensor. Pass the same value as
nvfp4_4over6=on the quantize call, orresolve_nvfp4_4over6()when that call leaves it unset.- Parameters:
input_tensor (torch.Tensor) – Tensor whose amax seeds the scale (per-tensor mode only).
per_token_activation (bool) – Per-token activation mode, where the scale is the pure function
1 / (e4m3_max * 6).global_scale (float, optional) – Explicit per-tensor scale, bypassing the amax reduction.
nvfp4_4over6_config (NVFP44Over6Config or None) – Resolved recipe;
None(the default) means 4over6 is off. Likenvfp4_e4m3_max(), omitting it does not read the environment.