flashinfer.quantization.make_nvfp4_global_scale

flashinfer.quantization.make_nvfp4_global_scale(input_tensor: Tensor, per_token_activation: bool, global_scale: float | None = None, nvfp4_4over6_config: NVFP44Over6Config | None = None) → Tensor

Build the NVFP4 global scale implied by a resolved 4over6 recipe.

The recipe must match the quantize call’s: the E4M3 clamp appears both in this scale and in the kernel’s candidate search, and a mismatch silently rescales the whole tensor. Pass the same value as nvfp4_4over6= on the quantize call, or resolve_nvfp4_4over6() when that call leaves it unset.

Parameters:
  • input_tensor (torch.Tensor) – Tensor whose amax seeds the scale (per-tensor mode only).

  • per_token_activation (bool) – Per-token activation mode, where the scale is the pure function 1 / (e4m3_max * 6).

  • global_scale (float, optional) – Explicit per-tensor scale, bypassing the amax reduction.

  • nvfp4_4over6_config (NVFP44Over6Config or None) – Resolved recipe; None (the default) means 4over6 is off. Like nvfp4_e4m3_max(), omitting it does not read the environment.