Quantization
Quantization stores and computes weights and activations at lower precision than bf16 (fp8, nvfp4, and so on) to save memory and raise throughput. The hard part is not how any single format does its math. It is that one engine has to accept several quantized checkpoints that come from different tools with different conventions, while leaving model code almost untouched. PhyAI splits the problem into two layers:- Semantic layer: what a tensor is quantized to. It describes only the mathematical properties (dtype, granularity, symmetry) and says nothing about scale layout, kernel, or device.
- Physical layer: what the tensor looks like in memory. Its dtype, the shape and layout of its scales, whether it needs post-load processing, and which kernel the forward pass runs.
QuantPlan, the QuantPlan resolves a QuantScheme by layer name, and materialize lowers that semantics into a WeightSpec. The first three stages never touch hardware. Only materialize reads the SM version, so the boundary falls between scheme and spec.
The four-stage pipeline
What each stage in the figure actually does:
Each stage reads only the output of the one before it and never reaches back. An importer collects assorted configs into a single rule table. The rule table answers one question by layer name: which scheme this layer uses.
materialize translates semantics into physical form and makes the hardware-dependent choices here (for example, NVFP4’s scale layout depends on the SM version).
Example: how three keys land
Zoom into theQuantPlan → WeightSpec span above. Take an fp8 checkpoint whose plan has two rules: skip lm_head, send everything else to fp8. Three weight prefixes run through the rule table and land on different physical formats.
qkv_proj and gate_up_proj are not named by any rule, so they fall through to default and get an Fp8Spec: the weight is stored as fp8_e4m3 with a fp32 weight_scale (per-channel, one per output channel), and the activation is quantized dynamically per token at runtime. lm_head hits the first skip rule, its scheme is None, so it stays bf16 with just a weight and no scale. Rules match top to bottom and stop at the first hit; everything else takes default.
Running a quantized checkpoint
Model entries already wire up quantized loading, so running a quantized checkpoint uses the same code as a normal one. Take pi0.5:QuantPlan, and sets it as the active plan:
load_quant_plan returns None, the model falls back to bf16, and behavior is identical to before. Wiring up quantization needs no branch on the model side. This loading is already connected in the pi0, pi0.5, and cosmos3 entries.
What’s supported today
On the config side, three frameworks are recognized:
For how to configure it or hand-write rules, see Configuring quantization. To change physical formats or add a framework, see Internals and extension.

