Skip to main content
MolmoAct2 predicts a chunk of robot actions from camera images, a task instruction, and robot state. PhyAI runs its vision-language backbone and continuous action expert in BF16. This guide uses the franka_droid robot configuration. You will first run a synthetic observation, then use a recorded sample and the Python API. Each example returns 15 actions with 8 values per action.
allenai/MolmoAct2 is a multi-embodiment foundation checkpoint intended for further robot fine-tuning. These examples check inference and input handling; they do not establish task success on a robot. Ai2 publishes deployment-specific checkpoints in its fine-tuned model collection.

Quickstart

Use Linux, Python 3.12 or newer, uv, and an NVIDIA GPU with BF16 support and enough free memory to load the model. See installation for the CUDA environment. The commands below run from the PhyAI repository root.
1

Install PhyAI

If you already have a checkout and environment, activate that environment and continue from its repository root.
2

Download the checkpoint

Download the whole repository. Alongside the weights, the processor needs the tokenizer, image processor, Python files, and norm_stats.json. PhyAI loads the processor from this local directory with trust_remote_code=True.For an existing download, set MOLMOACT2_CHECKPOINT to that directory instead. Keep this variable and the activated environment in the same shell for the rest of the guide.
3

Run one observation

Select an available GPU with nvidia-smi. This command uses GPU 0:
--synthetic creates camera images and uses the selected robot configuration’s mean state, so this first run needs no observation files. The script loads the model, predicts actions, and saves the postprocessed array. Its log includes:
4

Read the result

The first line should be (1, 15, 8) float32 True: one observation, 15 future actions, and 8 action values. For franka_droid, the values are joint_0 through joint_6, followed by gripper; the control mode in the checkpoint is absolute joint pose.

Run a recorded observation

The official DROID sample provides an exterior-camera image, a wrist-camera image, and the matching state. Download just those two images and save the sample’s state:
The sample supplies only one exterior view. Following the upstream example, pass it twice to fill the two exterior-camera slots, then pass the wrist view:
This still uses the foundation checkpoint downloaded above. It demonstrates file-based input with a real observation; downloading sample images does not switch the model to the DROID fine-tune.

Supply your own camera frames and state

Replace the sample files with one synchronized observation. For franka_droid, the order of repeated --image arguments must match this table: Provide RGB images; the CLI converts image files to RGB. In Python, use uint8 arrays shaped (height, width, 3). The processor handles resizing and crops. For your own recordings, use the corresponding second exterior view in slot 2. Save the measured state as a float32 NumPy .npy array shaped (8,) or (1, 8), ordered as [joint_0, joint_1, ..., joint_6, gripper]. Pass the original state values; the processor applies the checkpoint’s normalization and discrete state encoding. State and action conventions must match the selected checkpoint and robot configuration.

Select a different robot configuration

--norm-tag selects camera order, state/action statistics, control mode, and action horizon. It must match the observation. Inspect the available tags and their dimensions before preparing data for another embodiment:

Use the Python API

Save the following as run_molmoact2_api.py at the repository root. It uses the checkpoint variable and recorded-sample files prepared above. The camera dictionary makes the mapping explicit; an ordered image list is also accepted.
run_molmoact2_api.py
The first print returns (1, 15, 8) torch.float32 cpu. engine.step() itself returns normalized, padded actions shaped (1, 15, 32) for this configuration. processor.postprocess() removes padding, selects the requested action steps, and converts to the tag’s action scale. Keep the engine and processor alive across observations in an application; close the engine when the application finishes.

Inference options

Use python examples/molmoact2/run_molmoact2.py --help for all input options. In the Python API, set num_steps on MolmoAct2Request and n_action_steps on MolmoAct2Processor.from_pretrained(). CUDA Graph captures action velocity evaluation. Image/text preprocessing and prefill run outside the graph, and the first request includes capture work. Reuse the engine to reuse compatible captures. To disable capture in the Python example, set RuntimeConfig(use_cuda_graph=False).

Troubleshooting

Continuous actions require an action expert and action_mode="continuous" or "both". The low-level API also exposes MolmoAct2GenerationRequest for greedy token IDs; the action examples above use MolmoAct2Request and its processor. For deployment beyond a local process, see parallel serving.