franka_droid robot configuration. You will first run a
synthetic observation, then use a recorded sample and the Python API. Each example
returns 15 actions with 8 values per action.
allenai/MolmoAct2 is a multi-embodiment foundation checkpoint intended for
further robot fine-tuning. These examples check inference and input handling;
they do not establish task success on a robot. Ai2 publishes deployment-specific
checkpoints in its fine-tuned model collection.Quickstart
Use Linux, Python 3.12 or newer,uv, and an NVIDIA GPU with BF16 support and
enough free memory to load the model. See installation for the
CUDA environment. The commands below run from the PhyAI repository root.
1
Install PhyAI
2
Download the checkpoint
norm_stats.json. PhyAI loads the
processor from this local directory with trust_remote_code=True.For an existing download, set MOLMOACT2_CHECKPOINT to that directory instead.
Keep this variable and the activated environment in the same shell for the rest
of the guide.3
Run one observation
Select an available GPU with
nvidia-smi. This command uses GPU 0:--synthetic creates camera images and uses the selected robot configuration’s
mean state, so this first run needs no observation files. The script loads the
model, predicts actions, and saves the postprocessed array. Its log includes:4
Read the result
(1, 15, 8) float32 True: one observation, 15 future
actions, and 8 action values. For franka_droid, the values are
joint_0 through joint_6, followed by gripper; the control mode in the
checkpoint is absolute joint pose.Run a recorded observation
The official DROID sample provides an exterior-camera image, a wrist-camera image, and the matching state. Download just those two images and save the sample’s state:Supply your own camera frames and state
Replace the sample files with one synchronized observation. Forfranka_droid,
the order of repeated --image arguments must match this table:
Provide RGB images; the CLI converts image files to RGB. In Python, use
uint8
arrays shaped (height, width, 3). The processor handles resizing and crops.
For your own recordings, use the corresponding second exterior view in slot 2.
Save the measured state as a float32 NumPy .npy array shaped (8,) or (1, 8),
ordered as [joint_0, joint_1, ..., joint_6, gripper]. Pass the original state
values; the processor applies the checkpoint’s normalization and discrete state
encoding. State and action conventions must match the selected checkpoint and
robot configuration.
Select a different robot configuration
--norm-tag selects camera order, state/action statistics, control mode, and
action horizon. It must match the observation. Inspect the available tags and
their dimensions before preparing data for another embodiment:
Use the Python API
Save the following asrun_molmoact2_api.py at the repository root. It uses the
checkpoint variable and recorded-sample files prepared above. The camera
dictionary makes the mapping explicit; an ordered image list is also accepted.
run_molmoact2_api.py
(1, 15, 8) torch.float32 cpu. engine.step() itself
returns normalized, padded actions shaped (1, 15, 32) for this configuration.
processor.postprocess() removes padding, selects the requested action steps,
and converts to the tag’s action scale. Keep the engine and processor alive
across observations in an application; close the engine when the application
finishes.
Inference options
Use
python examples/molmoact2/run_molmoact2.py --help for all input options.
In the Python API, set num_steps on MolmoAct2Request and n_action_steps on
MolmoAct2Processor.from_pretrained().
CUDA Graph captures action velocity evaluation. Image/text preprocessing and
prefill run outside the graph, and the first request includes capture work.
Reuse the engine to reuse compatible captures. To disable capture in the Python
example, set RuntimeConfig(use_cuda_graph=False).
Troubleshooting
Continuous actions require an action expert and
action_mode="continuous" or
"both". The low-level API also exposes MolmoAct2GenerationRequest for greedy
token IDs; the action examples above use MolmoAct2Request and its processor.
For deployment beyond a local process, see parallel serving.
