phyai or phyai-gateway.
Prepare a server configuration
Install phyai and copy the PI0.5 example:pi05.yaml before starting it:
- Set
plugin_args.checkpoint_dirandadapter.args.tokenizer_dirto your local files. - Set
adapter.args.image_namesin checkpoint camera order, and set the correspondingplugin_args.inputs_image_shapeentries to the incoming image sizes. - Match
adapter.args.state_dimandadapter.args.action_dimto the checkpoint and client. The example’s two cameras, eight state values, and seven action values are example settings. - Choose
server.model_name,server.port, andconfig.device.target. The example usespi05, port50063, andcuda.
plugin, plugin_args, and adapter settings in the same YAML layout.
Load a separate kernel policy
Add this top-level field topi05.yaml:
kernel-policy.yaml beside it. A minimal policy is:
kernel_policy: /path/to/kernel-policy.yaml. Relative policy paths resolve from the directory containing pi05.yaml, regardless of where you run the command. For local checkpoint and tokenizer paths, use an absolute path or an explicit ./ or ../ prefix.
Start a model server
--check checks configuration fields and types without loading the model. It does not verify checkpoint contents or adapter startup.
Start two servers
Make one YAML file per server:
Run each command in a separate terminal:
50063 and their local GPU. Register each host’s address in the gateway.
Start the gateway
Repeat--backend MODEL HOST:PORT for each server:
50111 and HTTP port 30000. Each backend’s MODEL must match its YAML server.model_name. Use a reachable host address when the gateway runs on another machine.
To serve different models, give their server YAML files different model_name values and register both:
healthy field to become true before sending requests. Reapply registrations after restarting the gateway.
Use multiple replicas for stateless inference. Successive requests from one client can reach different servers; a model that keeps session state needs a single backend for that route.
Add or remove a backend while running
SERVER_ID with the server_id returned by POST or GET. POST returns 202; inspect GET to check readiness. DELETE stops new requests from using that registration and lets active requests finish.
Connect a LeRobot client
Install LeRobot support in the gateway’s environment. From the PhyAI repository root:pretrained_name_or_path. Set the server YAML’s server.model_name and the gateway’s registration name to that exact value. For example, if your client uses your-org/your-policy, set server.model_name: your-org/your-policy before starting the model server, then start the gateway:
10.0.0.11 with the model server’s address. For the PI05 adapter, match adapter.args.image_names to the client’s camera names after any rename_map. Set plugin_args.inputs_image_shape to the incoming image sizes, and match the state/action dimensions and ordering to your robot and checkpoint. The client’s actions_per_chunk must not exceed the checkpoint’s chunk size.
On the robot machine, use your existing LeRobot environment and client configuration. If that configuration is saved as robot-client.yaml, run:
10.0.0.10 with the gateway’s address. robot-client.yaml is your LeRobot RobotClientConfig, including robot, policy_type, pretrained_name_or_path, and actions_per_chunk. This command starts the original client’s robot control loop. If you already pass those settings as command-line arguments, keep them and change only --server_address. No client source changes or extra request headers are needed.
Match the gateway’s --lerobot-fps to the client’s fps; both default to 30. For a client configured at 20 Hz, add --lerobot-fps 20 to the gateway command. All LeRobot clients on one gateway must use the same rate to receive correctly spaced action timestamps. Use the LeRobot adapter with trusted clients.
Send one test observation
You can test a request without connecting a robot. This example uses the original PI05 YAML settings: model namepi05, eight state values, and the main_images and wrist_images cameras at 360 × 360 pixels. Start that model server, then start the gateway for this test with:
TimedAction because this example requests actions_per_chunk=1. Its action is a Torch tensor with seven values for the sample server’s action_dim: 7. Keep the same channel open for setup, observation uploads, and action requests.
