Author a policy
The @seq.policy class
A policy is a class decorated with @seq.policy: the arguments declare what it
needs (GPU, image, weights); the methods mark what runs when. seq init pi05-droid
clones a decorated pi0.5 template to repoint at your own checkpoint.
import seq @seq.policy( gpu="L4", # any card seq gpus lists max_gpus=2, # required: the most GPUs at once, one robot each image=( seq.Image.debian_slim() # a small, pinned Debian base .apt_install("libgl1", "libglib2.0-0") .run_commands("git clone https://github.com/Physical-Intelligence/openpi /opt/openpi", "cd /opt/openpi && uv pip install --system -e .") # openpi installs from source ), weights=seq.Volume.from_name("pi05-ft"), # your checkpoint: an uploaded volume, hf://, or a template's secrets=[seq.Secret.from_name("hf")], # optional, by name only — see Secrets ) class Pi05(seq.Policy): @seq.load # cold start, once: pull the checkpoint into VRAM def load(self, weights_dir): self.model = ... @seq.warmup # optional: one JIT-warming pass before serving def warmup(self): ... @seq.infer # per request: observation -> ActionChunk def infer(self, obs, *, seed=None): return self.model(obs) @seq.on_episode # optional: a stateful policy's episode boundary def on_episode(self, episode_id): ... @seq.shutdown # optional: cleanup before the worker stops def shutdown(self): ...
Decorator and lifecycle
The decorator arguments:
| argument | what it declares |
|---|---|
gpu | the card: "L4", "A100", "A100-40GB", "H100", "B200", "RTX-PRO-6000" or "cpu"; seq gpus lists them with their memory and prices. For a benchmark, "cpu" puts its physics on CPU machines, one environment per core; a benchmark that names a card must simulate on it, which registration checks. Any other name is refused at deploy with the list. A list is tried in order: a worker that has waited 60 s for one card tries the next, and is billed for the card it got. "H100:2" asks for two cards per worker, billed for each; past 8 — "H100:16" — the model runs across machines (see More than one machine) |
cpu, memory | the worker’s CPU cores and GiB of memory: a number (what it requests) or (request, limit). Unset, it requests the minimum — 0.125 core and 128 MiB — and may use more; a policy that loads a large model should request the memory it needs |
image | the container image, declared — see Images |
weights | an official template’s weights (seq.Weights.template("pi05-droid", version=1)), an hf://org/model[@revision] repo, or a volume you uploaded (seq.Volume.from_name(...) / vol://name) — copied to /data at deploy; or a dict of named parts, each from its own origin. See Volumes and weights |
action, observation | the I/O contract: what infer returns and reads, checked by the worker — see The I/O contract |
region | where its workers run, near your robots — see Scaling. Unpinned (the default), they run wherever a card is free |
timeouts | seq.Timeouts(startup=1800, request=300): the seconds @seq.load may take, and one inference. A request over its limit answers a timeout; the next one runs normally |
volumes | not supported yet — a deploy that declares it is refused rather than run without it |
secrets | credentials referenced by name — see Secrets |
max_gpus | the most GPUs the deployment holds at once, counted in cards — required for a GPU policy (or seq deploy --max-gpus N); see Scaling |
on_full, queue_timeout_s | what a robot does when every worker is full and no other may start: "queue" (the default) waits in line before its control loop starts, up to queue_timeout_s (120 s); "reject" is told at once |
min_workers, buffer_workers | workers kept running with no robot connected, and idle workers kept beside the busy ones while robots are connected: no cold start, billed while they run |
idle_timeout | how long an idle worker stays before it stops (default 600 s) |
@seq.concurrent(...) | copies of the model one worker runs — replicas_per_gpu per card, gpus_per_replica cards each (by default one copy sees every card) — and the robots it takes: target_inputs before another worker starts, max_inputs at most. Too many copies for a card fails the deploy with what one copy measured |
streaming | a stateful policy: every act() of one episode reaches the same replica (see Episodes) |
The methods bind to lifecycle stages. Only @seq.load and @seq.infer
are required; the rest are opt-in.
| decorator | when it runs |
|---|---|
@seq.load | once at cold start — build the model from weights_dir, where the declared weights were copied (see How load and infer are called) |
@seq.warmup | once after load — JIT-warm so the first real request is fast |
@seq.infer | every request — what the caller sent → what to send back: an Observation → an ActionChunk under an I/O contract, or your own inputs → your own outputs |
@seq.on_episode | at an episode boundary, for a stateful policy |
@seq.shutdown | before the worker stops: it is drained first — new requests go to another worker, those in flight finish (up to 20 s), then the hook runs |
Compilation is kept between cold starts. Each deployment has its own cache volume: the JAX,
torch.compile and Triton caches point at it, and what the first worker compiled is saved
once it is ready. A later worker reads it instead of compiling again — pi0.5 on an L4 loads in
53 s instead of 80 s. A worker keeps the platform code its version was built with, so a version built
before an improvement like this one gets it when you deploy again.
How load and infer are called
load and infer are your code; the platform decides when they run. It copies the weights
you declared into a directory, calls load once with that directory, then calls infer
for every request. Reading the files and building the model are load’s job.
| when | who | what |
|---|---|---|
seq deploy | you | the decorator’s declaration and your entry file go to the platform |
| a worker starts | the platform’s code in the worker | imports your file: the decorator registers the class, and the methods marked @seq.load, @seq.infer and the rest are found |
| cold start, once | the worker | load(weights_dir), the declared weights already copied there |
| every request | the worker | infer(inputs, seed=…), inputs being what the caller sent |
weights_diris/data/<policy name>: the directoryweights=is copied to before the worker starts. Weights in named parts (weights={"base": …, "lora": …}) arrive as{name: directory}instead.- With no weights declared, nothing is copied and
loadis still called once, with that path.toy-policykeeps it and builds nothing;pi05-droidbuilds the model from it with openpi (create_trained_policy(cfg, weights_dir, …)). loadmay take up to thestartuplimit oftimeouts=(1800 s by default; see the table above).- Which
inputsinfergets depends on how it is written: if its first argument is annotatedObservationor it returns anActionChunk, the worker parses each request into anObservationand checks it against the contract you declared (The I/O contract); otherwiseinfergets the request’s inputs as they were sent (Your own inputs and outputs).
Your own inputs and outputs
A policy that declares no contract defines its own inputs and outputs. The platform converts nothing: it hands
infer the dict the caller sent, unchanged, and sends back whatever infer returns. The format
is yours to choose, so write it where callers will look — infer’s docstring and your README.
toy-policy takes {"state": [x], "target": [t]} and returns {"delta": [dx]}.
- From a robot or a script, send that dict and get the outputs back as
inferreturned them:sequence_ai.connect("my-toy").predict({"state": [0.0], "target": [0.5]})returns{"delta": [0.1]}. A template’sclient.pybuilds the dict for you. - In an eval, an adapter turns the benchmark’s observation into your inputs and your outputs into
the benchmark’s action: one with
policy_io = "own"handsinferexactly the dict it takes (see Adapters). - A wrong key is not caught before
inferruns, because nothing knows your format:inputs["state"]on a request that sentarm_xraises inside your code, and the caller gets that error: a Python built-in exception comes back as itself — hereKeyError: 'state'— with your code’s traceback on.remote_traceback(any other type arrives asPolicyError, its name in.exception_type). It is a500 policy_error, so a caller can tell your code’s failure from the platform’s. Check the keys you need at the top ofinferand raise an error that names them, astoy-policydoes — or declare a contract and let the worker check for you.
The I/O contract
Declare what infer returns and what it reads. The worker then refuses an observation the
policy cannot read — a missing camera, too few frames, a state vector of the wrong length, an empty
instruction — with a 400 before any inference runs, and refuses a chunk that is not the
declared action space or dimension. When you deploy, the build sends one synthetic observation made from
the contract (black frames, zero vectors) to /act; anything but a 200 fails the
build with your policy’s own error. A policy that declares nothing is not checked.
@seq.policy(
gpu="L4", max_gpus=1, weights="hf://my-org/pi05-ft",
action=seq.ActionContract(
robot="franka",
action_space=seq.ActionSpace.JOINT_DELTA, # JOINT_ABSOLUTE, JOINT_DELTA, JOINT_VELOCITY, EE_ABSOLUTE, EE_DELTA
action_dim=8,
control_frequency_hz=15.0,
bounds=[(-0.05, 0.05)] * 7 + [(0.0, 1.0)], # optional: each dimension's [low, high]
max_delta=0.05, # optional: the most one step moves from the last
on_violation="clip", # "clip" (default) or "refuse"
),
observation=seq.ObservationContract(
cameras={"exterior": seq.Camera(width=224, height=224),
"wrist": seq.Camera(width=224, height=224)},
proprio={"joint_positions": 7, "gripper": 1},
history=1, # frames per camera the policy reads
),
)
- Limits. With
boundsormax_deltathe worker clips every chunk into them and lists each clipped step in the response’s warnings; the client clips again across chunks before an action reaches your robot.on_violation="refuse"answers an error instead of a clipped chunk. - Check before the arm moves.
connect(model, expect={"action_space": "joint_delta", "action_dim": 8})compares what you will execute with what the deployment declared, and raises before the first action if they differ.policy.contractholds the declared contract. - Depth.
seq.Camera(width=, height=, depth=True)declares that the view also sends a depth frame.sequence_ai.observation(images=..., depth={"wrist": depth_map})sends each map as a lossless 16-bit PNG of millimetres (metres in by default;depth_unit="mm"for a map already in millimetres). A declared depth frame that is missing is refused before inference.
Inference parameters
A keyword argument of infer with a default is an inference parameter. Annotate it
int, float, bool or str; seq validate lists
them. A robot sets them for its whole session with connect(model, params={"replan_steps": 3}); an eval sets
one with --param or compares values with --sweep (an adapter’s and a benchmark’s
too: Parameters in an eval); a raw /act request sets
them in its "params" object. A name the policy did not declare is
refused before inference, and each value is cast to its declared type.
A parameter can also say which values it takes: bound a number with
Annotated[int, seq.Range(1, 10)] (or seq.Range(min=0.0)), or list its values with
Annotated[str, seq.OneOf("fast", "careful")]. Any other value is refused before inference, saying
what is allowed — at connect, by seq policy params, by an eval before anything is
reserved, and by the worker on a raw /act; seq validate shows the range beside each
parameter (sequence-ai 0.33.27 and later).
A deployment’s defaults change without a redeploy: seq policy params <name> replan_steps=10.
Every session that connects from then on runs with them unless it names its own; a robot already connected
learns of them at its next lease renewal, within about five minutes, and switches at its next episode
(sequence-ai 0.33.27 and later; 0.33.21 to 0.33.26 switch at the renewal itself, an older client
runs with its own and the policy’s). An eval of the policy runs with them too unless it sets its own with
--param, and its report marks those values deployment’s default (reading such a report
needs sequence-ai 0.33.28 and later). --clear replan_steps removes one,
seq policy params <name> shows them beside what the policy declares, and
seq policy status <name> lists what each connected robot runs with.
@seq.infer def infer(self, obs, *, seed=None, replan_steps: Annotated[int, seq.Range(1, 10)] = 5, temperature: float = 0.0): ... # $ seq eval run my-pi05 --benchmark libero-custom --param replan_steps=3 # $ seq eval run my-pi05 --benchmark libero-custom --sweep replan_steps=3,5,8
Pacing on a real robot
Three numbers decide whether an arm keeps moving between chunks. The control rate — actions per
second — belongs to the model and the data it was trained on: each action means “move this much in
this step”, so it is not a knob. The actions run per call (replan_steps in the
templates) is: how many of a chunk’s actions the robot runs before it asks again. And one call’s
round trip is the policy’s inference plus the network. Each call buys
replan_steps ÷ control rate seconds of motion, and the SDK asks for the next chunk while they
run; when that is shorter than a round trip, the arm stops and waits on every chunk.
| pi0.5-LIBERO on an L4, 20 Hz | motion per call | one round trip | result |
|---|---|---|---|
replan_steps=5 (openpi’s LIBERO evaluation) | 250 ms | about 230 ms inference + the network: 300–350 ms measured from one robot | a pause on every chunk |
replan_steps=10 | 500 ms | the same | the next chunk arrives while the arm moves |
A simulator waits for each action, so an eval can ask often — more closed-loop, usually more successes; a
real robot does not wait. Run enough actions per call to cover a round trip, and measure what it costs in an eval
first: seq eval run <policy> --benchmark <name> --sweep replan_steps=5,10.
Batching requests
With @seq.batched the worker runs acts that are waiting together as one call — up to
max_batch_size: each argument of infer receives a list with one entry per request
(the observations, the seeds, every inference parameter), and it returns a list of chunks in the same order.
A batch raises throughput, not one robot’s speed: pi0.5 on an L4 answers one observation in 232 ms and
sixteen in 2270 ms — 1.6 times the requests per second, each of them slower.
- In an eval, many environments act at once, and their acts reach a policy instance spread over up
to a few hundred milliseconds: it gathers them until they stop coming — a batch runs once
eval_wait_ms(20 ms unless you name it) pass with no new act, once it is full, or 250 ms after its first act at the latest. - A deployment serving robots never waits: a robot’s acts come one at a time, so a wait would
only delay each of them. Acts that arrive while a copy is busy still run together when it frees, at no
extra wait. Only a deployment that crowds robots onto a copy (
@seq.concurrentwithmax_inputsabove its copies) may nameserve_wait_ms, at most 10 ms.
@seq.batched(max_batch_size=16, eval_wait_ms=20) # 2-256 requests; an eval's batch runs once 20 ms pass with no new act @seq.infer def infer(self, obs, *, seed=None): return self.model.batch(obs) # obs is a list; return one ActionChunk per entry
wait_ms= is the old name of eval_wait_ms=: still read, with a warning. Not for
streaming=True: a stateful episode owns its replica.
Images
Each seq.Image method returns a new image; the chain is a spec, and the step
order you write is the build order.
| method | what it declares |
|---|---|
seq.Image.debian_slim() | a small, pinned Debian base — the default when you declare no image |
seq.Image.from_registry("...") | start from any registry tag, carried through verbatim |
seq.Image.from_dockerfile("Dockerfile", context=".") | build from your own Dockerfile; its COPY lines read from context, uploaded when you deploy |
python="3.12" | on any of the three: the Python the image runs |
.apt_install(*pkgs) | system libraries (apt-get install) |
.uv_pip_install(*pkgs) | Python packages (uv pip install) |
.run_commands(*cmds, secrets=[...]) | arbitrary build commands, run in order. With secrets= (a private repo, a package index token) the credential is present for that step only, never in the running worker; the build’s copy of it is deleted when the build ends |
.env(...) | build-time environment variables (never secrets — those travel by name) |
Volumes and weights
weights= copies the checkpoint once, at deploy, into a volume the worker mounts at
/data; every cold start reads that copy. The volume is billed as storage at the size actually
copied. It names one of three things:
- An official template’s weights —
seq.Weights.template("pi05-droid", version=1)(the string form is"template://pi05-droid@v1"): the unmodified checkpoint a template ships with, plus what it loads at runtime (pi0.5’s tokenizer). A published version never changes; leaveversion=out to take the newest one at each deploy. No key is needed. This is whatseq init pi05-droidwrites. - Hugging Face —
hf://org/modelorhf://org/model@<revision>. The revision is resolved to a commit at deploy and exactly that commit is copied, so a re-deploy is deterministic. A private or gated repo needs a token:seq secret set hf --env HF_TOKEN, thensecrets=[seq.Secret.from_name("hf")]. The copy runs with only that token; declaring it also handsHF_TOKENto your policy at runtime. A repo that cannot be read is refused when you deploy, in seconds, with what to fix. - Your own files — upload a checkpoint directory once with
seq volume put pi05-ft ./checkpoints/pi05-ft(resumable: run it again after an interruption), then deploy it by name:weights=seq.Volume.from_name("pi05-ft"), orweights="vol://pi05-ft@v2"to pin a version. Every file’s md5 is checked when the upload completes; only then is the version active and billed as storage.seq volume lslists them;seq volume rmis refused while a deployment’s weights, or a registered benchmark’s assets, use the volume. - Several parts — a dict of named parts, each from its own origin and read with its
own declared secret;
@seq.loadreceives{name: directory}. Up to 8 parts, named[a-z][a-z0-9_-]:weights={"base": "hf://physical-intelligence/pi05-base", "lora": seq.Volume.from_name("my-lora"), "norm_stats": seq.Volume.from_name("droid-norm")}.
Deleted a volume, or replaced a version with a new seq volume put, by mistake?
Contact us within 7 days and we can bring it back.
Weights from your own storage bucket or an https URL are deprecated. This
SDK release still deploys them, with a warning; the next one refuses them. Move the files into a volume once
— download them, then seq volume put <name> <dir> — and deploy
weights=seq.Volume.from_name("<name>"), or point at the Hugging Face repo they came from.
Explicit volumes= mounts are coming — a deploy that declares volumes= now is
refused rather than run without it.
A policy with no checkpoint — a closed-form controller, say — leaves weights= out:
nothing is copied or mounted, and nothing is billed as storage. seq validate shows
weights none.