Docs

Author a policy

The @seq.policy class

A policy is a class decorated with @seq.policy: the arguments declare what it needs (GPU, image, weights); the methods mark what runs when. seq init pi05-droid clones a decorated pi0.5 template to repoint at your own checkpoint.

pi05_droid.pypython
import seq

@seq.policy(
    gpu="L4",                                  # any card seq gpus lists
    max_gpus=2,                                # required: the most GPUs at once, one robot each
    image=(
        seq.Image.debian_slim()                # a small, pinned Debian base
        .apt_install("libgl1", "libglib2.0-0")
        .run_commands("git clone https://github.com/Physical-Intelligence/openpi /opt/openpi",
                      "cd /opt/openpi && uv pip install --system -e .")   # openpi installs from source
    ),
    weights=seq.Volume.from_name("pi05-ft"),   # your checkpoint: an uploaded volume, hf://, or a template's
    secrets=[seq.Secret.from_name("hf")],      # optional, by name only — see Secrets
)
class Pi05(seq.Policy):
    @seq.load                                  # cold start, once: pull the checkpoint into VRAM
    def load(self, weights_dir):
        self.model = ...
    @seq.warmup                                # optional: one JIT-warming pass before serving
    def warmup(self):
        ...
    @seq.infer                                 # per request: observation -> ActionChunk
    def infer(self, obs, *, seed=None):
        return self.model(obs)
    @seq.on_episode                            # optional: a stateful policy's episode boundary
    def on_episode(self, episode_id):
        ...
    @seq.shutdown                              # optional: cleanup before the worker stops
    def shutdown(self):
        ...

Decorator and lifecycle

The decorator arguments:

argumentwhat it declares
gputhe card: "L4", "A100", "A100-40GB", "H100", "B200", "RTX-PRO-6000" or "cpu"; seq gpus lists them with their memory and prices. For a benchmark, "cpu" puts its physics on CPU machines, one environment per core; a benchmark that names a card must simulate on it, which registration checks. Any other name is refused at deploy with the list. A list is tried in order: a worker that has waited 60 s for one card tries the next, and is billed for the card it got. "H100:2" asks for two cards per worker, billed for each; past 8 — "H100:16" — the model runs across machines (see More than one machine)
cpu, memorythe worker’s CPU cores and GiB of memory: a number (what it requests) or (request, limit). Unset, it requests the minimum — 0.125 core and 128 MiB — and may use more; a policy that loads a large model should request the memory it needs
imagethe container image, declared — see Images
weightsan official template’s weights (seq.Weights.template("pi05-droid", version=1)), an hf://org/model[@revision] repo, or a volume you uploaded (seq.Volume.from_name(...) / vol://name) — copied to /data at deploy; or a dict of named parts, each from its own origin. See Volumes and weights
action, observationthe I/O contract: what infer returns and reads, checked by the worker — see The I/O contract
regionwhere its workers run, near your robots — see Scaling. Unpinned (the default), they run wherever a card is free
timeoutsseq.Timeouts(startup=1800, request=300): the seconds @seq.load may take, and one inference. A request over its limit answers a timeout; the next one runs normally
volumesnot supported yet — a deploy that declares it is refused rather than run without it
secretscredentials referenced by name — see Secrets
max_gpusthe most GPUs the deployment holds at once, counted in cards — required for a GPU policy (or seq deploy --max-gpus N); see Scaling
on_full, queue_timeout_swhat a robot does when every worker is full and no other may start: "queue" (the default) waits in line before its control loop starts, up to queue_timeout_s (120 s); "reject" is told at once
min_workers, buffer_workersworkers kept running with no robot connected, and idle workers kept beside the busy ones while robots are connected: no cold start, billed while they run
idle_timeouthow long an idle worker stays before it stops (default 600 s)
@seq.concurrent(...)copies of the model one worker runs — replicas_per_gpu per card, gpus_per_replica cards each (by default one copy sees every card) — and the robots it takes: target_inputs before another worker starts, max_inputs at most. Too many copies for a card fails the deploy with what one copy measured
streaminga stateful policy: every act() of one episode reaches the same replica (see Episodes)

The methods bind to lifecycle stages. Only @seq.load and @seq.infer are required; the rest are opt-in.

decoratorwhen it runs
@seq.loadonce at cold start — build the model from weights_dir, where the declared weights were copied (see How load and infer are called)
@seq.warmuponce after load — JIT-warm so the first real request is fast
@seq.inferevery request — what the caller sent → what to send back: an Observation → an ActionChunk under an I/O contract, or your own inputs → your own outputs
@seq.on_episodeat an episode boundary, for a stateful policy
@seq.shutdownbefore the worker stops: it is drained first — new requests go to another worker, those in flight finish (up to 20 s), then the hook runs

Compilation is kept between cold starts. Each deployment has its own cache volume: the JAX, torch.compile and Triton caches point at it, and what the first worker compiled is saved once it is ready. A later worker reads it instead of compiling again — pi0.5 on an L4 loads in 53 s instead of 80 s. A worker keeps the platform code its version was built with, so a version built before an improvement like this one gets it when you deploy again.

How load and infer are called

load and infer are your code; the platform decides when they run. It copies the weights you declared into a directory, calls load once with that directory, then calls infer for every request. Reading the files and building the model are load’s job.

whenwhowhat
seq deployyouthe decorator’s declaration and your entry file go to the platform
a worker startsthe platform’s code in the workerimports your file: the decorator registers the class, and the methods marked @seq.load, @seq.infer and the rest are found
cold start, oncethe workerload(weights_dir), the declared weights already copied there
every requestthe workerinfer(inputs, seed=…), inputs being what the caller sent

Your own inputs and outputs

A policy that declares no contract defines its own inputs and outputs. The platform converts nothing: it hands infer the dict the caller sent, unchanged, and sends back whatever infer returns. The format is yours to choose, so write it where callers will look — infer’s docstring and your README. toy-policy takes {"state": [x], "target": [t]} and returns {"delta": [dx]}.

The I/O contract

Declare what infer returns and what it reads. The worker then refuses an observation the policy cannot read — a missing camera, too few frames, a state vector of the wrong length, an empty instruction — with a 400 before any inference runs, and refuses a chunk that is not the declared action space or dimension. When you deploy, the build sends one synthetic observation made from the contract (black frames, zero vectors) to /act; anything but a 200 fails the build with your policy’s own error. A policy that declares nothing is not checked.

declare the contractpython
@seq.policy(
    gpu="L4", max_gpus=1, weights="hf://my-org/pi05-ft",
    action=seq.ActionContract(
        robot="franka",
        action_space=seq.ActionSpace.JOINT_DELTA,     # JOINT_ABSOLUTE, JOINT_DELTA, JOINT_VELOCITY, EE_ABSOLUTE, EE_DELTA
        action_dim=8,
        control_frequency_hz=15.0,
        bounds=[(-0.05, 0.05)] * 7 + [(0.0, 1.0)],  # optional: each dimension's [low, high]
        max_delta=0.05,                                # optional: the most one step moves from the last
        on_violation="clip",                           # "clip" (default) or "refuse"
    ),
    observation=seq.ObservationContract(
        cameras={"exterior": seq.Camera(width=224, height=224),
                 "wrist": seq.Camera(width=224, height=224)},
        proprio={"joint_positions": 7, "gripper": 1},
        history=1,                                     # frames per camera the policy reads
    ),
)

Inference parameters

A keyword argument of infer with a default is an inference parameter. Annotate it int, float, bool or str; seq validate lists them. A robot sets them for its whole session with connect(model, params={"replan_steps": 3}); an eval sets one with --param or compares values with --sweep (an adapter’s and a benchmark’s too: Parameters in an eval); a raw /act request sets them in its "params" object. A name the policy did not declare is refused before inference, and each value is cast to its declared type.

A parameter can also say which values it takes: bound a number with Annotated[int, seq.Range(1, 10)] (or seq.Range(min=0.0)), or list its values with Annotated[str, seq.OneOf("fast", "careful")]. Any other value is refused before inference, saying what is allowed — at connect, by seq policy params, by an eval before anything is reserved, and by the worker on a raw /act; seq validate shows the range beside each parameter (sequence-ai 0.33.27 and later).

A deployment’s defaults change without a redeploy: seq policy params <name> replan_steps=10. Every session that connects from then on runs with them unless it names its own; a robot already connected learns of them at its next lease renewal, within about five minutes, and switches at its next episode (sequence-ai 0.33.27 and later; 0.33.21 to 0.33.26 switch at the renewal itself, an older client runs with its own and the policy’s). An eval of the policy runs with them too unless it sets its own with --param, and its report marks those values deployment’s default (reading such a report needs sequence-ai 0.33.28 and later). --clear replan_steps removes one, seq policy params <name> shows them beside what the policy declares, and seq policy status <name> lists what each connected robot runs with.

declare, then set from an evalpython + shell
@seq.infer
def infer(self, obs, *, seed=None, replan_steps: Annotated[int, seq.Range(1, 10)] = 5,
          temperature: float = 0.0):
    ...

# $ seq eval run my-pi05 --benchmark libero-custom --param replan_steps=3
# $ seq eval run my-pi05 --benchmark libero-custom --sweep replan_steps=3,5,8

Pacing on a real robot

Three numbers decide whether an arm keeps moving between chunks. The control rate — actions per second — belongs to the model and the data it was trained on: each action means “move this much in this step”, so it is not a knob. The actions run per call (replan_steps in the templates) is: how many of a chunk’s actions the robot runs before it asks again. And one call’s round trip is the policy’s inference plus the network. Each call buys replan_steps ÷ control rate seconds of motion, and the SDK asks for the next chunk while they run; when that is shorter than a round trip, the arm stops and waits on every chunk.

pi0.5-LIBERO on an L4, 20 Hzmotion per callone round tripresult
replan_steps=5 (openpi’s LIBERO evaluation)250 msabout 230 ms inference + the network: 300–350 ms measured from one robota pause on every chunk
replan_steps=10500 msthe samethe next chunk arrives while the arm moves

A simulator waits for each action, so an eval can ask often — more closed-loop, usually more successes; a real robot does not wait. Run enough actions per call to cover a round trip, and measure what it costs in an eval first: seq eval run <policy> --benchmark <name> --sweep replan_steps=5,10.

Batching requests

With @seq.batched the worker runs acts that are waiting together as one call — up to max_batch_size: each argument of infer receives a list with one entry per request (the observations, the seeds, every inference parameter), and it returns a list of chunks in the same order. A batch raises throughput, not one robot’s speed: pi0.5 on an L4 answers one observation in 232 ms and sixteen in 2270 ms — 1.6 times the requests per second, each of them slower.

one forward pass for every environment acting nowpython
@seq.batched(max_batch_size=16, eval_wait_ms=20)  # 2-256 requests; an eval's batch runs once 20 ms pass with no new act
@seq.infer
def infer(self, obs, *, seed=None):
    return self.model.batch(obs)                 # obs is a list; return one ActionChunk per entry

wait_ms= is the old name of eval_wait_ms=: still read, with a warning. Not for streaming=True: a stateful episode owns its replica.

Images

Each seq.Image method returns a new image; the chain is a spec, and the step order you write is the build order.

methodwhat it declares
seq.Image.debian_slim()a small, pinned Debian base — the default when you declare no image
seq.Image.from_registry("...")start from any registry tag, carried through verbatim
seq.Image.from_dockerfile("Dockerfile", context=".")build from your own Dockerfile; its COPY lines read from context, uploaded when you deploy
python="3.12"on any of the three: the Python the image runs
.apt_install(*pkgs)system libraries (apt-get install)
.uv_pip_install(*pkgs)Python packages (uv pip install)
.run_commands(*cmds, secrets=[...])arbitrary build commands, run in order. With secrets= (a private repo, a package index token) the credential is present for that step only, never in the running worker; the build’s copy of it is deleted when the build ends
.env(...)build-time environment variables (never secrets — those travel by name)

Volumes and weights

weights= copies the checkpoint once, at deploy, into a volume the worker mounts at /data; every cold start reads that copy. The volume is billed as storage at the size actually copied. It names one of three things:

Deleted a volume, or replaced a version with a new seq volume put, by mistake? Contact us within 7 days and we can bring it back.

Weights from your own storage bucket or an https URL are deprecated. This SDK release still deploys them, with a warning; the next one refuses them. Move the files into a volume once — download them, then seq volume put <name> <dir> — and deploy weights=seq.Volume.from_name("<name>"), or point at the Hugging Face repo they came from.

Explicit volumes= mounts are coming — a deploy that declares volumes= now is refused rather than run without it.

A policy with no checkpoint — a closed-form controller, say — leaves weights= out: nothing is copied or mounted, and nothing is billed as storage. seq validate shows weights none.