The control loop, in Python
sequence_ai is the client behind Running a policy.
connect() opens a kept-alive line straight to the worker; act()
and run() play the action chunks it returns and raise typed errors. This is the
reference for that loop.
Install
One dependency (httpx), Python 3.9+. Authoring and deploying a policy add a
heavier extra (sequence-ai[authoring]); this page is
the client.
pip install sequence-ai # the client: connect / act / run pip install 'sequence-ai[image]' # adds Pillow, for ndarray / PIL observations
Install sequence-ai, import sequence_ai.
Authentication
One key identifies your tenant. The client reads it from api_key= first, then
$SEQUENCES_API_KEY, and resolves the gateway from it — no host to configure.
export SEQUENCES_API_KEY="seq_live_..."
Create one in the console.
sequence_ai.models() needs no key.
Quickstart
A complete control loop. connect() is a context manager, so the connection closes
even when the loop raises.
import sequence_ai with sequence_ai.connect(model="pi05-droid") as policy: out = policy.run( # required observe=robot.read_observation, # -> Observation act=robot.apply_action, # one step, a list of floats max_actions=250, # 16.7 s at 15 Hz # optional, and these three are the ones worth adding validate=robot.is_safe, # return False to stop; warns if omitted hold=robot.hold, # called once if it stops early on_underrun=robot.hold_position, # buffer ran dry — hold, do not coast ) print(out) # RunOutcome(250 actions over 17 chunks in 16.8s, completed)
Required: observe, act, max_actions. Everything
else has a default. validate is your bounds check and run() warns if
it is omitted; hold runs on any early ending; on_underrun fires when a
refill was late.
validate and hold.
connect() opens a direct line to the worker
connect() returns a Policy handle and keeps one connection open for
the life of the loop. A call is two phases; the gateway is in only the first.
| phase | reaches | purpose |
|---|---|---|
connect(model) | the gateway, once | admit, ensure a worker is up, return its worker_url, cert_fingerprint, direct_secret and lease |
act() / run() | the worker, directly | run inference on the reused, cert-pinned connection; Bearer <direct_secret> |
| reconnect | the gateway, on demand | on a rotated secret, expired lease or a worker that scaled to zero — fetch fresh coordinates and replay |
Direct connect and transparent fallback
You do not handle the 401 or 503 yourself. A rotated
direct_secret, an expired lease, a worker that scaled to zero, or a dropped
connection all surface as one transparent reconnect: the client fetches fresh
coordinates and replays the in-flight call. The same Policy handle keeps working.
A cold worker answers connect() with 503 and a
Retry-After while it boots; the client waits that out — see
cold starts and warming.
Four ways to drive it
The same request underneath, at four levels of control. The instruction is a field of
Observation, so the first three carry it with the observation; act()
takes it as an argument and overwrites whatever the observation holds.
| Call | Returns |
|---|---|
predict(obs) | Prediction — one whole action_chunk, plus perf, usage and warnings. You play it out. |
next_action(obs) | One action, a list[float]. Refills from an internal buffer when it empties. |
run(observe, act) | RunOutcome, after playing max_actions steps. |
act(instruction, observe, act) | RunOutcome, after playing one instruction until until, max_seconds or interrupt(). |
pred = policy.predict(observation) print(len(pred.action_chunk), "steps covering", pred.action_chunk.covers_seconds, "s at", pred.action_chunk.control_frequency_hz, "Hz")
while running:
action = policy.next_action(robot.read_observation())
robot.apply_action(action)
Pass observation= on every call. It is ignored while the buffer has actions, and is
what the refill is computed from when the buffer is empty. An empty buffer with no observation
raises ChunkExhausted.
run()
The full control-loop API. validate and hold are its parameters, not a
wrapper you build — attach your bounds check and your stop-behaviour here.
| Parameter | Default | What it does |
|---|---|---|
observe | required | Returns the current Observation. Called on the calling thread, never off it. |
act | required | Called once per action with one step, a list of floats. |
max_actions | required, keyword-only | Upper bound on steps executed. |
validate | None | Called before every act. Return False to stop. Omitting it warns. |
hold | None | Called once if the loop stops early, including on an exception. |
on_chunk | None | Called with each Prediction after a refill. |
on_underrun | None | Called with the stall in seconds when the buffer ran dry. |
on_status | None | Called with starting, running, waiting, done. |
progress | "auto" | Print status to stderr. "auto" means only when stderr is a terminal. |
prefetch | True | Overlap the next inference with playback. |
lookahead | 1 | Chunks to hold in hand or on the way. Above 1 the run is open loop over that depth. |
open_loop_horizon | None | Play only the first N steps of each chunk before refilling. |
max_seconds | None | Stop after this much wall-clock. reason becomes timeout. |
until | None | Checked before each act; returning True ends the run with reason until. |
startup_timeout_s | 1800.0 | Wait this long for a cold worker, before the first action only. 0 opts out. See how long it takes. |
smooth_seam | 0 | Blend N steps across a chunk boundary. Absolute action spaces only. |
max_actions is required and keyword-only. There is no run_forever().
hold runs in a finally block and fires on the exception path too.
With validate=None nothing is checked before actuation. Allowed for simulation and
bench work; it warns once per process, never silently on real hardware.
act() — one instruction, a time bound
act() pins one instruction and drives until it is done, a deadline, or an
interrupt — the shape a higher-level planner calls as a tool. It returns a
RunOutcome like run().
out = policy.act(
"pick up the cup",
observe=robot.read_observation, # cameras + joints
act=robot.apply_action, # one joint target
validate=robot.is_safe,
hold=robot.hold,
until=robot.cup_in_gripper, # optional: return True to finish early
max_seconds=30, # required — this drives hardware
)
act(instruction, observe, act, *, max_seconds, until=None). max_seconds
is required and has no default. act() overwrites the observation's instruction on
every call.
Wiring instructions together — retries, failure detection, choosing the next sub-goal — is yours. For a server-side planner, see a composite policy, where the VLM runs on your key.
Changing the instruction without stopping
policy.set_instruction("now put it on the plate") # returns immediately; the arm keeps moving policy.instruction # what is in force, or None
Thread-safe. The new sentence takes effect on the next request; the arm is not paused. Setting the same string again is ignored.
Call it from a thread that is not the one driving. Putting the planner call inside
observe() runs it on the driving thread and the arm stands still for its length;
on its own thread, the loop keeps prefetching under the instruction currently in force.
Use set_instruction() to refine or continue, and interrupt() to
abandon. The difference is what happens to the robot: the first keeps it moving, the second
stops it within one control period.
The new instruction is not instant: the in-flight request was built from the old one, so the
first action reflecting the call arrives at most lookahead + 1 chunks later.
RunOutcome.instruction_swaps and max_instruction_lag_s report it.
Changing your mind
policy.interrupt()
Thread-safe. Call it from any thread while run() or act() is driving.
Stops between actions, not at a chunk boundary — about 67 ms at 15 Hz.
Also interrupts a fetch that is still in flight, including during a cold start or a reconnect.
Unwinds through hold() and is not raised. Buffered and in-flight chunks are dropped.
reason comes back as interrupted.
The flag is cleared at the start of each run, so an interrupt arriving between runs does not cancel the next one.
RunOutcome
run() and act() return this. underruns and
underrun_s report how long the buffer was dry.
| Field | Meaning |
|---|---|
actions | Steps handed to act(). |
chunks | Inferences issued, including the first. |
elapsed_s | Wall clock for the whole run. |
underruns | Times the buffer ran dry before the next chunk arrived. |
underrun_s | Total seconds the arm spent with no command. |
stopped_early | Whether it ended before max_actions. |
reason | Which ending happened: completed (ran out of actions), until (your predicate said the subtask was done), timeout (max_seconds elapsed), interrupted, validate (a callback refused an action), or raised:<ExceptionClass> — the exception's own class name. |
max_seam_jump | Largest per-dimension jump across a chunk boundary, over the run. |
action_space | What the chunk said its numbers mean. Tells you how to read the field above. |
observe_s | Total seconds spent inside your observe(). Anything but a small fraction of elapsed_s means the loop is waiting on your code, not the network. |
instruction_swaps | How many times set_instruction() changed the sentence in force. |
max_instruction_lag_s | Longest gap between a set_instruction() call and the request that first carried it. |
Chunks are deadlines
One call returns a block of future actions. covers_seconds says how much motion it
holds, and the next chunk has to arrive before the current one finishes playing.
| Model | Steps | Hz | Chunk covers |
|---|---|---|---|
cosmos3-edge-policy-droid | 32 | 15 | 2.13 s |
groot-n1-7-3b | 40 | 30 | 1.33 s |
pi05-droid | 15 | 15 | 1.00 s |
pi05-yam | 16 | 30 | 0.53 s |
The client reuses one connection — now straight to the worker — for the life of a
Policy. Measure the difference on your own network with
sequence-ai doctor.
Prefetch and underruns
Without prefetch the arm stops once per chunk for a full round trip.
run() starts the next inference while the current chunk plays, timed from a moving
average of observed refill latency. Pass prefetch=False to disable it.
A prefetched chunk is computed from an observation taken before the previous one finished
playing, so that overlap is open loop. prefetch=False restores strictly
closed-loop behaviour, stall included.
Only the HTTP call runs on the background worker thread. observe,
validate and act stay on the thread that called run().
The seam
A new chunk replaces the buffer. lookahead=N extends it instead, holding N chunks.
out.max_seam_jump reports the largest such distance over a run, measured before any
blending. Read it against out.action_space.
smooth_seam=N blends the lead-in: the first N steps of each new chunk ramp linearly
out of the last executed action. Absolute action spaces only, and never the first chunk
of a run — a delta chunk is passed through untouched even when you ask. Off by default.
Cold starts and warming
Warming happens at connect(), not mid-loop. A worker that scaled to zero answers
connect with 503 and a retry_after_s; in Python that is
Unavailable with .warming set. connect(),
run() and act() wait it out for you.
try: policy.predict(observation) except sequence_ai.Unavailable as exc: if exc.warming: print(f"loading; ready in ~{exc.retry_after_s}s") # a fixed estimate, not a live countdown
run() and act() wait warming out by default
(startup_timeout_s=1800), and only before the first action; warming met
mid-run is raised and hold fires. Each attempt re-observes.
startup_timeout_s=0 opts out.
How long it actually takes
First request against a model with no GPU yet, measured end to end:
worker VM up 50 s weights fetched 274 s 12.43 GB, first deploy only — see below extracted 8 s policy loaded onto the GPU ~70 s ——— first action ~400 s about seven minutes, once every action after it ~0.56 s
The weights fetch is paid once: a restart after an idle stop finds the
checkpoint on disk and skips it — a couple of minutes, not seven. That is why
startup_timeout_s defaults to 1800, not 300: a 300 s budget expires before a
first deploy is ready.
Pass on_warming to show progress rather than block silently. It is called with the
seconds waited so far and the current estimate:
def warming(waited_s, eta_s): print(f"starting the model, ~{eta_s/60:.0f} min to go") with sequence_ai.connect(model="pi05-droid", on_warming=warming) as policy: ...
A worker is stopped 10–25 minutes after your last call; the next call then pays the cold start above. A session that pauses for less keeps its worker and its sub-second latency.
You can also check readiness before you call, instead of reacting to a 503.
policy.ready() returns (ready, eta_seconds) without blocking;
policy.wait_until_ready(timeout_s=300) blocks until the model can serve and returns
the seconds it waited. policy.buffered is how many actions are left in the local
buffer — zero means the next call is a round trip. policy.spec() is the
model's catalogue row (action space, image geometry, robot), fetched once.
ready, eta = policy.ready() # (False, 180.0) while warming, (True, None) when live policy.wait_until_ready() # block until it can serve; returns seconds waited policy.buffered # actions left locally; 0 means the next call hits the network policy.spec() # catalogue row: action space, image size, robot
Many robots at once
One Policy per control loop: it holds one action buffer, and a second thread
driving the same one raises SequencesError.
def drive(robot, model): with sequence_ai.connect(model=model) as policy: # one each return policy.run(observe=robot.read, act=robot.apply, validate=robot.is_safe, max_actions=250) with ThreadPoolExecutor() as pool: left, right = pool.map(drive, [arm_l, arm_r], ["pi05-droid"] * 2)
The handshake is paid once per Policy, not once per call.
Concurrent loops are independent: each has its own worker, buffer and connection. Only your
account's concurrency and quota bound them, enforced at connect() time, not per
step.
Observation
Plain dataclasses, with the wire field names kept verbatim. Validation happens on the worker; unknown fields on a response are preserved rather than dropped.
from sequence_ai import ImageFrame, Observation, Proprioception obs = Observation( images=[ ImageFrame(view="exterior_1", data=b64_png), ImageFrame(view="wrist_left", data=b64_png), ], proprioception=Proprioception(joint_positions=[...]), instruction="pick up the red cube", )
Observation is images, proprioception and
instruction. All three are required.
view must be one of the names the model declares. Ask the catalogue rather than
guessing — sequence_ai.models() returns each model's camera_views,
and it needs no API key. A name that is not in that list arrives as an empty view rather than an
error.
Proprioception is joint_positions (required),
joint_velocities, end_effector_pose and gripper. The
gripper is its own field, not an extra entry in joint_positions, and it is a
list — a dual-arm setup reports two. ImageFrame is
view, data and an optional timestamp_ms.
Let the catalogue fill it in
Four per-model numbers decide whether a request is correct, not merely accepted: which
views, the state-vector width, the training aspect ratio, and what the returned floats mean.
None is guessable, and each fails silently — an unexpected camera is
uploaded, counted in frames_used, then ignored; a short state vector is
zero-padded. The robot moves either way, on something other than what you meant.
observation() drops a camera the model does not declare and warns naming it,
rather than encoding and sending it. Nothing the model asked for is ever dropped, and a
missing view is still an error — a frame you do not have cannot be supplied by
discarding a different one.
model = {m.short_id: m for m in sequence_ai.models(endpoint="act")}["pi05-droid"]
obs = sequence_ai.observation(
model,
images={"exterior_1": cam.read(), "wrist_left": wrist.read()}, # ndarray or PIL
joints=arm.joint_positions(), gripper=[arm.gripper_position()],
instruction="pick up the cup",
)
It sizes each frame the way the model's own server would — resize_with_pad for
pi0.5, short edge for GR00T — encodes JPEG at quality 95, and raises locally,
before anything is billed, when the views, the state width, or the aspect ratio do not
match. policy.predict() runs the same check and warns if you assembled the
observation yourself; connect(validate="strict") makes that raise and
"off" skips it along with the single catalogue lookup it costs.
state_dim is the input width and is not action_dim. Across the
served catalogue it is 8, 17 and 8. Read it per model rather than from this sentence: the list
changes when the catalogue does.
state_layout says how that width splits on the wire, and the halves are not
interchangeable: {"joint_positions": 7, "gripper": 1} for pi0.5-DROID. Eight joints
with no gripper sums to 8 and is accepted — the eighth joint is then dropped and the
gripper zero-filled, commanding the hand open for the whole episode.
observation() checks each half.
Arrays and PIL images need Pillow: pip install 'sequence-ai[image]'. Passing base64
strings from your own pipeline needs nothing — but then the sizing is yours, and it is
required, not optional: resize each frame to the model's
client_resize and JPEG-encode it, or call encode(frame, model), which
does exactly that off the catalogue. A raw camera frame at any other size is rejected
with 400: the model resizes any input to client_resize anyway,
so a larger one only costs upload (5–16× the bytes) and can silently exceed the
worker's body limit — /act no longer forwards it.
What to do with the rows that come back
Send the rows raw. Every model reports pass_raw: true; the checkpoint's own
reference loop takes the row, binarises the gripper, clips, steps. Nothing added.
env = RobotEnv(action_space="joint_velocity", gripper_action_space="position") chunk = policy.predict(obs).action_chunk for i in range(chunk.open_loop_horizon or len(chunk)): env.step(sequence_ai.prepare(chunk, model, i))
model.action_semantics names the controller interface those rows are for, so you
learn it before the first request. The catalogue spans four action spaces, and reading
one as another is invisible in the shape.
| model | action_space | what a row is |
|---|---|---|
pi05-droid | joint_absolute | absolute joint targets in radians, gripper 0/1 — command as-is |
cosmos3-edge-policy-droid | joint_absolute | a joint target already |
Convert only if your controller has no interface for that space — applying a conversion a controller did not need breaks a working arm exactly as thoroughly as skipping one it did.
q_ref = arm.joint_positions() # ONCE, before the chunk is requested for target in sequence_ai.to_joint_positions(chunk, q_ref, model): arm.move_to_joint_positions(target)
Every row offsets the same q_ref. The values are cumulative displacements,
so re-reading the live pose each step turns the chunk into an integrator and the arm overshoots.
to_joint_positions() refuses any model that does not declare a verified conversion
rather than inventing one — it refuses cosmos3-edge-policy-droid, whose rows
are already targets.
sequence_ai.recipes holds one module per supported model —
pi05_droid and cosmos3_edge_policy_droid — each a worked loop
with a check() that compares it against the live catalogue. for_model()
returns None for the rest.
Image aspect ratio
Every frame must arrive at the model's client_resize size —
the size encode() / observation() produce. /act returns
400 for anything else, including a raw camera frame at the right aspect ratio: the
model resizes any input to client_resize anyway, so a larger one only costs upload.
A model that declares no client_resize is checked on aspect ratio only — that
path, and how the letterbox works, is below.
# models() returns a list, so key it yourself -- there is no dict form. m = {x.short_id: x for x in sequence_ai.models()}["pi05-droid"] m.training_frame_size # [320, 180] what the DATASET stored -- not what the model reads m.input_aspect_ratio # 1.7778 (16:9) m.input_resize # "pad" m.client_resize # {"size": [224, 224], "filter": "bilinear"} -- or {"shortest_edge": 256, ...}, or None
A model whose input transform pads to a square letterboxes your frame; the ratio decides how
much of the square the scene fills, and a ratio the model never trained on shifts the bars. If
your camera is not 16:9, crop to it — do not pad. The SDK will not do
this for you: it refuses the frame with a local ValueError naming the ratio, since
which part of the scene to give up is your decision.
The size is enforced, not just the ratio. /act rejects a frame
that is not the exact client_resize size with a 400; the model resizes
any input down anyway, so a larger one only costs upload and can exceed the worker's body limit.
Hand your native frame to encode() / observation() — the
codec is JPEG at quality 95 — and never hand-downscale to training_frame_size
first (a two-step resample that degrades the frame). A model with no client_resize
is capped at 640 on the long edge.
Only joint_positions is required. Omitting a field is not the same as sending zeros:
zeros claim the joints are stationary at the origin, omission says “not reported”,
and the model treats the two differently.
data is base64 PNG or JPEG. A URL is not fetched — the server
never dereferences a caller-supplied address, so a URL put here reaches the worker as an
unparseable image rather than as a picture.
Decide and perception
decide() is a one-shot call to a scoring model (jev): a
state (string, object, or array) and typed questions in, calibrated probabilities
out. It scores rather than generates — no max_tokens or
temperature, billed for input tokens only. Use it for routing, ranking and
verification rather than chat.
from sequence_ai import decide result = decide( "jev", "Help! My payouts have been failing for 3 days.", { "is_urgent": {"type": "noul", "instructions": "Does this convey urgency?", "criteria": {"true": "Time-sensitive", "false": "Not urgent"}}, "department": {"type": "choice", "instructions": "Which team handles this?", "criteria": {"billing": "Payments", "technical": "Bugs", "sales": "Pricing"}}, }, ) for key, ans in result.answers.items(): if ans["type"] == "noul": print(key, ans["noul"]) # is_urgent 0.96 elif ans["type"] == "choice": print(key, ans["choice"], ans["confidence"]) # department billing 0.81
Each question's type fixes the answer's shape: noul is a 0–1
probability, choice picks from options you name and carries the full distribution
beside a confidence, score is a position on an ordered rubric. The keys
are yours and each answer comes back under the same key, so branch on
answers[key]["type"] and gate on confidence rather than trusting the
single pick. Like chat(), decide() is answered by an external provider,
so there is no cold-start wait — a 503 means the provider is down, not warming.
decide(), detect(), depth() and embed() are
the SDK's calls that stand apart from the Policy loop — scoring and perception
rather than control. The models behind them live in the catalogue: the scoring model under
Decision, the perception endpoints under
Perception.
Errors
Everything raised inherits SequencesError, so catching that catches all of ours
and none of httpx's.
| Class | Means | Do |
|---|---|---|
AuthError | Key missing, revoked or expired | Stop. Retrying will not help. |
OutOfCredit | Balance exhausted | Stop and hold. Top up. |
InvalidRequest | Unknown model, malformed observation, body too large | Fix it. Deterministic — the identical request fails identically. |
Unavailable | Upstream blip, timeout, or a cold worker | Check .warming. Failed calls are not billed. |
ChunkExhausted | Asked for an action with an empty buffer and no observation | Pass observation= every call. |
Unavailable carries .warming, .retry_after_s and
.job_id. warming=True means retry after the stated interval;
warming=False means the upstream failed and retrying is a guess. A rotated secret or
an expired lease does not reach you as an error — it is the one case the client
resolves itself, by reconnecting through the gateway and replaying (see
fallback).
Wire endpoints
For a controller not written in Python. connect reaches the gateway;
/act reaches the worker directly.
Client → gateway, once. Admits the call and ensures a worker is up; returns
worker_url, cert_fingerprint, direct_secret and a
lease (ttl_s, expires_at, epoch). A cold
worker answers 503 with Retry-After while it boots.
Client → worker, directly — the gateway is not in this path. Runs inference and
returns an action_chunk. A 401 (rotated secret) or lease expiry sends
the client back to connect to fetch fresh coordinates.
Client → gateway, under your API key. The read-only status of the account the key belongs
to: balance (money left — the platform is prepaid), usage (active
policies and benchmarks) and limits (the per-account caps that are configured).
Creating keys and topping up stay in the console.
import sequence_ai acct = sequence_ai.account() acct["balance"]["balance"] # USD left on the prepaid account acct["usage"] # {"policies": 1, "benchmarks": 0}
The full three-contract wire — POST /v1/deployments to register a policy,
connect, and the direct /act — is laid out in the
platform docs.
Command line
Installing the package also installs sequence-ai, with seqai as an
alias.
sequence-ai 0.3.0
key SEQUENCES_API_KEY (seq_live_…a1b2)
gateway up auth=keystore # host resolved from the key; not configured
key valid yes
credit available
connect pi05-droid -> worker up, cert pinned, lease 300s
latency new connection each call <measured on your network>
one reused connection <measured on your network>
saved by keeping it open <the difference>
ready
doctor checks the key, the gateway, credit, and a real direct connect — it
opens one to a model, confirms the cert pins and the lease comes back, and times a reused
connection against a fresh one.
sequence-ai models lists the catalogue: how many dimensions come back, and which
cameras to send. It needs no key.