Python client

The control loop, in Python

sequence_ai is the client behind Running a policy. connect() opens a kept-alive line straight to the worker; act() and run() play the action chunks it returns and raise typed errors. This is the reference for that loop.

Install

One dependency (httpx), Python 3.9+. Authoring and deploying a policy add a heavier extra (sequence-ai[authoring]); this page is the client.

installshell
pip install sequence-ai                      # the client: connect / act / run
pip install 'sequence-ai[image]'                # adds Pillow, for ndarray / PIL observations

Install sequence-ai, import sequence_ai.

Authentication

One key identifies your tenant. The client reads it from api_key= first, then $SEQUENCES_API_KEY, and resolves the gateway from it — no host to configure.

set a keyshell
export SEQUENCES_API_KEY="seq_live_..."

Create one in the console. sequence_ai.models() needs no key.

Quickstart

A complete control loop. connect() is a context manager, so the connection closes even when the loop raises.

control looppython
import sequence_ai

with sequence_ai.connect(model="pi05-droid") as policy:
    out = policy.run(
        # required
        observe=robot.read_observation,   # -> Observation
        act=robot.apply_action,           # one step, a list of floats
        max_actions=250,               # 16.7 s at 15 Hz
        # optional, and these three are the ones worth adding
        validate=robot.is_safe,           # return False to stop; warns if omitted
        hold=robot.hold,                  # called once if it stops early
        on_underrun=robot.hold_position,  # buffer ran dry — hold, do not coast
    )

print(out)                                # RunOutcome(250 actions over 17 chunks in 16.8s, completed)

Required: observe, act, max_actions. Everything else has a default. validate is your bounds check and run() warns if it is omitted; hold runs on any early ending; on_underrun fires when a refill was late.

An action returned by any model is model output, not a safe robot command. This library checks no joint limit, reachability, collision or velocity. Put a bounds check, a watchdog and a physical e-stop between it and your motors — attach them at validate and hold.

connect() opens a direct line to the worker

connect() returns a Policy handle and keeps one connection open for the life of the loop. A call is two phases; the gateway is in only the first.

phasereachespurpose
connect(model)the gateway, onceadmit, ensure a worker is up, return its worker_url, cert_fingerprint, direct_secret and lease
act() / run()the worker, directlyrun inference on the reused, cert-pinned connection; Bearer <direct_secret>
reconnectthe gateway, on demandon a rotated secret, expired lease or a worker that scaled to zero — fetch fresh coordinates and replay

Direct connect and transparent fallback

You do not handle the 401 or 503 yourself. A rotated direct_secret, an expired lease, a worker that scaled to zero, or a dropped connection all surface as one transparent reconnect: the client fetches fresh coordinates and replays the in-flight call. The same Policy handle keeps working.

A cold worker answers connect() with 503 and a Retry-After while it boots; the client waits that out — see cold starts and warming.

Four ways to drive it

The same request underneath, at four levels of control. The instruction is a field of Observation, so the first three carry it with the observation; act() takes it as an argument and overwrites whatever the observation holds.

CallReturns
predict(obs)Prediction — one whole action_chunk, plus perf, usage and warnings. You play it out.
next_action(obs)One action, a list[float]. Refills from an internal buffer when it empties.
run(observe, act)RunOutcome, after playing max_actions steps.
act(instruction, observe, act)RunOutcome, after playing one instruction until until, max_seconds or interrupt().
predict — the whole chunkpython
pred = policy.predict(observation)
print(len(pred.action_chunk), "steps covering",
      pred.action_chunk.covers_seconds, "s at",
      pred.action_chunk.control_frequency_hz, "Hz")
next_action — we hold the bufferpython
while running:
    action = policy.next_action(robot.read_observation())
    robot.apply_action(action)

Pass observation= on every call. It is ignored while the buffer has actions, and is what the refill is computed from when the buffer is empty. An empty buffer with no observation raises ChunkExhausted.

run()

The full control-loop API. validate and hold are its parameters, not a wrapper you build — attach your bounds check and your stop-behaviour here.

ParameterDefaultWhat it does
observerequiredReturns the current Observation. Called on the calling thread, never off it.
actrequiredCalled once per action with one step, a list of floats.
max_actionsrequired, keyword-onlyUpper bound on steps executed.
validateNoneCalled before every act. Return False to stop. Omitting it warns.
holdNoneCalled once if the loop stops early, including on an exception.
on_chunkNoneCalled with each Prediction after a refill.
on_underrunNoneCalled with the stall in seconds when the buffer ran dry.
on_statusNoneCalled with starting, running, waiting, done.
progress"auto"Print status to stderr. "auto" means only when stderr is a terminal.
prefetchTrueOverlap the next inference with playback.
lookahead1Chunks to hold in hand or on the way. Above 1 the run is open loop over that depth.
open_loop_horizonNonePlay only the first N steps of each chunk before refilling.
max_secondsNoneStop after this much wall-clock. reason becomes timeout.
untilNoneChecked before each act; returning True ends the run with reason until.
startup_timeout_s1800.0Wait this long for a cold worker, before the first action only. 0 opts out. See how long it takes.
smooth_seam0Blend N steps across a chunk boundary. Absolute action spaces only.

max_actions is required and keyword-only. There is no run_forever().

hold runs in a finally block and fires on the exception path too.

With validate=None nothing is checked before actuation. Allowed for simulation and bench work; it warns once per process, never silently on real hardware.

act() — one instruction, a time bound

act() pins one instruction and drives until it is done, a deadline, or an interrupt — the shape a higher-level planner calls as a tool. It returns a RunOutcome like run().

act — the tool a planner callspython
out = policy.act(
    "pick up the cup",
    observe=robot.read_observation,   # cameras + joints
    act=robot.apply_action,           # one joint target
    validate=robot.is_safe,
    hold=robot.hold,
    until=robot.cup_in_gripper,       # optional: return True to finish early
    max_seconds=30,                   # required — this drives hardware
)

act(instruction, observe, act, *, max_seconds, until=None). max_seconds is required and has no default. act() overwrites the observation's instruction on every call.

Wiring instructions together — retries, failure detection, choosing the next sub-goal — is yours. For a server-side planner, see a composite policy, where the VLM runs on your key.

Changing the instruction without stopping

set_instruction — from the planner's threadpython
policy.set_instruction("now put it on the plate")   # returns immediately; the arm keeps moving
policy.instruction                                # what is in force, or None

Thread-safe. The new sentence takes effect on the next request; the arm is not paused. Setting the same string again is ignored.

Call it from a thread that is not the one driving. Putting the planner call inside observe() runs it on the driving thread and the arm stands still for its length; on its own thread, the loop keeps prefetching under the instruction currently in force.

Use set_instruction() to refine or continue, and interrupt() to abandon. The difference is what happens to the robot: the first keeps it moving, the second stops it within one control period.

The new instruction is not instant: the in-flight request was built from the old one, so the first action reflecting the call arrives at most lookahead + 1 chunks later. RunOutcome.instruction_swaps and max_instruction_lag_s report it.

Changing your mind

interrupt — from any thread, at any momentpython
policy.interrupt()

Thread-safe. Call it from any thread while run() or act() is driving.

Stops between actions, not at a chunk boundary — about 67 ms at 15 Hz.

Also interrupts a fetch that is still in flight, including during a cold start or a reconnect.

Unwinds through hold() and is not raised. Buffered and in-flight chunks are dropped. reason comes back as interrupted.

The flag is cleared at the start of each run, so an interrupt arriving between runs does not cancel the next one.

RunOutcome

run() and act() return this. underruns and underrun_s report how long the buffer was dry.

FieldMeaning
actionsSteps handed to act().
chunksInferences issued, including the first.
elapsed_sWall clock for the whole run.
underrunsTimes the buffer ran dry before the next chunk arrived.
underrun_sTotal seconds the arm spent with no command.
stopped_earlyWhether it ended before max_actions.
reasonWhich ending happened: completed (ran out of actions), until (your predicate said the subtask was done), timeout (max_seconds elapsed), interrupted, validate (a callback refused an action), or raised:<ExceptionClass> — the exception's own class name.
max_seam_jumpLargest per-dimension jump across a chunk boundary, over the run.
action_spaceWhat the chunk said its numbers mean. Tells you how to read the field above.
observe_sTotal seconds spent inside your observe(). Anything but a small fraction of elapsed_s means the loop is waiting on your code, not the network.
instruction_swapsHow many times set_instruction() changed the sentence in force.
max_instruction_lag_sLongest gap between a set_instruction() call and the request that first carried it.

Chunks are deadlines

One call returns a block of future actions. covers_seconds says how much motion it holds, and the next chunk has to arrive before the current one finishes playing.

ModelStepsHzChunk covers
cosmos3-edge-policy-droid32152.13 s
groot-n1-7-3b40301.33 s
pi05-droid15151.00 s
pi05-yam16300.53 s

The client reuses one connection — now straight to the worker — for the life of a Policy. Measure the difference on your own network with sequence-ai doctor.

Prefetch and underruns

Without prefetch the arm stops once per chunk for a full round trip.

run() starts the next inference while the current chunk plays, timed from a moving average of observed refill latency. Pass prefetch=False to disable it.

A prefetched chunk is computed from an observation taken before the previous one finished playing, so that overlap is open loop. prefetch=False restores strictly closed-loop behaviour, stall included.

Only the HTTP call runs on the background worker thread. observe, validate and act stay on the thread that called run().

The seam

A new chunk replaces the buffer. lookahead=N extends it instead, holding N chunks.

out.max_seam_jump reports the largest such distance over a run, measured before any blending. Read it against out.action_space.

smooth_seam=N blends the lead-in: the first N steps of each new chunk ramp linearly out of the last executed action. Absolute action spaces only, and never the first chunk of a run — a delta chunk is passed through untouched even when you ask. Off by default.

Cold starts and warming

Warming happens at connect(), not mid-loop. A worker that scaled to zero answers connect with 503 and a retry_after_s; in Python that is Unavailable with .warming set. connect(), run() and act() wait it out for you.

a warming workerpython
try:
    policy.predict(observation)
except sequence_ai.Unavailable as exc:
    if exc.warming:
        print(f"loading; ready in ~{exc.retry_after_s}s")   # a fixed estimate, not a live countdown

run() and act() wait warming out by default (startup_timeout_s=1800), and only before the first action; warming met mid-run is raised and hold fires. Each attempt re-observes. startup_timeout_s=0 opts out.

How long it actually takes

First request against a model with no GPU yet, measured end to end:

worker VM up                 50 s
weights fetched             274 s   12.43 GB, first deploy only — see below
extracted                     8 s
policy loaded onto the GPU  ~70 s
———
first action               ~400 s   about seven minutes, once
every action after it      ~0.56 s

The weights fetch is paid once: a restart after an idle stop finds the checkpoint on disk and skips it — a couple of minutes, not seven. That is why startup_timeout_s defaults to 1800, not 300: a 300 s budget expires before a first deploy is ready.

Pass on_warming to show progress rather than block silently. It is called with the seconds waited so far and the current estimate:

def warming(waited_s, eta_s):
    print(f"starting the model, ~{eta_s/60:.0f} min to go")

with sequence_ai.connect(model="pi05-droid", on_warming=warming) as policy:
    ...

A worker is stopped 10–25 minutes after your last call; the next call then pays the cold start above. A session that pauses for less keeps its worker and its sub-second latency.

You can also check readiness before you call, instead of reacting to a 503. policy.ready() returns (ready, eta_seconds) without blocking; policy.wait_until_ready(timeout_s=300) blocks until the model can serve and returns the seconds it waited. policy.buffered is how many actions are left in the local buffer — zero means the next call is a round trip. policy.spec() is the model's catalogue row (action space, image geometry, robot), fetched once.

check before you callpython
ready, eta = policy.ready()   # (False, 180.0) while warming, (True, None) when live
policy.wait_until_ready()     # block until it can serve; returns seconds waited
policy.buffered               # actions left locally; 0 means the next call hits the network
policy.spec()                 # catalogue row: action space, image size, robot

Many robots at once

One Policy per control loop: it holds one action buffer, and a second thread driving the same one raises SequencesError.

two armspython
def drive(robot, model):
    with sequence_ai.connect(model=model) as policy:   # one each
        return policy.run(observe=robot.read, act=robot.apply,
                          validate=robot.is_safe, max_actions=250)

with ThreadPoolExecutor() as pool:
    left, right = pool.map(drive, [arm_l, arm_r], ["pi05-droid"] * 2)

The handshake is paid once per Policy, not once per call.

Concurrent loops are independent: each has its own worker, buffer and connection. Only your account's concurrency and quota bound them, enforced at connect() time, not per step.

Observation

Plain dataclasses, with the wire field names kept verbatim. Validation happens on the worker; unknown fields on a response are preserved rather than dropped.

building an observationpython
from sequence_ai import ImageFrame, Observation, Proprioception

obs = Observation(
    images=[
        ImageFrame(view="exterior_1", data=b64_png),
        ImageFrame(view="wrist_left",  data=b64_png),
    ],
    proprioception=Proprioception(joint_positions=[...]),
    instruction="pick up the red cube",
)

Observation is images, proprioception and instruction. All three are required.

view must be one of the names the model declares. Ask the catalogue rather than guessing — sequence_ai.models() returns each model's camera_views, and it needs no API key. A name that is not in that list arrives as an empty view rather than an error.

Proprioception is joint_positions (required), joint_velocities, end_effector_pose and gripper. The gripper is its own field, not an extra entry in joint_positions, and it is a list — a dual-arm setup reports two. ImageFrame is view, data and an optional timestamp_ms.

Let the catalogue fill it in

Four per-model numbers decide whether a request is correct, not merely accepted: which views, the state-vector width, the training aspect ratio, and what the returned floats mean. None is guessable, and each fails silently — an unexpected camera is uploaded, counted in frames_used, then ignored; a short state vector is zero-padded. The robot moves either way, on something other than what you meant.

observation() drops a camera the model does not declare and warns naming it, rather than encoding and sending it. Nothing the model asked for is ever dropped, and a missing view is still an error — a frame you do not have cannot be supplied by discarding a different one.

one call, four numbers read off the cataloguepython
model = {m.short_id: m for m in sequence_ai.models(endpoint="act")}["pi05-droid"]

obs = sequence_ai.observation(
    model,
    images={"exterior_1": cam.read(), "wrist_left": wrist.read()},   # ndarray or PIL
    joints=arm.joint_positions(), gripper=[arm.gripper_position()],
    instruction="pick up the cup",
)

It sizes each frame the way the model's own server would — resize_with_pad for pi0.5, short edge for GR00T — encodes JPEG at quality 95, and raises locally, before anything is billed, when the views, the state width, or the aspect ratio do not match. policy.predict() runs the same check and warns if you assembled the observation yourself; connect(validate="strict") makes that raise and "off" skips it along with the single catalogue lookup it costs.

state_dim is the input width and is not action_dim. Across the served catalogue it is 8, 17 and 8. Read it per model rather than from this sentence: the list changes when the catalogue does.

state_layout says how that width splits on the wire, and the halves are not interchangeable: {"joint_positions": 7, "gripper": 1} for pi0.5-DROID. Eight joints with no gripper sums to 8 and is accepted — the eighth joint is then dropped and the gripper zero-filled, commanding the hand open for the whole episode. observation() checks each half.

Arrays and PIL images need Pillow: pip install 'sequence-ai[image]'. Passing base64 strings from your own pipeline needs nothing — but then the sizing is yours, and it is required, not optional: resize each frame to the model's client_resize and JPEG-encode it, or call encode(frame, model), which does exactly that off the catalogue. A raw camera frame at any other size is rejected with 400: the model resizes any input to client_resize anyway, so a larger one only costs upload (5–16× the bytes) and can silently exceed the worker's body limit — /act no longer forwards it.

What to do with the rows that come back

Send the rows raw. Every model reports pass_raw: true; the checkpoint's own reference loop takes the row, binarises the gripper, clips, steps. Nothing added.

a Franka on the DROID stackpython
env = RobotEnv(action_space="joint_velocity", gripper_action_space="position")

chunk = policy.predict(obs).action_chunk
for i in range(chunk.open_loop_horizon or len(chunk)):
    env.step(sequence_ai.prepare(chunk, model, i))

model.action_semantics names the controller interface those rows are for, so you learn it before the first request. The catalogue spans four action spaces, and reading one as another is invisible in the shape.

modelaction_spacewhat a row is
pi05-droidjoint_absoluteabsolute joint targets in radians, gripper 0/1 — command as-is
cosmos3-edge-policy-droidjoint_absolutea joint target already

Convert only if your controller has no interface for that space — applying a conversion a controller did not need breaks a working arm exactly as thoroughly as skipping one it did.

only for a position-only controllerpython
q_ref = arm.joint_positions()                  # ONCE, before the chunk is requested
for target in sequence_ai.to_joint_positions(chunk, q_ref, model):
    arm.move_to_joint_positions(target)

Every row offsets the same q_ref. The values are cumulative displacements, so re-reading the live pose each step turns the chunk into an integrator and the arm overshoots. to_joint_positions() refuses any model that does not declare a verified conversion rather than inventing one — it refuses cosmos3-edge-policy-droid, whose rows are already targets.

sequence_ai.recipes holds one module per supported model — pi05_droid and cosmos3_edge_policy_droid — each a worked loop with a check() that compares it against the live catalogue. for_model() returns None for the rest.

Image aspect ratio

Every frame must arrive at the model's client_resize size — the size encode() / observation() produce. /act returns 400 for anything else, including a raw camera frame at the right aspect ratio: the model resizes any input to client_resize anyway, so a larger one only costs upload. A model that declares no client_resize is checked on aspect ratio only — that path, and how the letterbox works, is below.

what the catalogue declarespython
# models() returns a list, so key it yourself -- there is no dict form.
m = {x.short_id: x for x in sequence_ai.models()}["pi05-droid"]
m.training_frame_size   # [320, 180]  what the DATASET stored -- not what the model reads
m.input_aspect_ratio    # 1.7778  (16:9)
m.input_resize          # "pad"
m.client_resize         # {"size": [224, 224], "filter": "bilinear"}  -- or {"shortest_edge": 256, ...}, or None

A model whose input transform pads to a square letterboxes your frame; the ratio decides how much of the square the scene fills, and a ratio the model never trained on shifts the bars. If your camera is not 16:9, crop to it — do not pad. The SDK will not do this for you: it refuses the frame with a local ValueError naming the ratio, since which part of the scene to give up is your decision.

The size is enforced, not just the ratio. /act rejects a frame that is not the exact client_resize size with a 400; the model resizes any input down anyway, so a larger one only costs upload and can exceed the worker's body limit. Hand your native frame to encode() / observation() — the codec is JPEG at quality 95 — and never hand-downscale to training_frame_size first (a two-step resample that degrades the frame). A model with no client_resize is capped at 640 on the long edge.

Only joint_positions is required. Omitting a field is not the same as sending zeros: zeros claim the joints are stationary at the origin, omission says “not reported”, and the model treats the two differently.

data is base64 PNG or JPEG. A URL is not fetched — the server never dereferences a caller-supplied address, so a URL put here reaches the worker as an unparseable image rather than as a picture.

Decide and perception

decide() is a one-shot call to a scoring model (jev): a state (string, object, or array) and typed questions in, calibrated probabilities out. It scores rather than generates — no max_tokens or temperature, billed for input tokens only. Use it for routing, ranking and verification rather than chat.

decidepython
from sequence_ai import decide

result = decide(
    "jev",
    "Help! My payouts have been failing for 3 days.",
    {
        "is_urgent":   {"type": "noul",   "instructions": "Does this convey urgency?",
                        "criteria": {"true": "Time-sensitive", "false": "Not urgent"}},
        "department":  {"type": "choice", "instructions": "Which team handles this?",
                        "criteria": {"billing": "Payments", "technical": "Bugs", "sales": "Pricing"}},
    },
)

for key, ans in result.answers.items():
    if ans["type"] == "noul":
        print(key, ans["noul"])                       # is_urgent 0.96
    elif ans["type"] == "choice":
        print(key, ans["choice"], ans["confidence"])   # department billing 0.81

Each question's type fixes the answer's shape: noul is a 0–1 probability, choice picks from options you name and carries the full distribution beside a confidence, score is a position on an ordered rubric. The keys are yours and each answer comes back under the same key, so branch on answers[key]["type"] and gate on confidence rather than trusting the single pick. Like chat(), decide() is answered by an external provider, so there is no cold-start wait — a 503 means the provider is down, not warming.

decide(), detect(), depth() and embed() are the SDK's calls that stand apart from the Policy loop — scoring and perception rather than control. The models behind them live in the catalogue: the scoring model under Decision, the perception endpoints under Perception.

Errors

Everything raised inherits SequencesError, so catching that catches all of ours and none of httpx's.

ClassMeansDo
AuthErrorKey missing, revoked or expiredStop. Retrying will not help.
OutOfCreditBalance exhaustedStop and hold. Top up.
InvalidRequestUnknown model, malformed observation, body too largeFix it. Deterministic — the identical request fails identically.
UnavailableUpstream blip, timeout, or a cold workerCheck .warming. Failed calls are not billed.
ChunkExhaustedAsked for an action with an empty buffer and no observationPass observation= every call.

Unavailable carries .warming, .retry_after_s and .job_id. warming=True means retry after the stated interval; warming=False means the upstream failed and retrying is a guess. A rotated secret or an expired lease does not reach you as an error — it is the one case the client resolves itself, by reconnecting through the gateway and replaying (see fallback).

Wire endpoints

For a controller not written in Python. connect reaches the gateway; /act reaches the worker directly.

POST/v1/act/connect

Client → gateway, once. Admits the call and ensures a worker is up; returns worker_url, cert_fingerprint, direct_secret and a lease (ttl_s, expires_at, epoch). A cold worker answers 503 with Retry-After while it boots.

POST{worker_url}/act

Client → worker, directly — the gateway is not in this path. Runs inference and returns an action_chunk. A 401 (rotated secret) or lease expiry sends the client back to connect to fetch fresh coordinates.

GET/v1/account

Client → gateway, under your API key. The read-only status of the account the key belongs to: balance (money left — the platform is prepaid), usage (active policies and benchmarks) and limits (the per-account caps that are configured). Creating keys and topping up stay in the console.

check money left before a big runpython
import sequence_ai

acct = sequence_ai.account()
acct["balance"]["balance"]   # USD left on the prepaid account
acct["usage"]                  # {"policies": 1, "benchmarks": 0}

The full three-contract wire — POST /v1/deployments to register a policy, connect, and the direct /act — is laid out in the platform docs.

Command line

Installing the package also installs sequence-ai, with seqai as an alias.

sequence-ai doctor
sequence-ai 0.3.0

  key           SEQUENCES_API_KEY  (seq_live_…a1b2)
  gateway       up   auth=keystore            # host resolved from the key; not configured
  key valid     yes
  credit        available

  connect       pi05-droid  ->  worker up, cert pinned, lease 300s
  latency       new connection each call      <measured on your network>
                one reused connection         <measured on your network>
                saved by keeping it open      <the difference>

ready

doctor checks the key, the gateway, credit, and a real direct connect — it opens one to a model, confirms the cert pins and the lease comes back, and times a reused connection against a fresh one.

sequence-ai models lists the catalogue: how many dimensions come back, and which cameras to send. It needs no key.