Reference
Wire endpoints
For a controller not written in Python. deploy and connect reach the
gateway; /act reaches the worker directly.
SDK → gateway. Registers a policy from its spec (image, weights, resources, autoscaler).
Returns 202 building; poll GET /v1/deployments/{id} until
ready.
Client → gateway, once. Admits the call and ensures a worker is up; returns
worker_url, cert_fingerprint, direct_secret and a
lease. A cold worker answers 503 with Retry-After while
it boots.
Client → worker, directly — the gateway is not in this path. Runs inference and
returns an action_chunk. A 401 (rotated secret) or lease expiry sends
the client back to connect to fetch fresh coordinates.
Client → gateway, under your API key. The read-only status of the account the key belongs
to: balance (money left — the platform is prepaid; an account with a contracted credit line
also gets credit_limit_usd and spendable, the two together), usage (active
policies and benchmarks) and limits (the per-account caps that are configured).
Creating keys and topping up stay in the console.
import sequence_ai acct = sequence_ai.account() acct["balance"]["balance"] # USD left on the prepaid account acct["usage"] # {"policies": 1, "benchmarks": 0}
Client → gateway, under your API key. Your GPU card-hours per hour or day: from and
to (ISO dates or times; the last 24 hours by default, at most 90 days), by
(hour up to a week, day beyond, unless given), tz (an IANA zone the
buckets start in, UTC by default) and group (none, kind,
deployment or gpu). The answer has total_card_hours, total_usd
and one entry per bucket and group in buckets: start, group,
card_hours, cost_usd and the deployments that ran. seq usage
prints it.
Client → gateway, under a full key of the account (a robot’s act key is refused). One row per robot
— per API key: name, last_four, role, state
(acting, connected, queued, offline, revoked),
deployment, version, last_seen, requests_1h,
errors_1h, p50_ms, p95_ms, labeled_24h,
success_rate_7d — in data, with total; limit and
offset page it. GET /v1/robots/{name}?hours=6 adds the robot’s latency per minute
(latency), its time budget (budget_ms, from chunk_frames and
control_hz), its sessions and labeled episodes;
GET /v1/robots/{name}/logs its lines; GET /v1/fleet the robots by state and each
deployment’s workers, robots and queued.
Official templates
The templates are in the model library: the physical AI templates and the
benchmark templates. seq init <template> [dir] writes
one into a folder of your own, from the platform; seq init --list is the live list, and each
template’s README says how to use it. A policy template deploys as your own policy (seq deploy),
a benchmark registers as your own benchmark, and an adapter registers as your own adapter, named in an eval with
--adapter.
The seq CLI
| command | what it does |
|---|---|
seq init <template> [dir] | start from a template, a policy, a benchmark or an adapter — see Official templates; seq init --list shows them all |
seq validate <file> | check a policy, benchmark or adapter against the platform contract, offline — nothing is built |
seq deploy <file> [--name <name>] | build and deploy a policy as <name>, or register a benchmark or an adapter; ready when the command returns. --dry-run prints the request instead |
seq eval run <policy> --benchmark <name> | score a deployed policy on a registered benchmark — --adapter name@vN, --suite, --trials N, --max-usd X; --rerun-failed EVAL_ID runs an eval’s errored episodes again into its report, --rerun-unsuccessful EVAL_ID (--metric NAME) runs the ones it did not succeed on as an eval of their own; prints each metric overall, per suite and per task, and the cost — every flag in Run an eval |
seq eval report <eval_id> | an eval’s report; --wait until it finishes; --videos DIR saves its recordings; --csv PATH one row per episode; --trajectories DIR every episode’s states and actions. seq eval cancel <eval_id> stops one |
seq eval status <eval_id> | where an eval is — waiting for GPUs, running or ended — with episodes done of the total and every metric so far; --watch until it ends |
seq eval logs <eval_id> | what an eval’s machines printed: the benchmark instances’ output and the policy instances’ lines; --side, --arm, --instance, --follow; kept 7 days |
seq eval ls <policy> | a deployed policy’s evals over time, and its schedules; seq eval ls --benchmark B lists every eval on that benchmark instead; seq eval schedule <policy> --benchmark B --every daily|weekly|deploy runs one on a schedule (--off) |
seq adapter ls | your adapters: latest version, versions, lossy or lossless, the last eval that ran each; seq adapter show <name>[@vN] says whether a version is lossy, seq adapter pull <name>[@vN] writes its file back to edit and deploy again, seq adapter check <file or name@vN> [--benchmark B] [--policy P] checks the labels and, against a benchmark or a policy, runs both pipelines on samples of what each side gives; such a check is kept with the version for its page in the console (--no-record keeps nothing), seq adapter rm <name> removes one (never while an eval runs it; its numbers are not given again) — what an adapter can be is in Adapters |
seq benchmark ls | your benchmarks, newest first: version, status, card, environments, latest eval — 50 a page (--limit, --cursor) |
seq benchmark evals <name> | every eval run on a benchmark, across your policies: when, policy and version, benchmark version, adapter, episodes, cost, the primary metric with its interval; --policy, --since |
seq benchmark logs <name> | a version’s build and check log, the check’s own output included (--version N) |
seq benchmark export | check-model | scaffold | compare | convert | port a classic-MuJoCo benchmark to MuJoCo Warp, or a SAPIEN one to SAPIEN’s GPU PhysX, and compare the two versions on the same episodes — Move a CPU benchmark to the GPU |
seq benchmark status <name> | a registered benchmark’s build state: building, ready or failed |
seq benchmark rm <name> | delete a benchmark: its name and quota place are free at once (--dry-run, --force cancels the evals running on it) |
seq policy ls | your deployments: the serving version, one building beside it, the card, the weights |
seq policy status <name> | a deployment’s state: its build (building, ready or failed), each worker with its robots and how long it has served, the queue, and what it costs an hour (--wait follows a build) |
seq policy versions <name> | its versions: building (or failed), serving, kept for rollback. A re-deploy builds beside the serving version, which keeps serving until the new one is ready — and if the new one fails |
seq policy rollback <name> [--to N] | serve a kept earlier version again, at once — nothing is rebuilt |
seq policy traffic <name> v5=10 v4=90 | split new sessions between ready versions; --pin KEY v5 / --unpin KEY for one robot key; --clear; no split shows the current one |
seq policy budget <name> --monthly USD | stop the deployment when this month’s spend reaches it; --off; without a flag shows it |
seq policy scale <name> --min N --during '…' | keep N workers warm in those hours (--tz, UTC by default); --clear |
seq policy alerts <name> --webhook URL | post signed alerts: --error-rate, --p95-ms, --budget-pct, failed builds; --off |
seq policy capture <name> --rate R | keep that share of episodes (--images for frames, --retention-days N); --off; without a flag shows what is kept |
seq policy data export <name> --out DIR | download the captured episodes as a LeRobot dataset (--format jsonl for the raw records) |
seq policy stop <name> / seq policy start <name> | emergency stop: every worker ends now and every connect is refused; start lets robots connect again |
seq policy rm <name> | delete it: workers stop now, the name and the quota slot are freed |
seq policy rename <name> <new> | give it a new name to connect by; its versions, keys and settings stay with it. Running workers end, and robots reconnect under the new name |
seq policy logs <name> [--rid r-…] | your workers’ own lines (load, every act with its request id and timings, errors) beside each connect and the timings your robots reported; --rid joins one request across them; --build shows the newest build instead. Kept 7 days |
seq policy doctor <name> | checks the deployment on the platform: its state and version, a running worker’s /ping, one /act on a synthetic observation from your contract, the round trip split into network and inference. Never starts a worker |
seq policy estimate <file> | price a policy before deploying it: the card per hour, the warm floor, use per day and month |
seq usage | your GPU card-hours per hour (or --by day) with each one’s cost and what ran: the last 24 hours, --days N, or --from/--to up to 90 days; --group kind|deployment|gpu; --tz (this machine’s zone by default) |
seq audit | who did what to your account (--limit, --before; seq --json audit for JSON) |
seq volume put <name> <dir> | upload a checkpoint or dataset directory (resumable; md5-checked); seq volume ls, seq volume rm <name> |
seq secret set <name> --env KEY | store a credential by name (value read from stdin) |
Secrets (BYOK)
Store a credential once by name and refer to it from a policy by that name; the value never enters your code, your logs, or any output. A secret is one name bound to a set of environment variables.
seq secret set hf --env HF_TOKEN # private hf:// weights (and your policy's own use) # there is no --value flag: pipe it in, or type it at the hidden prompt printf '%s' "$HF_TOKEN" | seq secret set hf --env HF_TOKEN seq secret ls # names + env-var NAMES + timestamps; the value is never shown
There is no command that reads a stored value back — a secret is
write-only, and re-running set rotates the whole bundle. Use it from a policy by
name:
@seq.policy(gpu="L4", max_gpus=1, secrets=[seq.Secret.from_name("hf")]) class MyPolicy(seq.Policy): ...
Error codes
Every response that is not 2xx carries one JSON body: detail (a sentence),
code, retryable, retry_after_s, request_id,
hint and details. Switch on code rather than the status, and
quote request_id when you report a problem — it is also in the
X-Seq-Request-Id header. 429 and 503 send
Retry-After too.
| status | code | what happened | what to do | retry | Python exception | seq exits |
|---|---|---|---|---|---|---|
| 400 | invalid_request | The request itself is wrong: a missing or mistyped field, an observation that does not match the deployment’s contract, bad params, or the wrong endpoint for the model (a chat model sent to /v1/decide). | Fix the request. | no | InvalidRequest | 2 |
| 413 | payload_too_large | The request body is over the size limit. | Send less: smaller or fewer images. | no | InvalidRequest | 2 |
| 401 | unauthenticated | No key, or a key that is invalid, revoked or expired. | Use a valid key. | no | AuthError | 3 |
| 403 | forbidden | The resource is yours to see, but this credential may not do this: an act key used for management or for a model its policy does not declare, or a key bound to another deployment. | Use a key that is allowed. | no | PermissionDenied | 3 |
| 402 | insufficient_credit | The account’s balance or credit line is used up. | Top up. | no | OutOfCredit | 6 |
| 402 | budget_exhausted | The deployment’s own monthly budget is reached. | Raise the budget or turn it off. | no | OutOfCredit | 6 |
| 404 | not_found | It does not exist, or it is not yours: another account’s deployment, volume, secret or eval reads as missing, and so do deleted deployments and retired routes. | Check the name. To start from a template: seq init <template>, then seq deploy. | no | NotFound | 4 |
| 409 | conflict | It conflicts with the current state: the name is taken, the deployment is building, failed or stopped, or the account is at its deployment limit. | Change the state first: another name, wait for the build, seq policy start, or delete a deployment. | no | Conflict | 5 |
| 410 | gone | The name existed and was retired on purpose; it will not come back. A catalogue model taken offline answers this, with the date, the reason and its replacement. | Use details.replacement; hint says how. | no | Gone | 4 |
| 429 | busy | Every replica of the worker is busy, or the key is calling faster than its rate limit. | Wait retry_after_s, then retry. | yes | Unavailable | 7 |
| 503 | warming | The worker is starting and will be ready shortly. | Wait retry_after_s, then retry. | yes | Unavailable | 7 |
| 503 | unavailable | Temporarily unavailable: no GPU capacity, an upstream briefly unreachable, or a worker draining. | Retry later; the SDK reconnects on its own. | yes | Unavailable | 7 |
| 502 | upstream_error | The provider behind a chat or decide model returned an error. | Retry once; if it persists, report it with the request_id. | yes | Unavailable | 7 |
| 504 | timeout | The request took longer than its time limit. | Retry. | yes | Unavailable | 7 |
| 500 | policy_error | Your deployed policy’s own code raised, or returned something its declared contract does not allow. | Read the traceback: seq policy logs <deployment>. | no | PolicyError | 1 |
| 500 | internal | A bug on our side. | Report it with the request_id. | no | InternalError | 1 |
| 424 | dependency_missing | A worker could not start: a secret its deployment references no longer exists. | Recreate the secret, or redeploy without it. | no | InternalError | 1 |
Errors in Python
In Python each code above is one exception class — the Python exception column — all
inheriting SequencesError: switch on .code, show .hint, quote
.request_id. Unavailable carries .warming and
.retry_after_s; Gone carries .replacement; a policy that raised comes back
as PolicyError with .remote_traceback, or as the built-in exception itself
(ValueError, …). ChunkExhausted is local: an action asked for with an empty
buffer and no observation. A failed seq command prints error[<code>]: …
and exits with the number in the last column, so a script can tell retry (7) from give up.
Set SEQUENCES_API_KEY in the environment, or pass api_key= to
connect(). Create a key in the
console; browse models in
the model library.
Keys
There are two kinds of key, and a robot never needs the first:
| key | what it may do |
|---|---|
| Full | everything the account can: deploy, manage, top up, run evals, call any model |
Robot (act) | connect to and act the deployment it is bound to (an unbound one: any of the account’s), and call on /v1/chat and /v1/decide the models that policy declares — nothing else. Any other model is refused 403 with the list it may call; managing anything is refused too. |
A policy declares the models it calls while it runs with models=; seq deploy refuses a
name the catalogue does not have. There is nothing else to set up: each time a robot connects, and while it keeps
its lease, the platform hands the policy’s worker a short-lived credential for exactly those models as
SEQUENCES_API_KEY, so the policy’s own chat() and decide() calls work
and no key is stored on the worker. The calls are billed to your account under the key the robot connected with
— when several robots share one worker, the one that connected or renewed last. Ten minutes after the last
robot leaves, the credential lapses; a call then is refused, and the robot that connects next gets a fresh one.
A SEQUENCES_API_KEY you give the policy yourself, as a secret, is used instead.
import sequence_ai, seq @seq.policy(gpu="L4", max_gpus=1, weights=seq.Weights.template("pi05-droid", version=1), models=["qwen3-8-27b"]) class PlannedPolicy(seq.Policy): @seq.plan(every_s=2.0) def plan(self, obs): r = sequence_ai.chat("qwen3-8-27b", [{"role": "user", "content": [ sequence_ai.image(obs.images[0].data), {"type": "text", "text": "Next step, in one sentence?"}]}], max_tokens=100, reasoning={"effort": "none"}) return r.content