Deploy
seq deploy
Deploy the decorated class; it is ready to call when the command returns.
seq deploy pi05_droid:Pi05 --name my-pi05 # ready when the command returns; then addressable as model="my-pi05" # --dry-run prints the request instead of sending it; seq validate pi05_droid:Pi05 checks it offline seq policy estimate pi05_droid:Pi05 # only price it: the card per hour, the warm floor, use per day and month seq policy logs my-pi05 --build # the newest version's build: each phase's time and its full log
While it builds, seq deploy shows each build phase and the newest log line. A deploy of a name
that already exists builds a new version beside the serving one — see
Versions and traffic.
Names. A deployment’s name is what robots connect to, and it is unique on the platform. Without
name= or --name it is the class name, lower-case with hyphens: ToyReacher
deploys as toy-reacher; seq init <template> without a directory names it after the
template and your account (toy-policy-<8 characters>). Give every policy robots connect to a
name= of its own. Deploying the same class to its name again is a new version; deploying
another class to a name you already use is refused unless you add --replace — its
robots then switch to the new policy once it builds — and a name another account uses is refused outright.
seq deploy --dry-run says which it would be before anything is sent.
What a deploy measures
Before a version serves, its build check starts one worker, sends it one act, then acts on it back to back as
one robot would. From that run seq deploy, seq policy status and the console’s
deployment page report the same three things:
- One copy’s GPU memory: the most the model’s framework held at once (JAX’s own peak, or PyTorch’s), plus what the process holds outside it. A JAX model alone on a card takes 75% of the card as it starts, so the card’s own figure shows that pool, not what the model needs.
- How many copies fit on the card, each given 10% more than it measured, and the declaration that puts
them there:
@seq.concurrent(replicas_per_gpu=N). Copies that share a card each take their share of it (JAX:XLA_PYTHON_CLIENT_MEM_FRACTION= 0.95/N − 0.02, unless your image sets it). Declaring more copies than fit fails the deploy with what one copy measured. With a list of cards (gpu=["L4", "A100-40GB"]) a worker starts on the first that has one free, so the copies are counted on the smallest of them, and the price is the first card’s with the later ones beside it. - One worker’s price per hour, two ways: at its requests —
cpu=andmemory=, or the minimum 0.125 core and 0.125 GiB when you declare none — and as billed: CPU and memory are billed on the greater of the request and actual use. When the use runs over the request, it names thecpu=andmemory=that make the two agree.
price: one worker, per hour GPU L4 $0.9600 CPU 0.125 cores requested, 0.93 used serving one robot $0.1584 memory 0.125 GiB requested, 9.92 GiB used $0.2867 total $0.9849/h at the requests; $1.4051/h as billed — CPU and memory are billed on the greater of the request and actual use declare @seq.policy(cpu=1, memory=11) and the price at the requests is what is billed one copy needs 8.0 GiB of GPU memory (its framework's own peak); one L4 (22.5 GiB) fits 2 copies with room to spare: @seq.concurrent(replicas_per_gpu=2) GPU cap: 2 (your account's: 8); with all of them busy (2 workers) about $2.81/h — change it any time: seq policy scale my-pi05 --max-gpus N
seq policy estimate <file> --gpu H100 shows, for a policy deployed before, how many copies fit
on a card and one worker’s price there; without --gpu, on every card. The card the build check ran
on is measured; the others are estimated from their size.
Control plane, data plane
A call is two phases: connect() goes through the gateway once and returns the
worker’s address, a pinned certificate, a short-lived secret and a lease; every
act() then connects directly to the worker.
| phase | reaches | purpose |
|---|---|---|
connect(model) | the gateway, once | admit, ensure a worker is up, return its address, cert, secret and lease |
act(...) | the worker, directly | run inference on a reused connection; the gateway is not in this path |
| reconnect | the gateway, on demand | on a rotated secret, expired lease or scaled-to-zero worker — fetch fresh coordinates and replay, transparently |
You do not handle the 401/503 yourself: a rotated secret, expired
lease, or worker that scaled to zero surfaces as a transparent reconnect, and the in-flight
call is replayed.
Scaling
A worker is stopped after idle_timeout seconds without a request (600 by default), and
is billed until it stops. A scaled-to-zero worker cold-starts in a couple of minutes;
connect() warms it before you drive it (see warming).
A deployment grows as robots connect, up to max_gpus, which a GPU policy must declare: each robot
gets a copy of the model of its own by default, so max_gpus=2 serves 2 robots at once, and
seq deploy prints what all of them cost per hour busy. Each
robot joins the least-busy worker with room — target_inputs robots per worker, from
@seq.concurrent, one by default — and keeps it across its lease renewals; when every worker
is full another starts. A robot is never crowded onto a full worker: when no worker has room and none may start,
it waits in line before its control loop begins (on_full="queue", up to
queue_timeout_s) and takes the next slot that frees — a robot that closes its connection
frees its slot at once. min_workers=N keeps N workers running even with no robot connected,
and buffer_workers=N keeps N idle beside the busy ones, so a robot never waits for a cold
start; they are billed while they run. seq policy scale <name> --min 1 --max-gpus 4
changes these live, with no redeploy; seq policy status <name> shows each worker’s robots and
the line.
How long a worker serves. A worker is replaced after 5 hours (seq policy scale <name>
--rotate-after SECONDS makes it sooner): a replacement starts beside it, and each of its robots moves to
the replacement at the end of its current episode — sequence-ai 0.33.17 and later do this by
themselves, so no episode is cut and no cold start is waited out (an older client moves at its next lease
renewal). Until the replacement serves, a robot that connects starts on the worker being replaced and moves with
the others, so it waits for no cold start either. The same moves happen when a new version starts serving (a
robot that connects then starts on the new version). While a worker is replaced, it and
its replacement both run, for a few minutes and at most 45, billed as usual; workers replaced at the same time
each get a replacement beside them, so a deployment can briefly run up to twice its max_gpus
(seq policy status <name> says what that costs an hour). Your account’s GPU limit still
counts them all, so for robots that run for hours keep a GPU of your account free for each worker replaced at
the same time — with none free a replacement cannot start, and its robots move through a cold start just
before the worker’s forced end. A worker that stops answering is replaced at once. An act still
running at your deployment’s request timeout (5 minutes unless you declare one) is answered with a timeout; if it
is still running one timeout later, the copy running it is ended and a fresh copy starts on the same worker
(seq policy logs says reason=hung), and a worker even that does not free is replaced
(12 minutes after the act came in, at the default). A worker still running 5 h 50 min after it started is ended
whatever it is doing, the last line of defence. Rarely, the provider ends a GPU machine under its robots; they
reconnect to another worker, through a cold start unless one is kept warm — for robots that must not wait,
seq policy scale <name> --buffer 1 keeps one spare while robots are connected. seq policy status <name> shows how long each worker has served
and when it is replaced.
region= keeps the workers near your robots. A broad region — us,
eu, ap — bills the card at 1.15× its rate; a narrow one —
us-east, us-central, us-west, eu-west,
eu-north, ap-northeast, ap-southeast, ap-south,
jp, uk, ca — at 1.75×. Any other name is refused at deploy;
seq policy estimate <file> shows the price.
More than one machine
Write the total: gpu="H100:16". Past 8 cards the platform runs the model on whole machines of 8
— 16 is two, at most 256 cards — started together, stopped together and billed together (every
machine’s cards, CPU and memory). The account needs a GPU limit that holds the whole group at once.
Your @seq.load and @seq.infer run on every machine, one process each, seeing that
machine’s 8 cards. Before your code is imported the platform sets RANK, WORLD_SIZE,
MASTER_ADDR, MASTER_PORT and LOCAL_RANK=0 and, when your image has torch,
initializes torch.distributed (NCCL) across the machines — check
torch.distributed.is_initialized() rather than initializing it again. A robot reaches machine 0;
every request, and every episode end, runs on every machine at once so your collectives meet, and machine 0’s
answer goes back to the robot.
The group is one copy of the model and holds one robot’s control loop: leave gpus_per_replica and
replicas_per_gpu unset, list one card type, and keep a policy with @seq.plan on one
machine. When one machine stops, the whole group stops (group_member_lost) and the next connect
starts a new one; seq policy status <name> shows each machine’s rank, state and last heartbeat.
Versions and traffic
Every deploy of a name is a new version. It builds beside the serving one, which keeps serving until the new
version is ready — and keeps serving if the new one fails. Earlier versions are kept, so
seq policy rollback serves one again at once, with nothing rebuilt.
seq policy versions my-pi05 # building / serving / kept for rollback seq policy traffic my-pi05 v5=10 v4=90 # one new session in ten goes to v5 seq policy traffic my-pi05 --pin arm-7 v5 # a robot key's new sessions all go to v5 (its id, name or prefix…last4) seq policy traffic my-pi05 --unpin arm-7 seq policy traffic my-pi05 --clear # every new session to the serving version again seq policy rollback my-pi05 --to 4
A split applies to new sessions only: a robot keeps its version across its lease renewals, so an episode never changes model halfway. A pin reaches every gateway within five minutes.
Budgets, alerts and warm hours
seq policy budget my-pi05 --monthly 300 # USD this month; --off removes it seq policy scale my-pi05 --min 1 --during "Mon-Fri 08:00-18:00" --tz America/Los_Angeles seq policy alerts my-pi05 --webhook https://ops.example.com/seq \ --error-rate 0.05 --p95-ms 250 --budget-pct 80
- Budget. Within minutes of this month’s spend reaching it, the deployment is
stopped: its workers end and a connect is refused with
402 budget_exhausted.seq policy startstays refused until you raise the budget or remove it. - Warm hours.
--duringtakes days and a time span (Mon-Fri,Sat,Sun,daily 22:00-06:00) in--tz(UTC by default). In those hours at least--minworkers run with no robot connected, billed while they run; outside them the deployment scales to zero as usual.--clearremoves the schedule. - Alerts. The platform posts JSON (
deployment,kind,detail,at) to your https webhook when an hour’s error rate (over at least 20 acts) or p95 latency crosses its threshold (error_rate,latency), when spend reaches--budget-pctof the budget or the budget itself (budget_threshold,budget_reached), when a new version fails to build (deploy_failed; off with--no-deploy-failed), and when the deployment reaches half of itsmax_gpus(gpu_ceiling_near) or a robot finds it full and waits or is refused (gpu_ceiling_reached). Each kind is sent at most once an hour, the budget kinds once a month. Setting the webhook prints a secret once: each body is signed with it inX-Seq-Signature: sha256=<HMAC-SHA256 of the body>. The webhook must be https on a public address, and redirects are not followed. - Rate limit. Each API key may make 600 calls a minute to the gateway; past that a call
answers
429withRetry-After. Acts sent directly to a worker are not counted.
Budget, schedule, alerts and capture belong to the deployment, not a version: a new deploy keeps them.