Docs

Deploy

seq deploy

Deploy the decorated class; it is ready to call when the command returns.

deployshell
seq deploy pi05_droid:Pi05 --name my-pi05
# ready when the command returns; then addressable as model="my-pi05"
# --dry-run prints the request instead of sending it; seq validate pi05_droid:Pi05 checks it offline

seq policy estimate pi05_droid:Pi05         # only price it: the card per hour, the warm floor, use per day and month
seq policy logs my-pi05 --build             # the newest version's build: each phase's time and its full log

While it builds, seq deploy shows each build phase and the newest log line. A deploy of a name that already exists builds a new version beside the serving one — see Versions and traffic.

Names. A deployment’s name is what robots connect to, and it is unique on the platform. Without name= or --name it is the class name, lower-case with hyphens: ToyReacher deploys as toy-reacher; seq init <template> without a directory names it after the template and your account (toy-policy-<8 characters>). Give every policy robots connect to a name= of its own. Deploying the same class to its name again is a new version; deploying another class to a name you already use is refused unless you add --replace — its robots then switch to the new policy once it builds — and a name another account uses is refused outright. seq deploy --dry-run says which it would be before anything is sent.

What a deploy measures

Before a version serves, its build check starts one worker, sends it one act, then acts on it back to back as one robot would. From that run seq deploy, seq policy status and the console’s deployment page report the same three things:

pi0.5 on an L4, nothing declaredoutput
price: one worker, per hour
  GPU     L4  $0.9600
  CPU     0.125 cores requested, 0.93 used serving one robot  $0.1584
  memory  0.125 GiB requested, 9.92 GiB used  $0.2867
  total   $0.9849/h at the requests; $1.4051/h as billed — CPU and memory are billed on the greater of the request and actual use
  declare @seq.policy(cpu=1, memory=11) and the price at the requests is what is billed
one copy needs 8.0 GiB of GPU memory (its framework's own peak); one L4 (22.5 GiB) fits 2 copies with room to spare: @seq.concurrent(replicas_per_gpu=2)
GPU cap: 2 (your account's: 8); with all of them busy (2 workers) about $2.81/h — change it any time: seq policy scale my-pi05 --max-gpus N

seq policy estimate <file> --gpu H100 shows, for a policy deployed before, how many copies fit on a card and one worker’s price there; without --gpu, on every card. The card the build check ran on is measured; the others are estimated from their size.

Control plane, data plane

A call is two phases: connect() goes through the gateway once and returns the worker’s address, a pinned certificate, a short-lived secret and a lease; every act() then connects directly to the worker.

phasereachespurpose
connect(model)the gateway, onceadmit, ensure a worker is up, return its address, cert, secret and lease
act(...)the worker, directlyrun inference on a reused connection; the gateway is not in this path
reconnectthe gateway, on demandon a rotated secret, expired lease or scaled-to-zero worker — fetch fresh coordinates and replay, transparently

You do not handle the 401/503 yourself: a rotated secret, expired lease, or worker that scaled to zero surfaces as a transparent reconnect, and the in-flight call is replayed.

Scaling

A worker is stopped after idle_timeout seconds without a request (600 by default), and is billed until it stops. A scaled-to-zero worker cold-starts in a couple of minutes; connect() warms it before you drive it (see warming).

A deployment grows as robots connect, up to max_gpus, which a GPU policy must declare: each robot gets a copy of the model of its own by default, so max_gpus=2 serves 2 robots at once, and seq deploy prints what all of them cost per hour busy. Each robot joins the least-busy worker with room — target_inputs robots per worker, from @seq.concurrent, one by default — and keeps it across its lease renewals; when every worker is full another starts. A robot is never crowded onto a full worker: when no worker has room and none may start, it waits in line before its control loop begins (on_full="queue", up to queue_timeout_s) and takes the next slot that frees — a robot that closes its connection frees its slot at once. min_workers=N keeps N workers running even with no robot connected, and buffer_workers=N keeps N idle beside the busy ones, so a robot never waits for a cold start; they are billed while they run. seq policy scale <name> --min 1 --max-gpus 4 changes these live, with no redeploy; seq policy status <name> shows each worker’s robots and the line.

How long a worker serves. A worker is replaced after 5 hours (seq policy scale <name> --rotate-after SECONDS makes it sooner): a replacement starts beside it, and each of its robots moves to the replacement at the end of its current episode — sequence-ai 0.33.17 and later do this by themselves, so no episode is cut and no cold start is waited out (an older client moves at its next lease renewal). Until the replacement serves, a robot that connects starts on the worker being replaced and moves with the others, so it waits for no cold start either. The same moves happen when a new version starts serving (a robot that connects then starts on the new version). While a worker is replaced, it and its replacement both run, for a few minutes and at most 45, billed as usual; workers replaced at the same time each get a replacement beside them, so a deployment can briefly run up to twice its max_gpus (seq policy status <name> says what that costs an hour). Your account’s GPU limit still counts them all, so for robots that run for hours keep a GPU of your account free for each worker replaced at the same time — with none free a replacement cannot start, and its robots move through a cold start just before the worker’s forced end. A worker that stops answering is replaced at once. An act still running at your deployment’s request timeout (5 minutes unless you declare one) is answered with a timeout; if it is still running one timeout later, the copy running it is ended and a fresh copy starts on the same worker (seq policy logs says reason=hung), and a worker even that does not free is replaced (12 minutes after the act came in, at the default). A worker still running 5 h 50 min after it started is ended whatever it is doing, the last line of defence. Rarely, the provider ends a GPU machine under its robots; they reconnect to another worker, through a cold start unless one is kept warm — for robots that must not wait, seq policy scale <name> --buffer 1 keeps one spare while robots are connected. seq policy status <name> shows how long each worker has served and when it is replaced.

region= keeps the workers near your robots. A broad region — us, eu, ap — bills the card at 1.15× its rate; a narrow one — us-east, us-central, us-west, eu-west, eu-north, ap-northeast, ap-southeast, ap-south, jp, uk, ca — at 1.75×. Any other name is refused at deploy; seq policy estimate <file> shows the price.

More than one machine

Write the total: gpu="H100:16". Past 8 cards the platform runs the model on whole machines of 8 — 16 is two, at most 256 cards — started together, stopped together and billed together (every machine’s cards, CPU and memory). The account needs a GPU limit that holds the whole group at once.

Your @seq.load and @seq.infer run on every machine, one process each, seeing that machine’s 8 cards. Before your code is imported the platform sets RANK, WORLD_SIZE, MASTER_ADDR, MASTER_PORT and LOCAL_RANK=0 and, when your image has torch, initializes torch.distributed (NCCL) across the machines — check torch.distributed.is_initialized() rather than initializing it again. A robot reaches machine 0; every request, and every episode end, runs on every machine at once so your collectives meet, and machine 0’s answer goes back to the robot.

The group is one copy of the model and holds one robot’s control loop: leave gpus_per_replica and replicas_per_gpu unset, list one card type, and keep a policy with @seq.plan on one machine. When one machine stops, the whole group stops (group_member_lost) and the next connect starts a new one; seq policy status <name> shows each machine’s rank, state and last heartbeat.

Versions and traffic

Every deploy of a name is a new version. It builds beside the serving one, which keeps serving until the new version is ready — and keeps serving if the new one fails. Earlier versions are kept, so seq policy rollback serves one again at once, with nothing rebuilt.

roll out a new version graduallyshell
seq policy versions my-pi05                # building / serving / kept for rollback
seq policy traffic my-pi05 v5=10 v4=90     # one new session in ten goes to v5
seq policy traffic my-pi05 --pin arm-7 v5  # a robot key's new sessions all go to v5 (its id, name or prefix…last4)
seq policy traffic my-pi05 --unpin arm-7
seq policy traffic my-pi05 --clear         # every new session to the serving version again
seq policy rollback my-pi05 --to 4

A split applies to new sessions only: a robot keeps its version across its lease renewals, so an episode never changes model halfway. A pin reaches every gateway within five minutes.

Budgets, alerts and warm hours

run a deployment on a budget and a scheduleshell
seq policy budget my-pi05 --monthly 300      # USD this month; --off removes it
seq policy scale  my-pi05 --min 1 --during "Mon-Fri 08:00-18:00" --tz America/Los_Angeles
seq policy alerts my-pi05 --webhook https://ops.example.com/seq \
                  --error-rate 0.05 --p95-ms 250 --budget-pct 80

Budget, schedule, alerts and capture belong to the deployment, not a version: a new deploy keeps them.