Benchmarks and evals
A benchmark is a decorated class, like a policy. act() serves a deployed policy;
eval() scores it.
| policy side | benchmark side |
|---|---|
@seq.policy — author a policy | @seq.benchmark — author a benchmark |
@seq.load — weights into VRAM | @seq.setup — load the sim and assets, once |
@seq.on_episode — episode boundary | @seq.reset — build one episode’s env |
@seq.infer — observation → action | @seq.check — read raw env state → done + your metrics |
seq deploy → act() serves it | registered → eval() scores it |
Author a benchmark
Implement the hooks; there is no for-loop to write. @seq.setup runs once;
@seq.reset builds one gym-style env per condition; @seq.check runs
once per sim step, reads the real env handle, and returns done plus your own
metrics; an optional @seq.score computes your own aggregates.
import seq @seq.benchmark( name="libero-custom", image=seq.Image.debian_slim().uv_pip_install("mujoco==3.2.3", "robosuite==1.4.1"), # + LIBERO gpu="cpu", cpu=16, memory=32, # classic MuJoCo: CPU machines, one environment per core conditions=seq.conditions(tasks=TASKS, seeds=[0, 1], suite_of={...}), # tasks × seeds, your suites max_steps=220, # cost backstop; done is decided by check assets=seq.Volume.from_name("libero-assets"), # your uploaded files, read-only at /assets secrets=[seq.Secret.from_name("bench-token")], # env vars for the sim, never the policy render=seq.Record(camera="agentview", fps=20), # an mp4 of each task's first seed ) class LiberoCustom(seq.Benchmark): @seq.setup def setup(self): self.settings = json.load(open("/assets/tasks.json")) @seq.reset def reset(self, seed, variant): return make_env(variant.task, seed) # a real env handle @seq.check def check(self, env): return seq.Step(done=env.done, metrics={"success": float(env.done)}) @seq.score def score(self, results): # once per group: overall, each suite, each task wins = [r for r in results if r.metrics["success"]] return {"fast_success": sum(r.steps <= 120 for r in wins) / len(results)}
- Score.
@seq.scoreruns after every episode, once per group — all episodes, each suite’s, each task’s — over the episodes that did not error. A key it returns replaces that group’s mean; a new key is added. If it raises, the report keeps the means and lists the error. - Assets. Upload once with
seq volume put libero-assets ./assets --kind dataset. Each run reads the volume’s current version read-only, so no eval can change what the next one sees. - Secrets.
seq secret set bench-token --env BENCH_TOKEN, then declare it. Only the sim gets it; error text in the report shows***in place of its value. - Recordings. The first seed of each task is recorded from your env’s
render(camera), or the observation’s<camera>_image. They are stored with your account aseval:<id>inseq volume ls; save them withseq eval report <eval_id> --videos ./videos. - Conditions.
seq.conditions(tasks, seeds, variants=, max_steps=, timeout_s=, suite_of=)builds every combination. For a sparse set, pass a list ofseq.Condition(task=, seed=, variation=, max_steps=, timeout_s=, suite=)— each with its own step and time limit. An episode past itstimeout_sends astimed_out. - On a GPU, or on the CPU. A benchmark that names a card simulates on it — an L4
by default, or
gpu="RTX-PRO-6000"— and registration steps it there and refuses a simulator that does not use the GPU; billed like any GPU time. To run a CPU simulator as it is, declaregpu="cpu": each instance is a CPU machine withcpu=cores (16 by default), one environment per core, andmemory=for all of them (2 GB a core by default), billed by the core-hour and the GiB-hour. Use it for tests, first runs and reproducing a paper’s numbers: at the same number of environments it costs roughly 2 to 7 times more per environment step than a GPU version, because every environment holds a core while it waits for the policy.seq eval run --dry-runshows the cost of both. The official benchmark templates come in both versions: LIBERO’sbenchmark.py(GPU,libero) andbenchmark_cpu.py(CPU,libero-cpu). - No adapter. A benchmark names none —
default_adapter=is refused when you register it. The adapter is registered on its own and the eval names it with--adapter(see Adapters). - Your benchmarks.
seq benchmark lslists them with their version, status, card, environments and latest eval.seq benchmark logs <name> [--version N]prints a version’s build and check: the image build, then the check’s own output — what your setup printed, such as which renderer it got — and what the check measured. - Delete one.
seq benchmark rm <name>deletes it: its name and its place in your account’s quota are free at once, a build in flight stops, and the evals that ran it keep their reports. An eval running on it refuses the delete unless--forcecancels it;--dry-runshows what it would touch, including eval schedules that name it, which wait while it is deleted.
Move a CPU benchmark to the GPU
A benchmark written for classic MuJoCo — robosuite, dm_control, a gymnasium MuJoCo environment or your own
MjModel — ports to MuJoCo Warp with these commands, and one written for SAPIEN (its CPU
PhysX) ports to SAPIEN’s GPU PhysX with the same ones. They read the CPU version through its own
@seq.reset, so the GPU version starts every episode where the CPU version does, and they end with a
comparison of the two versions on the same episodes.
export runs where your benchmark runs, on its own simulator and MuJoCo version.
check-model, scaffold and compare --replay run MuJoCo Warp (or SAPIEN’s
GPU PhysX): pip install "sequence-ai[port]" installs the versions they are tested with
(sequence-ai 0.33.59 and later, Python 3.10 or later; on Linux for ARM it leaves out SAPIEN, and MuJoCo Warp
there needs glibc 2.34 or later). When your benchmark needs another MuJoCo, give these an environment of their
own.
seq benchmark export my_bench.py:MyBench --seeds 2 # each task's scene and starts, through your reset() seq benchmark check-model scene/ # what MuJoCo Warp would compute differently seq benchmark scaffold scene/ --envs 64 # a batched GPU benchmark, with TODOs for you seq benchmark compare my-bench my-bench-gpu --policy my-policy --adapter my-adapter --trials 10 # both versions, the same episodes seq benchmark compare --replay scene/ --conditions 10 # the same controls on MuJoCo and MuJoCo Warp, step by step seq benchmark convert my_bench.py:MyBench # export, check-model and scaffold in one go
- export runs your benchmark’s reset for each task and seed and writes the compiled scene, every start’s state, what reset changes from one seed to the next (where a fixture stands, for example) and how many contacts each start holds. Tasks that compile to the same model share one scene, and so one set of GPU worlds; an episode that compiles its own scene with other numbers (robosuite’s Lift draws its cube’s size each episode) stays that one scene, those numbers kept per seed and set world by world (sequence-ai 0.33.56 and later). For a SAPIEN benchmark it writes each start’s bodies and articulations (poses, joints, drive targets) and groups the starts by what the scene is built of — shapes, materials, masses, joints — since one GPU system’s worlds cannot change them.
- check-model compiles the scene with the GPU engine’s MuJoCo and lists every compiled field that differs, by body and geom name (masses, inertias, geometry); every option or flag whose default changed between the two MuJoCo versions; and any feature the engine does not support. For a SAPIEN scene it checks for SAPIEN’s GPU PhysX instead: it fails a dynamic body with a triangle mesh, says how many GPU systems each task needs, and names the calls with no GPU form your reset made (the arms’ passive force, for example).
- scaffold writes a
@seq.benchmarkwith batched hooks: one world per environment, the per-seed changes applied world by world, the changed options pinned to your version’s values and the contact buffers sized from the busiest start. A scene too large to ship with the code (8 MB packed) is mounted from a volume instead, and the skeleton’s README gives theseq volume putthat uploads it (sequence-ai 0.33.56 and later). Yours to write, markedTODO: the controller that turns an action into actuator inputs, the observation, and the success check. A SAPIEN skeleton builds one GPU world per environment with your scene code, writes each start back after the GPU system starts, and leaves the same three to you. - compare runs both registered versions with your policy, and the adapter both use
(
--adapter), on the same tasks and seeds (or reads two finished evals,--evals), pairs the episodes, and reports each version’s rate with its 95% interval, the paired difference with Agresti and Min’s 95% interval and McNemar’s exact test, and whether the difference is within--threshold(0.05 by default) — when every pair agrees, that takes 38 pairs or more (sequence-ai 0.33.38 and later). With--replay scene/it compares the motion instead, on your machine: the same seeded controls drive the exported scene on MuJoCo and the port’s scene (the same one unless you name a second) on MuJoCo Warp, from each start, and it names the first step, joint and body more than--atolapart (0.001 by default, over 200 steps) — one link of robosuite’s Lift made half again as heavy shows within 50 steps (sequence-ai 0.33.56 and later). A scene exported with sequence-ai 0.33.57 or later also carries a run of each start on your benchmark’s own MuJoCo, and the replay holds this machine’s MuJoCo to it: a difference MuJoCo Warp shares is named as the MuJoCo version’s, not the port’s — Lift’s cube lands a millimetre apart on MuJoCo 3.3.1 and 3.15. - Engines. MuJoCo to MuJoCo Warp, and SAPIEN to SAPIEN’s GPU PhysX (sequence-ai 0.33.36 and later). The official LIBERO template’s GPU version was ported the same way; its README says what still differs from LIBERO, and by how much. So was robosuite’s Lift, which agrees with its CPU version on all 50 seeds.
Run an eval
Register the benchmark, then evaluate a deployed policy on it. The benchmark runs on its own GPU instances
(L4 or RTX-PRO-6000), each stepping several environments at once — or, for a gpu="cpu"
benchmark, on CPU machines with one core per environment — and the policy on its own instances, which
answer every environment’s requests in batches. You write how many of each; nothing is derived for you:
seq deploy libero_custom.py:LiberoCustom seq deploy libero_to_pi05.py # the adapter: registers libero-to-pi05@v1, nothing builds seq eval run my-pi05 --benchmark libero-custom --adapter libero-to-pi05 --envs 5 seq eval status <eval_id> --watch # episodes done of the total, every metric so far seq eval logs <eval_id> --follow # the benchmark's and the policy's lines as they print seq eval report <eval_id> --videos ./videos seq benchmark evals libero-custom # every eval on it, across your policies
| flag | what it does |
|---|---|
--adapter NAME[@vN] | the adapter between the benchmark and the policy: a registered one (libero-to-pi05@v1, or the name alone for its latest), or a local file, ./adapter.py:Class, registered first; none when the two already agree |
--tasks a,b, --suite S, --trials N, --variants AXIS=V1,V2 | run a subset: these tasks, one suite, the first N seeds of each task, these variant values |
--policy-instances N, --policy-gpu CARD | policy instances serving the eval (default 1), on the deployment’s cards unless a card is named |
--benchmark-instances N, --benchmark-gpu L4|RTX-PRO-6000, --envs N | benchmark instances per arm, each on its own GPU (default 1, on the benchmark’s card), and the environments each runs at once (default: the benchmark’s envs=, else 1). A gpu="cpu" benchmark’s instances are CPU machines: --envs N gives each N cores (default its cpu=, at most 64), and --benchmark-gpu does not apply |
--max-usd X | stop once the eval’s machines have cost this much: no new episode starts a minute before, and the eval ends with it, its running episodes too; the report covers what ran |
--rerun-failed EVAL_ID | run only that eval’s errored episodes again, with the adapter version it ran; one report covers both runs |
--rerun-unsuccessful EVAL_ID, --metric NAME | run the episodes that eval ran without error and did not succeed on again, as an eval of their own — with --record all to watch them — never merged into its report. --metric names the metric that says success (default success; 0, false or not reported is not a success) |
--compare MODEL | also run MODEL on the same episodes and seeds, as another arm of this eval, and print the paired difference (repeatable) |
--gate "success>=0.8" | exit with code 8 unless the overall metric passes — for CI (repeatable) |
--param NAME=V, --sweep NAME=V1,V2 | set an inference parameter for every act; or an arm per value, side by side in this eval |
--record per-task|all|none | which episodes are recorded (default: the first seed of each task) |
--dry-run, --no-wait | show the layout, the cards it needs against your account and the estimate, starting nothing; or return once it starts and read it later with seq eval report (stop it with seq eval cancel <eval_id>) |
The report gives every metric key overall, per suite and per task with n and a 95% confidence
interval (Wilson for a 0/1 metric, bootstrap otherwise; fewer than 10 episodes is flagged), a table per arm with
the paired differences, the cost of the time every machine actually ran, a reproducibility manifest (model and benchmark versions, code sha256, adapter,
seeds, parameters, platform revision), the recordings, and any episode errors.
seq eval report <eval_id> --csv runs.csv writes one row per episode;
--trajectories DIR saves every episode’s states and actions as gzip JSONL.
Each act carries a seed of its own, hashed from the episode’s seed and the act’s index: a rerun of the
episode carries the same ones and replays exactly, and a policy that draws its noise from seed
(flow matching, diffusion) starts every replan from fresh noise rather than from the noise its previous one used,
which keeps a stuck state stuck. A trajectory records the seed each act carried.
While an eval runs, seq eval status <eval_id> — and its page in the console — shows
how many episodes are done of the total, per arm, and every metric so far with its interval;
--watch follows it to the end. An eval that is canceled or fails keeps what ran: its report says
the results are partial and how many episodes they cover. seq eval logs <eval_id> prints what
its machines printed — each benchmark instance’s output and each policy instance’s lines;
--side benchmark|policy, --arm and --instance narrow them and
--follow streams them. Lines are kept 7 days. seq benchmark evals <name> (the same as
seq eval ls --benchmark <name>) lists every eval run on a benchmark across your policies, newest
first, 50 a page (--limit, --cursor; --policy, --since 7d);
seq.benchmarks() and seq.benchmark_evals(name) return the same pages in Python.
Adapters
A @seq.adapter bridges a benchmark’s observation and action to a policy’s own inputs and outputs.
It is the third kind you register, and the only one without a machine: it belongs to neither the benchmark nor
the policy but to the pairing an eval runs. seq deploy libero_to_pi05.py registers it as a version
— its file is sent, nothing is built or run — under @seq.adapter(name=...), else its class
name in lower case with hyphens. The version is the content: the same file again is the version it already is, a
changed one the next. The eval names it, --adapter libero-to-pi05@v1; when the eval starts, the
benchmark’s container fetches that version, checks every label again and runs it between the simulation and
the policy. Changing it changes neither the benchmark’s version nor the policy’s, and the report’s first line
reads benchmark vN + adapter vN + policy vN. seq adapter ls lists yours,
seq adapter show <name>[@vN] says whether a version is lossy, and
seq adapter pull <name>[@vN] writes its file back exactly as it was registered
(seq.adapter_source(name, version) in Python): edit it and seq deploy it again for the
next version. check and rm check and remove them.
Label every transform lossless or lossy(reason) — you cannot self-certify a
lossy transform as clean, and lossy transforms are recorded in the report.
| lossless (keep as default) | lossy (declare + recorded in the report) |
|---|---|
| key rename, dimension reorder, dtype-preserving reshape, identity, unit-preserving coordinate relabel | image resize / crop / downsample, dtype quantize, dropping a camera or channel, action-space projection (abs↔delta, mismatched dims), interpolating coordinate transforms |
When shapes or spaces do not match and no adapter is given, eval refuses rather than guessing.
What an adapter can be: one file of at most 256 KB, run inside the benchmark’s container — it imports
what the benchmark’s image installs (numpy, say) and the platform’s seq.tf transforms, and
nothing else: no file beside it, no package of its own. It feeds one kind of policy, the robot-arm preset or a
policy’s own inputs and outputs, and an eval pairing it with the other kind is refused. A version an eval is
running cannot be removed. Numbers are never given twice: an adapter removed and registered again under its name
goes on from its next number, and --rerun-failed runs the version its eval ran, refused (409) if that
number no longer holds the same code.
Parameters in an eval
Every side of an eval declares what an eval may set the way infer does — keyword arguments
with defaults, typed int, float, bool or str — and
registering it sends them: the policy’s @seq.infer, the adapter’s
__init__, the benchmark’s @seq.setup (after the environment count the platform
passes). The eval loop has one of its own, eval.steps_per_query: how many actions of each chunk run
before the policy is asked again. Set any of them for the whole eval with
--param side.name=value — the bare name when only one side declares it — and sweep any
of them with --sweep. A name no side declares, a wrong type or a steps_per_query below
1 is refused before anything is reserved, with every parameter the eval can set.
@seq.adapter(name="to-my-policy") class ToMyPolicy(seq.Adapter): obs_transforms = [seq.tf.resize("image", (256, 256), (224, 224), lossy="the policy takes 224x224")] def __init__(self, *, steps_per_query: int = 5, image_size: int = 224): # built from the parameter, so the eval’s value reaches the transform (and the report) self.obs_transforms = [seq.tf.resize("image", (256, 256), (image_size, image_size), lossy=f"the policy takes {image_size}x{image_size}")] # $ seq eval run my-policy --benchmark my-bench --adapter to-my-policy --param adapter.image_size=192 # $ seq eval run my-policy --benchmark my-bench --sweep eval.steps_per_query=5,10
Who decides. The eval’s value, else the default of the side that declares it.
steps_per_query is one setting shared with an adapter that declares an int of that name
(it is handed the value used): the eval’s, else the adapter’s default, else each chunk’s
open_loop_horizon, else the whole chunk. One longer than a chunk stops the eval at the first query
and names the most that fits. --dry-run and the report print every arm’s values with where
each came from, and how many of its steps asked the policy: a batched benchmark renders only at those steps, so
steps_per_query=10 renders half as often as 5.
Composite policies
A composite runs a planner model beside pi0.5: @seq.plan asks the planner for the next step about
every every_s seconds, for each robot on its own and off the per-step path; @seq.infer
runs pi0.5 every step with that robot’s latest step, which the platform passes in as
instruction — the robot’s own task until the first plan arrives. The planner is a model of
the catalogue, called through /v1/chat with sequence_ai.chat: declare it in
models=, and that is all — each time a robot connects, the platform hands the policy a
short-lived credential for exactly those models (see Keys). Each planner call is billed to your
account like any chat call. A model that reasons by default, as qwen3-8-27b does, is told not to with
reasoning={"effort": "none"}: a planner wants its one short line now, and thinking would spend the
reply’s few tokens and leave it empty (see Reasoning).
@seq.policy(name="my-robot", gpu="L4", max_gpus=1, weights=seq.Weights.template("pi05-droid", version=1), models=["qwen3-8-27b"]) # the planner, from the catalogue class Composite(seq.Policy): @seq.load def load(self, weights_dir): self.low = load_pi05(weights_dir) # runs on your deployment's GPU @seq.plan(every_s=2.0) # off the per-step path, per robot def plan(self, obs): reply = sequence_ai.chat("qwen3-8-27b", [ {"role": "system", "content": "Reply with the one next step the arm should do now."}, {"role": "user", "content": [{"type": "text", "text": f"Task: {obs.instruction}"}, *[sequence_ai.image(f.data) for f in obs.images]]}], max_tokens=32, temperature=0, reasoning={"effort": "none"}) return reply.content # the next step; None keeps the current one @seq.infer def infer(self, obs, *, seed=None, instruction=None): return self.low(obs, instruction or obs.instruction) # every step: this robot's latest step
plan returns the step instead of storing it on self. A
@seq.batched infer receives instruction as a list, one entry per request. An error in
plan keeps the current step, and the planner asks again on its next turn.
Scheduled evals
Evaluate a deployment on a schedule — daily, weekly, or after every deploy — and read its evals as a
trend: version, primary metric and its 95% interval. Each run is an ordinary eval, billed like one and stopped at
--max-usd; a deployment stopped at its budget runs none.
seq eval schedule my-pi05 --benchmark libero-custom --every weekly --trials 2 --max-usd 20 seq eval schedule my-pi05 --benchmark libero-custom --every deploy# each time the serving version changes seq eval ls my-pi05 # its evals over time seq eval schedule my-pi05 --benchmark libero-custom --off