Docs

Benchmarks and evals

A benchmark is a decorated class, like a policy. act() serves a deployed policy; eval() scores it.

policy sidebenchmark side
@seq.policy — author a policy@seq.benchmark — author a benchmark
@seq.load — weights into VRAM@seq.setup — load the sim and assets, once
@seq.on_episode — episode boundary@seq.reset — build one episode’s env
@seq.infer — observation → action@seq.check — read raw env state → done + your metrics
seq deploy → act() serves itregistered → eval() scores it

Author a benchmark

Implement the hooks; there is no for-loop to write. @seq.setup runs once; @seq.reset builds one gym-style env per condition; @seq.check runs once per sim step, reads the real env handle, and returns done plus your own metrics; an optional @seq.score computes your own aggregates.

libero_custom.py — LIBERO is one example benchmarkpython
import seq

@seq.benchmark(
    name="libero-custom",
    image=seq.Image.debian_slim().uv_pip_install("mujoco==3.2.3", "robosuite==1.4.1"),   # + LIBERO
    gpu="cpu", cpu=16, memory=32,               # classic MuJoCo: CPU machines, one environment per core
    conditions=seq.conditions(tasks=TASKS, seeds=[0, 1], suite_of={...}),   # tasks × seeds, your suites
    max_steps=220,                                # cost backstop; done is decided by check
    assets=seq.Volume.from_name("libero-assets"),    # your uploaded files, read-only at /assets
    secrets=[seq.Secret.from_name("bench-token")], # env vars for the sim, never the policy
    render=seq.Record(camera="agentview", fps=20),   # an mp4 of each task's first seed
)
class LiberoCustom(seq.Benchmark):
    @seq.setup
    def setup(self):
        self.settings = json.load(open("/assets/tasks.json"))
    @seq.reset
    def reset(self, seed, variant):
        return make_env(variant.task, seed)                 # a real env handle
    @seq.check
    def check(self, env):
        return seq.Step(done=env.done, metrics={"success": float(env.done)})
    @seq.score
    def score(self, results):                        # once per group: overall, each suite, each task
        wins = [r for r in results if r.metrics["success"]]
        return {"fast_success": sum(r.steps <= 120 for r in wins) / len(results)}

Move a CPU benchmark to the GPU

A benchmark written for classic MuJoCo — robosuite, dm_control, a gymnasium MuJoCo environment or your own MjModel — ports to MuJoCo Warp with these commands, and one written for SAPIEN (its CPU PhysX) ports to SAPIEN’s GPU PhysX with the same ones. They read the CPU version through its own @seq.reset, so the GPU version starts every episode where the CPU version does, and they end with a comparison of the two versions on the same episodes.

export runs where your benchmark runs, on its own simulator and MuJoCo version. check-model, scaffold and compare --replay run MuJoCo Warp (or SAPIEN’s GPU PhysX): pip install "sequence-ai[port]" installs the versions they are tested with (sequence-ai 0.33.59 and later, Python 3.10 or later; on Linux for ARM it leaves out SAPIEN, and MuJoCo Warp there needs glibc 2.34 or later). When your benchmark needs another MuJoCo, give these an environment of their own.

terminalshell
seq benchmark export my_bench.py:MyBench --seeds 2     # each task's scene and starts, through your reset()
seq benchmark check-model scene/                       # what MuJoCo Warp would compute differently
seq benchmark scaffold scene/ --envs 64                # a batched GPU benchmark, with TODOs for you
seq benchmark compare my-bench my-bench-gpu --policy my-policy --adapter my-adapter --trials 10   # both versions, the same episodes
seq benchmark compare --replay scene/ --conditions 10  # the same controls on MuJoCo and MuJoCo Warp, step by step
seq benchmark convert my_bench.py:MyBench              # export, check-model and scaffold in one go

Run an eval

Register the benchmark, then evaluate a deployed policy on it. The benchmark runs on its own GPU instances (L4 or RTX-PRO-6000), each stepping several environments at once — or, for a gpu="cpu" benchmark, on CPU machines with one core per environment — and the policy on its own instances, which answer every environment’s requests in batches. You write how many of each; nothing is derived for you:

register, run, read the reportshell
seq deploy libero_custom.py:LiberoCustom
seq deploy libero_to_pi05.py                # the adapter: registers libero-to-pi05@v1, nothing builds
seq eval run my-pi05 --benchmark libero-custom --adapter libero-to-pi05 --envs 5
seq eval status <eval_id> --watch          # episodes done of the total, every metric so far
seq eval logs <eval_id> --follow            # the benchmark's and the policy's lines as they print
seq eval report <eval_id> --videos ./videos
seq benchmark evals libero-custom           # every eval on it, across your policies
flagwhat it does
--adapter NAME[@vN]the adapter between the benchmark and the policy: a registered one (libero-to-pi05@v1, or the name alone for its latest), or a local file, ./adapter.py:Class, registered first; none when the two already agree
--tasks a,b, --suite S, --trials N, --variants AXIS=V1,V2run a subset: these tasks, one suite, the first N seeds of each task, these variant values
--policy-instances N, --policy-gpu CARDpolicy instances serving the eval (default 1), on the deployment’s cards unless a card is named
--benchmark-instances N, --benchmark-gpu L4|RTX-PRO-6000, --envs Nbenchmark instances per arm, each on its own GPU (default 1, on the benchmark’s card), and the environments each runs at once (default: the benchmark’s envs=, else 1). A gpu="cpu" benchmark’s instances are CPU machines: --envs N gives each N cores (default its cpu=, at most 64), and --benchmark-gpu does not apply
--max-usd Xstop once the eval’s machines have cost this much: no new episode starts a minute before, and the eval ends with it, its running episodes too; the report covers what ran
--rerun-failed EVAL_IDrun only that eval’s errored episodes again, with the adapter version it ran; one report covers both runs
--rerun-unsuccessful EVAL_ID, --metric NAMErun the episodes that eval ran without error and did not succeed on again, as an eval of their own — with --record all to watch them — never merged into its report. --metric names the metric that says success (default success; 0, false or not reported is not a success)
--compare MODELalso run MODEL on the same episodes and seeds, as another arm of this eval, and print the paired difference (repeatable)
--gate "success>=0.8"exit with code 8 unless the overall metric passes — for CI (repeatable)
--param NAME=V, --sweep NAME=V1,V2set an inference parameter for every act; or an arm per value, side by side in this eval
--record per-task|all|nonewhich episodes are recorded (default: the first seed of each task)
--dry-run, --no-waitshow the layout, the cards it needs against your account and the estimate, starting nothing; or return once it starts and read it later with seq eval report (stop it with seq eval cancel <eval_id>)

The report gives every metric key overall, per suite and per task with n and a 95% confidence interval (Wilson for a 0/1 metric, bootstrap otherwise; fewer than 10 episodes is flagged), a table per arm with the paired differences, the cost of the time every machine actually ran, a reproducibility manifest (model and benchmark versions, code sha256, adapter, seeds, parameters, platform revision), the recordings, and any episode errors. seq eval report <eval_id> --csv runs.csv writes one row per episode; --trajectories DIR saves every episode’s states and actions as gzip JSONL.

Each act carries a seed of its own, hashed from the episode’s seed and the act’s index: a rerun of the episode carries the same ones and replays exactly, and a policy that draws its noise from seed (flow matching, diffusion) starts every replan from fresh noise rather than from the noise its previous one used, which keeps a stuck state stuck. A trajectory records the seed each act carried.

While an eval runs, seq eval status <eval_id> — and its page in the console — shows how many episodes are done of the total, per arm, and every metric so far with its interval; --watch follows it to the end. An eval that is canceled or fails keeps what ran: its report says the results are partial and how many episodes they cover. seq eval logs <eval_id> prints what its machines printed — each benchmark instance’s output and each policy instance’s lines; --side benchmark|policy, --arm and --instance narrow them and --follow streams them. Lines are kept 7 days. seq benchmark evals <name> (the same as seq eval ls --benchmark <name>) lists every eval run on a benchmark across your policies, newest first, 50 a page (--limit, --cursor; --policy, --since 7d); seq.benchmarks() and seq.benchmark_evals(name) return the same pages in Python.

Adapters

A @seq.adapter bridges a benchmark’s observation and action to a policy’s own inputs and outputs. It is the third kind you register, and the only one without a machine: it belongs to neither the benchmark nor the policy but to the pairing an eval runs. seq deploy libero_to_pi05.py registers it as a version — its file is sent, nothing is built or run — under @seq.adapter(name=...), else its class name in lower case with hyphens. The version is the content: the same file again is the version it already is, a changed one the next. The eval names it, --adapter libero-to-pi05@v1; when the eval starts, the benchmark’s container fetches that version, checks every label again and runs it between the simulation and the policy. Changing it changes neither the benchmark’s version nor the policy’s, and the report’s first line reads benchmark vN + adapter vN + policy vN. seq adapter ls lists yours, seq adapter show <name>[@vN] says whether a version is lossy, and seq adapter pull <name>[@vN] writes its file back exactly as it was registered (seq.adapter_source(name, version) in Python): edit it and seq deploy it again for the next version. check and rm check and remove them.

Label every transform lossless or lossy(reason) — you cannot self-certify a lossy transform as clean, and lossy transforms are recorded in the report.

lossless (keep as default)lossy (declare + recorded in the report)
key rename, dimension reorder, dtype-preserving reshape, identity, unit-preserving coordinate relabel image resize / crop / downsample, dtype quantize, dropping a camera or channel, action-space projection (abs↔delta, mismatched dims), interpolating coordinate transforms

When shapes or spaces do not match and no adapter is given, eval refuses rather than guessing.

What an adapter can be: one file of at most 256 KB, run inside the benchmark’s container — it imports what the benchmark’s image installs (numpy, say) and the platform’s seq.tf transforms, and nothing else: no file beside it, no package of its own. It feeds one kind of policy, the robot-arm preset or a policy’s own inputs and outputs, and an eval pairing it with the other kind is refused. A version an eval is running cannot be removed. Numbers are never given twice: an adapter removed and registered again under its name goes on from its next number, and --rerun-failed runs the version its eval ran, refused (409) if that number no longer holds the same code.

Parameters in an eval

Every side of an eval declares what an eval may set the way infer does — keyword arguments with defaults, typed int, float, bool or str — and registering it sends them: the policy’s @seq.infer, the adapter’s __init__, the benchmark’s @seq.setup (after the environment count the platform passes). The eval loop has one of its own, eval.steps_per_query: how many actions of each chunk run before the policy is asked again. Set any of them for the whole eval with --param side.name=value — the bare name when only one side declares it — and sweep any of them with --sweep. A name no side declares, a wrong type or a steps_per_query below 1 is refused before anything is reserved, with every parameter the eval can set.

adapter.pypython
@seq.adapter(name="to-my-policy")
class ToMyPolicy(seq.Adapter):
    obs_transforms = [seq.tf.resize("image", (256, 256), (224, 224), lossy="the policy takes 224x224")]
    def __init__(self, *, steps_per_query: int = 5, image_size: int = 224):
        # built from the parameter, so the eval’s value reaches the transform (and the report)
        self.obs_transforms = [seq.tf.resize("image", (256, 256), (image_size, image_size),
                                             lossy=f"the policy takes {image_size}x{image_size}")]

# $ seq eval run my-policy --benchmark my-bench --adapter to-my-policy --param adapter.image_size=192
# $ seq eval run my-policy --benchmark my-bench --sweep eval.steps_per_query=5,10

Who decides. The eval’s value, else the default of the side that declares it. steps_per_query is one setting shared with an adapter that declares an int of that name (it is handed the value used): the eval’s, else the adapter’s default, else each chunk’s open_loop_horizon, else the whole chunk. One longer than a chunk stops the eval at the first query and names the most that fits. --dry-run and the report print every arm’s values with where each came from, and how many of its steps asked the policy: a batched benchmark renders only at those steps, so steps_per_query=10 renders half as often as 5.

Composite policies

A composite runs a planner model beside pi0.5: @seq.plan asks the planner for the next step about every every_s seconds, for each robot on its own and off the per-step path; @seq.infer runs pi0.5 every step with that robot’s latest step, which the platform passes in as instruction — the robot’s own task until the first plan arrives. The planner is a model of the catalogue, called through /v1/chat with sequence_ai.chat: declare it in models=, and that is all — each time a robot connects, the platform hands the policy a short-lived credential for exactly those models (see Keys). Each planner call is billed to your account like any chat call. A model that reasons by default, as qwen3-8-27b does, is told not to with reasoning={"effort": "none"}: a planner wants its one short line now, and thinking would spend the reply’s few tokens and leave it empty (see Reasoning).

policy.pypython
@seq.policy(name="my-robot", gpu="L4", max_gpus=1, weights=seq.Weights.template("pi05-droid", version=1),
            models=["qwen3-8-27b"])                       # the planner, from the catalogue
class Composite(seq.Policy):
    @seq.load
    def load(self, weights_dir):
        self.low = load_pi05(weights_dir)                 # runs on your deployment's GPU
    @seq.plan(every_s=2.0)                               # off the per-step path, per robot
    def plan(self, obs):
        reply = sequence_ai.chat("qwen3-8-27b", [
            {"role": "system", "content": "Reply with the one next step the arm should do now."},
            {"role": "user", "content": [{"type": "text", "text": f"Task: {obs.instruction}"},
                                         *[sequence_ai.image(f.data) for f in obs.images]]}],
            max_tokens=32, temperature=0, reasoning={"effort": "none"})
        return reply.content                             # the next step; None keeps the current one
    @seq.infer
    def infer(self, obs, *, seed=None, instruction=None):
        return self.low(obs, instruction or obs.instruction)   # every step: this robot's latest step
The platform keeps the step for you, one per robot — one per environment in an eval — and hands each request its own, so plan returns the step instead of storing it on self. A @seq.batched infer receives instruction as a list, one entry per request. An error in plan keeps the current step, and the planner asks again on its next turn.

Scheduled evals

Evaluate a deployment on a schedule — daily, weekly, or after every deploy — and read its evals as a trend: version, primary metric and its 95% interval. Each run is an ordinary eval, billed like one and stopped at --max-usd; a deployment stopped at its budget runs none.

a weekly regression checkshell
seq eval schedule my-pi05 --benchmark libero-custom --every weekly --trials 2 --max-usd 20
seq eval schedule my-pi05 --benchmark libero-custom --every deploy# each time the serving version changes
seq eval ls my-pi05                                               # its evals over time
seq eval schedule my-pi05 --benchmark libero-custom --off