Exit codes
The commands that gate a pipeline report their verdict through the exit code, and the codes are not uniform across Soup. This page lists what each command returns at v0.75.2. The rows come from the call sites in the v0.75.2 source and, wherever it could be done without a model or a network service, were checked by running the command against a small local fixture. Rows for commands that need a model, a GPU or a network service (soup data canary check, soup data doctor, soup draft measure, the verdict of soup shrink, soup eval unlearning, soup eval behavior, soup eval checklist, soup probe, soup can run, soup ui) were read from the call site and not run, and so were the rarer error branches elsewhere, such as the hardware-fit refusal in soup train.
The PR gate that soup ci init writes runs soup data validate, soup expect and soup ship --evidence in that order, and a non-zero exit from any step fails the job. Everything below is what those exits, and the rest, mean.
Three rules that hold for every command
- 0 is success.
--helpexits 0. - Click rejects a bad invocation with exit 2 before the command's own code runs. That covers an unknown command or flag, a missing required option, and a value that fails Click's own type check:
soup ship --bogusandsoup ship --forgetting-threshold abcboth exit 2. This applies to every command, including the ones that use 2 for a verdict. - An unexpected exception prints a short error and exits 1, and Ctrl+C exits 130.
soup --verbose <command>shows the traceback. A malformed YAML suite passed tosoup eval gateis one example.soup traininstalls no Ctrl+C handler of its own at v0.75.2, so interrupting a run is an ordinary 130, see Observability.
Past those three, a code means what the command says it means. In the gates, 2 is usually a bad verdict (MAJOR, DON'T SHIP, gameable) and 3 a failed check or drift; in many validators 2 is a rejected argument; soup ship uses 3 for usage errors. Read the row for the command you gate on and do not carry it over to a sibling.
Where one code means two things
soup ship: exit 2 is DON'T SHIP, and it is also what Click returns for an unknown flag or a malformed number. A typo in the CI file reads as a regression. Soup's own usage errors are 3, but Click's are not.soup diagnose,soup data doctor,soup data lint,soup eval unlearning,soup eval behavior,soup eval checklistandsoup probe: exit 2 is a MAJOR verdict, and each of them also uses 2 for some rejected inputs (a missing--evidencefile or an invalid--citation-styleonsoup diagnose, preference data handed tosoup data doctor).soup eval gate: exit 1 is a failed gate and also an unloadable suite, baseline or model.soup adapters scan: exit 1 is a WARN verdict and also an adapter that is missing or unreadable.soup adapters verifyandsoup adapters check-safetensors: a failed check is 1 by default and 3 with--strict. Both also exit 3, with or without--strict, when the directory is outside the working directory or is a symlink, andsoup adapters verifyexits 3 on a.soup-signature.jsonthat cannot be parsed.soup adapters merge: exit 2 is a rejected argument and, with--strict-verdict, a MAJOR canary verdict; exit 3 is a refused gate, not a verdict.soup env check: exit 3 is ABI drift against the lock and also an installed package outside the bounds Soup declares for itself.soup train: exit 2 can arrive after a finished run, and a run that a gate stopped early exits 0. See the train rows below.
Release and quality gates
Ship and the eval commands
| Command | Code | Means |
|---|---|---|
soup ship | 0 | SHIP. |
soup ship | 2 | DON'T SHIP: the task leg or the forgetting leg said no. |
soup ship | 3 | A usage or validation error Soup detected itself: no mode chosen, --base without --tuned or --adapter, --noise-floor together with --evidence, an unreadable --config, or evidence whose provenance does not match the committed config. |
soup ship | 1 | A runtime error: an evidence file that is missing, is not JSON or has a malformed score block, a model that fails to load, an unwritable --output. |
soup eval gate | 0 | Every task passed. An empty suite passes, and without --model the run uses a stub generator that returns empty text, so it checks the wiring and not a model. |
soup eval gate | 1 | A task failed, or anything else went wrong (suite, baseline or model could not be loaded, --regression-threshold out of range). There is no separate code for a broken setup. |
soup eval against | 0 | No regression on the metric. |
soup eval against | 1 | A regression, or an error (a locked suite that is missing or invalid, an empty metric series, a run id that cannot be read). |
soup eval quant-check | 0 | Always, once it has run. The OK, MINOR and MAJOR verdicts are printed and do not change the exit code, and a model that fails to load falls back to a stub with a yellow line and still exits 0. Read the --format json output if you need to gate on it. |
soup eval quant-check | 1 | An unknown --format, or a --before, --after or --tasks path that is missing, outside the working directory or an unresolvable registry reference. A tasks file that is not valid JSONL also exits 1. |
soup eval unlearning | 0 | Overall OK or MINOR. Without --evidence every metric is scored a neutral OK and the command exits 0. |
soup eval unlearning | 2 | Overall MAJOR, or a rejected input. |
soup eval behavior | 0 | Overall OK or MINOR. Without --base-model or --evidence it prints "No --evidence supplied; emitting neutral OK report." and exits 0 without scoring anything. |
soup eval behavior | 2 | Overall MAJOR, or a rejected input. |
soup eval checklist | 0 | Overall OK or MINOR. |
soup eval checklist | 2 | Overall MAJOR, or a rejected input. |
soup diagnose | 0 | Overall OK or MINOR. The worst mode decides: a score of 0.85 or above is OK, 0.60 or above is MINOR, anything lower is MAJOR. Without --base-model or --evidence every mode is scored a neutral OK and the command exits 0, so a pipeline has to pass one of them. |
soup diagnose | 2 | Overall MAJOR. Also an --evidence file that does not exist and an invalid --citation-style. |
soup diagnose | 1 | An --evidence file that is not valid JSON, a live run that failed, or an unwritable --output. |
soup eval capability and soup eval citation return no verdict code. A failed --push PR comment on soup ship prints a warning and leaves the verdict's exit code alone.
Several of these commands score only what you give them, and a missing input is a pass. soup diagnose, soup eval behavior, soup eval unlearning and soup probe sleeper exit 0 when run with no evidence (and, for the first two, no --base-model): they print a neutral OK report and score nothing. A CI step that drops the flag stays green.
Data
| Command | Code | Means |
|---|---|---|
soup data validate | 0 | At least one row is usable and any --min-valid-fraction is met. Issues on some rows do not fail the run when no minimum is given. An empty file also exits 0 when --format is given explicitly (it reports 0/0 rows valid); without --format it exits 1. |
soup data validate | 2 | No usable rows, or the valid fraction is below --min-valid-fraction. |
soup data validate | 1 | A missing file, an unknown --format, or a format that cannot be detected. |
soup data doctor | 0 | Overall OK or MINOR. |
soup data doctor | 2 | Overall MAJOR, or data of the wrong kind: preference data (use soup data lint), or RAFT data without --show-mask. |
soup data doctor | 1 | A missing path, an empty dataset, a tokenizer that fails to load, or an unwritable --output. |
soup data lint | 0 | Overall OK or MINOR. |
soup data lint | 2 | Overall MAJOR, or data that is not preference data (use soup data doctor). |
soup data lint | 1 | A missing path, an empty dataset, or an unwritable --output. |
soup data canary check | 0 | OK or MINOR. |
soup data canary check | 2 | MAJOR: a canary was memorized. |
soup data canary check | 1 | A runtime error: manifest unreadable, model load failed, or the [train] extra is missing. |
soup data brain-rot | 0 | Always, unless --strict is set. |
soup data brain-rot | 3 | With --strict, the fraction of MAJOR rows exceeds --max-major-fraction. |
soup data brain-rot | 2 | A rejected input. |
soup expect | 0 | Every expectation passed. |
soup expect | 3 | At least one expectation failed. |
soup expect | 2 | The suite or the data was rejected: a missing or invalid suite, an unreadable dataset, a path outside the working directory. |
soup data decontaminate, soup data toxicity, soup data langdetect, soup data pii, soup data educational and soup data score filter or score a file and write the result. They exit 0 whatever they find.
Models, adapters and supply chain
| Command | Code | Means |
|---|---|---|
soup adapters scan | 0 | OK. |
soup adapters scan | 1 | WARN, or the adapter is missing or unreadable. |
soup adapters scan | 3 | FAIL: the spectral scan found a likely backdoor pattern (rank-1 dominance, energy concentration, NaN or Inf, norm outliers). |
soup adapters scan | 2 | An unknown --format, or an adapter path outside the working directory. |
soup adapters verify | 0 | A signature record is present and the files match it. An unsigned record is a hash manifest with no signer; an ed25519 record is checked against its own embedded key unless --public-key is given. |
soup adapters verify | 1 | The signature is absent or does not match, or the directory is missing. |
soup adapters verify | 3 | The same failure as 1, with --strict. Also without --strict: a directory outside the working directory or a symlink, or a signature file that cannot be parsed. |
soup adapters check-safetensors | 0 | Every weight file is safetensors. |
soup adapters check-safetensors | 1 | A pickle or PyTorch-classic weight file was found, or the directory is missing. |
soup adapters check-safetensors | 3 | The same finding as 1, with --strict (a missing directory stays 1). Also without --strict: a directory outside the working directory or a symlink. |
soup adapters merge | 0 | Merged. |
soup adapters merge | 3 | Refused by a gate: the spectral scan failed or could not run on an input (--allow-unscanned skips it), or the licenses conflict (--license-override with a reason of at least 8 characters proceeds and is recorded to the audit log). |
soup adapters merge | 2 | Rejected arguments: an unknown --strategy, fewer than two adapters, a --license count that does not match the adapters, --strategy cmaes without --eval (or with --canary or --strict-verdict), or a --canary file that cannot be read or scored. Also a MAJOR canary verdict when --strict-verdict is given: the merged adapter is already written when it exits. |
soup adapters bisect | 0 | ALL_OK: the endpoints both pass. |
soup adapters bisect | 2 | BROKEN_AT, or rejected arguments. The --eval-command is judged by its own exit code: 0 passes a checkpoint and anything else fails it. |
soup probe sleeper | 2 | A MAJOR verdict, or a rejected input. The same holds for soup probe truth and soup probe harm. |
soup probe interference | 2 | The worst pair scores 0.20 or more in absolute value, or a rejected input. |
soup shrink | 0 | SHIP: perplexity regression within --tolerance. Also --plan-only. |
soup shrink | 2 | DON'T SHIP: perplexity regression above --tolerance. |
soup shrink | 1 | Anything else, bad arguments included. Unlike soup ship, a usage error is 1 here. |
soup draft measure | 0 | The report was produced. Without --min-acceptance no acceptance rate fails the run. |
soup draft measure | 2 | --min-acceptance was given and the measured acceptance rate is below it. |
soup draft measure | 1 | An error: unreadable prompts, a model pair that fails to load, or a target that generated no tokens. |
soup reward synth | 0 | A verifier was emitted, or --plan-only printed the plan. |
soup reward synth | 2 | Refused: the verifier could not tell the references from perturbed negatives. Nothing is left at -o. |
soup reward synth | 1 | Usage or runtime error: no -o, an existing output without --force, an output that is not a .py file. |
soup reward stress | 0 | Robust: the junk completions stayed under --max-gameable. |
soup reward stress | 2 | Gameable. |
soup reward stress | 1 | Usage or runtime error: the target was not found or could not be loaded, a bad --attacks list. |
Provenance, locks and plans
| Command | Code | Means |
|---|---|---|
soup attest emit | 0 | Statement printed or written. --sign sigstore is accepted, falls back to an unsigned statement and still exits 0. |
soup attest emit | 2 | Rejected input: an invalid stage or SHA, a signing failure (ed25519 without a usable key), a write failure, or --attach-to-registry without --output. |
soup attest emit | 1 | --attach-to-registry failed after the statement was written. |
soup attest verify | 0 | The ed25519 signature is valid against the key embedded in the sidecar. That shows the statement matches its sidecar, not who signed it: pass --public-key trusted.pem to require a specific signer. The subject digest is asserted by the signer and is not re-checked against an artifact. |
soup attest verify | 3 | Invalid: a tampered statement or wrong key, an unsigned sidecar, no public key to verify with, or an embedded key that does not match --public-key. |
soup attest verify | 2 | An input that cannot be read, or a path outside the working directory. |
soup bom emit | 0 | BOM printed or written. |
soup bom emit | 2 | Rejected input: an unsupported --format, a SHA that is not 64 hex characters, a bad --energy file, or --attach-to-registry without --output. |
soup bom emit | 1 | --attach-to-registry failed after the BOM was written. |
soup lock check | 0 | The closure matches. |
soup lock check | 3 | Drift. |
soup lock check | 2 | The lock file is missing or unreadable, or a SHA is malformed. |
soup env check | 0 | No drift and every installed package is inside the bounds Soup declares. |
soup env check | 3 | ABI drift against the lock, or a package outside the declared bounds. The bounds check needs no lock file and comes first. |
soup env check | 1 | There is no lock file. |
soup env check | 2 | The lock file is unreadable. |
soup plan | 0 | The plan was written to soup.tfstate. |
soup plan | 1 | The config file was not found. |
soup plan | 2 | The config or the state was rejected (for example a config that is not a mapping). A YAML syntax error is not caught there and exits 1. |
soup apply | 0 | No drift. --dry-run stops there. Without it the state is marked applied and the soup train command is printed: soup apply does not start training. |
soup apply | 3 | The config has drifted from the plan. |
soup apply | 1 | The config or the state file does not exist. |
soup apply | 2 | The config or the state was rejected. A YAML syntax error in the config exits 1. |
soup license-advisor | 0 | Advice printed. A warn-level risk does not fail the run. |
soup license-advisor | 3 | --license was given and its risk for the --target is at block severity. |
soup license-advisor | 2 | An unknown --target. |
soup drift-alarm | 0 | The KL divergence is within --threshold. Also when either file has no usable rows (a row needs a string output, response or text field): the divergence is then reported as 0 and nothing is compared, so check n_reference and n_live in the output. |
soup drift-alarm | 3 | Drift. Any Slack or Discord webhook is posted first. |
soup drift-alarm | 1 | A missing file, or a path outside the working directory. |
soup drift-alarm | 2 | A --threshold that is not above zero or is above 100, or a webhook URL that is refused. |
soup apple-adapter | 0 | Converted, or --plan-only. |
soup apple-adapter | 2 | Rejected input or a missing source. |
soup apple-adapter | 3 | The conversion is gated upstream and was refused. |
soup apple-adapter | 1 | A dependency is missing. |
Commands that start something
| Command | Code | Means |
|---|---|---|
soup train | 0 | Training finished and the adapter is saved. Also --dry-run on a valid config, and --cloud without --cloud-submit (it prints a plan and starts nothing). A run stopped early by --gate or training.eval_gate with on_regression: stop, by the loss watchdog, or by a reward-hack or echo-trap halt also exits 0: soup train does not turn the stop into a failure, so check the run's output and metrics. |
soup train | 1 | The config is missing, unreadable or invalid (an unknown key is refused), the hardware-fit gate refuses the run (--allow-oom-attempt overrides it), a --gate suite cannot be loaded, a --replay file does not exist, --gpus above 1 was given with --no-reexec (it prints the launch command and stops), or training raised an error. The SFT trainer also refuses to save the final model, and exits 1, when the last logged loss or the trainable weights are non-finite. |
soup train | 2 | A flag override that fails validation (the --replay options), an invalid tracker option, an invalid --cloud or --gpu (--cloud runpod is recognized and refused), or --diagnose-gate finding a MAJOR mode. The last one fires after training: the adapter is already saved and the run is recorded as finished. |
soup train | the submit call's code | With --cloud modal or --cloud lambda plus --cloud-submit, the return code of the submit call (1 when the backend's submit is unavailable). With --gpus above 1 the run is handed to accelerate launch, so the code is the launcher's rather than Soup's. |
soup quickstart | the child's code | Exits with the code of the soup train subprocess it starts, or 0 with --dry-run. |
soup llama | the child's code | Exits with the code of the llama.cpp program it runs, 2 when that program is not found on PATH, 1 when it cannot be launched. |
soup serve | 2 | A refused option: a non-loopback --host without --tool-auth-token, --bank, --mole or --steer with a backend other than transformers or a path outside the working directory, --kv-cache-type fp8, or an unusable --hub. |
soup serve | 1 | Other startup failures, for example a model path that does not exist. |
soup mcp serve | 2 | A refused option combination: an unknown --transport, --host, --port or --auth-token with stdio, --allow-execute with sse or http, or an invalid --auth-token. |
soup mcp serve | 1 | The mcp SDK is not installed, or is too old for the transport. |
soup mcp runs reconcile | 2 | --expunge-launching was not passed, and no run record changed. |
soup mcp runs reconcile | 1 | A launching run whose process is still alive was refused. |
soup ui | 2 | Reserved: a non-loopback host (--public or --host) without a valid auth token. The command's own options cannot reach it, because a token is generated at startup and an invalid --auth-token exits 1 first, so soup ui --host 0.0.0.0 starts the server. |
soup ui | 1 | An invalid --auth-token, or the [ui] extra is not installed. |
soup can run | 1 | Without --yes it prints the confirmation panel, exits 1 and runs nothing. Also an unreadable can. |
soup can run | the child's code | With --yes, the exit code of the training subprocess is passed through, then that of the deploy step when --deploy is set. |
soup doctor | 0 | Advisory findings (a missing optional extra, a torchvision mismatch) and settings your config writes that its backend never reads do not change the exit code. |
soup doctor | 1 | A required core dependency is missing or incompatible. |
soup doctor | 2 | --config named a file that is missing, unparseable or fails the schema. It is reported after the environment summary, so "All checks passed" can print above it. |
Commands with no verdict code of their own
Besides Click's exit 2 for a bad invocation, these exit only 0 or 1: there is no other exit code anywhere in their source at v0.75.2. Gate on 0 against everything else.
soup init,soup cost,soup card,soup diff,soup history,soup why,soup profile,soup sweep,soup autopilot,soup ci initsoup advise(run, explain and compare),soup recipes,soup registrysoup eval benchmark,soup eval custom,soup eval judge,soup eval auto,soup eval compare,soup eval leaderboard,soup eval human,soup eval aidersoup datainspect, convert, merge, dedup, filter, stats, sample, split, search, preview, from-traces, review, ingest, preprocess, generate and topicssoup runsshow, compare, replay and deletesoup canpack, inspect, verify, publish and forksoup deploy ollama,soup deploy hf-space,soup agent synth,soup agent eval
soup eval custom never exits non-zero on a low score, so it cannot act as a checkpoint test for soup adapters bisect on its own. Several validating commands outside the tables above, among them soup export, soup merge, soup quantize, soup push, soup infer, soup fetch, soup ingest, soup tunability, soup spectrum scan, soup deploy autopilot and soup loop, use 2 for a rejected argument and have no verdict.
Soup is free and Apache-2.0. If it saved you a training run, starring the repo costs nothing and helps most. You can also fund the GPU time behind the work a 4 GB laptop cannot reach.