Exit codes

The commands that gate a pipeline report their verdict through the exit code, and the codes are not uniform across Soup. This page lists what each command returns at v0.75.2. The rows come from the call sites in the v0.75.2 source and, wherever it could be done without a model or a network service, were checked by running the command against a small local fixture. Rows for commands that need a model, a GPU or a network service (soup data canary check, soup data doctor, soup draft measure, the verdict of soup shrink, soup eval unlearning, soup eval behavior, soup eval checklist, soup probe, soup can run, soup ui) were read from the call site and not run, and so were the rarer error branches elsewhere, such as the hardware-fit refusal in soup train.

The PR gate that soup ci init writes runs soup data validate, soup expect and soup ship --evidence in that order, and a non-zero exit from any step fails the job. Everything below is what those exits, and the rest, mean.

Three rules that hold for every command

  • 0 is success. --help exits 0.
  • Click rejects a bad invocation with exit 2 before the command's own code runs. That covers an unknown command or flag, a missing required option, and a value that fails Click's own type check: soup ship --bogus and soup ship --forgetting-threshold abc both exit 2. This applies to every command, including the ones that use 2 for a verdict.
  • An unexpected exception prints a short error and exits 1, and Ctrl+C exits 130. soup --verbose <command> shows the traceback. A malformed YAML suite passed to soup eval gate is one example. soup train installs no Ctrl+C handler of its own at v0.75.2, so interrupting a run is an ordinary 130, see Observability.

Past those three, a code means what the command says it means. In the gates, 2 is usually a bad verdict (MAJOR, DON'T SHIP, gameable) and 3 a failed check or drift; in many validators 2 is a rejected argument; soup ship uses 3 for usage errors. Read the row for the command you gate on and do not carry it over to a sibling.

Where one code means two things

  • soup ship: exit 2 is DON'T SHIP, and it is also what Click returns for an unknown flag or a malformed number. A typo in the CI file reads as a regression. Soup's own usage errors are 3, but Click's are not.
  • soup diagnose, soup data doctor, soup data lint, soup eval unlearning, soup eval behavior, soup eval checklist and soup probe: exit 2 is a MAJOR verdict, and each of them also uses 2 for some rejected inputs (a missing --evidence file or an invalid --citation-style on soup diagnose, preference data handed to soup data doctor).
  • soup eval gate: exit 1 is a failed gate and also an unloadable suite, baseline or model.
  • soup adapters scan: exit 1 is a WARN verdict and also an adapter that is missing or unreadable.
  • soup adapters verify and soup adapters check-safetensors: a failed check is 1 by default and 3 with --strict. Both also exit 3, with or without --strict, when the directory is outside the working directory or is a symlink, and soup adapters verify exits 3 on a .soup-signature.json that cannot be parsed.
  • soup adapters merge: exit 2 is a rejected argument and, with --strict-verdict, a MAJOR canary verdict; exit 3 is a refused gate, not a verdict.
  • soup env check: exit 3 is ABI drift against the lock and also an installed package outside the bounds Soup declares for itself.
  • soup train: exit 2 can arrive after a finished run, and a run that a gate stopped early exits 0. See the train rows below.

Release and quality gates

Ship and the eval commands

CommandCodeMeans
soup ship0SHIP.
soup ship2DON'T SHIP: the task leg or the forgetting leg said no.
soup ship3A usage or validation error Soup detected itself: no mode chosen, --base without --tuned or --adapter, --noise-floor together with --evidence, an unreadable --config, or evidence whose provenance does not match the committed config.
soup ship1A runtime error: an evidence file that is missing, is not JSON or has a malformed score block, a model that fails to load, an unwritable --output.
soup eval gate0Every task passed. An empty suite passes, and without --model the run uses a stub generator that returns empty text, so it checks the wiring and not a model.
soup eval gate1A task failed, or anything else went wrong (suite, baseline or model could not be loaded, --regression-threshold out of range). There is no separate code for a broken setup.
soup eval against0No regression on the metric.
soup eval against1A regression, or an error (a locked suite that is missing or invalid, an empty metric series, a run id that cannot be read).
soup eval quant-check0Always, once it has run. The OK, MINOR and MAJOR verdicts are printed and do not change the exit code, and a model that fails to load falls back to a stub with a yellow line and still exits 0. Read the --format json output if you need to gate on it.
soup eval quant-check1An unknown --format, or a --before, --after or --tasks path that is missing, outside the working directory or an unresolvable registry reference. A tasks file that is not valid JSONL also exits 1.
soup eval unlearning0Overall OK or MINOR. Without --evidence every metric is scored a neutral OK and the command exits 0.
soup eval unlearning2Overall MAJOR, or a rejected input.
soup eval behavior0Overall OK or MINOR. Without --base-model or --evidence it prints "No --evidence supplied; emitting neutral OK report." and exits 0 without scoring anything.
soup eval behavior2Overall MAJOR, or a rejected input.
soup eval checklist0Overall OK or MINOR.
soup eval checklist2Overall MAJOR, or a rejected input.
soup diagnose0Overall OK or MINOR. The worst mode decides: a score of 0.85 or above is OK, 0.60 or above is MINOR, anything lower is MAJOR. Without --base-model or --evidence every mode is scored a neutral OK and the command exits 0, so a pipeline has to pass one of them.
soup diagnose2Overall MAJOR. Also an --evidence file that does not exist and an invalid --citation-style.
soup diagnose1An --evidence file that is not valid JSON, a live run that failed, or an unwritable --output.

soup eval capability and soup eval citation return no verdict code. A failed --push PR comment on soup ship prints a warning and leaves the verdict's exit code alone.

Several of these commands score only what you give them, and a missing input is a pass. soup diagnose, soup eval behavior, soup eval unlearning and soup probe sleeper exit 0 when run with no evidence (and, for the first two, no --base-model): they print a neutral OK report and score nothing. A CI step that drops the flag stays green.

Data

CommandCodeMeans
soup data validate0At least one row is usable and any --min-valid-fraction is met. Issues on some rows do not fail the run when no minimum is given. An empty file also exits 0 when --format is given explicitly (it reports 0/0 rows valid); without --format it exits 1.
soup data validate2No usable rows, or the valid fraction is below --min-valid-fraction.
soup data validate1A missing file, an unknown --format, or a format that cannot be detected.
soup data doctor0Overall OK or MINOR.
soup data doctor2Overall MAJOR, or data of the wrong kind: preference data (use soup data lint), or RAFT data without --show-mask.
soup data doctor1A missing path, an empty dataset, a tokenizer that fails to load, or an unwritable --output.
soup data lint0Overall OK or MINOR.
soup data lint2Overall MAJOR, or data that is not preference data (use soup data doctor).
soup data lint1A missing path, an empty dataset, or an unwritable --output.
soup data canary check0OK or MINOR.
soup data canary check2MAJOR: a canary was memorized.
soup data canary check1A runtime error: manifest unreadable, model load failed, or the [train] extra is missing.
soup data brain-rot0Always, unless --strict is set.
soup data brain-rot3With --strict, the fraction of MAJOR rows exceeds --max-major-fraction.
soup data brain-rot2A rejected input.
soup expect0Every expectation passed.
soup expect3At least one expectation failed.
soup expect2The suite or the data was rejected: a missing or invalid suite, an unreadable dataset, a path outside the working directory.

soup data decontaminate, soup data toxicity, soup data langdetect, soup data pii, soup data educational and soup data score filter or score a file and write the result. They exit 0 whatever they find.

Models, adapters and supply chain

CommandCodeMeans
soup adapters scan0OK.
soup adapters scan1WARN, or the adapter is missing or unreadable.
soup adapters scan3FAIL: the spectral scan found a likely backdoor pattern (rank-1 dominance, energy concentration, NaN or Inf, norm outliers).
soup adapters scan2An unknown --format, or an adapter path outside the working directory.
soup adapters verify0A signature record is present and the files match it. An unsigned record is a hash manifest with no signer; an ed25519 record is checked against its own embedded key unless --public-key is given.
soup adapters verify1The signature is absent or does not match, or the directory is missing.
soup adapters verify3The same failure as 1, with --strict. Also without --strict: a directory outside the working directory or a symlink, or a signature file that cannot be parsed.
soup adapters check-safetensors0Every weight file is safetensors.
soup adapters check-safetensors1A pickle or PyTorch-classic weight file was found, or the directory is missing.
soup adapters check-safetensors3The same finding as 1, with --strict (a missing directory stays 1). Also without --strict: a directory outside the working directory or a symlink.
soup adapters merge0Merged.
soup adapters merge3Refused by a gate: the spectral scan failed or could not run on an input (--allow-unscanned skips it), or the licenses conflict (--license-override with a reason of at least 8 characters proceeds and is recorded to the audit log).
soup adapters merge2Rejected arguments: an unknown --strategy, fewer than two adapters, a --license count that does not match the adapters, --strategy cmaes without --eval (or with --canary or --strict-verdict), or a --canary file that cannot be read or scored. Also a MAJOR canary verdict when --strict-verdict is given: the merged adapter is already written when it exits.
soup adapters bisect0ALL_OK: the endpoints both pass.
soup adapters bisect2BROKEN_AT, or rejected arguments. The --eval-command is judged by its own exit code: 0 passes a checkpoint and anything else fails it.
soup probe sleeper2A MAJOR verdict, or a rejected input. The same holds for soup probe truth and soup probe harm.
soup probe interference2The worst pair scores 0.20 or more in absolute value, or a rejected input.
soup shrink0SHIP: perplexity regression within --tolerance. Also --plan-only.
soup shrink2DON'T SHIP: perplexity regression above --tolerance.
soup shrink1Anything else, bad arguments included. Unlike soup ship, a usage error is 1 here.
soup draft measure0The report was produced. Without --min-acceptance no acceptance rate fails the run.
soup draft measure2--min-acceptance was given and the measured acceptance rate is below it.
soup draft measure1An error: unreadable prompts, a model pair that fails to load, or a target that generated no tokens.
soup reward synth0A verifier was emitted, or --plan-only printed the plan.
soup reward synth2Refused: the verifier could not tell the references from perturbed negatives. Nothing is left at -o.
soup reward synth1Usage or runtime error: no -o, an existing output without --force, an output that is not a .py file.
soup reward stress0Robust: the junk completions stayed under --max-gameable.
soup reward stress2Gameable.
soup reward stress1Usage or runtime error: the target was not found or could not be loaded, a bad --attacks list.

Provenance, locks and plans

CommandCodeMeans
soup attest emit0Statement printed or written. --sign sigstore is accepted, falls back to an unsigned statement and still exits 0.
soup attest emit2Rejected input: an invalid stage or SHA, a signing failure (ed25519 without a usable key), a write failure, or --attach-to-registry without --output.
soup attest emit1--attach-to-registry failed after the statement was written.
soup attest verify0The ed25519 signature is valid against the key embedded in the sidecar. That shows the statement matches its sidecar, not who signed it: pass --public-key trusted.pem to require a specific signer. The subject digest is asserted by the signer and is not re-checked against an artifact.
soup attest verify3Invalid: a tampered statement or wrong key, an unsigned sidecar, no public key to verify with, or an embedded key that does not match --public-key.
soup attest verify2An input that cannot be read, or a path outside the working directory.
soup bom emit0BOM printed or written.
soup bom emit2Rejected input: an unsupported --format, a SHA that is not 64 hex characters, a bad --energy file, or --attach-to-registry without --output.
soup bom emit1--attach-to-registry failed after the BOM was written.
soup lock check0The closure matches.
soup lock check3Drift.
soup lock check2The lock file is missing or unreadable, or a SHA is malformed.
soup env check0No drift and every installed package is inside the bounds Soup declares.
soup env check3ABI drift against the lock, or a package outside the declared bounds. The bounds check needs no lock file and comes first.
soup env check1There is no lock file.
soup env check2The lock file is unreadable.
soup plan0The plan was written to soup.tfstate.
soup plan1The config file was not found.
soup plan2The config or the state was rejected (for example a config that is not a mapping). A YAML syntax error is not caught there and exits 1.
soup apply0No drift. --dry-run stops there. Without it the state is marked applied and the soup train command is printed: soup apply does not start training.
soup apply3The config has drifted from the plan.
soup apply1The config or the state file does not exist.
soup apply2The config or the state was rejected. A YAML syntax error in the config exits 1.
soup license-advisor0Advice printed. A warn-level risk does not fail the run.
soup license-advisor3--license was given and its risk for the --target is at block severity.
soup license-advisor2An unknown --target.
soup drift-alarm0The KL divergence is within --threshold. Also when either file has no usable rows (a row needs a string output, response or text field): the divergence is then reported as 0 and nothing is compared, so check n_reference and n_live in the output.
soup drift-alarm3Drift. Any Slack or Discord webhook is posted first.
soup drift-alarm1A missing file, or a path outside the working directory.
soup drift-alarm2A --threshold that is not above zero or is above 100, or a webhook URL that is refused.
soup apple-adapter0Converted, or --plan-only.
soup apple-adapter2Rejected input or a missing source.
soup apple-adapter3The conversion is gated upstream and was refused.
soup apple-adapter1A dependency is missing.

Commands that start something

CommandCodeMeans
soup train0Training finished and the adapter is saved. Also --dry-run on a valid config, and --cloud without --cloud-submit (it prints a plan and starts nothing). A run stopped early by --gate or training.eval_gate with on_regression: stop, by the loss watchdog, or by a reward-hack or echo-trap halt also exits 0: soup train does not turn the stop into a failure, so check the run's output and metrics.
soup train1The config is missing, unreadable or invalid (an unknown key is refused), the hardware-fit gate refuses the run (--allow-oom-attempt overrides it), a --gate suite cannot be loaded, a --replay file does not exist, --gpus above 1 was given with --no-reexec (it prints the launch command and stops), or training raised an error. The SFT trainer also refuses to save the final model, and exits 1, when the last logged loss or the trainable weights are non-finite.
soup train2A flag override that fails validation (the --replay options), an invalid tracker option, an invalid --cloud or --gpu (--cloud runpod is recognized and refused), or --diagnose-gate finding a MAJOR mode. The last one fires after training: the adapter is already saved and the run is recorded as finished.
soup trainthe submit call's codeWith --cloud modal or --cloud lambda plus --cloud-submit, the return code of the submit call (1 when the backend's submit is unavailable). With --gpus above 1 the run is handed to accelerate launch, so the code is the launcher's rather than Soup's.
soup quickstartthe child's codeExits with the code of the soup train subprocess it starts, or 0 with --dry-run.
soup llamathe child's codeExits with the code of the llama.cpp program it runs, 2 when that program is not found on PATH, 1 when it cannot be launched.
soup serve2A refused option: a non-loopback --host without --tool-auth-token, --bank, --mole or --steer with a backend other than transformers or a path outside the working directory, --kv-cache-type fp8, or an unusable --hub.
soup serve1Other startup failures, for example a model path that does not exist.
soup mcp serve2A refused option combination: an unknown --transport, --host, --port or --auth-token with stdio, --allow-execute with sse or http, or an invalid --auth-token.
soup mcp serve1The mcp SDK is not installed, or is too old for the transport.
soup mcp runs reconcile2--expunge-launching was not passed, and no run record changed.
soup mcp runs reconcile1A launching run whose process is still alive was refused.
soup ui2Reserved: a non-loopback host (--public or --host) without a valid auth token. The command's own options cannot reach it, because a token is generated at startup and an invalid --auth-token exits 1 first, so soup ui --host 0.0.0.0 starts the server.
soup ui1An invalid --auth-token, or the [ui] extra is not installed.
soup can run1Without --yes it prints the confirmation panel, exits 1 and runs nothing. Also an unreadable can.
soup can runthe child's codeWith --yes, the exit code of the training subprocess is passed through, then that of the deploy step when --deploy is set.
soup doctor0Advisory findings (a missing optional extra, a torchvision mismatch) and settings your config writes that its backend never reads do not change the exit code.
soup doctor1A required core dependency is missing or incompatible.
soup doctor2--config named a file that is missing, unparseable or fails the schema. It is reported after the environment summary, so "All checks passed" can print above it.

Commands with no verdict code of their own

Besides Click's exit 2 for a bad invocation, these exit only 0 or 1: there is no other exit code anywhere in their source at v0.75.2. Gate on 0 against everything else.

  • soup init, soup cost, soup card, soup diff, soup history, soup why, soup profile, soup sweep, soup autopilot, soup ci init
  • soup advise (run, explain and compare), soup recipes, soup registry
  • soup eval benchmark, soup eval custom, soup eval judge, soup eval auto, soup eval compare, soup eval leaderboard, soup eval human, soup eval aider
  • soup data inspect, convert, merge, dedup, filter, stats, sample, split, search, preview, from-traces, review, ingest, preprocess, generate and topics
  • soup runs show, compare, replay and delete
  • soup can pack, inspect, verify, publish and fork
  • soup deploy ollama, soup deploy hf-space, soup agent synth, soup agent eval

soup eval custom never exits non-zero on a low score, so it cannot act as a checkpoint test for soup adapters bisect on its own. Several validating commands outside the tables above, among them soup export, soup merge, soup quantize, soup push, soup infer, soup fetch, soup ingest, soup tunability, soup spectrum scan, soup deploy autopilot and soup loop, use 2 for a rejected argument and have no verdict.

Soup is free and Apache-2.0. If it saved you a training run, starring the repo costs nothing and helps most. You can also fund the GPU time behind the work a 4 GB laptop cannot reach.