Contributing to Soup

This page is a contributor guide for a first pull request. It restates upstream's CONTRIBUTING.md and AGENTS.md, CODEOWNERS, CODE_OF_CONDUCT.md, CONTRIBUTORS.md, the changelog fragment README, the pull request template and the CI workflows, all read at the v0.75.2 tag, and it adds what the test suite enforces. Where this page and those files differ, the files win.

The short version. Claim an issue by commenting on it. Branch, write the tests first, run ruff check and pytest, add one changelog fragment, and open a pull request. A merged pull request puts your name in CONTRIBUTORS.md.

Why contribute, and where work is handed out

Soup is Apache-2.0, and three recent releases came almost entirely from outside. All 24 pull requests in v0.73.3 came from people other than the maintainer, 116 of the 120 merged into v0.74.0 did, and so did all 60 in v0.75.0. The two releases after that, v0.75.1 and v0.75.2, were security backports the maintainer wrote alone.

Work comes from four places:

  • GitHub issues. Start with the good first issue label. The help wanted filter is where upstream's README lists the work that is blocked on hardware, so it suits you if you have a bigger GPU. Issues labelled infra-blocked are waiting on hardware or infrastructure that the person picking them up may not have.
  • Soup Tasters, the Telegram channel where unclaimed issues, hardware still to be tested on and benchmarks still to be run are posted.
  • Discord, which upstream recommends for quick questions and for pairing on a pull request. Anything worth finding later still belongs in Issues or Discussions, and a security report never goes in Discord.
  • GitHub Discussions, for questions.

Upstream names four good first contributions: a new recipe (a ready-made config for a popular model), documentation (docstrings, README examples, example configs), tests (more coverage for existing commands) and bug fixes (issues labelled bug). Measurements count as well: two of the measurement records published in v0.74.0 came from contributors, on hardware the project does not own.

Before you write code

  • Read the feature's page first. The README and the matching page under docs/ in the repository describe each feature; docs/commands.md is the full command list. The Pydantic schema in src/soup_cli/config/schema.py is the single source of truth for every config field; Configuration is the user-facing reference.
  • Read the code of conduct. It is the Contributor Covenant, and it applies in every project space, Discord included. Report a violation to team@trysoup.dev or makazanalpamys@gmail.com.
  • Know who reviews. CODEOWNERS names the maintainer, @MakazhanAlpamys, as the default owner of every path. It lists src/soup_cli/config/ and src/soup_cli/trainer/ separately, and src/soup_cli/ui/, src/soup_cli/data/providers/ and src/soup_cli/utils/ollama.py under a heading called "Security-sensitive areas".
  • AI coding agents. AGENTS.md is upstream's tool-agnostic entry point for Codex, Cursor, Aider, Claude Code and similar tools. It repeats the build commands and the conventions on this page in a shorter form. Its lint line, ruff check src/soup_cli/ tests/, leaves out scripts/ and benchmarks/, which CI also lints, so use the four-directory form from Set up your environment.

Set up your environment

You need Python 3.10, 3.11 or 3.12. Upstream declares requires-python = ">=3.10,<3.13", CI runs exactly those three versions, and tests/test_requires_python_bound.py derives the declared bound from the CI matrix, so widening one without the other fails the suite.

Fork the repository on GitHub and clone it. Upstream's guide shows the main repository's URL; use your fork's URL if you want to push to it. Then work inside a virtual environment and install in editable mode with the dev extra:

bash
git clone https://github.com/MakazhanAlpamys/Soup.git
cd Soup
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

These are upstream's commands, quoted from its guide and not run for this page. On Windows, activate with .venv\Scripts\activate. The virtual environment is not optional on Debian 12 or Ubuntu 23.04 and later: without it pip refuses with error: externally-managed-environment (PEP 668). pipx is the right tool for installing Soup and the wrong one here, because the test suite needs an editable checkout it can import.

In pyproject.toml at the tag, dev is the train, mcp and data extras plus cryptography (so the ed25519 signing tests run), reportlab (so the Annex XI and XII PDF tests run), pytest, pytest-cov, httpx, mypy, ruff, pre-commit and fastapi. That includes the full training stack, torch among it, so expect a large install. Every soup command writes under ~/.soup/; Environment variables and state files says how to keep a scratch session out of your real home.

Confirm the setup by running the suite and the linter. If both pass, you are ready:

bash
pytest tests/ -v --tb=short
ruff check src/soup_cli/ scripts/ tests/ benchmarks/

Some unit tests load small public models from the Hugging Face Hub (CI pre-downloads a handful of small repositories such as sshleifer/tiny-gpt2), so the first full run needs network access. CI forces UTF-8 on every platform with PYTHONUTF8=1 and PYTHONIOENCODING=utf-8, because Windows defaults to cp1252 and some upstream packages read their own files without naming an encoding. If a test dies on Windows with a charmap error, set the same two variables.

Git hooks

Upstream's .pre-commit-config.yaml installs with one command (the dev extra already provides pre-commit):

bash
pre-commit install

The hooks run ruff --fix, actionlint over the workflow files, and checks for end-of-file newlines, trailing whitespace, YAML and TOML syntax, merge-conflict markers and files over 5 MB. The actionlint hook builds with Go; the config says to swap it for the actionlint-docker hook id if you have no Go toolchain. The ruff-format hook is disabled on purpose and CI never runs ruff format --check: the config explains that enabling it would rewrite nearly every file and destroy git blame. Keep your diff to the lines you change.

How the repository is laid out

This map is built from the tree at v0.75.2. Upstream's own directory listing in CONTRIBUTING.md is older and omits several of these packages.

Inside src/soup_cli

PathWhat lives there
cli.pyThe entry point. It imports every command module, registers each on the Typer app, and defines the root options (--verbose, --log-level, --no-audit-log, --no-telemetry).
commands/One module per command or command group. commands/train.py is where a task is routed to its trainer.
config/schema.py, the Pydantic v2 models that define every config field and default, and loader.py, which loads a soup.yaml and refuses unknown keys.
trainer/One wrapper per training task around the TRL, Hugging Face or MLX trainer (sft.py, dpo.py, grpo.py and so on), plus stream_setup.py, the shared layer-streaming setup.
data/Loading, format detection and conversion (formats.py, loader.py), validation, collators, chat templates, loss masks, the synthetic-data providers and trace harvesting.
recipes/catalog.py: the ready-made configs, one RecipeMeta entry each.
templates/The YAML files behind soup init --template, and manifest.json, which lists them.
eval/The eval platform: custom tasks, the LLM judge, the eval gate, the leaderboard, forgetting detection, quantization checks and the Aider benchmark.
experiment/The SQLite experiment tracker.
monitoring/Training callbacks and the live display.
registry/ and cans/The model registry (hashing, store, diff, attach) and the Soup Can artifact format (pack, unpack, verify, publish, run).
migrate/Config conversion from LLaMA-Factory, Axolotl and Unsloth.
autopilot/The zero-config decision engine behind soup autopilot.
cloud/The Lambda and Modal backends for soup train --cloud.
envs/Bundled toy rollout environments (calculator, guess_number, retrieval_qa), used through training.rollout_func.
mcp_server/The MCP server behind soup mcp serve: server.py, the tool registry in registry.py, and execution.py.
plugins/The plugin and hook registry.
ui/The Web UI: a FastAPI app and its static single-page front end.
tui_app.pyThe terminal UI.
autodistill/The artifact contract for AutoDistill. Upstream's docs/autodistill.md calls it a design-and-contract slice that adds no task or training.
utils/Shared helpers: GPU detection, path containment (paths.py), hardware fit, quantization, layer streaming and the rest.

Everywhere else

PathWhat lives there
tests/The test suite, with tests/conftest.py for shared fixtures and tests/fixtures/ for data.
docs/Upstream's per-topic feature guides, with docs/commands.md as the command list.
examples/Example configs and datasets.
benchmarks/The published measurement records and a harness/ directory. CI lints it.
scripts/Among others, assemble_changelog.py, generate_recipe_snapshot.py and check_recipe_repo_ids.py.
changelog.d/Pending changelog fragments.
notebooks/Notebooks, including the reproduction notebook.
.github/The workflows, the issue forms, the pull request template and the pinned constraints file for the transformers-floor job.

Pick work and claim it

Comment on the issue saying you are taking it. GitHub only lets upstream set the formal assignee badge for repository collaborators, so the comment is the claim: the issue will look unassigned either way, and the thread is the record.

Upstream makes two commitments in return:

  • If a maintainer ends up doing the work instead, you are told in the thread, before or when it lands.
  • A claim that goes quiet is released with a comment, not silently. A maintainer checks in first, and "life got in the way" needs no explanation.

If an issue needs hardware you do not have, say so in the claim. Several open issues carry the infra-blocked label for exactly that reason, and knowing early is more useful than a stalled branch.

To file an issue, use one of the two forms. The bug report asks for a description, steps to reproduce, your soup.yaml if relevant, the error output, and your environment. The description, the steps and the environment are required, and the environment is the output of soup doctor. For the error output the form asks for the full error with a traceback, which comes from the root option --verbose, so it goes before the command name: soup --verbose train --config soup.yaml. The feature request asks for the problem, your proposed solution, alternatives you considered and an area; the problem, the solution and the area are required. Blank issues are allowed, and the form chooser links questions to Discussions.

Make the change

  1. Create a branch. Upstream's examples are feature/your-feature-name and fix/your-bug-fix.
  2. Write the tests first, then the code that passes them. Keep commits focused and logical.
  3. Add a changelog fragment if the change is visible to users. The next section has the rules.
  4. Lint, then test. Upstream's guide omits benchmarks/ from its --fix line, but CI lints it, so the form below covers all four directories.
  5. Commit with a Conventional Commits message. The types are feat, fix, refactor, docs, test, chore, perf and ci. Stage specific files rather than everything.
  6. Push and open a pull request. CI runs on pull requests to main.

The commands for steps 4 to 6:

bash
ruff check --fix src/soup_cli/ scripts/ tests/ benchmarks/
pytest tests/ -v --tb=short
git add <specific-files>
git commit -m "feat: add support for X"
git push origin feature/your-feature-name

Give the pull request a clear title, a description of what changed and why, any related issue (for example "Closes #123"), and your test results. The template then asks you to pick a type of change (bug fix, new feature, refactor or cleanup, documentation, tests) and to tick four boxes:

  • ruff check passes. The template's command omits benchmarks/; use the CI form above.
  • pytest tests/ -v passes.
  • The relevant docs are updated (README.md and the matching page under docs/), if needed.
  • A changelog fragment is added, or the change has no user-visible impact.

CONTRIBUTING.md keeps its own version of the checklist, which adds two items: new tests for new functionality, and no breaking changes, or any that are documented in the pull request description.

Changelog fragments

A user-visible pull request adds one Markdown file instead of editing CHANGELOG.md:

text
changelog.d/<latest-release>/<number>.<category>.md
  • <latest-release> is the newest released version heading at the top of CHANGELOG.md on the branch you are building on, so read it there and do not copy a version from this page. A fragment filed under any other version fails validation, and tests/test_issue487_changelog_fragments.py runs that validation against the repository, so it fails the test suite and not only the release. After a release, an older fragment forces its pull request to rebase and move the file.
  • <number> should be the pull request number rather than the issue it closes, so two pull requests for the same issue stay distinct. The validator rejects a second category with the same number.
  • <category> is one of added, changed, deprecated, removed, fixed or security.
  • The content is the complete changelog list item: it must start with - , mention the same #<number>, be non-empty UTF-8, and may include paragraphs, tables and fenced code blocks, which assembly copies verbatim. Symlinks are not allowed in changelog.d/.

A fragment for pull request 1234 that fixes a bug would be changelog.d/<latest-release>/1234.fixed.md, starting like this:

text
- Fix the crash when a config names a missing dataset (#1234).

Before a release, a maintainer runs python scripts/assemble_changelog.py. It validates every fragment, inserts the entries under [Unreleased] and consumes the files. Publishing to PyPI independently refuses to continue while any fragment remains.

Code standards

Ruff is the style gate in CI. Its settings in pyproject.toml are a line length of 100, the rule sets E, F, I, N and W, and a Python 3.10 target.

  • Imports are sorted (ruff's I rules).
  • Names: no single-letter variables. Ruff's E741 rejects l, O and I; use entry, part or length.
  • Lazy imports. Import heavy dependencies (torch, transformers, peft, trl, mlx) inside functions, never at module level, so the CLI stays fast. A test guards the result for the training stack: tests/test_cli_startup_is_light.py imports soup_cli.cli in a fresh interpreter and fails if torch, transformers, accelerate, peft, trl, datasets or bitsandbytes came with it.
  • Config validation uses Pydantic v2 with BaseModel and Field.
  • Output goes through rich.console.Console, never a bare print().
  • Type hints on every parameter and return value. CI runs mypy as a non-blocking baseline: it reports issues as a warning and does not fail the job.
  • Path containment uses os.path.realpath plus os.path.commonpath, not Path.resolve() plus relative_to(), which breaks on Windows 8.3 short names. The shared helpers in src/soup_cli/utils/paths.py are is_under, is_under_cwd and enforce_under_cwd_and_no_symlink. A test scans src/soup_cli/ and fails on the banned idiom used as a containment check.
  • No third-party licence headers. The project is Apache-2.0 and ships to PyPI, so an SPDX-License-Identifier or a copyright line naming anyone else is a licensing claim it cannot make. If you generate a file from a template, including with automated tooling, strip its header before you push. tests/test_no_foreign_license_headers.py fails the build on it.

Upstream's own example of the lazy-import and output rules, lightly adapted:

python
# WRONG: a heavy import at module level and a bare print
import transformers

def train():
    print("Starting training")

# CORRECT
from rich.console import Console

def train():
    import transformers

    console = Console()
    console.print("Starting training")

Tests

bash
pytest tests/ -v --tb=short
pytest tests/test_config.py -v
pytest tests/test_data.py::test_detect_alpaca_format -v
pytest tests/ --cov=soup_cli --cov-report=html

The four commands run everything, one file, one test, and a coverage report. The default options in pyproject.toml deselect the smoke marker and measure coverage of soup_cli against a floor of 77 percent. CI's own subset runs pass --no-cov, and some also clear the defaults with -o addopts=; the Recipe Validation workflow runs its two files with --no-cov. The markers are smoke (slow tests that download models and train; run them with pytest -m smoke), unit (fast and isolated: no subprocess, network, filesystem or real model load) and integration (real subprocess, SQLite, filesystem or HTTP).

Conventions that the tree shows:

  • File names. A test file is named after what it covers (tests/test_config.py, tests/test_formats.py); a fix for a numbered issue is often tests/test_issue<number>_<topic>.py.
  • Shared helpers live in tests/conftest.py. Import strip_ansi from it rather than writing another copy. An autouse fixture points SOUP_DB_PATH at a per-test temporary file, so a test never sees runs in your real ~/.soup/experiments.db. tmp_data_dir, sample_alpaca_data and sample_config give you sample data and a config.
  • Strip ANSI before asserting on --help output. Typer renders --help through Rich, which on a colour-capable stream puts escape codes inside the flag name, so "--noise-floor" in result.output fails. Windows passes because Rich turns colour off there, which is how an assertion that passed on a Windows machine has turned CI red before.
  • The Web UI escaping tests run safe_html.js under node and skip when node is absent. CI installs node 22 so they cannot vanish silently.

Tests that guard repository-wide rules

A pull request can break one of these without touching the code a test is named after.

Test fileIt fails when
test_cli_startup_is_light.pyImporting soup_cli.cli loads the training stack.
test_requires_python_bound.pyThe declared requires-python and the CI Python matrix disagree.
test_no_foreign_license_headers.pyA file carries a third-party licence header.
test_issue775_path_containment_ratchet.pyrelative_to or is_relative_to is used as a containment check under src/soup_cli/.
test_issue748_config_fields_reach_a_consumer.pyA field is declared in config/schema.py and nothing under src/ reads it, unless it is allowlisted with a reason. The check matches names globally, so a field named like a common attribute can pass without being wired.
test_cli_help_assertions_are_ansi_safe.pyA test asserts on raw --help output.
test_readme_translation_sync.pyREADME.md changed and a README.<lang>.md translation no longer matches it. Update each translation, then replace its stamp with the line the failure prints.
test_recipe_count_is_synced.pyThe recipe count written in the documentation differs from the catalog.
test_issue621_recipe_snapshot.pyA recipe's resolved config, or a schema default, differs from tests/fixtures/recipe_config_snapshots.json.

What CI runs on your pull request

The CI workflow runs on pushes to main and release/** and on pull requests to main. Its jobs are:

  • lint (Ubuntu, Python 3.11): ruff check src/soup_cli/ scripts/ tests/ benchmarks/, then actionlint over the workflow files. The workflow's own comments call this a required check.
  • type-check: mypy, non-blocking.
  • test: pip install -e ".[dev]" and the whole suite with coverage, on Python 3.10, 3.11 and 3.12 across Ubuntu, Windows and macOS. CONTRIBUTING.md says this must pass.
  • pytorch-smoke (Ubuntu, Python 3.11): the smoke-marked tests against the real PyTorch pipeline, which the job's comment describes as SFT and DPO training.
  • mlx-smoke (macOS 14, Python 3.12): installs only the mlx extra, checks that the PyTorch training stack is absent, and runs one MLX SFT step through the real CLI.
  • transformers-floor (Ubuntu, Python 3.11): installs .[dev] under .github/constraints/transformers-floor.txt, which pins exact versions of torch, transformers, trl, peft and plotext, then runs pip check, asserts that the installed versions match the pins, and runs a compatibility test set.

Which checks are required for a merge is a repository setting, not a file in the tree, so the list above is what runs rather than what blocks. The other workflows are:

  • Recipe Validation runs the two recipe test files on pull requests and pushes that touch recipes, config, the data formats or loader, or those tests.
  • dependency drift runs weekly, on Mondays at 06:00 UTC or on demand. It installs the latest resolvable stack, writes a resolved-versus-declared table into the run summary and runs the suite against it.
  • recipe repo ids runs weekly, on Mondays at 07:00 UTC or on demand, and resolves every recipe's model id on the Hub anonymously. It is deliberately not a pull request check.
  • Publish to PyPI runs on a v* tag, and Build and Publish Docker Image runs when a GitHub release is published.

How credit works

Upstream credits every contributor in three places:

  • CONTRIBUTORS.md. Every merged pull request, or one adopted in-tree, adds you there with your work and the pull request number. The file is maintained by hand and listed by first contribution. If you merged a pull request and are not in it, open a pull request adding your line; that counts too.
  • CHANGELOG.md. From the release that ships your change, the entry is tagged (#NNN by @you).
  • Commit trailers. Pull requests land with a Co-Authored-By: line for you, so GitHub's contributors graph reflects your work even after a squash merge.

If you contributed and are not credited somewhere, upstream calls that a bug: open a pull request or an issue. A security reporter is credited by name in the release notes unless they ask for anonymity.

Report a vulnerability

Do not open a public issue for anything security-sensitive, and do not report it in Discord. Report it privately through GitHub Security Advisories, which is the preferred route, or, if you cannot use GitHub, by email to team@trysoup.dev. The supported versions, the scope, what to expect and the fact that there is no bug bounty are in Reporting a vulnerability.

Security expectations for code

None of CONTRIBUTING.md, AGENTS.md or SECURITY.md contains a separate security checklist for pull requests. What they state is:

  • Path containment follows the rule in Code standards, through the helpers in utils/paths.py.
  • The classes of bug SECURITY.md treats as in scope are the ones to keep in mind when code reads or writes anything a user names: path traversal or arbitrary file read or write from config, dataset or artifact paths; SSRF in the synthetic-data providers, the inference server or the hub and endpoint validators; command, Modelfile, Jinja chat-template or systemd and launchd unit injection; secret leakage in logs, crash bundles or generated artifacts; and sandbox escape in the code-execution reward path.
  • The paths CODEOWNERS marks as security-sensitive are src/soup_cli/ui/, src/soup_cli/data/providers/ and src/soup_cli/utils/ollama.py.

How to add a recipe

  1. Add a RecipeMeta entry to the RECIPES dict in src/soup_cli/recipes/catalog.py. Its fields are model, task, size, tags, description and yaml_str, and the dict key is the recipe's name. yaml_str is a complete soup.yaml that must load as a SoupConfig. In the existing recipes the model id and the YAML base: both name a Hugging Face repository, and a weekly workflow resolves every id.
  2. Add tests in tests/test_recipes.py.
  3. Regenerate the config snapshot with python scripts/generate_recipe_snapshot.py and review the diff. The new recipe's resolved config has to be in the committed snapshot, or test_fixture_and_catalog_have_the_same_recipe_names fails.
  4. Move the recipe count everywhere it is written. tests/test_recipe_count_is_synced.py checks five places: a comment in catalog.py, CONTRIBUTING.md, docs/commands.md and two lines in docs/serving-and-export.md. Several tests also pin the exact catalog size, in tests/test_recipes.py and in three older release-milestone test files. Pull request 864, which added two recipes, touched all of them.
  5. Update the recipes section of README.md, and add a changelog fragment in the added category.

You can check the result with soup recipes show <name>, which prints the recipe's YAML, and soup recipes list. A new recipe is one of the first contributions upstream suggests; see Recipes for what the catalog holds.

How to add a command

  1. Create src/soup_cli/commands/your_command.py with a handler function. A command group is a typer.Typer() app in the module.
  2. Register it in src/soup_cli/cli.py. Import the module in cli.py (the oldest commands share one from soup_cli.commands import (...) block; newer ones are imported one at a time, often under an alias such as as _ship_cmd), then register it the way its neighbours are: app.command()(your_command.your_command) for one command (pass name= to set the CLI name), or app.add_typer(your_command.app, name="your-group", help="...") for a group. Upstream's guide says @app.command(); the file registers by calling app.command()(function).
  3. Import heavy dependencies inside the handler, so soup --help stays fast.
  4. Add tests in tests/test_your_command.py, using strip_ansi for any --help assertion.
  5. Update the help text and README.md, and add the command to docs/commands.md, the full command list. Add a changelog fragment.

How to add a training task

  1. Write src/soup_cli/trainer/your_trainer.py, a wrapper class around the appropriate trainer. commands/train.py builds it with config, device, report_to, deepspeed_config, fsdp_config and trust_remote_code, then calls setup(dataset) and train(...). trainer/orpo.py shows the shape.
  2. Add the task name and its config fields to src/soup_cli/config/schema.py: the name goes in the task literal on SoupConfig. Every new field needs a consumer, or the repository-wide test above fails, and since v0.75.0 a key that is not in the schema is refused when a config loads.
  3. Add a template: a YAML file under src/soup_cli/templates/ and its entry in manifest.json.
  4. Route it in src/soup_cli/commands/train.py with an elif cfg.task == "your_task": branch that imports the wrapper inside the branch. Runs on the MLX backend go through trainer/mlx_routing.py first.
  5. Add tests in tests/test_your_trainer.py. Upstream's guide asks for at least thirty for a new trainer.
  6. Update README.md, the relevant page under docs/, and add a changelog fragment.

For a new data format, the steps are shorter: add detection and conversion in src/soup_cli/data/formats.py, add the name to the format literal in config/schema.py, update data/loader.py if needed, add tests in tests/test_formats.py, and document it in the relevant docs/ page. See Data formats for the formats that exist.

What happens after your pull request merges

Release steps are the maintainer's side of the process, and upstream's guide lists them under "Releases". The project uses semantic versioning, MAJOR.MINOR.PATCH.

  1. Refresh the dev stack with pip install -U -e ".[dev]", or read the last dependency drift run.
  2. Run python scripts/assemble_changelog.py and review the assembled CHANGELOG.md.
  3. Update the version in pyproject.toml and src/soup_cli/__init__.py.
  4. Run the full test suite and the linter, including the recipe snapshot test.
  5. Update README.md, the relevant page under docs/, and SECURITY.md if the release is security-related.
  6. Commit as Release v0.X.0 and tag v0.X.0.
  7. GitHub Actions verifies that no fragment was left behind, then publishes to PyPI on the v* tag through a Trusted Publisher (OIDC).

See also

Soup is free and Apache-2.0. If it saved you a training run, starring the repo costs nothing and helps most. You can also fund the GPU time behind the work a 4 GB laptop cannot reach.