Recipe catalog
Soup ships 167 ready-made configs at v0.75.2, and this page lists every one of them. The Recipes page explains how to use them; this one is the catalog, grouped by training task.
soup recipes list # all of them
soup recipes search qwen --task grpo # keyword, --task and --size filters
soup recipes show llama3.1-8b-sft # print the YAML
soup recipes use llama3.1-8b-sft -o soup.yaml # write it, then: soup train --config soup.yamlWhat a recipe is, and is not
A recipe is a complete soup.yaml for one base model and one task. All 167 of them load through the v0.75.2 config schema with no unknown keys: that was checked by loading every one. That says the file is valid. It does not say the run works end to end, and it does not say it was trained to a quality target: the data paths are placeholders you replace, a few recipes stop at trainer setup or carry a key the trainer ignores at v0.75.2 (the list below names them), and this page loaded every recipe through the schema without training any of them. Treat every hyperparameter as a starting point, and read the base model's license before you train on it.
The Size column is the base model's parameter size as the catalog states it, not a memory figure, and it is not reliable for the largest and mixture-of-experts entries: a few look like placeholders (deepseek-v3-7b-sft lists 7B for DeepSeek-V3-0324, which deepseek-v3-pipeline below lists at 671B, and glm-4.6-sft, glm-5-sft and minimax-m2-sft list 9B). N/A means the catalog gives no size. Check the model card before you size a GPU. Quantization is the training.quantization the recipe sets (4bit is 4-bit NF4 loading, the QLoRA setup; none means no quantization; 8bit is 8-bit bitsandbytes loading; not set leaves the schema default, which is 4bit; bitnet_1.58 appears once and does not train, see below). Recipes that run on Apple Silicon say mlx in the notes.
Before you run one
Loading as a config and running are different things. These are the recipes where the YAML is valid and the run still stops, needs something you have to supply first, or carries a key the trainer does not read at v0.75.2. Each was checked against the v0.75.2 source; none was trained to confirm it.
- Sharding is a launch flag, not a recipe key. Where a description names DeepSpeed ZeRO-2 or ZeRO-3 or FSDP2 (for example
llama3.1-70b-sft,qwen2.5-72b-sft,gemma3-27b-sft,qwen3-32b-sft,deepseek-r1-32b-grpoandllama3-70b-fsdp2) or says multi-node, the YAML does not select it. Pass--deepspeed zero2,--deepspeed zero3or--fsdp full_shard, plus--gpus N, tosoup train, as the multi-GPU page shows.llama3-70b-fsdp2setsuse_fsdp2_compile: true, andsoup train --config soup.yamlwithout--fsdpexits 1. - Some recipes need something first.
llama3.1-8b-ppoloadstraining.reward_model: ./output_rm, a reward model you train beforehand.llama3.1-8b-longctxsetsuse_flash_attn: true, which exits 1 without CUDA and theflash-attnpackage.baichuan-sftsetstraining.hub: modelscope, so the base model is fetched through ModelScope into./.soup_hub_cache. - Keys the trainer ignores. Eight MoE DPO and GRPO recipes (
qwen3.5-35b-a3b-dpo,deepseek-v4-flash-dpo,glm-5.1-dpo,kimi-k2.6-dpo,qwen3-30b-a3b-reasoning-grpo,deepseek-v4-flash-grpo,glm-5.1-grpoandkimi-k2.6-grpo) setmoe_lora: true, and five of them setmoe_aux_loss_coeff: 0.01. Only the SFT and pretrain trainers read those two keys, so on these tasks they are accepted and ignored.deepseek-v3-reasoning,kimi-k2-thinking-grpoandqwq-32b-grposetverifiable_domain: mathbeside areward_fnthat is notverifiable, andverifiable_domainis read only by theverifiablereward, so it does nothing there. - Multimodal, speech and BitNet caveats. The DPO and GRPO trainers are text-only, so
pixtral-dpoandllama3.2-vision-grpodo not use images. Of the five TTS recipes onlyorpheus-tts-sfthas a live codec: the other four setdata.format: audioand raise.falcon-e-bitnet-sftdoes not train, because the quantization loader rejectsbitnet_1.58.whisper-large-v3-ftgoes through the generic audio SFT path; for Whisper speech-to-text usewhisper-large-v3-asr.deepseek-v3-pipelinevalidates its pipeline setting and then runs data-parallel.
SFT: supervised fine-tuning
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
baichuan-sft | baichuan-inc/Baichuan2-13B-Chat | 13B | 4bit | Baichuan 2 13B chat SFT |
biomistral-7b-sft | BioMistral/BioMistral-7B | 7B | 4bit | BioMistral 7B medical/biomedical domain SFT |
codellama-13b-sft | codellama/CodeLlama-13b-Instruct-hf | 13B | 4bit | Code Llama 13B code-specialist SFT |
codellama-70b-sft | codellama/CodeLlama-70b-Instruct-hf | 70B | 4bit | Code Llama 70B code-specialist SFT (multi-GPU) |
cogito-v2-sft | deepcogito/cogito-v2-preview | 14B | 4bit | Cogito v2 preview SFT |
deepseek-ocr-sft | deepseek-ai/DeepSeek-OCR | N/A | 4bit | DeepSeek-OCR vision OCR SFT |
deepseek-r1-distill-llama-8b-sft | deepseek-ai/DeepSeek-R1-Distill-Llama-8B | 8B | 4bit | DeepSeek-R1-Distill Llama 8B reasoning SFT |
deepseek-r1-distill-qwen-1.5b-sft | deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B | 1.5B | 4bit | DeepSeek-R1-Distill Qwen 1.5B reasoning SFT |
deepseek-r1-distill-qwen-7b-sft | deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | 7B | 4bit | DeepSeek-R1-Distill Qwen 7B reasoning SFT |
deepseek-v3-7b-sft | deepseek-ai/DeepSeek-V3-0324 | 7B | 4bit | DeepSeek V3 SFT with MoE LoRA |
deepseek-v3-pipeline | deepseek-ai/DeepSeek-V3 | 671B | 4bit | DeepSeek V3 SFT scaffold. Its parallelism: pipeline (4 stages) is checked (it needs CUDA and at least 4 GPUs) and then the run is data-parallel at v0.75.2, see Pipeline parallelism |
deepseek-v4-flash-sft | deepseek-ai/DeepSeek-V4-Flash | N/A | 4bit | DeepSeek V4 Flash MoE SFT (MIT, efficiency-tier) |
deepseek-v4-pro-sft | deepseek-ai/DeepSeek-V4-Pro | N/A | 4bit | DeepSeek V4 Pro flagship MoE SFT (MIT, 1.6T-class). Requires multi-node DeepSpeed. |
devstral-sft | mistralai/Devstral-Small | 24B | 4bit | Devstral Small code/agent SFT |
falcon-e-bitnet-sft | tiiuae/Falcon-E-1B-Instruct | 1B | bitnet_1.58 | Falcon-E BitNet 1.58-bit SFT. Does not train at v0.75.2: the config validates, then the quantization loader rejects bitnet_1.58, see BitNet |
gemma2-2b-sft | google/gemma-2-2b-it | 2B | 4bit | Gemma 2 2B SFT (edge-friendly) |
gemma3-12b-sft | google/gemma-3-12b-it | 12B | 4bit | Gemma 3 12B instruction tuning |
gemma3-27b-sft | google/gemma-3-27b-it | 27B | 4bit | Gemma 3 27B SFT with DeepSpeed ZeRO-2 |
gemma3-4b-sft-mlx | mlx-community/gemma-3-4b-it-4bit | 4B | 4bit | Gemma 3 4B SFT on Apple Silicon via MLX (M1+ 16GB) (mlx) |
gemma3-9b-sft | google/gemma-3-9b-it | 9B | 4bit | Gemma 3 9B instruction tuning |
glm-4.6-sft | THUDM/glm-4.6 | 9B | 4bit | GLM 4.6 instruction tuning with LoRA |
glm-5-sft | zai-org/GLM-5 | 9B | 4bit | GLM 5 SFT (next-gen GLM family) |
glm-5.1-sft | zai-org/GLM-5.1 | 754B | 4bit | GLM 5.1 MoE SFT (MIT, 754B). Multi-GPU / multi-node recommended. |
gpt-oss-120b-sft | openai/gpt-oss-120b | 120B | 4bit | GPT-OSS 120B SFT (multi-GPU recommended) |
gpt-oss-20b-sft | openai/gpt-oss-20b | 20B | 4bit | GPT-OSS 20B SFT (reasoning_effort=medium) |
granite-4-sft | ibm-granite/granite-4.0-tiny-base | 3B | 4bit | IBM Granite 4.0 tiny SFT |
internvl-2.5-8b-sft | OpenGVLab/InternVL2_5-8B | 8B | 4bit | InternVL 2.5 8B vision-language SFT |
internvl-3-5-sft | OpenGVLab/InternVL3-5 | 8B | 4bit | InternVL 3.5 vision SFT |
kimi-k2-sft | moonshotai/Kimi-K2 | N/A | 4bit | Kimi K2 SFT (Moonshot MoE, long-context-aware) |
kimi-k2.5-sft | moonshotai/Kimi-K2.5 | 1T | 4bit | Kimi K2.5 MoE SFT (Modified MIT, ~1T / 32B active). Requires multi-node DeepSpeed. |
kimi-k2.6-sft | moonshotai/Kimi-K2.6 | 1T | 4bit | Kimi K2.6 MoE SFT (Modified MIT, ~1T / 32B active). Requires multi-node DeepSpeed. |
lfm2-sft | LiquidAI/LFM2-1.2B | 1.2B | 4bit | Liquid LFM2 1.2B SFT (edge-optimised) |
llama2-13b-finance-sft | meta-llama/Llama-2-13b-hf | 13B | 4bit | Llama 2 13B finance-domain SFT (FinGPT-style starter recipe) |
llama3-70b-fsdp2 | meta-llama/Llama-3.1-70B-Instruct | 70B | 4bit | Llama 3.1 70B SFT with FSDP2 full shard + torch.compile. Requires 8 x A100/H100 80GB. |
llama3.1-70b-sft | meta-llama/Llama-3.1-70B-Instruct | 70B | 4bit | Llama 3.1 70B SFT with DeepSpeed ZeRO-3 |
llama3.1-8b-longctx | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | Llama 3.1 8B long-context (32k) with YaRN RoPE scaling |
llama3.1-8b-sft | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | Llama 3.1 8B instruction tuning with LoRA |
llama3.1-8b-sft-mlx | mlx-community/Llama-3.1-8B-Instruct-4bit | 8B | 4bit | Llama 3.1 8B SFT on Apple Silicon via MLX (M2+ 16GB) (mlx) |
llama3.2-11b-vision | meta-llama/Llama-3.2-11B-Vision-Instruct | 11B | 4bit | Llama 3.2 11B Vision multimodal fine-tuning |
llama3.2-1b-sft | meta-llama/Llama-3.2-1B-Instruct | 1B | 8bit | Llama 3.2 1B instruction tuning (mobile-friendly) |
llama3.2-3b-sft | meta-llama/Llama-3.2-3B-Instruct | 3B | 4bit | Llama 3.2 3B instruction tuning (edge-friendly) |
llama3.2-vision-90b-sft | meta-llama/Llama-3.2-90B-Vision-Instruct | 90B | 4bit | Llama 3.2 90B Vision multimodal SFT (8 x A100/H100 80GB) |
llama4-scout-17b-sft | meta-llama/Llama-4-Scout-17B-16E-Instruct | 17B | 4bit | Llama 4 Scout 17B SFT with LoRA (4bit) |
llama4-scout-tools | meta-llama/Llama-4-Scout-17B-16E-Instruct | 17B | 4bit | Llama 4 Scout 17B tool-calling / function-calling SFT |
llava-next-sft | llava-hf/llava-v1.6-mistral-7b-hf | 7B | 4bit | LLaVA-Next 7B vision SFT |
magicoder-7b-sft | ise-uiuc/Magicoder-S-DS-6.7B | 6.7B | 4bit | Magicoder S-DS 6.7B code-specialist SFT |
magistral-small-sft | mistralai/Magistral-Small | 24B | 4bit | Magistral Small reasoning SFT |
mathstral-7b-sft | mistralai/Mathstral-7B-v0.1 | 7B | 4bit | Mathstral 7B math/STEM-specialist SFT |
medgemma-sft | google/medgemma-4b-it | 4B | 4bit | MedGemma 4B medical SFT |
meditron-7b-sft | epfl-llm/meditron-7b | 7B | 4bit | Meditron 7B medical / clinical domain SFT |
minicpm-v-2.6-sft | openbmb/MiniCPM-V-2_6 | 8B | 4bit | MiniCPM-V 2.6 vision-language SFT (edge-friendly multimodal) |
minimax-m2-sft | MiniMaxAI/MiniMax-M2 | 9B | 4bit | MiniMax M2 SFT instruction tuning |
minimax-m3-sft | MiniMaxAI/MiniMax-M3 | 428B | 4bit | MiniMax M3 MoE SFT (428B / 23B active). MiniMax Community License - commercial use requires a separate agreement. Multi-GPU recommended. |
ministral-sft | mistralai/Ministral-8B-Instruct-2410 | 8B | 4bit | Ministral 8B SFT |
mistral-7b-sft | mistralai/Mistral-7B-Instruct-v0.3 | 7B | 4bit | Mistral 7B instruction tuning |
mistral-large-3-sft | mistralai/Mistral-Large-3-675B-Instruct-2512 | 675B | 4bit | Mistral Large 3 MoE SFT (Apache-2.0, 675B / 41B active, multimodal). Requires multi-node DeepSpeed. |
mistral-medium-3-5-sft | mistralai/Mistral-Medium-3.5 | N/A | 4bit | Mistral Medium 3.5 SFT |
mistral-small-3-sft | mistralai/Mistral-Small-24B-Instruct-2501 | 24B | 4bit | Mistral Small 3 24B SFT |
nemotron-4-340b-sft | nvidia/Nemotron-4-340B-Instruct | 340B | 4bit | Nemotron-4 340B SFT (massive multi-node deployment) |
paddle-ocr-sft | PaddlePaddle/PaddleOCR-VL | N/A | 4bit | Paddle-OCR-VL OCR SFT |
phi3.5-mini-sft | microsoft/Phi-3.5-mini-instruct | 3.8B | 4bit | Phi-3.5-mini 3.8B SFT (small / edge) |
phi4-14b-sft | microsoft/phi-4 | 14B | 4bit | Phi-4 14B instruction tuning |
pixtral-12b-sft | mistralai/Pixtral-12B-2409 | 12B | 4bit | Pixtral 12B vision-language SFT with LoRA |
qvq-72b-sft | Qwen/QVQ-72B-Preview | 72B | 4bit | QVQ 72B vision-reasoning SFT |
qwen-image-sft | Qwen/Qwen-Image | N/A | 4bit | Qwen-Image image-output multimodal SFT |
qwen2-audio-7b-sft | Qwen/Qwen2-Audio-7B-Instruct | 7B | 4bit | Qwen2-Audio 7B audio-language SFT |
qwen2-vl-72b-sft | Qwen/Qwen2-VL-72B-Instruct | 72B | 4bit | Qwen2-VL 72B vision-language SFT (multi-GPU recommended) |
qwen2-vl-7b-sft | Qwen/Qwen2-VL-7B-Instruct | 7B | 4bit | Qwen2-VL 7B vision-language SFT |
qwen2.5-0.5b-sft | Qwen/Qwen2.5-0.5B-Instruct | 0.5B | 8bit | Qwen 2.5 0.5B SFT (mobile / edge) |
qwen2.5-1.5b-sft | Qwen/Qwen2.5-1.5B-Instruct | 1.5B | 8bit | Qwen 2.5 1.5B SFT (edge-friendly) |
qwen2.5-3b-sft | Qwen/Qwen2.5-3B-Instruct | 3B | 4bit | Qwen 2.5 3B SFT (small / edge) |
qwen2.5-72b-sft | Qwen/Qwen2.5-72B-Instruct | 72B | 4bit | Qwen 2.5 72B SFT with DeepSpeed ZeRO-3 |
qwen2.5-7b-sft | Qwen/Qwen2.5-7B-Instruct | 7B | 4bit | Qwen 2.5 7B instruction tuning with LoRA |
qwen2.5-coder-1.5b-sft | Qwen/Qwen2.5-Coder-1.5B-Instruct | 1.5B | 4bit | Qwen 2.5 Coder 1.5B instruction tuning with LoRA |
qwen2.5-coder-14b-sft | Qwen/Qwen2.5-Coder-14B-Instruct | 14B | 4bit | Qwen 2.5 Coder 14B instruction tuning with LoRA |
qwen2.5-coder-32b-sft | Qwen/Qwen2.5-Coder-32B-Instruct | 32B | 4bit | Qwen 2.5 Coder 32B instruction tuning with LoRA |
qwen2.5-coder-7b-sft | Qwen/Qwen2.5-Coder-7B-Instruct | 7B | 4bit | Qwen 2.5 Coder 7B instruction tuning with LoRA |
qwen2.5-math-1.5b-sft | Qwen/Qwen2.5-Math-1.5B-Instruct | 1.5B | 4bit | Qwen 2.5 Math 1.5B instruction tuning with LoRA |
qwen2.5-math-7b-sft | Qwen/Qwen2.5-Math-7B-Instruct | 7B | 4bit | Qwen 2.5 Math 7B instruction tuning with LoRA |
qwen3-14b-sft | Qwen/Qwen3-14B | 14B | 4bit | Qwen 3 14B instruction tuning with LoRA |
qwen3-30b-a3b-sft | Qwen/Qwen3-30B-A3B | 30B | 4bit | Qwen 3 30B-A3B MoE instruction tuning |
qwen3-32b-sft | Qwen/Qwen3-32B | 32B | 4bit | Qwen 3 32B SFT with DeepSpeed ZeRO-2 |
qwen3-32b-zeropp | Qwen/Qwen3-32B | 32B | 4bit | Qwen3 32B SFT with DeepSpeed ZeRO++ (quantized gradients + hierarchical partitioning). Launch with --deepspeed zero++. |
qwen3-8b-sft | Qwen/Qwen3-8B | 8B | 4bit | Qwen 3 8B instruction tuning |
qwen3-8b-sft-mlx | mlx-community/Qwen3-8B-4bit | 8B | 4bit | Qwen 3 8B SFT on Apple Silicon via MLX (M2+ 16GB) (mlx) |
qwen3-8b-tools | Qwen/Qwen3-8B | 8B | 4bit | Qwen 3 8B tool-calling / function-calling SFT |
qwen3-coder-30b-sft | Qwen/Qwen3-Coder-30B-A3B-Instruct | 30B | 4bit | Qwen3-Coder 30B (A3B MoE) code-specialist SFT |
qwen3.5-0.8b-sft | Qwen/Qwen3.5-0.8B | 0.8B | 4bit | Qwen 3.5 0.8B SFT (Apache-2.0, tiny / mobile) |
qwen3.5-122b-a10b-sft | Qwen/Qwen3.5-122B-A10B | 122B | 4bit | Qwen 3.5 122B-A10B MoE SFT (Apache-2.0, 10B active). Multi-GPU recommended (8 x A100/H100 80GB). |
qwen3.5-27b-sft | Qwen/Qwen3.5-27B | 27B | 4bit | Qwen 3.5 27B SFT (Apache-2.0) with DeepSpeed ZeRO-2 |
qwen3.5-2b-sft | Qwen/Qwen3.5-2B | 2B | 4bit | Qwen 3.5 2B SFT (Apache-2.0, small / edge) |
qwen3.5-35b-a3b-sft | Qwen/Qwen3.5-35B-A3B | 35B | 4bit | Qwen 3.5 35B-A3B MoE SFT (Apache-2.0, 3B active) |
qwen3.5-397b-a17b-sft | Qwen/Qwen3.5-397B-A17B | 397B | 4bit | Qwen 3.5 397B-A17B flagship MoE SFT (Apache-2.0, 17B active). Requires multi-node DeepSpeed / FSDP. |
qwen3.5-4b-sft | Qwen/Qwen3.5-4B | 4B | 4bit | Qwen 3.5 4B SFT (Apache-2.0, 262K context) |
qwen3.5-9b-sft | Qwen/Qwen3.5-9B | 9B | 4bit | Qwen 3.5 9B SFT (Apache-2.0, 262K context) |
qwen3.6-27b-sft | Qwen/Qwen3.6-27B | 27B | 4bit | Qwen 3.6 27B SFT (Apache-2.0) with DeepSpeed ZeRO-2 |
qwen3.6-35b-a3b-sft | Qwen/Qwen3.6-35B-A3B | 35B | 4bit | Qwen 3.6 35B-A3B MoE SFT (Apache-2.0, 3B active) |
qwen3.8-27b-sft | Qwen/Qwen3.8-27B | 27B | 4bit | Qwen 3.8 27B text-only SFT (Apache-2.0) with DeepSpeed ZeRO-2 |
ra-dit-llama3-8b | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | RA-DIT stage 2 (Meta 2023) — RAFT-style SFT on the generator. Pairs with ra-dit-retriever (stage 1). Uses RAFT data format. |
raft-llama3-8b | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | RAFT (Retrieval-Augmented Fine-Tuning, Stanford 2024) — train an 8B Llama 3.1 to answer queries given a golden doc + distractor docs. Rows: {query, golden_doc, distractor_docs, answer}. |
seamlessm4t-v2-sft | facebook/seamless-m4t-v2-large | 2.3B | 4bit | SeamlessM4T v2 multilingual speech-to-text SFT |
smollm2-1.7b-sft | HuggingFaceTB/SmolLM2-1.7B-Instruct | 1.7B | 8bit | SmolLM2 1.7B SFT (small / edge) |
smollm2-135m-sft | HuggingFaceTB/SmolLM2-135M-Instruct | 135M | none | SmolLM2 135M SFT (ultra-tiny / mobile) |
smollm2-360m-sft | HuggingFaceTB/SmolLM2-360M-Instruct | 360M | none | SmolLM2 360M SFT (tiny / mobile) |
smollm3-3b-sft | HuggingFaceTB/SmolLM3-3B | 3B | 8bit | SmolLM3 3B SFT (small / edge) |
smolvlm-256m-sft | HuggingFaceTB/SmolVLM-256M-Instruct | 256M | none | SmolVLM 256M vision SFT (llava format) — a tiny VLM. SmolVLM uses an Idefics3 processor; Soup mirrors its nested tokenizer surface and performs processor-aware vision collation so image-token expansion and pixel_values reach the model. target_modules are pinned to q_proj/v_proj (auto cannot infer them for Idefics3). |
voxtral-sft | mistralai/Voxtral-Mini-3B | 3B | 4bit | Voxtral Mini 3B audio SFT |
whisper-large-v3-ft | openai/whisper-large-v3 | 1.5B | 8bit | Whisper Large v3 through the generic audio SFT path (modality: audio), which loads the model with AutoModel, the bare Whisper encoder-decoder without a language-model head. For Whisper speech-to-text use the ASR task, see ASR fine-tuning and whisper-large-v3-asr below |
DPO: preference optimization
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
deepseek-r1-distill-llama-8b-dpo | deepseek-ai/DeepSeek-R1-Distill-Llama-8B | 8B | 4bit | DeepSeek-R1-Distill Llama 8B DPO alignment |
deepseek-r1-distill-qwen-1.5b-dpo | deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B | 1.5B | 4bit | DeepSeek-R1-Distill Qwen 1.5B DPO alignment |
deepseek-r1-distill-qwen-7b-dpo | deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | 7B | 4bit | DeepSeek-R1-Distill Qwen 7B DPO alignment |
deepseek-v4-flash-dpo | deepseek-ai/DeepSeek-V4-Flash | N/A | 4bit | DeepSeek V4 Flash MoE DPO alignment (MIT, efficiency-tier) |
gemma3-27b-dpo | google/gemma-3-27b-it | 27B | 4bit | Gemma 3 27B DPO alignment |
glm-5.1-dpo | zai-org/GLM-5.1 | 754B | 4bit | GLM 5.1 MoE DPO alignment (MIT, 754B). Multi-GPU / multi-node recommended. |
kimi-k2.6-dpo | moonshotai/Kimi-K2.6 | 1T | 4bit | Kimi K2.6 MoE DPO alignment (Modified MIT, ~1T / 32B active). Requires multi-node DeepSpeed. |
llama3.1-8b-dpo | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | Llama 3.1 8B DPO alignment |
llama4-scout-17b-dpo | meta-llama/Llama-4-Scout-17B-16E-Instruct | 17B | 4bit | Llama 4 Scout 17B DPO alignment |
mistral-7b-dpo | mistralai/Mistral-7B-Instruct-v0.3 | 7B | 4bit | Mistral 7B DPO alignment |
pixtral-dpo | mistralai/Pixtral-12B-2409 | 12B | 4bit | Pixtral 12B DPO. Not usable as shipped at v0.75.2: the DPO trainer is text-only (no vision processor), and data.format: llava produces messages and image columns without the chosen and rejected that DPO reads |
qwen2.5-7b-dpo | Qwen/Qwen2.5-7B-Instruct | 7B | 4bit | Qwen 2.5 7B DPO alignment |
qwen3.5-35b-a3b-dpo | Qwen/Qwen3.5-35B-A3B | 35B | 4bit | Qwen 3.5 35B-A3B MoE DPO alignment (Apache-2.0, 3B active) |
qwen3.5-9b-dpo | Qwen/Qwen3.5-9B | 9B | 4bit | Qwen 3.5 9B DPO alignment (Apache-2.0) |
GRPO: reinforcement learning with a verifier
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
deepseek-r1-32b-grpo | deepseek-ai/DeepSeek-R1-Distill-Qwen-32B | 32B | 4bit | DeepSeek R1 32B GRPO with DeepSpeed |
deepseek-r1-8b-grpo | deepseek-ai/DeepSeek-R1-Distill-Llama-8B | 8B | 4bit | DeepSeek R1 8B GRPO reasoning |
deepseek-v3-reasoning | deepseek-ai/DeepSeek-V3 | N/A | 4bit | DeepSeek V3 (671B-parameter MoE) GRPO reasoning recipe. Requires multi-node DeepSpeed. Scores with the accuracy and format rewards; its verifiable_domain: math is only read by reward_fn: verifiable and has no effect here |
deepseek-v4-flash-grpo | deepseek-ai/DeepSeek-V4-Flash | N/A | 4bit | DeepSeek V4 Flash MoE GRPO reasoning training (MIT, efficiency-tier) |
glm-5.1-grpo | zai-org/GLM-5.1 | 754B | 4bit | GLM 5.1 MoE GRPO reasoning training (MIT, 754B). Multi-GPU / multi-node recommended. |
grpo-env-calculator | HuggingFaceTB/SmolLM2-135M-Instruct | 135M | not set | SmolLM2-135M GRPO on the bundled calculator env (openenv rollout) |
grpo-env-guess-number | HuggingFaceTB/SmolLM2-135M-Instruct | 135M | not set | SmolLM2-135M GRPO on the bundled number-deduction env (openenv rollout) |
grpo-env-retrieval-qa | HuggingFaceTB/SmolLM2-135M-Instruct | 135M | not set | SmolLM2-135M GRPO on the bundled retrieval-QA env (openenv rollout) |
kimi-k2-thinking-grpo | moonshotai/Kimi-K2-Thinking | N/A | 4bit | Kimi K2 Thinking GRPO reasoning |
kimi-k2.6-grpo | moonshotai/Kimi-K2.6 | 1T | 4bit | Kimi K2.6 MoE GRPO reasoning (Modified MIT, ~1T / 32B active). Requires multi-node DeepSpeed. |
llama3.1-8b-grpo | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | Llama 3.1 8B GRPO reasoning training |
llama3.2-vision-grpo | meta-llama/Llama-3.2-11B-Vision-Instruct | 11B | 4bit | Llama 3.2 11B Vision GRPO. The GRPO trainer is text-only at v0.75.2, so the images are not used, see GRPO Plus |
llama4-scout-17b-grpo | meta-llama/Llama-4-Scout-17B-16E-Instruct | 17B | 4bit | Llama 4 Scout 17B GRPO reasoning training |
phi4-reasoning-grpo | microsoft/phi-4 | 14B | 4bit | Phi-4 14B GRPO reasoning training |
qwen2.5-7b-grpo | Qwen/Qwen2.5-7B-Instruct | 7B | 4bit | Qwen 2.5 7B GRPO reasoning training |
qwen3-30b-a3b-reasoning-grpo | Qwen/Qwen3-30B-A3B | 30B | 4bit | Qwen3 30B-A3B GRPO reasoning training (MoE thinking model) |
qwen3-8b-grpo | Qwen/Qwen3-8B | 8B | 4bit | Qwen 3 8B GRPO reasoning training |
qwen3.5-0.8b-grpo | Qwen/Qwen3.5-0.8B | 0.8B | 4bit | Qwen 3.5 0.8B GRPO reasoning training (Apache-2.0, tiny / mobile) |
qwen3.5-2b-grpo | Qwen/Qwen3.5-2B | 2B | 4bit | Qwen 3.5 2B GRPO reasoning training (Apache-2.0, small / edge) |
qwen3.5-9b-grpo | Qwen/Qwen3.5-9B | 9B | 4bit | Qwen 3.5 9B GRPO reasoning training (Apache-2.0, 262K context) |
qwq-32b-grpo | Qwen/QwQ-32B | 32B | 4bit | QwQ 32B GRPO reasoning training |
r1-distill-llama-70b-grpo | deepseek-ai/DeepSeek-R1-Distill-Llama-70B | 70B | 4bit | DeepSeek-R1-Distill Llama 70B GRPO reasoning (multi-GPU) |
r1-distill-qwen-1.5b-grpo | deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B | 1.5B | 4bit | DeepSeek-R1-Distill Qwen 1.5B GRPO reasoning training |
r1-distill-qwen-14b-grpo | deepseek-ai/DeepSeek-R1-Distill-Qwen-14B | 14B | 4bit | DeepSeek-R1-Distill Qwen 14B GRPO reasoning training |
r1-distill-qwen-7b-grpo | deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | 7B | 4bit | DeepSeek-R1-Distill Qwen 7B GRPO reasoning training |
KTO
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
llama3.1-8b-kto | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | Llama 3.1 8B KTO unpaired preference alignment |
ORPO
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
llama3.1-8b-orpo | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | Llama 3.1 8B ORPO reference-free alignment |
SimPO
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
llama3.1-8b-simpo | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | Llama 3.1 8B SimPO length-normalized alignment |
IPO
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
llama3.1-8b-ipo | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | Llama 3.1 8B IPO regularized preference alignment |
PPO
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
llama3.1-8b-ppo | meta-llama/Llama-3.1-8B-Instruct | 8B | 4bit | Llama 3.1 8B PPO (RLHF stage 3) |
Online DPO
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
online-dpo-smollm2-135m | HuggingFaceTB/SmolLM2-135M-Instruct | 135M | none | SmolLM2 135M Online DPO — on-policy generation judged by a pairwise LLM judge (the recipe sets training.online_dpo_judge: ollama://llama3.1, which needs a local Ollama server; point it at your own judge) |
Reward model
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
qwen2.5-7b-reward | Qwen/Qwen2.5-7B-Instruct | 7B | 4bit | Qwen 2.5 7B reward model (RLHF stage 2) |
Embedding
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
embedding-gemma-sft | google/embeddinggemma-300m | 300M | 4bit | EmbeddingGemma 300M sentence-embedding SFT |
llama3.1-8b-embed | meta-llama/Llama-3.1-8B | 8B | 4bit | Llama 3.1 8B sentence embedding with cosine loss |
ra-dit-retriever | sentence-transformers/all-mpnet-base-v2 | N/A | not set | RA-DIT stage 1 (Meta 2023) — contrastive retriever training. Pairs with ra-dit-llama3-8b (stage 2) for the full pipeline. |
Continued pre-training
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
qwen2.5-7b-pretrain | Qwen/Qwen2.5-7B | 7B | 4bit | Qwen 2.5 7B continued pre-training |
qwen3.5-4b-pretrain | Qwen/Qwen3.5-4B-Base | 4B | 4bit | Qwen 3.5 4B Base continued pre-training (Apache-2.0) |
ASR: speech recognition
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
whisper-base-asr | openai/whisper-base | 74M | none | Whisper base (74M) ASR fine-tune with LoRA on q/v projections — fits a 4 GB GPU |
whisper-large-v3-asr | openai/whisper-large-v3 | 1.5B | none | Whisper large-v3 (1.5B) ASR fine-tune with LoRA — needs a larger GPU (>= 16 GB); the tiny/base recipes fit a 4 GB card |
whisper-tiny-asr | openai/whisper-tiny | 39M | none | Whisper tiny (39M) ASR fine-tune with LoRA on q/v projections — fits a 4 GB GPU. Rows: {"audio": path, "text": transcript} |
TTS: text to speech
| Recipe | Base model | Size | Quantization | What it is |
|---|---|---|---|---|
llasa-tts | HKUSTAudio/Llasa-1B | 1B | not set | Llasa-TTS. As shipped it sets data.format: audio, the live-codec path, which raises at v0.75.2 because only the Orpheus encoder exists; pre-encode your audio to codec tokens and train on chatml rows instead, see TTS |
orpheus-tts-sft | canopylabs/orpheus-3b-0.1-ft | 3B | not set | Orpheus emotional TTS. Its data.format: audio is the live-codec path and needs pip install snac; the pre-encoded path needs no extra package, see TTS |
oute-tts | OuteAI/OuteTTS-0.3-500M | 0.5B | not set | Oute-TTS with emotion conditioning. As shipped it sets data.format: audio, the live-codec path, which raises at v0.75.2 because only the Orpheus encoder exists; pre-encode your audio to codec tokens and train on chatml rows instead, see TTS |
sesame-csm-tts | sesame/csm-1b | 1B | not set | Sesame CSM conversational TTS. As shipped it sets data.format: audio, the live-codec path, which raises at v0.75.2 because only the Orpheus encoder exists; pre-encode your audio to codec tokens and train on chatml rows instead, see TTS |
spark-tts | SparkAudio/Spark-TTS-0.5B | 0.5B | not set | Spark-TTS. As shipped it sets data.format: audio, the live-codec path, which raises at v0.75.2 because only the Orpheus encoder exists; pre-encode your audio to codec tokens and train on chatml rows instead, see TTS |
See also
- Recipes: how to list, search, show and use them.
- Configuration: every key a recipe can set.
- Quant Menu: what each
quantizationvalue does.
Soup is free and Apache-2.0. If it saved you a training run, starring the repo costs nothing and helps most. You can also fund the GPU time behind the work a 4 GB laptop cannot reach.