Recipe catalog

Soup ships 167 ready-made configs at v0.75.2, and this page lists every one of them. The Recipes page explains how to use them; this one is the catalog, grouped by training task.

bash
soup recipes list                      # all of them
soup recipes search qwen --task grpo   # keyword, --task and --size filters
soup recipes show llama3.1-8b-sft      # print the YAML
soup recipes use llama3.1-8b-sft -o soup.yaml   # write it, then: soup train --config soup.yaml

What a recipe is, and is not

A recipe is a complete soup.yaml for one base model and one task. All 167 of them load through the v0.75.2 config schema with no unknown keys: that was checked by loading every one. That says the file is valid. It does not say the run works end to end, and it does not say it was trained to a quality target: the data paths are placeholders you replace, a few recipes stop at trainer setup or carry a key the trainer ignores at v0.75.2 (the list below names them), and this page loaded every recipe through the schema without training any of them. Treat every hyperparameter as a starting point, and read the base model's license before you train on it.

The Size column is the base model's parameter size as the catalog states it, not a memory figure, and it is not reliable for the largest and mixture-of-experts entries: a few look like placeholders (deepseek-v3-7b-sft lists 7B for DeepSeek-V3-0324, which deepseek-v3-pipeline below lists at 671B, and glm-4.6-sft, glm-5-sft and minimax-m2-sft list 9B). N/A means the catalog gives no size. Check the model card before you size a GPU. Quantization is the training.quantization the recipe sets (4bit is 4-bit NF4 loading, the QLoRA setup; none means no quantization; 8bit is 8-bit bitsandbytes loading; not set leaves the schema default, which is 4bit; bitnet_1.58 appears once and does not train, see below). Recipes that run on Apple Silicon say mlx in the notes.

Before you run one

Loading as a config and running are different things. These are the recipes where the YAML is valid and the run still stops, needs something you have to supply first, or carries a key the trainer does not read at v0.75.2. Each was checked against the v0.75.2 source; none was trained to confirm it.

  • Sharding is a launch flag, not a recipe key. Where a description names DeepSpeed ZeRO-2 or ZeRO-3 or FSDP2 (for example llama3.1-70b-sft, qwen2.5-72b-sft, gemma3-27b-sft, qwen3-32b-sft, deepseek-r1-32b-grpo and llama3-70b-fsdp2) or says multi-node, the YAML does not select it. Pass --deepspeed zero2, --deepspeed zero3 or --fsdp full_shard, plus --gpus N, to soup train, as the multi-GPU page shows. llama3-70b-fsdp2 sets use_fsdp2_compile: true, and soup train --config soup.yaml without --fsdp exits 1.
  • Some recipes need something first. llama3.1-8b-ppo loads training.reward_model: ./output_rm, a reward model you train beforehand. llama3.1-8b-longctx sets use_flash_attn: true, which exits 1 without CUDA and the flash-attn package. baichuan-sft sets training.hub: modelscope, so the base model is fetched through ModelScope into ./.soup_hub_cache.
  • Keys the trainer ignores. Eight MoE DPO and GRPO recipes (qwen3.5-35b-a3b-dpo, deepseek-v4-flash-dpo, glm-5.1-dpo, kimi-k2.6-dpo, qwen3-30b-a3b-reasoning-grpo, deepseek-v4-flash-grpo, glm-5.1-grpo and kimi-k2.6-grpo) set moe_lora: true, and five of them set moe_aux_loss_coeff: 0.01. Only the SFT and pretrain trainers read those two keys, so on these tasks they are accepted and ignored. deepseek-v3-reasoning, kimi-k2-thinking-grpo and qwq-32b-grpo set verifiable_domain: math beside a reward_fn that is not verifiable, and verifiable_domain is read only by the verifiable reward, so it does nothing there.
  • Multimodal, speech and BitNet caveats. The DPO and GRPO trainers are text-only, so pixtral-dpo and llama3.2-vision-grpo do not use images. Of the five TTS recipes only orpheus-tts-sft has a live codec: the other four set data.format: audio and raise. falcon-e-bitnet-sft does not train, because the quantization loader rejects bitnet_1.58. whisper-large-v3-ft goes through the generic audio SFT path; for Whisper speech-to-text use whisper-large-v3-asr. deepseek-v3-pipeline validates its pipeline setting and then runs data-parallel.

SFT: supervised fine-tuning

RecipeBase modelSizeQuantizationWhat it is
baichuan-sftbaichuan-inc/Baichuan2-13B-Chat13B4bitBaichuan 2 13B chat SFT
biomistral-7b-sftBioMistral/BioMistral-7B7B4bitBioMistral 7B medical/biomedical domain SFT
codellama-13b-sftcodellama/CodeLlama-13b-Instruct-hf13B4bitCode Llama 13B code-specialist SFT
codellama-70b-sftcodellama/CodeLlama-70b-Instruct-hf70B4bitCode Llama 70B code-specialist SFT (multi-GPU)
cogito-v2-sftdeepcogito/cogito-v2-preview14B4bitCogito v2 preview SFT
deepseek-ocr-sftdeepseek-ai/DeepSeek-OCRN/A4bitDeepSeek-OCR vision OCR SFT
deepseek-r1-distill-llama-8b-sftdeepseek-ai/DeepSeek-R1-Distill-Llama-8B8B4bitDeepSeek-R1-Distill Llama 8B reasoning SFT
deepseek-r1-distill-qwen-1.5b-sftdeepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B1.5B4bitDeepSeek-R1-Distill Qwen 1.5B reasoning SFT
deepseek-r1-distill-qwen-7b-sftdeepseek-ai/DeepSeek-R1-Distill-Qwen-7B7B4bitDeepSeek-R1-Distill Qwen 7B reasoning SFT
deepseek-v3-7b-sftdeepseek-ai/DeepSeek-V3-03247B4bitDeepSeek V3 SFT with MoE LoRA
deepseek-v3-pipelinedeepseek-ai/DeepSeek-V3671B4bitDeepSeek V3 SFT scaffold. Its parallelism: pipeline (4 stages) is checked (it needs CUDA and at least 4 GPUs) and then the run is data-parallel at v0.75.2, see Pipeline parallelism
deepseek-v4-flash-sftdeepseek-ai/DeepSeek-V4-FlashN/A4bitDeepSeek V4 Flash MoE SFT (MIT, efficiency-tier)
deepseek-v4-pro-sftdeepseek-ai/DeepSeek-V4-ProN/A4bitDeepSeek V4 Pro flagship MoE SFT (MIT, 1.6T-class). Requires multi-node DeepSpeed.
devstral-sftmistralai/Devstral-Small24B4bitDevstral Small code/agent SFT
falcon-e-bitnet-sfttiiuae/Falcon-E-1B-Instruct1Bbitnet_1.58Falcon-E BitNet 1.58-bit SFT. Does not train at v0.75.2: the config validates, then the quantization loader rejects bitnet_1.58, see BitNet
gemma2-2b-sftgoogle/gemma-2-2b-it2B4bitGemma 2 2B SFT (edge-friendly)
gemma3-12b-sftgoogle/gemma-3-12b-it12B4bitGemma 3 12B instruction tuning
gemma3-27b-sftgoogle/gemma-3-27b-it27B4bitGemma 3 27B SFT with DeepSpeed ZeRO-2
gemma3-4b-sft-mlxmlx-community/gemma-3-4b-it-4bit4B4bitGemma 3 4B SFT on Apple Silicon via MLX (M1+ 16GB) (mlx)
gemma3-9b-sftgoogle/gemma-3-9b-it9B4bitGemma 3 9B instruction tuning
glm-4.6-sftTHUDM/glm-4.69B4bitGLM 4.6 instruction tuning with LoRA
glm-5-sftzai-org/GLM-59B4bitGLM 5 SFT (next-gen GLM family)
glm-5.1-sftzai-org/GLM-5.1754B4bitGLM 5.1 MoE SFT (MIT, 754B). Multi-GPU / multi-node recommended.
gpt-oss-120b-sftopenai/gpt-oss-120b120B4bitGPT-OSS 120B SFT (multi-GPU recommended)
gpt-oss-20b-sftopenai/gpt-oss-20b20B4bitGPT-OSS 20B SFT (reasoning_effort=medium)
granite-4-sftibm-granite/granite-4.0-tiny-base3B4bitIBM Granite 4.0 tiny SFT
internvl-2.5-8b-sftOpenGVLab/InternVL2_5-8B8B4bitInternVL 2.5 8B vision-language SFT
internvl-3-5-sftOpenGVLab/InternVL3-58B4bitInternVL 3.5 vision SFT
kimi-k2-sftmoonshotai/Kimi-K2N/A4bitKimi K2 SFT (Moonshot MoE, long-context-aware)
kimi-k2.5-sftmoonshotai/Kimi-K2.51T4bitKimi K2.5 MoE SFT (Modified MIT, ~1T / 32B active). Requires multi-node DeepSpeed.
kimi-k2.6-sftmoonshotai/Kimi-K2.61T4bitKimi K2.6 MoE SFT (Modified MIT, ~1T / 32B active). Requires multi-node DeepSpeed.
lfm2-sftLiquidAI/LFM2-1.2B1.2B4bitLiquid LFM2 1.2B SFT (edge-optimised)
llama2-13b-finance-sftmeta-llama/Llama-2-13b-hf13B4bitLlama 2 13B finance-domain SFT (FinGPT-style starter recipe)
llama3-70b-fsdp2meta-llama/Llama-3.1-70B-Instruct70B4bitLlama 3.1 70B SFT with FSDP2 full shard + torch.compile. Requires 8 x A100/H100 80GB.
llama3.1-70b-sftmeta-llama/Llama-3.1-70B-Instruct70B4bitLlama 3.1 70B SFT with DeepSpeed ZeRO-3
llama3.1-8b-longctxmeta-llama/Llama-3.1-8B-Instruct8B4bitLlama 3.1 8B long-context (32k) with YaRN RoPE scaling
llama3.1-8b-sftmeta-llama/Llama-3.1-8B-Instruct8B4bitLlama 3.1 8B instruction tuning with LoRA
llama3.1-8b-sft-mlxmlx-community/Llama-3.1-8B-Instruct-4bit8B4bitLlama 3.1 8B SFT on Apple Silicon via MLX (M2+ 16GB) (mlx)
llama3.2-11b-visionmeta-llama/Llama-3.2-11B-Vision-Instruct11B4bitLlama 3.2 11B Vision multimodal fine-tuning
llama3.2-1b-sftmeta-llama/Llama-3.2-1B-Instruct1B8bitLlama 3.2 1B instruction tuning (mobile-friendly)
llama3.2-3b-sftmeta-llama/Llama-3.2-3B-Instruct3B4bitLlama 3.2 3B instruction tuning (edge-friendly)
llama3.2-vision-90b-sftmeta-llama/Llama-3.2-90B-Vision-Instruct90B4bitLlama 3.2 90B Vision multimodal SFT (8 x A100/H100 80GB)
llama4-scout-17b-sftmeta-llama/Llama-4-Scout-17B-16E-Instruct17B4bitLlama 4 Scout 17B SFT with LoRA (4bit)
llama4-scout-toolsmeta-llama/Llama-4-Scout-17B-16E-Instruct17B4bitLlama 4 Scout 17B tool-calling / function-calling SFT
llava-next-sftllava-hf/llava-v1.6-mistral-7b-hf7B4bitLLaVA-Next 7B vision SFT
magicoder-7b-sftise-uiuc/Magicoder-S-DS-6.7B6.7B4bitMagicoder S-DS 6.7B code-specialist SFT
magistral-small-sftmistralai/Magistral-Small24B4bitMagistral Small reasoning SFT
mathstral-7b-sftmistralai/Mathstral-7B-v0.17B4bitMathstral 7B math/STEM-specialist SFT
medgemma-sftgoogle/medgemma-4b-it4B4bitMedGemma 4B medical SFT
meditron-7b-sftepfl-llm/meditron-7b7B4bitMeditron 7B medical / clinical domain SFT
minicpm-v-2.6-sftopenbmb/MiniCPM-V-2_68B4bitMiniCPM-V 2.6 vision-language SFT (edge-friendly multimodal)
minimax-m2-sftMiniMaxAI/MiniMax-M29B4bitMiniMax M2 SFT instruction tuning
minimax-m3-sftMiniMaxAI/MiniMax-M3428B4bitMiniMax M3 MoE SFT (428B / 23B active). MiniMax Community License - commercial use requires a separate agreement. Multi-GPU recommended.
ministral-sftmistralai/Ministral-8B-Instruct-24108B4bitMinistral 8B SFT
mistral-7b-sftmistralai/Mistral-7B-Instruct-v0.37B4bitMistral 7B instruction tuning
mistral-large-3-sftmistralai/Mistral-Large-3-675B-Instruct-2512675B4bitMistral Large 3 MoE SFT (Apache-2.0, 675B / 41B active, multimodal). Requires multi-node DeepSpeed.
mistral-medium-3-5-sftmistralai/Mistral-Medium-3.5N/A4bitMistral Medium 3.5 SFT
mistral-small-3-sftmistralai/Mistral-Small-24B-Instruct-250124B4bitMistral Small 3 24B SFT
nemotron-4-340b-sftnvidia/Nemotron-4-340B-Instruct340B4bitNemotron-4 340B SFT (massive multi-node deployment)
paddle-ocr-sftPaddlePaddle/PaddleOCR-VLN/A4bitPaddle-OCR-VL OCR SFT
phi3.5-mini-sftmicrosoft/Phi-3.5-mini-instruct3.8B4bitPhi-3.5-mini 3.8B SFT (small / edge)
phi4-14b-sftmicrosoft/phi-414B4bitPhi-4 14B instruction tuning
pixtral-12b-sftmistralai/Pixtral-12B-240912B4bitPixtral 12B vision-language SFT with LoRA
qvq-72b-sftQwen/QVQ-72B-Preview72B4bitQVQ 72B vision-reasoning SFT
qwen-image-sftQwen/Qwen-ImageN/A4bitQwen-Image image-output multimodal SFT
qwen2-audio-7b-sftQwen/Qwen2-Audio-7B-Instruct7B4bitQwen2-Audio 7B audio-language SFT
qwen2-vl-72b-sftQwen/Qwen2-VL-72B-Instruct72B4bitQwen2-VL 72B vision-language SFT (multi-GPU recommended)
qwen2-vl-7b-sftQwen/Qwen2-VL-7B-Instruct7B4bitQwen2-VL 7B vision-language SFT
qwen2.5-0.5b-sftQwen/Qwen2.5-0.5B-Instruct0.5B8bitQwen 2.5 0.5B SFT (mobile / edge)
qwen2.5-1.5b-sftQwen/Qwen2.5-1.5B-Instruct1.5B8bitQwen 2.5 1.5B SFT (edge-friendly)
qwen2.5-3b-sftQwen/Qwen2.5-3B-Instruct3B4bitQwen 2.5 3B SFT (small / edge)
qwen2.5-72b-sftQwen/Qwen2.5-72B-Instruct72B4bitQwen 2.5 72B SFT with DeepSpeed ZeRO-3
qwen2.5-7b-sftQwen/Qwen2.5-7B-Instruct7B4bitQwen 2.5 7B instruction tuning with LoRA
qwen2.5-coder-1.5b-sftQwen/Qwen2.5-Coder-1.5B-Instruct1.5B4bitQwen 2.5 Coder 1.5B instruction tuning with LoRA
qwen2.5-coder-14b-sftQwen/Qwen2.5-Coder-14B-Instruct14B4bitQwen 2.5 Coder 14B instruction tuning with LoRA
qwen2.5-coder-32b-sftQwen/Qwen2.5-Coder-32B-Instruct32B4bitQwen 2.5 Coder 32B instruction tuning with LoRA
qwen2.5-coder-7b-sftQwen/Qwen2.5-Coder-7B-Instruct7B4bitQwen 2.5 Coder 7B instruction tuning with LoRA
qwen2.5-math-1.5b-sftQwen/Qwen2.5-Math-1.5B-Instruct1.5B4bitQwen 2.5 Math 1.5B instruction tuning with LoRA
qwen2.5-math-7b-sftQwen/Qwen2.5-Math-7B-Instruct7B4bitQwen 2.5 Math 7B instruction tuning with LoRA
qwen3-14b-sftQwen/Qwen3-14B14B4bitQwen 3 14B instruction tuning with LoRA
qwen3-30b-a3b-sftQwen/Qwen3-30B-A3B30B4bitQwen 3 30B-A3B MoE instruction tuning
qwen3-32b-sftQwen/Qwen3-32B32B4bitQwen 3 32B SFT with DeepSpeed ZeRO-2
qwen3-32b-zeroppQwen/Qwen3-32B32B4bitQwen3 32B SFT with DeepSpeed ZeRO++ (quantized gradients + hierarchical partitioning). Launch with --deepspeed zero++.
qwen3-8b-sftQwen/Qwen3-8B8B4bitQwen 3 8B instruction tuning
qwen3-8b-sft-mlxmlx-community/Qwen3-8B-4bit8B4bitQwen 3 8B SFT on Apple Silicon via MLX (M2+ 16GB) (mlx)
qwen3-8b-toolsQwen/Qwen3-8B8B4bitQwen 3 8B tool-calling / function-calling SFT
qwen3-coder-30b-sftQwen/Qwen3-Coder-30B-A3B-Instruct30B4bitQwen3-Coder 30B (A3B MoE) code-specialist SFT
qwen3.5-0.8b-sftQwen/Qwen3.5-0.8B0.8B4bitQwen 3.5 0.8B SFT (Apache-2.0, tiny / mobile)
qwen3.5-122b-a10b-sftQwen/Qwen3.5-122B-A10B122B4bitQwen 3.5 122B-A10B MoE SFT (Apache-2.0, 10B active). Multi-GPU recommended (8 x A100/H100 80GB).
qwen3.5-27b-sftQwen/Qwen3.5-27B27B4bitQwen 3.5 27B SFT (Apache-2.0) with DeepSpeed ZeRO-2
qwen3.5-2b-sftQwen/Qwen3.5-2B2B4bitQwen 3.5 2B SFT (Apache-2.0, small / edge)
qwen3.5-35b-a3b-sftQwen/Qwen3.5-35B-A3B35B4bitQwen 3.5 35B-A3B MoE SFT (Apache-2.0, 3B active)
qwen3.5-397b-a17b-sftQwen/Qwen3.5-397B-A17B397B4bitQwen 3.5 397B-A17B flagship MoE SFT (Apache-2.0, 17B active). Requires multi-node DeepSpeed / FSDP.
qwen3.5-4b-sftQwen/Qwen3.5-4B4B4bitQwen 3.5 4B SFT (Apache-2.0, 262K context)
qwen3.5-9b-sftQwen/Qwen3.5-9B9B4bitQwen 3.5 9B SFT (Apache-2.0, 262K context)
qwen3.6-27b-sftQwen/Qwen3.6-27B27B4bitQwen 3.6 27B SFT (Apache-2.0) with DeepSpeed ZeRO-2
qwen3.6-35b-a3b-sftQwen/Qwen3.6-35B-A3B35B4bitQwen 3.6 35B-A3B MoE SFT (Apache-2.0, 3B active)
qwen3.8-27b-sftQwen/Qwen3.8-27B27B4bitQwen 3.8 27B text-only SFT (Apache-2.0) with DeepSpeed ZeRO-2
ra-dit-llama3-8bmeta-llama/Llama-3.1-8B-Instruct8B4bitRA-DIT stage 2 (Meta 2023) — RAFT-style SFT on the generator. Pairs with ra-dit-retriever (stage 1). Uses RAFT data format.
raft-llama3-8bmeta-llama/Llama-3.1-8B-Instruct8B4bitRAFT (Retrieval-Augmented Fine-Tuning, Stanford 2024) — train an 8B Llama 3.1 to answer queries given a golden doc + distractor docs. Rows: {query, golden_doc, distractor_docs, answer}.
seamlessm4t-v2-sftfacebook/seamless-m4t-v2-large2.3B4bitSeamlessM4T v2 multilingual speech-to-text SFT
smollm2-1.7b-sftHuggingFaceTB/SmolLM2-1.7B-Instruct1.7B8bitSmolLM2 1.7B SFT (small / edge)
smollm2-135m-sftHuggingFaceTB/SmolLM2-135M-Instruct135MnoneSmolLM2 135M SFT (ultra-tiny / mobile)
smollm2-360m-sftHuggingFaceTB/SmolLM2-360M-Instruct360MnoneSmolLM2 360M SFT (tiny / mobile)
smollm3-3b-sftHuggingFaceTB/SmolLM3-3B3B8bitSmolLM3 3B SFT (small / edge)
smolvlm-256m-sftHuggingFaceTB/SmolVLM-256M-Instruct256MnoneSmolVLM 256M vision SFT (llava format) — a tiny VLM. SmolVLM uses an Idefics3 processor; Soup mirrors its nested tokenizer surface and performs processor-aware vision collation so image-token expansion and pixel_values reach the model. target_modules are pinned to q_proj/v_proj (auto cannot infer them for Idefics3).
voxtral-sftmistralai/Voxtral-Mini-3B3B4bitVoxtral Mini 3B audio SFT
whisper-large-v3-ftopenai/whisper-large-v31.5B8bitWhisper Large v3 through the generic audio SFT path (modality: audio), which loads the model with AutoModel, the bare Whisper encoder-decoder without a language-model head. For Whisper speech-to-text use the ASR task, see ASR fine-tuning and whisper-large-v3-asr below

DPO: preference optimization

RecipeBase modelSizeQuantizationWhat it is
deepseek-r1-distill-llama-8b-dpodeepseek-ai/DeepSeek-R1-Distill-Llama-8B8B4bitDeepSeek-R1-Distill Llama 8B DPO alignment
deepseek-r1-distill-qwen-1.5b-dpodeepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B1.5B4bitDeepSeek-R1-Distill Qwen 1.5B DPO alignment
deepseek-r1-distill-qwen-7b-dpodeepseek-ai/DeepSeek-R1-Distill-Qwen-7B7B4bitDeepSeek-R1-Distill Qwen 7B DPO alignment
deepseek-v4-flash-dpodeepseek-ai/DeepSeek-V4-FlashN/A4bitDeepSeek V4 Flash MoE DPO alignment (MIT, efficiency-tier)
gemma3-27b-dpogoogle/gemma-3-27b-it27B4bitGemma 3 27B DPO alignment
glm-5.1-dpozai-org/GLM-5.1754B4bitGLM 5.1 MoE DPO alignment (MIT, 754B). Multi-GPU / multi-node recommended.
kimi-k2.6-dpomoonshotai/Kimi-K2.61T4bitKimi K2.6 MoE DPO alignment (Modified MIT, ~1T / 32B active). Requires multi-node DeepSpeed.
llama3.1-8b-dpometa-llama/Llama-3.1-8B-Instruct8B4bitLlama 3.1 8B DPO alignment
llama4-scout-17b-dpometa-llama/Llama-4-Scout-17B-16E-Instruct17B4bitLlama 4 Scout 17B DPO alignment
mistral-7b-dpomistralai/Mistral-7B-Instruct-v0.37B4bitMistral 7B DPO alignment
pixtral-dpomistralai/Pixtral-12B-240912B4bitPixtral 12B DPO. Not usable as shipped at v0.75.2: the DPO trainer is text-only (no vision processor), and data.format: llava produces messages and image columns without the chosen and rejected that DPO reads
qwen2.5-7b-dpoQwen/Qwen2.5-7B-Instruct7B4bitQwen 2.5 7B DPO alignment
qwen3.5-35b-a3b-dpoQwen/Qwen3.5-35B-A3B35B4bitQwen 3.5 35B-A3B MoE DPO alignment (Apache-2.0, 3B active)
qwen3.5-9b-dpoQwen/Qwen3.5-9B9B4bitQwen 3.5 9B DPO alignment (Apache-2.0)

GRPO: reinforcement learning with a verifier

RecipeBase modelSizeQuantizationWhat it is
deepseek-r1-32b-grpodeepseek-ai/DeepSeek-R1-Distill-Qwen-32B32B4bitDeepSeek R1 32B GRPO with DeepSpeed
deepseek-r1-8b-grpodeepseek-ai/DeepSeek-R1-Distill-Llama-8B8B4bitDeepSeek R1 8B GRPO reasoning
deepseek-v3-reasoningdeepseek-ai/DeepSeek-V3N/A4bitDeepSeek V3 (671B-parameter MoE) GRPO reasoning recipe. Requires multi-node DeepSpeed. Scores with the accuracy and format rewards; its verifiable_domain: math is only read by reward_fn: verifiable and has no effect here
deepseek-v4-flash-grpodeepseek-ai/DeepSeek-V4-FlashN/A4bitDeepSeek V4 Flash MoE GRPO reasoning training (MIT, efficiency-tier)
glm-5.1-grpozai-org/GLM-5.1754B4bitGLM 5.1 MoE GRPO reasoning training (MIT, 754B). Multi-GPU / multi-node recommended.
grpo-env-calculatorHuggingFaceTB/SmolLM2-135M-Instruct135Mnot setSmolLM2-135M GRPO on the bundled calculator env (openenv rollout)
grpo-env-guess-numberHuggingFaceTB/SmolLM2-135M-Instruct135Mnot setSmolLM2-135M GRPO on the bundled number-deduction env (openenv rollout)
grpo-env-retrieval-qaHuggingFaceTB/SmolLM2-135M-Instruct135Mnot setSmolLM2-135M GRPO on the bundled retrieval-QA env (openenv rollout)
kimi-k2-thinking-grpomoonshotai/Kimi-K2-ThinkingN/A4bitKimi K2 Thinking GRPO reasoning
kimi-k2.6-grpomoonshotai/Kimi-K2.61T4bitKimi K2.6 MoE GRPO reasoning (Modified MIT, ~1T / 32B active). Requires multi-node DeepSpeed.
llama3.1-8b-grpometa-llama/Llama-3.1-8B-Instruct8B4bitLlama 3.1 8B GRPO reasoning training
llama3.2-vision-grpometa-llama/Llama-3.2-11B-Vision-Instruct11B4bitLlama 3.2 11B Vision GRPO. The GRPO trainer is text-only at v0.75.2, so the images are not used, see GRPO Plus
llama4-scout-17b-grpometa-llama/Llama-4-Scout-17B-16E-Instruct17B4bitLlama 4 Scout 17B GRPO reasoning training
phi4-reasoning-grpomicrosoft/phi-414B4bitPhi-4 14B GRPO reasoning training
qwen2.5-7b-grpoQwen/Qwen2.5-7B-Instruct7B4bitQwen 2.5 7B GRPO reasoning training
qwen3-30b-a3b-reasoning-grpoQwen/Qwen3-30B-A3B30B4bitQwen3 30B-A3B GRPO reasoning training (MoE thinking model)
qwen3-8b-grpoQwen/Qwen3-8B8B4bitQwen 3 8B GRPO reasoning training
qwen3.5-0.8b-grpoQwen/Qwen3.5-0.8B0.8B4bitQwen 3.5 0.8B GRPO reasoning training (Apache-2.0, tiny / mobile)
qwen3.5-2b-grpoQwen/Qwen3.5-2B2B4bitQwen 3.5 2B GRPO reasoning training (Apache-2.0, small / edge)
qwen3.5-9b-grpoQwen/Qwen3.5-9B9B4bitQwen 3.5 9B GRPO reasoning training (Apache-2.0, 262K context)
qwq-32b-grpoQwen/QwQ-32B32B4bitQwQ 32B GRPO reasoning training
r1-distill-llama-70b-grpodeepseek-ai/DeepSeek-R1-Distill-Llama-70B70B4bitDeepSeek-R1-Distill Llama 70B GRPO reasoning (multi-GPU)
r1-distill-qwen-1.5b-grpodeepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B1.5B4bitDeepSeek-R1-Distill Qwen 1.5B GRPO reasoning training
r1-distill-qwen-14b-grpodeepseek-ai/DeepSeek-R1-Distill-Qwen-14B14B4bitDeepSeek-R1-Distill Qwen 14B GRPO reasoning training
r1-distill-qwen-7b-grpodeepseek-ai/DeepSeek-R1-Distill-Qwen-7B7B4bitDeepSeek-R1-Distill Qwen 7B GRPO reasoning training

KTO

RecipeBase modelSizeQuantizationWhat it is
llama3.1-8b-ktometa-llama/Llama-3.1-8B-Instruct8B4bitLlama 3.1 8B KTO unpaired preference alignment

ORPO

RecipeBase modelSizeQuantizationWhat it is
llama3.1-8b-orpometa-llama/Llama-3.1-8B-Instruct8B4bitLlama 3.1 8B ORPO reference-free alignment

SimPO

RecipeBase modelSizeQuantizationWhat it is
llama3.1-8b-simpometa-llama/Llama-3.1-8B-Instruct8B4bitLlama 3.1 8B SimPO length-normalized alignment

IPO

RecipeBase modelSizeQuantizationWhat it is
llama3.1-8b-ipometa-llama/Llama-3.1-8B-Instruct8B4bitLlama 3.1 8B IPO regularized preference alignment

PPO

RecipeBase modelSizeQuantizationWhat it is
llama3.1-8b-ppometa-llama/Llama-3.1-8B-Instruct8B4bitLlama 3.1 8B PPO (RLHF stage 3)

Online DPO

RecipeBase modelSizeQuantizationWhat it is
online-dpo-smollm2-135mHuggingFaceTB/SmolLM2-135M-Instruct135MnoneSmolLM2 135M Online DPO — on-policy generation judged by a pairwise LLM judge (the recipe sets training.online_dpo_judge: ollama://llama3.1, which needs a local Ollama server; point it at your own judge)

Reward model

RecipeBase modelSizeQuantizationWhat it is
qwen2.5-7b-rewardQwen/Qwen2.5-7B-Instruct7B4bitQwen 2.5 7B reward model (RLHF stage 2)

Embedding

RecipeBase modelSizeQuantizationWhat it is
embedding-gemma-sftgoogle/embeddinggemma-300m300M4bitEmbeddingGemma 300M sentence-embedding SFT
llama3.1-8b-embedmeta-llama/Llama-3.1-8B8B4bitLlama 3.1 8B sentence embedding with cosine loss
ra-dit-retrieversentence-transformers/all-mpnet-base-v2N/Anot setRA-DIT stage 1 (Meta 2023) — contrastive retriever training. Pairs with ra-dit-llama3-8b (stage 2) for the full pipeline.

Continued pre-training

RecipeBase modelSizeQuantizationWhat it is
qwen2.5-7b-pretrainQwen/Qwen2.5-7B7B4bitQwen 2.5 7B continued pre-training
qwen3.5-4b-pretrainQwen/Qwen3.5-4B-Base4B4bitQwen 3.5 4B Base continued pre-training (Apache-2.0)

ASR: speech recognition

RecipeBase modelSizeQuantizationWhat it is
whisper-base-asropenai/whisper-base74MnoneWhisper base (74M) ASR fine-tune with LoRA on q/v projections — fits a 4 GB GPU
whisper-large-v3-asropenai/whisper-large-v31.5BnoneWhisper large-v3 (1.5B) ASR fine-tune with LoRA — needs a larger GPU (>= 16 GB); the tiny/base recipes fit a 4 GB card
whisper-tiny-asropenai/whisper-tiny39MnoneWhisper tiny (39M) ASR fine-tune with LoRA on q/v projections — fits a 4 GB GPU. Rows: {"audio": path, "text": transcript}

TTS: text to speech

RecipeBase modelSizeQuantizationWhat it is
llasa-ttsHKUSTAudio/Llasa-1B1Bnot setLlasa-TTS. As shipped it sets data.format: audio, the live-codec path, which raises at v0.75.2 because only the Orpheus encoder exists; pre-encode your audio to codec tokens and train on chatml rows instead, see TTS
orpheus-tts-sftcanopylabs/orpheus-3b-0.1-ft3Bnot setOrpheus emotional TTS. Its data.format: audio is the live-codec path and needs pip install snac; the pre-encoded path needs no extra package, see TTS
oute-ttsOuteAI/OuteTTS-0.3-500M0.5Bnot setOute-TTS with emotion conditioning. As shipped it sets data.format: audio, the live-codec path, which raises at v0.75.2 because only the Orpheus encoder exists; pre-encode your audio to codec tokens and train on chatml rows instead, see TTS
sesame-csm-ttssesame/csm-1b1Bnot setSesame CSM conversational TTS. As shipped it sets data.format: audio, the live-codec path, which raises at v0.75.2 because only the Orpheus encoder exists; pre-encode your audio to codec tokens and train on chatml rows instead, see TTS
spark-ttsSparkAudio/Spark-TTS-0.5B0.5Bnot setSpark-TTS. As shipped it sets data.format: audio, the live-codec path, which raises at v0.75.2 because only the Orpheus encoder exists; pre-encode your audio to codec tokens and train on chatml rows instead, see TTS

See also

Soup is free and Apache-2.0. If it saved you a training run, starring the repo costs nothing and helps most. You can also fund the GPU time behind the work a 4 GB laptop cannot reach.