LoRA keeps the base model frozen: read, never written. So Soup keeps it in system RAM and streams it into the GPU one decoder layer at a time. Peak VRAM becomes one layer instead of the whole model. Measured on an RTX 3050 Laptop 4 GB: Llama-3.1-8B trains at 119.6 tok/s in 3.32 GB peak. One YAML, one command. SFT, DPO, GRPO, KTO, plus eval, gating and export. Apache-2.0. Every number is published, including the ones I measured and threw away.
Exploring Soup CLI for tiny-GPU fine-tuning? Try Ollama for the easiest local model setup, or pair with Llama weights. Tap Hugging Face for datasets and trainers, sprint with Groq Chat for blazing inference, and track evals with Langfuse.