Launching today

Soup CLI
Fine-tune an 8B LLM on a 4 GB laptop GPU
27 followers
Fine-tune an 8B LLM on a 4 GB laptop GPU
27 followers
LoRA keeps the base model frozen: read, never written. So Soup keeps it in system RAM and streams it into the GPU one decoder layer at a time. Peak VRAM becomes one layer instead of the whole model. Measured on an RTX 3050 Laptop 4 GB: Llama-3.1-8B trains at 119.6 tok/s in 3.32 GB peak. One YAML, one command. SFT, DPO, GRPO, KTO, plus eval, gating and export. Apache-2.0. Every number is published, including the ones I measured and threw away.
Products used by Soup CLI
Explore the tech stack and tools that power Soup CLI. See what products Soup CLI uses for development, design, marketing, analytics, and more.
LLMs 1
LLMs 1
Engineering & Development 2
Engineering & Development 2

