Deterministic VRAM math and a validated throughput model for LLM inference GPU sizing. Pick a model, quant, and context length — get a real number for fit, speed, and cost across 100+ GPUs, validated to 11.3% median error against 19 published benchmarks. Free, no card required.