"Can we self-host this model, and how much?" One number never ends the discussion. Pick a model (KIMI, GLM, DeepSeek…) and quantization, choose on-prem or cloud, set assumptions, and get a VRAM-matched hardware plan plus a monthly cost split. Every result shows its price sources and assumptions, refreshed on a regular schedule and never made up. Share one URL so the team debates the same inputs. A planning estimate, not a quote.
Hey Product Hunt! 👋
This work started with a boss asking whether we could self-host a new open-weight model, and what it would cost. A quick total didn't settle it. It turned into "what exactly are we buying, how many GPUs, and does quantizing or renting first change the number?"
So I built a calculator that answers with its work shown. Pick a model (KIMI, GLM, DeepSeek…), a quantization, on-prem or cloud, and monthly runtime; it returns the VRAM-matched hardware and a monthly cost split, with the price sources and assumptions right there. Electricity rate, ops labor, depreciation. Edit any of them and the budget moves. One URL carries the whole scenario, so the review argues about the inputs, not the total.
The first version only gave a price. Review meetings are where that fell apart, because people wanted to see where the money went. So the tool grew to show its assumptions and sources instead of hiding them. It's a planning estimate you can defend, with no benchmark claims and no fake precision.
Happy to talk through the math or the limits. Feedback welcome.
Report
No reviews yetBe the first to leave a review for How Much to Run AI