OpenAI's new model costs $2 in, $10 out per million tokens. The cheap part was never the tokens.

Cheaper near-frontier models make it tempting to send more context on every call. For anything personal, that is the wrong place to spend the savings.

OpenAI announced GPT-6.1 Sol at DevDay on 2 October. Its API page lists $2.00 per million input tokens, $0.10 for cached input and $10.00 for output, with a 1,050,000-token context window. Prompts over 272K input tokens are billed at double the input rate and 1.5x output for the whole request. OpenAI says the model gives near-Astra intelligence at a fifth of Astra's price. That comparison is OpenAI's own claim about its own product, and I haven't seen an independent benchmark of it, so treat it as marketing until someone you trust has run it on your workload.

The same event added computer use to the Agents API, so an agent can operate software through its interface, available through the API and in Codex and ChatGPT Work on Pro 500 and Enterprise plans.

Here is what I think a small team should take from it. When the per-token price drops, the first instinct is to spend the savings on more context: send the whole history, the whole document, the whole account. The model can take a million tokens, so why not use them.

The 272K line on the pricing page is a good reminder that the price isn't flat anyway. But the bigger issue is not the bill. Every extra token you send is data you now have to justify, disclose and protect. If your product handles anything personal, the question isn't "can we afford to send this?" It's "does this answer get better enough to be worth sending it?"

We learned this the hard way at Murror. We used to keep every journal entry forever because more history looked like better insight. When we capped retention at 30 days, the insights barely changed and people started writing more openly. The context we thought we needed was mostly comfort for us, not value for them.

So my rule for a cheaper model is this. Take the savings and spend them on evaluation instead. Pick twenty real cases, run them with a small context and a large one, and have someone read the outputs side by side. If the large context doesn't win clearly, you've just found a privacy and cost win in the same afternoon.

On computer use, I'd go one step at a time. Agents that click through interfaces mean the way your app is labelled, and whether it can be operated cleanly, starts to matter to software as well as to people. I haven't tested this on our own app yet, so this is a hypothesis, not a finding. The cheap first check is the one you should be doing anyway: do your buttons and fields have real accessible labels?

If you're choosing a model this week, what are you measuring besides price per token?

7 views

Add a comment

Replies

Be the first to comment