OpenAI's new model costs $2 in, $10 out per million tokens. The cheap part was never the tokens.

Cheaper near-frontier models make it tempting to send more context on every call. For anything personal, that is the wrong place to spend the savings.

OpenAI announced GPT-6.1 Sol at DevDay on 2 October. Its API page lists $2.00 per million input tokens, $0.10 for cached input and $10.00 for output, with a 1,050,000-token context window. Prompts over 272K input tokens are billed at double the input rate and 1.5x output for the whole request. OpenAI says the model gives near-Astra intelligence at a fifth of Astra's price. That comparison is OpenAI's own claim about its own product, and I haven't seen an independent benchmark of it, so treat it as marketing until someone you trust has run it on your workload.

The same event added computer use to the Agents API, so an agent can operate software through its interface, available through the API and in Codex and ChatGPT Work on Pro 500 and Enterprise plans.

Here is what I think a small team should take from it. When the per-token price drops, the first instinct is to spend the savings on more context: send the whole history, the whole document, the whole account. The model can take a million tokens, so why not use them.

The 272K line on the pricing page is a good reminder that the price isn't flat anyway. But the bigger issue is not the bill. Every extra token you send is data you now have to justify, disclose and protect. If your product handles anything personal, the question isn't "can we afford to send this?" It's "does this answer get better enough to be worth sending it?"

We learned this the hard way at Murror. We used to keep every journal entry forever because more history looked like better insight. When we capped retention at 30 days, the insights barely changed and people started writing more openly. The context we thought we needed was mostly comfort for us, not value for them.

So my rule for a cheaper model is this. Take the savings and spend them on evaluation instead. Pick twenty real cases, run them with a small context and a large one, and have someone read the outputs side by side. If the large context doesn't win clearly, you've just found a privacy and cost win in the same afternoon.

On computer use, I'd go one step at a time. Agents that click through interfaces mean the way your app is labelled, and whether it can be operated cleanly, starts to matter to software as well as to people. I haven't tested this on our own app yet, so this is a hypothesis, not a finding. The cheap first check is the one you should be doing anyway: do your buttons and fields have real accessible labels?

If you're choosing a model this week, what are you measuring besides price per token?

35 views

Add a comment

Replies

Best

the "30 day cap barely changed the insights" result is the real finding here, everything after that is commentary. we run into the same temptation at Dial with call transcripts - it's tempting to feed the agent the customer's entire call history on every turn because the context window allows it, but most of what actually matters to route a call correctly is the last interaction plus two or three structured facts, not the full narrative. what I measure besides price per token is cost per resolved call without a human handoff, because a cheaper model that needs more retries or escalates more often isn't actually cheaper. latency variance under load matters more to us than the sticker price too - a model that's fast on average but spikes to 4 seconds on 1 in 20 calls breaks a phone conversation in a way the benchmark never shows.

 Cost per resolved call without a handoff is a much better yardstick than anything on a pricing page, thank you for putting it that way. The latency spike point is a good one too: averages hide exactly the failure a live conversation can't tolerate. We saw a similar pattern with the 30-day cap, where the last few entries plus a handful of structured themes did nearly all the work. Do you keep those structured facts as a separate store, or rebuild them from the transcript each turn?

 separate store. we keep a small structured record per customer (last outcome, open items, a couple of preferences) that gets updated after the call, not rebuilt from the transcript each time. rebuilding from raw transcript on every turn was our first version and it was too slow and too inconsistent - the model would phrase the same fact differently each time and occasionally drop it. the structured version also makes it easy to spot when a fact is stale, which a transcript replay doesn't give you for free.

I generate flashcards between language pairs. On the cheaper tier the overall error rate looked fine, but pairs where neither language is English came out consistently worse on translation. So those requests go to the pricier model automatically and everything else stays cheap. Same price per token, two very different costs depending on who's asking.

The disadvantage: every new model release means two evals instead of one

 That is a great concrete example: an aggregate error rate that looks fine while one slice quietly fails is exactly what a single eval hides. Routing by language pair is a neat way to pay for quality only where it shows up. The two-evals-per-release cost is real, but a small fixed set of pairs you re-run each time probably keeps it manageable. Did you find the weak pairs by reading outputs yourself, or did users flag them first?

 Neither, honestly. It came out of an eval run. I keep a fixed sweep of language pairs, five of them with no English on either side, and a second model grades every card.

The funny part is that on the newest cheap model the gap mostly closed in my sample. I kept the routing anyway, because the sample is small and the saving is pennies

Well fair point! The eval comparison is what I’d spend the savings on too, I’ve seen a bigger context hide the part that actually needs review. Do you keep the small and large runs side by side, or only compare summaries?

Honestly, the point that extra tokens are extra data to protect is the one I'd keep. People talk about the bill, but a million tokens of someone's history is a lot to be responsible for.