GPT-5.6 Luna's Price-to-Performance Is Absurd

·4 min read

GPT-5.6 Luna is $0.053 per task.

Let that number sink in.

Luna-Max — at that price — scores higher than models that cost 6 to 8 times more to run. Not "close." Higher. On real benchmarks.

That's not a incremental improvement. That's a structural collapse in what AI costs per useful output.

The Numbers From the Chart

The pricing data is stark:

Luna-Max vs Opus 5: Luna scores higher. Opus 5 costs nearly 6× more per task.

Luna-Max vs Sonnet 5 High: Luna scores higher. Sonnet 5 High costs over 6× more per task.

Luna-Max vs Gemini 3.6 Flash: Luna scores higher. Gemini 3.6 Flash costs over 8× more per task.

Luna-Max vs 3.5 Flash-Lite: Luna scores higher. 3.5 Flash-Lite costs nearly 1.5× more per task.

Every comparison is the same story: Luna wins on performance AND on price. Not by a little. By margins that make the economics of using anything else at scale genuinely hard to justify.

Why $0.053 Per Task Matters

The previous generation of "good enough" AI was: use Sonnet or Opus for anything serious, use Flash or Lite for everything else, watch your budget.

The new generation is: Luna at $0.053 does more of the serious work at higher performance than the models you were paying 6–8× more for.

At $0.053 per task, you can run 19 Luna tasks for the cost of one Opus 5 task. For the cost of one Sonnet 5 High task, you can run 19 Luna tasks and have change left over.

If you're running any kind of AI-assisted workflow at scale — coding, research, drafting, analysis — the unit economics just moved. Tasks that required you to be selective about which model you used can now use Luna by default and get better results for less money.

The "Good Enough" Threshold Just Moved

The phrase "good enough" in AI has always been a budget constraint in disguise. When Sonnet 5 was the best mid-tier option, "good enough" meant "what I can afford to run frequently."

Luna at this price point redefines the threshold. "Good enough" is no longer a compromise — it's outperforming the expensive options on the metrics that matter.

For daily use — the kind of consistent, high-volume AI assistance that actually changes how you work — Luna is now the obvious default. Not because it's cheap. Because it's better AND cheap.

What This Means For Your Stack

If you're on a Codex plan, Luna's pricing already flows through your usage count. Every task you run on Luna instead of Opus or Sonnet stretches your budget further and — based on these benchmarks — potentially delivers better results.

The practical implication: audit your default model selection. If Luna is beating the models you've been using as your "serious work" tier, your serious work tier should probably be Luna now.

The one caveat: frontier reasoning tasks — novel mathematical problems, complex multi-step science, the kind of work where you're genuinely pushing the outer edge of what current models can do — still favor the highest-capability models. Luna is extraordinary at this price point. It's not yet replacing the absolute frontier on every dimension.

But for the 80% of real daily work — drafting, coding, research, analysis, iteration — the benchmark data says Luna should be your default. Not a budget option. The primary option.

The Daily Driver Implication

"Everyone's daily driver" isn't hyperbole here. At $0.053 per task with higher benchmark performance than models costing 6–8× more, Luna has crossed the threshold where the price-performance ratio is so good that not using it as your default requires a specific reason — not the other way around.

The old logic: "Sonnet/Opus for the important stuff, Flash for everything else."

The new logic: "Luna for everything. Scale up to the frontier models only when Luna hits its ceiling."

That shift — that's what "everyone's daily driver" actually means. Not a lesser model for casual use. A model so good at this price that it replaces the expensive options as your first call.

Use it.