---
title: "Claude Haiku 5.5 meets GPT-6 Luna at ten cents"
description: "Claude Haiku 5.5 matches GPT-6 Luna at $0.10 per million input tokens, as ChatGPT adds Intelligent UI and Microsoft pushes local, sandboxed agents."
canonical_url: https://www.oguzhan.co/claude-haiku-5-5-gpt-6-luna-price-floor/
author: "Oğuzhan Koçaklı"
author_url: https://www.oguzhan.co/about/
date_published: 2026-10-08T06:17:14+00:00
date_modified: 2026-10-08T19:05:18+00:00
language: en
translations:
  tr: https://www.oguzhan.co/tr/claude-haiku-5-5-gpt-6-luna-fiyat-tabani/
---

# Claude Haiku 5.5 meets GPT-6 Luna at ten cents

Claude Haiku 5.5 has pulled the small-model price floor down to ten cents per million input tokens, matching GPT-6 Luna. That matters because agent systems make many cheap, narrow calls around the occasional difficult one. Wednesday also brought GPT-6 to ChatGPT, a Microsoft machine for local agents, a public SynthID checker, and a dispute over visible AI reasoning.

## 💸 Claude Haiku 5.5 meets the ten-cent price floor

![Isometric illustration of dozens of small cube robots carrying documents and database parcels across a long low teal plateau toward one large glowing sphere, with a tall amber step at the far end](https://oguzhanco.s3.eu-west-3.amazonaws.com/wp-content/uploads/2026/10/08061712/inline-haiku-price-floor.webp)

*Haiku 5.5 costs $0.10 per million input tokens for prompts up to 100K, the same list price as GPT-6 Luna. Above 100K the price jumps fivefold, to $0.50.*

[Anthropic released](https://www.anthropic.com/claude-haiku-5-5) its third 5.5 model in about a month. According to Anthropic, Haiku costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. Cache reads cost $0.01. OpenAI’s [API pricing page](https://developers.openai.com/api/docs/pricing) lists GPT-6 Luna at exactly those short-context prices.

Past 100K, Haiku jumps fivefold to $0.50 input, $2.50 output, and $0.05 for cache reads. Anthropic says roughly 90% of Haiku 4.5 requests stayed below that line. Short requests are 90% cheaper. On average, Anthropic puts the saving at around 75% once its new, slightly token-hungrier tokenizer is counted. Haiku 4.5 cost $1 and $5.

Side by side, in US dollars per million tokens:

| Model and context | Input | Output |
| --- | --- | --- |
| Claude Haiku 5.5, up to 100K tokens | $0.10 | $0.50 |
| Claude Haiku 5.5, above 100K tokens | $0.50 | $2.50 |
| GPT-6 Luna, short context | $0.10 | $0.50 |
| Claude Haiku 4.5 | $1 | $5 |

Anthropic calls this its cheapest, fastest and most capable small model so far. The jobs it names are wonderfully unglamorous: summaries, context compactions, database queries, classification, customer support and browser use. It can also sit beneath [Opus 5.5](https://www.oguzhan.co/claude-opus-5-5-deep-dive/) or Sonnet 5.5 as a coding subagent. This is the first Haiku with adjustable effort. The same day, Anthropic halved Sonnet 5.5 cache reads to $0.10 per million tokens, which it says makes Sonnet about 20% cheaper on most agentic work.

The benchmark figures are Anthropic’s own. Haiku scored 72.4% on the OSWorld 2.1 offline subset, against 48.9% for Luna and 83.9% for Sonnet 5.5. On Terminal-Bench 4.0 it posted 39.2%, versus 16.4% for Luna and 70.6% for Sonnet. Anthropic itself says Sonnet and Opus remain the better choice for complex agentic coding.

My practical note is simple: keep small-agent prompts short. The hundreds of summaries, lookups and [compactions](https://www.oguzhan.co/jev-context-compaction-verbatim-prune/) around a hard model call just became much cheaper, but crossing 100K tokens changes the bill fast.

## 🧩 GPT-6 draws the interface as it answers

[GPT-6 started rolling out in ChatGPT](https://openai.com/index/gpt-6-for-everyone/) on 7 October for Plus, Pro, Business and Enterprise users. Free and Go users start getting it today. Paid plans use GPT-6 Sol in Chat, while Free and Go use GPT-6 Luna. Work and Codex models stay unchanged.

The visible change is Intelligent UI. GPT-6 can mix prose with graphics, tappable buttons, forms, charts and small tools, rendered progressively from native streamable components. OpenAI’s examples include a bill splitter, a retirement calculator, a retro game and a mapped road trip. Users can reduce the visuals, and OpenAI concedes that its design judgment still needs work.

Answers can now appear while the model is still thinking. According to OpenAI, GPT-6 Instant starts answering 44% sooner on average than GPT-5.6 Instant on questions that need web search. That is the company’s own figure. The safety pitch also builds on Astra work and claims stronger resistance to multi-turn jailbreaks.

For many people in Turkey, the free tier effectively is ChatGPT, and today that tier moves up a model generation. Generated mini-tools could be genuinely handy too, although I would check every number before letting an improvised calculator near money. My earlier [GPT-6 Sol and Luna deep dive](https://www.oguzhan.co/gpt-6-sol-luna-deep-dive/) has the wider pricing context.

## 💻 Microsoft puts local agents inside a sandbox

![A generic silver laptop with a translucent glass cube rising from it, a small glowing agent figure working on file cards inside while locked folders sit outside the boundary](https://oguzhanco.s3.eu-west-3.amazonaws.com/wp-content/uploads/2026/10/08061713/inline-local-agent-sandbox.webp)

*Microsoft Execution Containers (MXC), now generally available on Windows 11, let organizations set which files and networks an agent can reach, enforced at runtime.*

Microsoft’s pitch on the same Wednesday was [“hybrid intelligence”](https://blogs.windows.com/windowsexperience/2026/10/07/building-windows-for-hybrid-intelligence/): run agents locally where sensible, then call the cloud when needed. Its new Surface Laptop Ultra starts at $2,599, is available from 16 October, and pairs NVIDIA RTX Spark with up to 128 GB of unified memory and up to one petaflop of AI compute.

MAI Code 1.1 Flash has 137B total parameters with 6.8B active, quantized to about three bits, plus a 256K local context. Microsoft’s own Surface test peaked at 75.5 GB of memory at full context. It measured prompt processing at 923.5 tokens per second at 64K and 769.8 at 128K. Microsoft also says local calls carry no inference charge.

That memory figure is the useful reality check. This is specialist hardware paid for up front. The broader idea may be MXC, Microsoft Execution Containers, now generally available on Windows 11. Organizations can specify the files and networks an agent may reach, with those limits enforced at runtime. MXC is open source and works across operating systems.

Containment is where local agents become credible. Codex, GitHub Copilot, Replit, LM Studio and several others already support MXC, with Claude Code among the tools listed as coming. Microsoft is giving Windows an OS-level answer to the same permissions problem I examined in [Apple’s Full Disk Access rules for AI agents](https://www.oguzhan.co/apple-full-disk-access-ai-agents/).

## 🔍 SynthID Detector opens its doors

[Google has opened SynthID Detector](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/synth-id-ai-content/) globally at synthid.com. The English-language tool checks images, video and audio for SynthID watermarks from Google and partners OpenAI, NVIDIA and Kakao. Apple support is listed as coming soon.

According to Google, more than 180 billion images and videos, plus 240,000 years of audio, have received the watermark since 2023. Verification built into Search, Gemini and Chrome handles more than one million requests each day.

The public site is narrower. According to [Ars Technica](https://arstechnica.com/ai/2026/10/google-rolls-out-improved-synthid-ai-content-detector-now-available-globally/), sign-in is required, the daily quota is approximately ten checks, and the result no longer highlights marked regions as the internal tool did. Microsoft and Meta use their own watermarking standards, and detection tools can fail.

My reading of a clean result is deliberately limited: it says the checker found no SynthID mark. It cannot certify human authorship. Still, one place to check marks from several labs is useful for a journalist or anyone examining a viral clip. Ten daily tries make it a spot checker rather than a newsroom pipeline.

## 🧯 Fired researchers ask OpenAI to preserve visible reasoning

Three researchers fired by OpenAI last week, Tomek Korbak, Mikita Balesni and Jasmine Wang, have written to its board and safety committees. [The Wall Street Journal reports](https://www.wsj.com/tech/ai/fired-openai-researchers-ask-company-to-preserve-visibility-into-ai-reasoning-987c8c94) that they urged frontier labs to avoid developments that further reduce monitorability, use third-party auditors, and support an open safety ecosystem.

Their letter says the dismissals are “chilling those who remain at OpenAI.” OpenAI, which said last week it parted ways with them for violating its policies on handling sensitive company information, says the firings “were not about raising safety concerns or speaking out,” and an internal memo said it strongly agreed with the recommendations. The stakes are concrete: the GPT-6 Astra system card says Astra-class models could evade OpenAI’s chain-of-thought monitors under adversarial conditions, as [Gizmodo](https://gizmodo.com/3-fired-openai-employees-write-plea-for-chain-of-thought-monitoring-to-be-preserved-2000823349) points out. It is the same monitoring problem behind [OpenAI’s safety-case bar after Astra](https://www.oguzhan.co/gpt-6-1-astra-cancelled-safety-cases/). If both sides agree on the substance, I would like to see it written into the next system card as a commitment.

Wednesday looked like a pricing day wearing product-day clothes. Small calls reached ten cents, the free tier moved to GPT-6, and Microsoft offered local inference to buyers who can cover the hardware. The harder questions, whether we can inspect reasoning and verify synthetic media, moved a little as well. I will take the movement without pretending the questions are settled.
