Anthropic's 'Honeycomb' Leak: Claude Opus 5 Spotted in Cursor With 1M Context and Fable-Tier Performance
Claude Honeycomb (Opus 5) leaked in Cursor with 1M context and xHigh reasoning. Full specs, benchmark projections, pricing rumors, and July 23 launch timeline.
On July 9, 2026, a developer opened Cursor IDE, scrolled through the model picker, and spotted something that should not have been there. Sitting alongside the familiar Claude Opus 4.8 and Sonnet 4.5 entries was a new name: Claude Honeycomb EAP. The description read: "Anthropic research model with per-turn controls and safety fallbacks. Early access preview. 1M context window. Version: extra high effort."
Within hours, it was gone. But the two prompts that went through before removal — plus a cascade of leaks from data-center filings, Vertex AI label strings, and Polymarket betting odds — have painted the clearest picture yet of Anthropic's next frontier model. That model, widely understood to be Claude Opus 5, could launch as soon as today.
This is the most significant Anthropic release since Fable 5 shipped in June. It arrives at a moment when the frontier AI leaderboard has never been more crowded: GPT-5.6 Sol, Kimi K3, Qwen3.8, and GLM-5.2 have all launched or announced within the past month. Opus 5 is Anthropic's answer to all of them — and the leaks suggest it will land closer to Fable-tier performance than anyone expected from a non-Mythos model.
This article assembles everything verifiable about Honeycomb / Opus 5: the leak timeline, confirmed specifications, performance projections, pricing rumors, and what the model means for developers, enterprises, and the broader AI market.
The Leak: How Honeycomb Appeared — and Vanished
The chain of events moved fast.
| Date | Event |
|---|---|
| July 9, 2026 | Claude Honeycomb EAP appears in Cursor IDE's model picker. Developer @chetaslua screenshots the listing and posts it on X. |
| July 9 (hours later) | Honeycomb is removed from Cursor. Only two prompts were processed before the entry disappeared. |
| July 9 | @chetaslua reports: hard prompts that Honeycomb could not handle were routed to Opus 4.8 as a fallback — confirming Honeycomb sits above 4.8 in Anthropic's capability hierarchy. |
| July 12 | @LuminaXspace publishes a detailed leak summary, adding that Honeycomb is "widely predicted to be Opus 5" and listing other Anthropic codenames: Capybara, Fennec, Numbat, KAIROS, ULTRAPLAN. |
| July 14 | Multiple leakers converge on a July 20–21 launch window. @Mr_Salio reports a 2M context variant may also exist. |
| July 19 | Anthropic tightens Fable 5 access on certain subscription tiers — a move widely interpreted as clearing capacity for Opus 5 deployment. DeepSeek V4 GA launches the same day. |
| July 20–21 | The rumored window passes without a launch. Polymarket odds shift. |
| July 22 | Polymarket shows 31% probability for a July 23 launch, 82% for July 31. @pankajkumar_dev reports Anthropic is targeting Thursday, July 23. |
| July 23 | Launch day target. Early indicators suggest the model may appear first on Vertex AI and in Cursor before hitting api.anthropic.com. |
The Cursor leak is especially significant because of precedent. Every previous Claude EAP model that appeared in Cursor — including early versions of Sonnet 4.5 and Opus 4.8 — shipped publicly within two to four weeks. If that pattern holds, Honeycomb's public release is imminent.
But a leak timeline only tells you when something might ship. The more pressing question is what it actually is — and whether the specifications justify the anticipation.
What Honeycomb / Opus 5 Actually Is
Unlike Fable 5, which was a step-change leap into Anthropic's top-tier architecture, Opus 5 represents the maturation of the Generation 5 platform. It is not a research prototype. The leak describes it as an "Anthropic research model" but the language around per-turn controls, fallbacks, and production-ready tooling suggests this is a shipping-grade product — not a concept demo.
Here is what the leak and corroborating sources confirm.
Confirmed Specifications
| Specification | Detail | Source |
|---|---|---|
| Model codename | Honeycomb (internal), Opus 5 (expected public name) | Cursor leak, multiple X sources |
| Context window | 1M tokens (2M variant possibly in development) | Cursor listing, @Mr_Salio, @LuminaXspace |
| Reasoning modes | Extra High Effort (xHigh), configurable per-request | Cursor listing, @adxtyahq |
| Safety architecture | Per-turn controls; hard cases fall back to Opus 4.8 | Cursor listing, @chetaslua test results |
| Architecture generation | Generation 5 (shared with Fable 5) | @ReadFuturist |
| Primary focus areas | Long-running autonomous agents, advanced coding, multi-step tool use, extended reasoning | Multiple sources |
The Safety Fallback Design
One of the most revealing technical details came from the two live prompts that went through Honeycomb before it was pulled offline. @chetaslua reported that prompts Honeycomb could not handle were automatically forwarded to Opus 4.8 as a fallback.
This is the inverse of most fallback architectures. In a typical system, a weaker model escalates to a stronger one. Here, Honeycomb escalates down — implying it sits above Opus 4.8 in Anthropic's internal capability ranking. Combined with the "per-turn controls" language, the fallback mechanism suggests Anthropic is building Opus 5 as a production-safe model that prioritizes reliability: when it encounters a prompt it cannot address with high confidence, it hands off rather than hallucinating.
For enterprise deployments where output quality is more important than latency, this architecture is a meaningful differentiator. It is the opposite of the "always respond" philosophy baked into most API endpoints.
Specifications describe what a model can do on paper. Benchmarks describe what it can do in practice. The gap between the two is where most deployment decisions are made — and lost.
Performance: Where Opus 5 Lands on the Leaderboard
No official benchmarks for Honeycomb / Opus 5 have been published. All performance projections come from leakers, inference-time behavior, and analyst modeling. The consensus across five independent sources is remarkably consistent.
Projected Performance Tier
| Source | Projection |
|---|---|
| @pankajkumar_dev | "Benchmark performance will be comparable to Fable 5, but won't surpass it." |
| @LuminaXspace | "Expected to push performance much closer to Fable and Mythos levels." |
| @goodworse | "This model will be very close to Fable 5 in quality and better than GPT-5.6." |
| @ReadFuturist | Opus 5 is positioned as the production-grade flagship below Mythos-class models. |
| @adxtyahq | "Significantly stronger coding and agent performance, with rumors placing it much closer to the latest frontier models." |
For context, here is where Opus 4.8 (the current Opus flagship) sits relative to the frontier:
| Model | Intelligence Index (Artificial Analysis) | Relative Position |
|---|---|---|
| Fable 5 (Max) | ~60 | #1 overall |
| GPT-5.6 Sol (Max) | ~59 | #2 |
| Kimi K3 | ~57 | #3 (best open-weight) |
| GLM-5.2 (Max) | ~51 | Best open-weights below Kimi |
| Opus 4.8 | ~44 (est.) | Current Opus ceiling |
If Opus 5 reaches Fable-adjacent performance as projected, it would represent a roughly 15-point jump on the Intelligence Index from Opus 4.8 — the largest single-generation improvement in the Opus line since Claude 3.5 Opus launched in early 2025.
Three specific capability improvements are consistently mentioned across leak sources:
-
Agentic coding: The model is expected to handle multi-file codebases, long-horizon debugging, and autonomous refactoring with significantly fewer errors than Opus 4.8. The 1M context window makes entire repository-scale prompts practical without chunking.
-
Multi-step tool use: Configurable reasoning effort means developers can request xHigh reasoning for complex function-calling chains and standard effort for simple completions — paying only for the compute they need per request, rather than a fixed premium for every token.
-
Extended autonomous runs: Several leakers mention Opus 5 is being designed for agents that run for hours or days, not minutes. The safety fallback architecture supports this: when an agent drifts into uncertain territory, it escalates to a known-safe model rather than compounding errors.
Raw performance decides whether a model gets attention. Pricing decides whether anyone actually builds on it. For Opus 5, the rumored price point may matter more than the benchmark scores.
Pricing: The Strategic Lever
No official pricing has been published, but the strategic context makes the direction clear.
Anthropic is under pressure from three directions:
- OpenAI released GPT-5.6 Sol at aggressive pricing and is gaining B2B traction through Codex, which now has over 10 million paying users.
- Moonshot AI released Kimi K3 as an open-weight model (free to self-host) with performance above Opus 4.8. Its API pricing is approximately $0.94 per task — roughly half the cost of equivalent US models.
- Alibaba shipped Qwen3.8 at 2.4T parameters, open-weight, second only to Fable 5 on internal benchmarks.
In this environment, Opus 5 cannot launch at Fable 5 prices and compete. The leaked positioning — "comparable to Fable 5, but won't surpass it" — suggests Anthropic is deliberately creating a clear price-performance tier below Fable and Mythos.
@goodworse summarized the thesis: "This model will be significantly cheaper than Fable. The price-to-quality ratio will be one of the best on the AI market."
If that holds, Opus 5 could occupy the same strategic slot that Claude 3.5 Sonnet held in 2024: not the absolute best, but the best value at the frontier. For most production workloads, that is a larger market than the absolute performance crown.
Subscription Impact
Fable 5 access was tightened on Max-tier subscriptions starting July 19 — the same day DeepSeek V4 GA launched. Anthropic framed this as capacity management, but the timing with Opus 5's rumored launch window is difficult to dismiss as coincidence.
The most likely scenario: Opus 5 becomes the default high-capability model on Max and Team plans, with Fable 5 reserved for Premium/Enterprise tiers or usage above certain rate limits. This would mirror the tiering strategy Anthropic used when Fable 5 first shipped with Opus 4.8 as the "included" flagship.
Developers currently relying on Opus 4.8 should expect Opus 5 to slot into the same API endpoint (claude-opus-5-*) at a similar or modestly higher price point — not Fable-tier pricing. The dramatic 15-point capability jump from Opus 4.8 to Opus 5 suggests the value proposition is to make Opus 4.8 obsolete overnight rather than to charge a premium.
Pricing strategy is one lever. Timing is the other — and Opus 5 is not shipping into a vacuum. To understand why this launch matters, you have to look at what every competitor is doing right now.
Strategic Context: Why Opus 5 Matters Now
The timing of Opus 5 is not arbitrary. Every major frontier lab has shipped or announced a new model in July 2026:
| Lab | Model | Date | Key Detail |
|---|---|---|---|
| OpenAI | GPT-5.6 Sol | July 2026 | Matches Fable 5 on most benchmarks |
| Moonshot AI | Kimi K3 | July 2026 | 2.8T params, open-weight (free July 27) |
| DeepSeek | V4 GA | July 19 | 80.6% SWE-bench, peak-valley pricing |
| Alibaba | Qwen3.8 | July 19 | 2.4T params, open-weight, Fable-adjacent |
| Zhipu AI | GLM-5.2 | June/July 2026 | Best open-weight model before Kimi K3 |
| Anthropic | Opus 5 (Honeycomb) | July 23 (projected) | Fable-adjacent at lower cost |
Anthropic's competitive position is unique. It holds the absolute performance crown with Fable 5 and Mythos, but the gap between those models and the rest of the market is shrinking fast. Kimi K3 is open-weight and free to self-host. Qwen3.8 is open-weight. DeepSeek V4 GA is MIT-licensed. Every month that passes without a new Anthropic model below the Fable tier costs ground to competitors who offer 80-90% of the quality at 10-50% of the price.
Opus 5 is the counterweight. If it can deliver Fable-adjacent performance at Sonnet-tier pricing — or even Opus-tier pricing that undercuts GPT-5.6 Sol — it becomes the default choice for the vast majority of production workloads that do not require absolute frontier reasoning.
The broader market context is worth noting. Bank of America told clients that Kimi K3 proves "Chinese labs can keep making big leaps even with limited chips." The pressure from open-weight models is structural, not cyclical. Every US frontier lab now faces the same question: if a free model does most of what your paid model does, what is the value of the paid model?
Opus 5's answer appears to be: production safety, per-turn control, and enterprise-grade reliability — attributes that self-hosted open-weight models struggle to match at scale.
Strategy is useful context, but most developers reading this want one thing: a checklist they can act on in the next ten minutes.
What Developers Should Do Right Now
If you build on Claude APIs or use Cursor / Claude Code, here is the practical checklist:
-
Monitor
api.anthropic.comand Vertex AI model lists. Opus 5 will likely appear on Vertex AI first (as previous Claude models have) before the general API. Cursor integration typically follows within days. -
Do not hard-code model strings. If your application pins
claude-opus-4-8-*, prepare a fallback path toclaude-opus-5-*or a model-family wildcard. Anthropic's deprecation timelines for Opus 4.8 are not yet announced, but Opus 4.5 was sunset within weeks of 4.8 shipping. -
Warm up a backup model. During the Fable 5 launch, Anthropic's API experienced significant rate-limiting as demand spiked. Keep DeepSeek V4 Pro, Kimi K3 API, or GPT-5.6 Sol warmed as a fallback if Opus 5 capacity tightens on launch day.
-
Audit your system prompts. Anthropic engineers have recommended lean system prompts for Generation 5 models — removing examples and hard constraints that can confuse the model when they collide with user instructions. The same advice will almost certainly apply to Opus 5.
-
Budget for per-request reasoning costs. If Opus 5 follows the same pattern as Fable 5's configurable reasoning effort, costs will scale with the reasoning level selected. Standard effort may be 2-3x cheaper than xHigh effort for the same prompt. Profile your workload before committing to a default effort tier.
Rule of Thumb: Bet on Opus 5 for the next 6–12 months, but never bet on a single model. Keep an open-weight fallback (Kimi K3 or DeepSeek V4) warmed at all times. The gap between proprietary and open-weight models shrinks weekly, and the team that ships faster is the team that can switch faster.
The 60-Second Validation Test
Before you migrate anything, open a terminal and run one prompt through whatever Opus 5 endpoint appears first (Vertex AI model list or Cursor model picker). Do not test with "hello world." Use a real prompt from your production workload — ideally one that Opus 4.8 got wrong or answered slowly. If Opus 5 handles it correctly and fast, migration is worth the effort. If it does not, you have lost 60 seconds and learned exactly where the ceiling sits for your use case.
Expert Pitfall: A common mistake during model transitions is comparing old-model-xHigh to new-model-standard and concluding the new model has regressed. Generation 5 models respond non-linearly to effort levels — a standard-effort Opus 5 prompt can look noticeably worse than an xHigh-effort Opus 4.8 prompt, even though the underlying model is stronger. Profile Opus 5 at xHigh reasoning effort against Opus 4.8's best-effort output before making any performance comparison. Match the effort tier to the task before drawing conclusions.
Frequently Asked Questions
What is Claude Honeycomb?
Claude Honeycomb EAP is the internal codename and early-access preview name for Anthropic's Claude Opus 5 model. It briefly appeared in Cursor IDE's model picker on July 9, 2026, before being removed. The leak described it as an "Anthropic research model with per-turn controls and safety fallbacks."
When will Claude Opus 5 be released?
Multiple leakers point to Thursday, July 23, 2026 as the target launch date, with Polymarket showing 31% probability for that date and 82% probability for a launch before July 31. Anthropic has not made an official announcement. If the July 23 window passes without a launch, the next likely dates are July 24 or July 30-31.
How does Opus 5 compare to Fable 5?
Opus 5 is projected to deliver benchmark performance comparable to Fable 5, but not surpassing it. It is positioned as a production-grade alternative to Fable 5 at a lower price point, with particular strength in coding, agentic workflows, and multi-step reasoning. Fable 5 and Mythos remain Anthropic's absolute-performance leaders.
How does Opus 5 compare to GPT-5.6 Sol?
Both models represent the frontier just below the absolute performance ceiling (Fable 5 / Mythos). Opus 5's per-turn safety controls and 1M+ context window are architectural differentiators. GPT-5.6 Sol has the advantage of being already available and battle-tested in production. If Opus 5 ships at a lower price than GPT-5.6 Sol, it could shift developer preference rapidly.
What is the context window for Opus 5?
The Cursor leak confirms a 1M token context window. Some sources report a 2M variant may exist, likely for enterprise or Vertex AI deployments. At 1M tokens, Opus 5 can ingest entire codebases, multi-hundred-page documents, or days of agent conversation history in a single prompt.
Will Opus 5 be open-source?
No. Anthropic has not released any Claude model as open-weight. Opus 5 will be available through Anthropic's API (api.anthropic.com), Google Cloud Vertex AI, Amazon Bedrock, and Cursor IDE. Unlike Kimi K3 and Qwen3.8 (both open-weight), Opus 5 is a proprietary, API-access-only model.
How much will Opus 5 cost?
Official pricing has not been released. Based on leaks and strategic positioning, Opus 5 is expected to price below Fable 5 — likely in the $3-$8 per million output tokens range (compared to Fable 5's estimated $15-$36/M tokens and Opus 4.8's $15/M tokens). The configurable reasoning effort will likely create a tiered cost structure within the model itself.
What are the other Anthropic codenames?
Beyond Honeycomb (Opus 5), leaked Anthropic codenames include: Fennec (possibly Sonnet 5), Numbat (unknown model in testing), Capybara (unknown), KAIROS (possibly a persistent background agent), and ULTRAPLAN (possibly an extended planning and reasoning mode). None of these have been officially confirmed.
What happens to Opus 4.8 when Opus 5 launches?
Based on Anthropic's historical pattern, Opus 4.8 will likely remain available for a transition period of 2-4 weeks, then be deprecated. Developers should prepare to migrate model strings and test prompt compatibility with Opus 5 as soon as the model is available.
These questions cover the tactical details. The bigger picture is what matters next.
Core Summary
The Claude Honeycomb leak is a rare window into how frontier AI labs stage their releases — from internal codenames, to accidental IDE exposure, to capacity-clearing maneuvers on subscription plans, to coordinated launch-day rollouts across cloud providers.
- The leak is credible. Cursor IDE model picker exposure, two live prompts processed, and a safety fallback to Opus 4.8 confirm Honeycomb is a real, production-grade model in final pre-release testing.
- The performance is projected to be Fable-adjacent. Five independent sources converge on Opus 5 landing near Fable 5 in benchmarks — a ~15-point jump from Opus 4.8 and the largest single-generation Opus improvement in over a year.
- The pricing is the strategic play. Opus 5 is not being positioned to beat Fable 5. It is being positioned to make Opus 4.8 obsolete and to compete with open-weight models (Kimi K3, Qwen3.8) on value, not on absolute performance.
- The launch window is today. Polymarket odds (82% by July 31) and historical Cursor EAP-to-launch timelines (2–4 weeks) both point to an imminent release. July 23 is the most likely date.
If Opus 5 delivers on the projected performance at the rumored price point, July 2026 will be remembered as the month the frontier AI market fundamentally restructured. A model with Fable-adjacent reasoning, a 1M-token context window, and production-grade safety controls — priced below the absolute frontier — resets the value equation for every team building on language model APIs.
For developers, the practical takeaway is simple: the model that ships this week could be the one you build on for the next six months. Your fastest move: open a terminal, pull the model list from api.anthropic.com, and run your hardest production prompt through Opus 5 the moment it appears. You will know in 60 seconds whether migration is worth it.
Last updated: July 23, 2026. This article will be updated with official benchmarks, pricing, and access details as Anthropic makes its public announcement.
Explore more AI model coverage on wan27.org — from DeepSeek V4 GA benchmarks to Kimi K3 explained and the latest in open-weight model releases.
Author
Categories
Seedance 2.0
ByteDance latest video model. Text & image to video, up to 1080p.
Try Seedance 2.0 →Wan Video
Wan 2.7 series — text, image, reference to video & video editing.
Try Wan Video →AI Image Generator
Nano Banana Pro, GPT Image 2 & more. Generate stunning images in seconds.
Try Image Generator →More Posts

Wan 2.7 Text-to-Image Pro: Up to 4K AI Image Generation With Thinking Mode
Wan 2.7 Text-to-Image Pro generates images up to 4K resolution from text prompts with thinking mode, superior text rendering, and magazine-cover quality. Generate directly at wan27.org.

25 Wan 2.7 Prompt Templates (Text-to-Video + Image-to-Video)
Copy-friendly Wan 2.7 prompt templates for T2V and I2V: camera moves, motion patterns, product shots, portraits, and cinematic scenes — plus how to customize them without breaking output quality.

Wan 2.2 vs Wan 2.7: Which One Should You Use on wan27.org?
A practical Wan 2.2 vs Wan 2.7 comparison using the actual workflows available on wan27.org, including modes, resolution, clip length, pricing, and when each model makes sense.
Newsletter
Join the community
Subscribe to our newsletter for the latest news and updates