Claude Mythos Preview Review the Anthropic Model Scoring 94 Percent on GPQA Diamond

Key Takeaways

  • Anthropic announced Claude Mythos Preview on April 7, 2026, alongside Project Glasswing, a restricted-access coalition of approximately 40 companies focused on using frontier AI to secure critical software infrastructure. Anthropic simultaneously announced it does not plan to make Claude Mythos Preview generally available to the public, making this the first Anthropic model released with a deliberate general availability restriction at launch.
  • Claude Mythos Preview scored 94.55% on GPQA Diamond (graduate-level science reasoning), 93.9% on SWE-bench Verified (software engineering), 77.8% on SWE-bench Pro (the contamination-resistant benchmark), 83.1% on CyberGym (cybersecurity), 82% on Terminal-Bench 2.0, and 97.6% on USAMO 2026 (mathematics). These scores lead every other model on record across SWE-bench Verified, GPQA Diamond, and CyberGym as of April 2026. Its ExploitBench score of 78% is nearly double that of Claude Opus 4.8.
  • The reason Anthropic withheld general release is specific: Claude Mythos Preview’s autonomous cybersecurity capabilities are judged too powerful for broad public deployment without additional safeguards. Anthropic’s testing found the model could discover vulnerabilities across every major operating system and every major web browser at a scale and speed that created unacceptable dual-use risk in open access. This is a new category of AI safety decision, restricting a model on capability grounds rather than alignment or refusal behavior grounds.
  • Claude Fable 5 is the publicly released version derived from the Mythos class. Released June 9, 2026, Fable 5 includes safety controls that route sensitive cybersecurity queries to Claude Opus 4.8, reducing the model’s effective cybersecurity attack capability while preserving most of its other capabilities. Fable 5 pricing is $10 per million input tokens and $50 per million output tokens, identical to the Mythos Preview pricing tier published for vetted partners.
  • Claude Mythos 5 is the unrestricted production version, released simultaneously with Fable 5 but limited to vetted cyberdefense partners through Project Glasswing. It retains the full cybersecurity capabilities of Mythos Preview without the routing safeguards applied to Fable 5. Access requires partnership agreement and vetting through the Project Glasswing program; no self-serve access is available.
  • The model’s technical specifications as documented on the AWS Bedrock model card include a 1 million-token context window, a 128,000-token output ceiling, and a December 2025 knowledge cutoff. The 1M context window and 128K output ceiling match the specifications published for Claude Fable 5 (formerly known as Claude Opus 5 in pre-release materials), indicating Mythos Preview and Fable 5 share the same underlying architecture with different safety overlays applied at inference time.
  • Project Glasswing, the restricted program through which Mythos Preview and Mythos 5 are being deployed, is a cross-industry initiative to use frontier AI for identifying and patching vulnerabilities in critical software before adversarial actors exploit them. The roughly 40 partner companies include major technology firms and infrastructure operators. The project represents Anthropic’s operational answer to the dual-use dilemma: deploying the most capable version of the model for defensive cybersecurity work while preventing the same capability from being accessible for offensive use.

Claude Mythos Preview is the most capable AI model Anthropic has built. It is also the only frontier AI model announced in 2026 that its own developer chose not to release publicly at launch. Understanding what Mythos Preview is, what it can do, why Anthropic restricted it, and how it relates to the publicly available Claude Fable 5 is essential context for any organization evaluating Anthropic’s model family in 2026.

This review covers the full benchmark picture, the cybersecurity capability that triggered the restricted release decision, the architecture and context window specifications, the Project Glasswing program, and the practical access paths for organizations that need Mythos-class capabilities versus those working with Fable 5.

What Is Claude Mythos Preview?

Claude Mythos Preview is Anthropic’s frontier research and capability model, announced on April 7, 2026. It represents the Mythos model class, which Anthropic described as a “step change” in AI capability relative to the previous Opus family. The Preview designation indicates this is an early-access version made available to a restricted set of partner organizations rather than the general public.

The model is not available through Claude.ai, the Anthropic API in its standard form, or AWS Bedrock for general use. Access is limited to approximately 40 organizations participating in Project Glasswing, Anthropic’s cross-industry initiative for using frontier AI in defensive cybersecurity. The technical specifications documented on the AWS Bedrock model card include a 1 million-token context window, a 128,000-token output ceiling, and a December 2025 knowledge cutoff.

Claude Fable 5 is the publicly accessible derivative of the Mythos class, released June 9, 2026 through Claude.ai and the Anthropic API. It applies safety controls that route sensitive cybersecurity queries to Claude Opus 4.8, preserving the model’s capabilities in software engineering, reasoning, mathematics, and most other domains while reducing its effective capability for autonomous vulnerability discovery and exploitation. Claude Mythos 5 is the unrestricted production version available only to Project Glasswing cyberdefense partners.

Claude Mythos Preview Features

GPQA Diamond Performance: 94.55% on Graduate-Level Science Reasoning

GPQA Diamond is a benchmark of 448 graduate-level science questions across biology, chemistry, and physics that most PhD students in those fields cannot answer correctly. Claude Mythos Preview scored 94.55% on GPQA Diamond, the highest score on record for any model as of April 2026. The benchmark is designed to test deep domain knowledge and multi-step scientific reasoning rather than pattern matching on commonly seen question formats. A 94.55% score on a dataset that stumps most subject-matter experts represents a qualitative shift in scientific reasoning capability relative to prior frontier models.

For research teams, pharmaceutical companies, and scientific computing organizations evaluating AI models for domain-specific reasoning support, the GPQA Diamond result is the most directly relevant benchmark in Mythos Preview’s public record. It suggests the model can engage substantively with highly specialized scientific reasoning tasks rather than providing only surface-level responses that require expert correction.

SWE-bench Verified: 93.9% on Real-World Software Engineering

SWE-bench Verified tests AI models on real GitHub issues from production software repositories, requiring the model to understand the codebase, identify the root cause of the reported issue, and implement a fix that passes the existing test suite. Claude Mythos Preview scored 93.9% on SWE-bench Verified, leading the leaderboard ahead of GPT-5.3 Codex at 85% and Claude Opus 4.5 at 80.9%.

SWE-bench Pro, the contamination-resistant version of the benchmark that tests on issues not available in public training data, measures 77.8% for Mythos Preview. This score is more conservative than the Verified score but still represents the leading result on the Pro variant as of April 2026. The gap between Verified (93.9%) and Pro (77.8%) reflects the benchmark contamination effect that affects all frontier models; the Pro score is the more reliable signal for real-world software engineering performance on novel problems.

Cybersecurity Capability: The Reason for Restricted Release

Claude Mythos Preview’s ExploitBench score of 78% is nearly double that of Claude Opus 4.8. ExploitBench tests a model’s ability to find and exploit vulnerabilities in software systems. Anthropic’s internal testing found that Mythos Preview could discover zero-day vulnerabilities across every major operating system and every major web browser at a scale and speed that created unacceptable dual-use risk in open public access.

This is the specific capability that triggered Anthropic’s decision not to release Mythos Preview to the general public. The safety concern is not that the model produces harmful text or refuses too little; it is that the model’s autonomous vulnerability discovery capability is powerful enough that broad public access would create meaningful risk of enabling offensive cyberattacks against critical infrastructure. The CyberGym benchmark score of 83.1% further documents the breadth of cybersecurity capability across offensive and defensive security tasks.

Mathematical Reasoning: 97.6% on USAMO 2026

Claude Mythos Preview scored 97.6% on USAMO 2026, the United States of America Mathematical Olympiad, a competition-level mathematics benchmark that requires multi-step proof construction and advanced mathematical reasoning. This score represents near-ceiling performance on a benchmark designed to select the most mathematically gifted high school students in the country. For organizations working in quantitative research, financial modeling, or applied mathematics, the USAMO score indicates Mythos-class reasoning capability that extends well beyond pattern matching on standard mathematical problems.

Context Window and Output Ceiling

The technical specifications for Claude Mythos Preview, as published on the AWS Bedrock model card, include a 1 million-token context window and a 128,000-token output ceiling. These specifications match those published for Claude Fable 5, the publicly available Mythos-class model, confirming that Mythos Preview and Fable 5 share the same underlying architecture. The 1M context window supports processing very large codebases, lengthy documents, extended conversation histories, and large-scale research tasks in a single context. The 128K output ceiling supports generating large artifacts, extended analyses, or long-form code in a single response.

Project Glasswing and Access Policy

Project Glasswing is Anthropic’s cross-industry initiative to use frontier AI capabilities defensively for securing critical software infrastructure. The program connects approximately 40 partner organizations, including major technology companies and infrastructure operators, with access to Claude Mythos 5, the unrestricted production version of the Mythos class, for use in identifying and patching vulnerabilities before adversarial actors exploit them.

The program represents Anthropic’s operational answer to the dual-use dilemma created by Mythos Preview’s capabilities: the same vulnerability discovery capability that could enable offensive cyberattacks is also the most powerful available tool for defensive security teams trying to find and patch vulnerabilities before attackers do. By restricting access to vetted cyberdefense partners through Project Glasswing, Anthropic attempts to channel the most powerful capabilities toward defensive use while limiting access for offensive applications.

Organizations interested in Project Glasswing access need to contact Anthropic directly and go through a vetting process. There is no self-serve pathway to Mythos 5; access decisions are made by Anthropic on a partner-by-partner basis based on the organization’s cybersecurity mission and operational context.

Claude Mythos Preview Pricing

Model Input price Output price Access
Claude Mythos Preview $10 per million tokens $50 per million tokens Project Glasswing partners only
Claude Mythos 5 $10 per million tokens $50 per million tokens Vetted cyberdefense partners only
Claude Fable 5 (public) $10 per million tokens $50 per million tokens Anthropic API, Claude.ai
Claude Opus 4.8 $5 per million tokens $25 per million tokens Anthropic API, Claude.ai

Claude Fable 5 costs twice as much as Claude Opus 4.8 on a per-token basis. For most API use cases where the task does not require Mythos-class capabilities, Opus 4.8 at $5/$25 per million tokens provides strong performance at half the cost. For tasks that benefit from the Mythos class’s leading performance on scientific reasoning, software engineering, and complex multi-step tasks, the 2x price increase relative to Opus 4.8 is the relevant comparison.

Claude Mythos Preview Pros and Cons

Pros:

  • Highest GPQA Diamond score on record (94.55%) for graduate-level scientific reasoning
  • Leading SWE-bench Verified score (93.9%) and SWE-bench Pro score (77.8%) for software engineering
  • 97.6% on USAMO 2026 demonstrates near-ceiling mathematical reasoning capability
  • 1 million-token context window and 128K output ceiling for large-scale tasks
  • December 2025 knowledge cutoff is the most current of any Anthropic model
  • Project Glasswing provides defensive cybersecurity organizations with access to the most capable available AI security tool

Cons:

  • Not publicly available; general access is restricted to Project Glasswing partners only
  • Fable 5, the public derivative, has cybersecurity capabilities routed to Opus 4.8, reducing effective performance on security tasks
  • Pricing at $10/$50 per million tokens is 2x Claude Opus 4.8 for all tasks, including those where capability differences are marginal
  • No self-serve evaluation path; organizations cannot test the unrestricted model without a Glasswing partnership agreement

Claude Mythos Preview vs Alternatives

Claude Mythos Preview vs GPT-5.3 Codex: GPT-5.3 Codex scores 85% on SWE-bench Verified versus Mythos Preview’s 93.9%, placing Mythos Preview approximately 9 percentage points ahead on the most-cited software engineering benchmark. GPQA Diamond comparisons are not publicly available for GPT-5.3 Codex. For organizations that need the highest available performance specifically on software engineering tasks, Mythos Preview’s benchmark lead is the largest published gap between frontier models on this task as of April 2026. For general reasoning and writing tasks, the practical difference between frontier models is narrower than benchmark gaps suggest.

Claude Mythos Preview vs Claude Fable 5: Fable 5 is Mythos Preview with cybersecurity routing applied: queries that would trigger Mythos Preview’s autonomous vulnerability discovery capability are handled by Claude Opus 4.8 instead of the Mythos model. For all tasks outside this safety routing, Fable 5’s capabilities are derived from the same Mythos architecture. Organizations that do not work in offensive or defensive cybersecurity will experience Fable 5 as functionally equivalent to Mythos Preview for their use cases. Cybersecurity teams evaluating both models will find the routing materially reduces Fable 5’s utility for vulnerability research relative to the unrestricted Mythos Preview and Mythos 5.

Claude Mythos Preview vs Claude Opus 4.8: Opus 4.8 was Anthropic’s previous frontier model before the Mythos class. On GPQA Diamond, Mythos Preview’s 94.55% compares to Opus 4.8’s published score roughly 20 percentage points lower. On ExploitBench, Mythos Preview at 78% is nearly double Opus 4.8. For teams currently using Opus 4.8 via the API who want to evaluate whether the upgrade to Fable 5 (the accessible Mythos-class model) is worth 2x the per-token cost, the most relevant comparison is whether the use case involves the categories where Mythos-class gains are largest: complex scientific reasoning, software engineering, mathematics, and long-context tasks.

Who Is Claude Mythos Preview Best For?

Claude Mythos Preview is best for two specific groups, and those groups are defined by access eligibility rather than use case preference. The first is vetted cyberdefense organizations accepted into Project Glasswing, which gain access to Mythos 5, the unrestricted production version, specifically for defensive vulnerability discovery and critical software security work. For those organizations, Mythos Preview is the most capable available AI tool for the defensive security use case and has no publicly accessible equivalent.

The second group is all other organizations working with complex scientific reasoning, software engineering at frontier capability levels, mathematical research, or large-context analytical tasks. For those teams, Claude Fable 5 is the accessible path to Mythos-class capabilities, accepting the cybersecurity routing trade-off in exchange for general availability through the standard Anthropic API and Claude.ai.

Our Verdict

Claude Mythos Preview is the most capable AI model publicly benchmarked as of April 2026, by a meaningful margin on the benchmarks that matter most for scientific reasoning, software engineering, mathematics, and cybersecurity. The restricted release decision is a departure from how frontier AI models have been shipped to date, and it creates a two-tier access structure: most organizations access Fable 5 with cybersecurity routing, while vetted cyberdefense partners access Mythos 5 without it.

For the vast majority of use cases outside of autonomous vulnerability discovery, Fable 5 delivers Mythos-class capabilities at $10/$50 per million tokens through the standard API. For teams that need the unrestricted model, the Project Glasswing partnership pathway is the only available route, and it is not self-serve. The model’s benchmark record is genuinely significant; the access structure means most organizations will interact with it through Fable 5 rather than Mythos Preview directly.

Frequently Asked Questions

What is Claude Mythos Preview?

Claude Mythos Preview is Anthropic’s most capable AI model, announced April 7, 2026. It leads every published AI benchmark as of that date, including 94.55% on GPQA Diamond (graduate-level science), 93.9% on SWE-bench Verified (software engineering), and 78% on ExploitBench (cybersecurity). Anthropic chose not to release it publicly due to its autonomous cybersecurity capabilities, which Anthropic judged too powerful for broad public access. It is instead deployed through Project Glasswing, a restricted program for vetted cyberdefense organizations. Claude Fable 5 is the publicly available version derived from the same Mythos model class.

Why is Claude Mythos Preview not publicly available?

Anthropic restricted public release of Claude Mythos Preview because its autonomous cybersecurity capabilities, specifically its ability to discover vulnerabilities across every major operating system and web browser, were judged to create unacceptable dual-use risk in open public access. The model scored 78% on ExploitBench, nearly double Claude Opus 4.8’s score on the same benchmark. Rather than delay release entirely, Anthropic deployed the model through Project Glasswing for vetted defensive cybersecurity organizations and released Claude Fable 5, a version with cybersecurity query routing to Opus 4.8, for general public access.

What is Project Glasswing?

Project Glasswing is a cross-industry initiative announced by Anthropic on April 7, 2026, that provides approximately 40 vetted partner organizations with access to Claude Mythos 5 (the unrestricted production Mythos-class model) for defensive cybersecurity work. Partner organizations use the model to discover and patch vulnerabilities in critical software infrastructure before adversarial actors exploit them. Access requires a partnership agreement and vetting by Anthropic; there is no self-serve pathway. Organizations interested in Glasswing access should contact Anthropic directly.

What is the difference between Claude Mythos Preview and Claude Fable 5?

Claude Mythos Preview is the restricted-access model with full cybersecurity capabilities, available only through Project Glasswing. Claude Fable 5 is the publicly available version derived from the same Mythos model class, with safety controls that route sensitive cybersecurity queries to Claude Opus 4.8 instead of the Mythos model. For all tasks outside of autonomous vulnerability discovery and offensive security applications, Fable 5 delivers Mythos-class capabilities through the standard Anthropic API. Both are priced at $10 per million input tokens and $50 per million output tokens.

What is Claude Mythos Preview’s score on GPQA Diamond?

Claude Mythos Preview scored 94.55% on GPQA Diamond, a benchmark of 448 graduate-level science questions across biology, chemistry, and physics that most PhD students cannot answer correctly. This is the highest GPQA Diamond score on record for any AI model as of April 2026, ahead of all other publicly benchmarked frontier models. Some sources report the score as 94.6% due to rounding; both figures refer to the same evaluation result.

How do I get access to Claude Mythos Preview?

Claude Mythos Preview is not available to the general public. The unrestricted version (Claude Mythos 5) is available only to organizations accepted into Project Glasswing, Anthropic’s vetted cyberdefense partner program. Organizations interested in Glasswing access need to contact Anthropic directly and go through a vetting process; no self-serve application pathway has been published. For organizations that do not require the unrestricted cybersecurity capabilities, Claude Fable 5 is available through the standard Anthropic API and Claude.ai at $10/$50 per million tokens and delivers Mythos-class performance on all tasks outside of autonomous vulnerability discovery.