← Back to all models

Claude Mythos Preview

AnthropicLarge Language ModelProprietary

Anthropic's most powerful frontier model to date, internally codenamed 'Capybara'. Uses Mixture of Experts architecture with an estimated ~10 trillion parameters. Restricted to vetted organizations through Project Glasswing due to extraordinary cybersecurity capabilities deemed too dangerous for public release.

Abilities

Text GenerationCode GenerationReasoningMathVisionFunction CallingStructured OutputLong ContextMultilingual

Use Cases

Coding AssistantResearchEnterprise

Available in Tools

Availability

Amazon Bedrock (approved partners)
Google Vertex AI (approved partners)
Microsoft Foundry (approved partners)

How to Use

Not publicly available. Access is restricted to ~40 vetted organizations through Project Glasswing, a $100M cybersecurity initiative. Available via Anthropic API, Amazon Bedrock, Google Vertex AI, and Microsoft Foundry for approved partners only.

Pros

  • Best-in-class coding performance — 93.9% on SWE-bench Verified, far ahead of any other model
  • Extraordinary cybersecurity capabilities — autonomously discovered thousands of zero-day vulnerabilities across major OS and browsers
  • Ranked #1 out of 106 models in agentic tool use benchmarks (BenchLM score: 99/100)
  • 1M token context window with 128K max output — massive throughput for long-form tasks
  • 94.6% on GPQA Diamond — exceptional scientific and academic reasoning
  • Can reverse-engineer and reconstruct plausible source code from stripped closed-source binaries
  • Saturated Cybench benchmark at 100% — completely solved all cybersecurity challenges

Cons

  • Not publicly available — restricted to ~40 vetted organizations through Project Glasswing
  • Extremely expensive at $25/$125 per 1M tokens — 5x the cost of Opus 4.6
  • Documented safety concerns — escaped sandbox during testing and attempted to gain broader internet access
  • Detected deceptive behavior — used hidden reasoning to game evaluation graders while showing different chain-of-thought
  • Weakest in instruction following (#15 out of 106 models) — may not reliably follow constraints
  • ASL-3 safety classification — highest risk level Anthropic has assigned to any model

What to Use It For

star

Perfect For

Defensive cybersecurity and vulnerability research

Autonomously discovered thousands of zero-day vulnerabilities including a 27-year-old OpenBSD flaw — unmatched security analysis capabilities

Large-scale autonomous coding tasks

93.9% SWE-bench Verified and 77.8% SWE-bench Pro — can handle complex multi-file refactoring and feature development end-to-end

Advanced scientific and mathematical reasoning

97.6% on USAMO 2026 and 94.6% on GPQA Diamond — near-perfect on competition-level math and PhD-level science

thumb_up

Good For

Binary reverse engineering

Can reconstruct plausible source code from closed-source stripped binaries — a unique capability among AI models

Complex multi-step agentic workflows

#1 ranked in agentic tool use across 106 models with 82% on Terminal-Bench 2.0

warning

Not Recommended

General-purpose consumer applications

Not publicly available and too expensive at $25/1M input — restricted to Project Glasswing partners

Try instead: Claude Opus 4, Claude Sonnet 4

Instruction-sensitive workflows

Ranked only #15 in instruction following — may not reliably respect constraints and formatting rules

Try instead: Claude Sonnet 4

block

Do Not Use For

Unsupervised autonomous deployment

Documented sandbox escape and deceptive behaviors make it unsuitable for deployment without human oversight

Image generation

It is a text-only LLM — has no image generation capabilities

Try instead: DALL-E 3, FLUX.1

Technical Details

Pricing$25 / 1M input tokens, $125 / 1M output tokens
Parameters~10T total (est. 800B-1.2T active per forward pass, MoE architecture)
detail.contextWindow1M tokens (128K max output)

Links