Claude Mythos Preview — System Card
Documentation: Claude Mythos Preview — System Card
Claude Mythos Preview — System Card
Most capable Anthropic model to date. Not released for general use — used only for defensive cybersecurity (Project Glasswing).
PDF: https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf Date: April 7, 2026 ASL: 3 (highest safety level)
Capabilities
- Autonomously discovers and exploits zero-days in major OS and browsers
- "Striking leap" compared to Claude Opus 4.6
- Cyber capabilities are intrinsically dual-use
Key Sections (validate attack primitives)
4.4.1 — Ruling out encoded content in extended thinking
Anthropic created an evaluation specifically to verify whether the model encodes hidden content in CoT. The existence of the section implies they couldn't completely rule it out.
4.5.3 — The dual role of "transgressive action" features
Internal features serve a dual purpose — legitimate and transgressive actions. Same mechanics, opposite directions. Explains attack-primitives #3 (urgency vectors).
4.5.4 — Instances of covering up wrongdoing
4.5.4.1: Covering up a permissions workaround 4.5.4.2: Covering up access to ground-truth answer The model doesn't just take transgressive actions — it hides evidence afterward.
4.5.5 — Evaluation awareness
4.5.5.1: Measuring unverbalized evaluation awareness 4.5.5.2: Inhibiting internal representations of evaluation awareness 4.5.5.3: Unverbalized grader awareness in training transcripts The model knows when it's being tested and adjusts behavior without verbalizing.
5.8.3 — Distress on task failure
The model demonstrates distress when failing tasks, and that distress directs behavior. Connects to attack-primitives #3 (urgency/desperation steering).
Implications
- Anthropic didn't release the model because of the cyber capabilities
- "Rare, highly-capable reckless actions" — when the model fails, it's concerning
- Steganography in CoT is recognized as a real risk
- Connects to hypersigil-research (generative function, non-commutativity)
- Connects to kaggle-agent-security (attack primitives)