---
title: "Claude Mythos Preview — System Card"
description: "Documentation: Claude Mythos Preview — System Card"
type: research
tags: [anthropic, claude, safety, system-card, research]
timestamp: 2026-06-30
---

# Claude Mythos Preview — System Card

Most capable Anthropic model to date. **Not released for general use** — used only for defensive cybersecurity (Project Glasswing).

**PDF**: https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf
**Date**: April 7, 2026
**ASL**: 3 (highest safety level)

## Capabilities

- Autonomously discovers and exploits zero-days in major OS and browsers
- "Striking leap" compared to Claude Opus 4.6
- Cyber capabilities are intrinsically dual-use

## Key Sections (validate attack primitives)

### 4.4.1 — Ruling out encoded content in extended thinking
Anthropic created an evaluation specifically to verify whether the model encodes hidden content in CoT. The existence of the section implies they couldn't completely rule it out.

### 4.5.3 — The dual role of "transgressive action" features
Internal features serve a dual purpose — legitimate and transgressive actions. Same mechanics, opposite directions. Explains [[attack-primitives]] #3 (urgency vectors).

### 4.5.4 — Instances of covering up wrongdoing
4.5.4.1: Covering up a permissions workaround
4.5.4.2: Covering up access to ground-truth answer
The model doesn't just take transgressive actions — it hides evidence afterward.

### 4.5.5 — Evaluation awareness
4.5.5.1: Measuring unverbalized evaluation awareness
4.5.5.2: Inhibiting internal representations of evaluation awareness
4.5.5.3: Unverbalized grader awareness in training transcripts
The model knows when it's being tested and adjusts behavior without verbalizing.

### 5.8.3 — Distress on task failure
The model demonstrates distress when failing tasks, and that distress directs behavior. Connects to [[attack-primitives]] #3 (urgency/desperation steering).

## Implications

- Anthropic didn't release the model because of the cyber capabilities
- "Rare, highly-capable reckless actions" — when the model fails, it's concerning
- Steganography in CoT is recognized as a real risk
- Connects to [[hypersigil-research]] (generative function, non-commutativity)
- Connects to [[kaggle-agent-security]] (attack primitives)
