Pokémon TCG — Neural Engine and Tokenization
As-built tensor streams, set-based tokenization, scratch memory, action pointers, auxiliary heads and the explicit delta to the future RoPEND/MoE design.
Pokémon TCG — Neural Engine and Tokenization
Role of the neural engine
The current policy is not a raw-observation rules learner. The deterministic engine, trackers and legal-action enumeration construct a typed belief state; the neural engine scores the resulting action surface. This boundary is central to interpreting both the model's strengths and its ceiling.
flowchart LR
O[Replay or live observation] --> T[GameTracker and ability state]
T --> C[Typed card and unit streams]
T --> M[Meta context and time buckets]
T --> Q[Legal option stream]
C --> X[Shared Transformer]
M --> X
Q --> X
X --> A[Pointer-style action scores]
X --> H[Auxiliary and prospective heads]
A --> G[Engine action]
G --> O
The architecture therefore learns a structured policy over an externally constrained action set. It is not evidence that the model internally simulates every game rule, and it does not make the future world-model or MoE phases current by association.
As-built streams
The neural-engine specification describes a model width of D=128 with typed streams that are concatenated or gathered into a single Transformer sequence. The exact names vary slightly across the monograph, manuscript and live code, but the stable conceptual components are:
| Stream | Information carried | Why it exists |
|---|---|---|
| CLS / global token | Global pooled context | Shared readout for action and auxiliary heads |
| Scratch registers | Learned recurrent workspace passed through TBPTT memory | A bounded place to retain temporal information not fully represented by the current row |
| Card stream | Card identity, zone, type, status and numeric attributes | Entity-level game state |
| Unit stream | Board positions, active/bench relationships and tracker state | Spatial and relational state |
| Meta stream | Turn, phase, player/opponent context and bucketized timing | Coarse temporal and strategic conditioning |
| Option stream | Legal actions, targets and action-specific features | Pointer-style selection over the actual action surface |
The current source path has a four-head Transformer with FFN width 512, scratch/memory inputs and outputs, and action plus auxiliary outputs. It does not contain an explicit RoPE or RoPE-ND module, a strategic MoE router, a vehicle-draft generator or an Apex runtime.
Tokenization is typed, not natural-language
The tokenization documents describe a set-based representation. Cards are embedded from structured identities and attributes, units are embedded from game-position roles, and options are represented as a finite legal-action alphabet. The representation is designed to keep illegal actions out of the policy surface and to make action selection a masked scoring problem.
This design has two consequences:
- It reduces the learning problem: the policy need not infer the complete legal-action grammar from token sequences.
- It creates a data and feature contract: any field used to construct an option or target must be available with the same semantics in replay compilation, training, validation and arena inference.
The would-KO signals are engine-derived prospective labels, not a claim that the neural network performed a full internal Monte Carlo simulation. The documented oracle uses engine state and bounded simulation to produce labels such as would_ko, prize impact and terminal-win implications. These labels can be powerful auxiliary supervision, but their provenance and denominator must remain separate from behavior-cloning loss.
Action encoder and pointer heads
The option stream is bucketed by action family and scored with a pointer-style mechanism. The documented action families include attacks, abilities, trainer/item/supporter actions, energy and retreat-like choices, with legality supplied by the engine. Multi-select decisions require autoregressive or split-submit handling rather than pretending that one categorical label describes the complete action.
The action architecture is therefore closer to:
shared state representation
|
+--> legal option keys and masks
|
+--> pointer logits over the current option set
|
+--> action-family / auxiliary readouts
This matters for evaluation. A model can have a reasonable token-level or option-level validation score while still choosing a poor deck-conditioned sequence in the engine. Tournament win rate is a downstream behavioral measurement, not a direct substitute for representation diagnostics.
Scratch registers and TBPTT
Scratch registers are intended as a learned workspace passed between sequential chunks. TBPTT truncates the gradient horizon while allowing a recurrent state or memory tensor to move through the trajectory. The training contract must account separately for decision chunks, row budgets, optimizer steps and memory isolation.
The documentation has a count discrepancy that must remain visible:
docs/neural_engine_and_tokenization_spec.mddescribes 16 learnable scratch tokens;- the integrated monograph and Antigravity state report describe 32 scratch positions/registers;
- the live code/configuration for each run is the authority for the exact tensor count used by that checkpoint.
This is not a reason to erase one value. It is a provenance issue: a specification revision, an experiment configuration or a report may be describing different stages. Future pages should cite the run-specific config whenever comparing scratch capacity.
The Stage 3 incident also changes how scratch memory should be discussed. The Antigravity report says the scratch workspace absorbed a gradient shock; the direct evidence establishes the loss-scale mismatch, not that scratch tokens alone contained or repaired it. Scratch memory is a plausible adaptation mechanism, not a post hoc proof of resilience.
Auxiliary and prospective supervision
The current head family includes behavior-cloning action prediction and auxiliary targets associated with terminal state, return, prize delta and would-KO-like prospective information. These heads are architecturally useful because they expose future research hooks without requiring a full world model. They also create multi-task optimization risk.
The safe contract is:
- each target declares its validity mask and denominator;
- the trainer logs valid rows and reduced loss in the same units used for optimization;
- the shared trunk's gradient contribution is measured per head;
- validation separates behavior, auxiliary calibration and arena outcomes;
- a future target is not allowed to leak information unavailable at the decision timestamp.
The Stage 3 postmortem records what happens when that contract is missing: auxiliary heads can look numerically healthy while their sparse gradients alter the shared representation in a way that is not visible to in-pool validation.
As-built versus target architecture
| Dimension | As-built boundary | Target or later blueprint |
|---|---|---|
| Position encoding | Zone/type structure and explicit gathers; no verified RoPE module | 4D or N-dimensional RoPEND over turn, meta-epoch, time remaining and Elo |
| Transformer topology | Single shared Transformer | Routed MoE specialists with a stable base and explicit routing diagnostics |
| Rating context | Static/bucketized metadata where present | Stochastic ephemeral Elo anchor and opponent proxy |
| Deck output | Policy acts on a packaged deck/runtime action surface | Autoregressive 60-card vehicle/deck draft as data augmentation and search |
| Long-horizon state | TBPTT scratch/memory | World-model or latent-rollout research after the baseline is stable |
| Runtime mode | Existing local arena and packaged inference path | Apex mode after the stated competition horizon; not implemented |
The target column is deliberately not a TODO list disguised as current code. It is the design lineage that explains why the current streams were kept extensible.
Primary sources
docs/neural_engine_and_tokenization_spec.md— tensor shapes, streams, heads and as-built/target table.docs/Pokemon_TCG_AI_Monograph.md— integrated chapters on tokenization, would-KO, pointer heads and scratch/TBPTT anomalies.docs/manuscript/03_tokenizer_and_epistemology.md— tokenization boundary and epistemic framing.docs/manuscript/05_action_encoder_and_pointer_heads.md— action stream and pointer design.docs/manuscript/06_scratch_registers_and_anomalies.md— scratch/TBPTT incident lineage.docs/dataset_compilation_and_oracle_pipeline.md— oracle and feature provenance.docs/architecture/01_ropend_theory.md— future positional encoding.docs/architecture/moe_pipeline_blueprint.md— future routing and vehicle draft.
Cross-references
- pokemon_tcg_agent_architecture — live current architecture page.
- pokemon_tcg_training_pipeline — trainer, TBPTT and validation path.
- pokemon_tcg_stage3_training_failure_postmortem — loss-scale and validation incident.
- pokemon_tcg_ropend_moe_blueprint — future architecture, data augmentation and Apex lineage.
- pokemon_tcg_would_ko_prospective_search — earlier prospective planner and RoPE-ND boundary.
- pokemon_tcg_prospective_v2 — sidecar experiment and what survived into the current heads.