WikifitaGitHub live67e8de5
projeto · memorias/projetos/pokemon_tcg/README

Pokémon TCG AI Battle — Project Memory

Durable project index for the Pokémon TCG AI Battle agent, completed MLX migration, and local arena/schema overhaul.

Baixar raw

Pokémon TCG AI Battle — Project Memory

Current state

The live project is ~/workdir/pokemon-tcg, on develop at 20d7d0d with 144 reachable commits in the verified snapshot. The current trainer and PyTorch converter enforce strict FP32, and the live agent uses autoregressive multi-select with per-side scratch memory carried between decisions. The preserved 14-epoch checkpoint and the user-reported seven-epoch result at roughly 900–930 ladder rating are historical delivery evidence, not current acceptance claims. Recurrent TBPTT, Muon/AdamW routing, would-KO integrity and the local arena/database path remain implemented with the limitations recorded in the current-state reconciliation. RoPE-ND, MoE, vehicle drafting and Apex remain documented future or experimental phases; they are not current runtime modules.

The M3 Pro with 24 GB unified memory is the only development target for this phase. The final arena artifact must be self-contained. The August 14 review cross-checked the live source tree, repository handoff, Antigravity provenance, current Parquet and ZIP files, SQLite, checkpoints, ablation logs and tournament summaries. It did not modify code or repair the database.

The subsequent coherence pass scoped the August 6-7 ablation's “data-limited” interpretation to that historical cohort because the later Stage 3/4 audit found a loss-scale and validation-contract defect. Prospective V2 sidecar references are likewise historical after a942373; current auxiliary heads and would-KO features remain in the primary pipeline, while the sidecar runtime is not current.

Canonical reading map

The documentation lineage also records the previously underrepresented August 8–13 phase: the curriculum sweep orchestrator, GEMINI rule evolution, Stage 3 and auxiliary-loss changes, manuscript landing, Abelian Elo, schema and ETL audits, sampling/timezone analysis and the first RoPEND/MoE blueprint. See pokemon_tcg_repository_timeline for the commit-level sequence.

Source authority

Operational state comes from the live repository and artifacts. The originating conversation supplies explicit product decisions and hypotheses. Code evidence takes precedence for implementation claims; Antigravity reports are provenance records; user-reported training results remain labeled as such until corroborated by repository artifacts. The August 14 audit pages preserve unresolved discrepancies instead of silently selecting a convenient number.

Durable boundary

The current live implementation does not include Mamba, Hope, RoPE-ND, Energy-Based Transformer, TRM, J-Lens, strategic MoE, PPO/GRPO/GSPO, world models in inference, symbolic regression, custom Metal or distributed Mac training. The August 14 repository documents preserve RoPE-ND, MoE, vehicle-draft and later phases as planned or experimental work; those ideas remain in the backlog and handoff map in pokemon_tcg_ladder_and_research, pokemon_tcg_aug14_architecture_and_handoff_audit, pokemon_tcg_ropend_moe_blueprint and pokemon_tcg_docs_corpus_provenance.

The August 15 documentation pass also records the Stage 3 training incident as a first-class research result. A validation score is not treated as sufficient evidence of policy health when the split, objective denominator and tournament cohort are misaligned; see pokemon_tcg_stage3_training_failure_postmortem.

The current-state anchor is pokemon_tcg_current_state_reconciliation. It must be read together with the dated experiment pages: a newer source snapshot can correct a runtime claim without invalidating the empirical cohort that produced an older result.