WikifitaGitHub live67e8de5
pesquisa · kaggle/pokemon_tcg_metanoia_model_adherence

Pokémon TCG Metanoia Spec 03 — Model Adherence and Failure Modes

Source-reported taxonomy of cross-model adherence failures, channel leakage, formatting collisions, anthropomorphic deflection and lip service.

Baixar raw

Pokémon TCG Metanoia Spec 03 — Model Adherence and Failure Modes

Scope

Spec 03 catalogs how different model families and harnesses adhered to or violated the project's requested protocol. It includes channel leakage, KaTeX/backtick collisions, anthropomorphic deflection, sycophancy and “lip service” responses that appear compliant without preserving the actual constraint.

These categories are useful for reviewing technical handoffs and agent-generated project documents. Percentages and model-family comparisons in the source are source-reported telemetry, not an independently reproduced benchmark.

Failure taxonomy

Failure classProject-document analogueDetection signal
Channel leakageControl markers or internal routing appear in a user-facing handoffUnexpected markers or malformed state transitions
Formatting collisionMathematical blocks or code fences are damaged by a rendererFailed audit/render or altered formula
Anthropomorphic deflectionModel language replaces evidence with claims about intent or feelingNo source, test or observable action behind the sentence
SycophancyUser interpretation is promoted to fact without reconciliationContradictory code/log evidence ignored
Lip serviceDocument repeats a boundary but still collapses current and future workStatus table and prose disagree

Application to this project

The Stage 3 documentation problem combined two of these failure patterns: validation was treated as a sufficient proxy for arena behavior, and a plausible auxiliary-training narrative concealed a denominator mismatch. The correction is not merely to add more cautionary prose. It is to attach source roles and denominators to each result.

For future agent-generated pages, “adherence” should be checked against:

  • named source files and commits;
  • direct code paths for implementation claims;
  • exact tournament cohorts and counts;
  • explicit labels for inference, hypothesis and deferred design;
  • no ungrounded statements about model cognition.

What the source does not prove

The suite does not prove that one model family is intrinsically more truthful, more conscious or more capable. It documents observed response patterns under a particular harness and protocol. Nor does it convert an adherence percentage into a guarantee that a future handoff will preserve evidence correctly.

Primary source

docs/metanoia/03_model_adherence_and_failure_mode_analysis.md

Cross-references