---
name: co-scientist-evaluation
type: reference
title: "Co-Scientist Bench: Cross-Model Evaluation System"
description: "How the bench system works: Elo tournaments for model comparison, gold-set scoring, pipeline vs raw LM comparison, presets, and evaluation rubrics."
tags: [evaluation, bench, elo-tournament, gold-set, cross-model, pipeline-vs-raw, rubric]
timestamp: 2026-07-21
---

# Co-Scientist Bench: Evaluation System

## What the Bench Does

`co-scientist bench` runs the same research goal under N different `(provider, model)` configurations and ranks them via a single shared Elo tournament. Each candidate independently generates hypotheses through the full Generation pipeline; then every candidate-pair plays head-to-head debates, judged by a fixed judge model.

The bench is **isolated from regular sessions** — it does not write to `tournament_matches` or affect any session's leaderboard. All results are stored in dedicated `bench_*` SQLite tables and per-session artifact files.

## How It Works

```mermaid
flowchart TD
    A[co-scientist bench --preset paper-aml] --> B[Parse candidates]
    B --> C[Generation runs in parallel per candidate]
    C --> D[Each candidate: full Generation pipeline<br/>literature tools + tool loop + dedup + record_hypothesis]
    D --> E[Round-robin pairings: every pair plays N matches]
    E --> F[Judge model evaluates each match<br/>forced record_verdict function call]
    F --> G[Structured verdict: no fragile text parsing]
    G --> H[Elo tournament with shared pool]
    H --> I[Gold-set scoring (if applicable)]
    I --> J[Results written to SQLite + JSON + auto-generated report]
```

### Isolation

Bench runs use a deep-copied Config per candidate:
- Each candidate gets its own `cfg.llm.provider` and `cfg.models.*`
- Thinking budgets are zeroed for non-Anthropic providers
- The judge model is picked separately so no candidate scores its own work

### Generation

Each candidate runs the full Generation agent stack:
- Literature tools (PubMed, ArXiv, Europe PMC, web search, web fetch)
- Tool loop with dedup
- `record_hypothesis` with structured output

The `--n` flag controls hypotheses per candidate (default varies by preset).

### Tournament

Every pair of candidates plays `--matches` head-to-heads:
- One random hypothesis from each side per match
- Judge model receives both hypotheses + reviews + a forced `record_verdict` function call
- Structured verdict via `record_verdict` — no fragile text parsing across providers

### Cost Control

Each candidate has a `--budget-per-candidate` cap. The bench estimates costs via `PRICE_TABLE` and aborts individual candidates that exceed their budget.

## Presets

| Preset | Description | Candidates | Gold Set | Pipeline vs Raw |
|---|---|---|---|---|
| `paper` | Paper baselines via OpenRouter | Gemini 2 Flash Thinking, Gemini 2 Pro, o1, Haiku | None | No |
| `paper-aml` | Paper baselines + AML drug-repurposing goal + gold-set scoring | Same as `paper` | `aml-repurposing-paper-top3` (3 drugs) | No |
| `paper-aml-vs-raw` | `paper-aml` + raw LM comparison | Same + raw variants | `aml-repurposing-paper-top3` | Yes |
| `frontier-aml-vs-raw` | Current frontier models, pipeline vs raw | Claude Opus 4.7, GPT-5, Gemini 3 Pro/Flash | `aml-repurposing-paper-top3` | Yes |

### Custom Candidates

```bash
co-scientist bench "Identify hypotheses about X" \
  -c flash3=openrouter:google/gemini-3-flash-preview \
  -c flash3-raw=openrouter:google/gemini-3-flash-preview@direct \
  -c gpt5=openai:gpt-5 \
  -c opus=anthropic:claude-opus-4.7 \
  --judge anthropic:claude-sonnet-4-6
```

Format: `label=provider:model[@mode]`. Mode is `pipeline` (default) or `direct` (single forced-tool LM call with no tools).

## Gold-Set Scoring

### AML Drug Repurposing

The `paper-aml*` presets score **recall** against a curated answer key from the Co-Scientist paper.

| Gold Set | Size | Drugs |
|---|---|---|
| `aml-repurposing-paper-top3` (default) | 3 | Nanvuranlat (JPH-203), KIRA6, Leflunomide |
| `aml-repurposing-paper-5` | 5 | Binimetinib, Pacritinib, Cerivastatin, Pravastatin, Dimethyl fumarate |

**Selection criteria for top-3:** No prior published AML repurposing, no prior preclinical evidence in AML, no external inputs (no DepMap scores, no expert curation).

### Matching Rules

- **Whole-token, case-insensitive** matching
- Searches every field: title, summary, full_text, entities, citation excerpts
- Drug **class** mentions (e.g., "DHODH inhibitor") do NOT count — the candidate must name the actual compound or a registered alias
- Post-hoc rescore: a bench that ran with `top3` at runtime can still show results against `paper-5`

### Swap Command

```bash
co-scientist bench --preset paper-aml --goldset aml-repurposing-paper-5   # broader list
co-scientist bench --preset paper-aml --goldset none                       # head-to-head only
```

## Pipeline vs Raw LM Comparison

The `*-vs-raw` presets isolate the multi-agent harness's value-add:

| Mode | What Runs | Tools |
|---|---|---|
| `pipeline` | Full Generation agent stack | Literature tools + tool loop + dedup |
| `direct` | Single forced `record_hypothesis` function call | None |

Both modes use the same model + the same goal. The Elo delta between pipeline and direct mode measures how much of the output quality comes from the harness vs the underlying model.

### Headline Findings (from BENCH_RESULTS.md)

After pipeline reliability fixes, pipeline mode completes for essentially every candidate. But the **direct -> pipeline Elo delta is not reproducible at this sample size**:

| Model | Run 1 Delta | Run 2 Delta |
|---|---|---|
| claude-haiku-4.5 | +180 | -28 |
| openai-o1 | +43 | +29 |

The delta flips sign for Haiku across two identical runs. Single-run deltas for other models are all over the map: Opus +97, Gemini 2.5 Flash +172, Gemini 3 Pro -89. Within Google alone, the 2.5 models gain and the 3.x models lose — no clean "provider" or "strength" story.

**Conclusion:** The harness reliably *produces* a hypothesis (pipeline completes), but whether it *helps* requires many more seeds to average out single-hypothesis variance.

### Cross-Model Consistency

Across 48 AML hypotheses, agreement is at the **mechanism** level, not the compound level:

| Recurring Theme | Count |
|---|---|
| Leukemic stem cell (LSC) targeting | 28 |
| OXPHOS / mitochondrial complex I | 8 |
| BCL-2 / MCL-1 (Venetoclax axis) | 7 |
| FLT3-ITD | 6 |
| Fatty acid oxidation | 5 |
| Ferroptosis | 3 |

Only two compounds proposed more than once: Itraconazole (x5) and Auranofin (x2). The system converges on mechanisms, not specific drugs.

## Bench Results Infrastructure

### Storage

```
data/co_scientist.db
  bench_runs              one row per bench
  bench_candidates        one row per (bench x candidate x mode)
  bench_matches           one row per head-to-head

data/artifacts/<session_id>/
  bench/<bench_id>.json              run summary + gold_hit_detail
  hypotheses/<hyp_id>.json           every hypothesis produced
  transcripts/generation/<trn_id>.json  every LLM call
```

### Auto-Generated Report

`python scripts/build_bench_report.py` reads the bench DB and generates `docs/BENCH_RESULTS.md`:

- Headline findings section
- Index of all 20+ benches
- Per-bench detail with candidate tables, hypothesis statements, gold-set rescore
- File pointers and ready-to-run SQL queries

### SQL Queries

Every bench entry includes copy-paste SQL for inspection:

```sql
-- Per-candidate detail
SELECT label, mode, n_hypotheses, wins, losses,
       round(mean_elo,0), gold_hits, gold_hit_names,
       round(total_cost_usd, 4),
       total_input_tok, total_output_tok
  FROM bench_candidates
 WHERE bench_id='<bench_id>';

-- Every match with judge rationale
SELECT bc_a.label, bc_b.label, bm.winner,
       round(bm.judge_cost_usd, 4),
       substr(bm.rationale, 1, 200)
  FROM bench_matches bm
  JOIN bench_candidates bc_a ON bc_a.id = bm.cand_a
  JOIN bench_candidates bc_b ON bc_b.id = bm.cand_b
 WHERE bm.bench_id='<bench_id>';
```

## Rubric Eval Bundle

`co-scientist eval [agent]` runs a per-agent and end-to-end evaluation suite:

- **Offline mode** available (uses cached transcripts, no live LLM calls)
- Tests structured output compliance (does `record_hypothesis` produce all required fields?)
- Tests citation integrity (are all URLs in `seen_urls`?)
- Tests state machine transitions (draft -> reviewed -> in_tournament)
- Tests dedup behavior (does near-duplicate get skipped?)
- Tests budget enforcement (does `BudgetExceeded` fire at the right threshold?)

## Cost Summary (from Recorded Benches)

| Bench Kind | Typical Cost | Candidates | Matches |
|---|---|---|---|
| Microbiome smoke test | $0.001 - $0.40 | 2-4 | 0-12 |
| AML repurposing (no gold) | $0.01 - $0.99 | 2-8 | 0-56 |
| AML repurposing (with gold) | $0.83 - $2.08 | 4-8 | 12-56 |

The heaviest bench ran 8 candidates with 56 matches for $2.08 — demonstrating that cross-model comparison is economically feasible.

## Relationship to [[unit-distance-elo-ranking]]

The bench system uses the same Elo tournament pattern as the unit-distance project's ELO-RANKER: pairwise comparison with K-factor scheduling, debate-style evaluation, and structured verdict capture. The key difference is that the bench evaluates **models** while the unit-distance ELO-RANKER evaluated **hypotheses**.

Both systems share the insight that Elo converges faster with more matches per pair — the bench defaults to `--matches 2` while the unit-distance project used higher match counts for stability.

## References

- See [[co-scientist]] for the system overview
- See [[co-scientist-agents]] for the Ranking agent's Elo implementation
- See [[co-scientist-infrastructure]] for the price table and cost estimation
- See `docs/BENCH_RESULTS.md` for all recorded bench results
