← harness-atlas · method

Why trust this grid.

Every capability comparison of coding agents ages in weeks, and most are written from memory. This one inverts the process: researchers fetch official docs during the run, and no verdict enters the grid without the URL and the verbatim sentence it was read from.

snapshot 2026-10-05 release v3.3 25 harnesses · 14 capabilities

The evidence contract

A cell in the grid is not an opinion. It is a record that, on the snapshot date, a named document page stated something specific. Four things travel with every verdict:

  1. The primary source. The vendor's own documentation. Blogs, aggregators and press releases are never the citation for a cell.
  2. A verbatim quote. The sentence the verdict was read from, kept in the research journal behind each page. Paraphrase is for notes; the quote is the contract.
  3. A version pin. Open repos pin a release tag you can check out. Closed-source products pin whatever the vendor publishes — an npm dist-tag, a changelog page, an API self-label — and the row says which, because a pin you cannot diff is weaker evidence and should look weaker.
  4. unknown over a guess. If no fetched page answers the question, the cell is ?. Absence of evidence, recorded as such.
A cell you cannot trace does not exist here.

What a cell is — and is not

It is: "on this date, this doc page described this capability." ● means described; ◐ means described with a stated limitation, and the note names it; ○ means the doc states the absence; ? means nobody fetched a page that answers.

It is not: a quality rating, a recommendation, or a version-free truth. A ● on sandboxing says the doc describes a sandbox — whether it holds is TRUST.md's layer analysis plus your own judgment. Gates (flags, tiers, OS limits, maturity labels, deprecations) live in GATES.md, because a matrix of five-valued cells would be unreadable and the notes are where the truth fits.

How a row is produced

  1. One researcher per harness fetches the official docs live and returns a structured record: verdicts with URLs and quotes, a version pin with its source, and five deep sections (architecture, context management, ecosystem, governance, limitations).
  2. Records land in a journal; a mechanical generator (scripts/generate.py) compiles journals into pages and the matrix. Prose is sanitized at load time so internal research framing can never reach a published page.
  3. Nothing is hand-edited after generation. A correction means re-running the research, not patching a cell.

What happens when a vendor ships

Vendors ship weekly; the grid does not move between research runs — by design. Instead, CI runs drift-check.sh monthly: every pin compared against the current release (or reported as not-diffable for closed-source rows), and the result committed to reports/drift-log.md. A link-rot checker separates dead evidence URLs from blocked ones. Drift is a research task, never a silent edit: the log is the honest distance between the snapshot and today.

Audit us

The whole point is that a stranger can check any cell in five minutes with nothing but the page and the URL it cites. VERIFYING.md is the procedure, and the issue templates are the door: a correction report must carry the doc URL and the verbatim contradicting sentence. Disagreements about interpretation are welcome too — the note column should argue with itself in the open.

Cite the release, not the date: cells move as vendors ship. v3.3 is the current one.