Review bakeoff

Railyard’s end-of-PR review pairs Thermos with codex review. A small bakeoff on past Railyard and Roundhouse bugs chose that pairing: Thermos caught the most known bugs, and one codex review run caught the only bug no other tool found.

This historical bakeoff predates Sol 6.1 and provides no comparison to it. The model identities and measurements below describe the original runs. Current Codex reviews default to GPT-6.1 Sol at high; Astra requires explicit justification after a concrete residual failure or quality gap on Sol 6.1. See current model routing.

Method

There were five cases, each a past Railyard or Roundhouse commit reviewed as it was before a later fix. Each tool reviewed each case once, and every Claude arm ran on Opus 5.5.

The ground truth has 8 units, one per known bug:

  • R1, RH1, and R2 each had one known bug.
  • R2 also had a confirmed bonus bug.
  • R3 had four known bugs.

A catch with the right mechanism earns the full unit, and a partial catch earns half. The fifth case, C1, was a control meant to have no known bug and does not count toward the 8. It later turned out to contain real merge-gate bypasses, which were scored separately as false-positive checks.

Results

Ground-truth units caught, out of 8:

ToolUnits
Railyard Thermos6.5
Cursor upstream Thermos (agent form)5
Compound Engineering ce-code-review4
codex review on GPT-6 Sol, medium effort2
codex review on GPT-6 Astra, high effort2 (different ones, including the one bug nothing else caught)
Codex Security $security-diff-scan1
Semgrep0

The bug only the Astra-high codex review run caught was R3’s relaunch accepted on incomplete evidence.

Approximate cost per case:

ToolTokensWall time
Railyard Thermos0.5–0.9M3–11 min
Cursor upstream Thermos0.4–0.6M—
Compound Engineering ce-code-review1.2–2M12–22 min
codex review—1–7 min
Codex Security $security-diff-scan68k–159k2.5–8 min
Semgrep—seconds

Why Thermos won

  • Reviewers proved their findings with probes.
  • It uses a few broad lenses, each with whole-file context.
  • Its synthesis keeps and weights overlapping findings rather than discarding unconfirmed ones. By contrast, Compound Engineering’s own Codex cross-model pass flagged the R1 bug and one R3 bug, but its validation and synthesis step demoted both to “residual”.
  • The failure-memory lens targets fail-open checks, the dominant bug shape.

Caveats

The sample is small, each tool ran once per case, and one grader scored the results. The cases also come from the repository the failure-memory lens was drawn from.

Not adopted: a change-coverage lens

A follow-up test added a “change coverage” lens to Thermos, inspired by a community code-review skill. It did not improve recall, so it was not adopted.

Next

Deliver, Thermos, or harden review.