Skip to content
Pairon

Pairon · private beta

Benchmarks

How Pairon measured project memory with Codex: 360 graded runs, the cost ratios, the method, the features that did not save money, and the limits.

Updated 27 September 2026

What was measured

On 23 September 2026, Codex with Pairon’s project memory solved the same pinned tasks for less than Codex alone. Cost per successful answer was 33% lower on Simulation 1, 22% lower on Simulation 2, and 19% lower on Simulation 3.

The run was Codex · gpt-5.6-luna · low effort, 360 paired runs. Every 90% confidence interval for those ratios was below 1. No significant difference in hand-graded answer quality on any set. Cost uses the harness’s fixed notional token prices, the same for both arms, not what Codex bills.

Source: Pairon-Backend docs/benchmarks.md, report BENCHMARK-2026-09-23-v3-codex-luna-low.md, dated 2026-09-23.

The three simulations

  • Simulation 1. Ratio 0.674 (33% lower cost per success). 90% CI 0.538–0.835.
  • Simulation 2. Ratio 0.778 (22% lower cost per success). 90% CI 0.619–0.944.
  • Simulation 3. Ratio 0.811 (19% lower cost per success). 90% CI 0.702–0.936.

One of the three repositories is not maintained by Pairon, so the result does not rest only on Pairon’s own code.

How a run is graded

  • Two identical copies. Each task runs on a plain checkout and on a Pairon checkout of the same pinned commit. Both reset before every run, and the arm order alternates.
  • Answers stay hidden. Benchmark material is excluded from both copies, so neither agent can read the expected answer.
  • People grade the work. Every automatic failure and a sample of passes are checked by hand: cited lines, completeness, exact locations, clean edits.
  • Cost per success. Total cost divided by the answers that passed grading, so a cheap wrong answer never counts as a saving.

Ratio means cost per successful answer with Pairon, divided by cost per successful answer without it. A number below 1 is a saving. A cheap wrong answer does not count as one.

What did not save money

These were measured on Codex and left off by default. Publishing only the memory result would hide them.

  • Hand-off package. Ratio 0.97 [0.87, 1.09]. Its facts were change notes rather than how the code works. Opt-in only.
  • Procedures. −14%, +20%, +14%. Missed the −30% target and stayed within run-to-run noise. Opt-in only.
  • Output compressor. +3%, +33%, +80%. On the three valid pairs; the test output here was too small to compress. Off by default.
  • pairon_ask tool. +10% cost. Codex never called it. No value on Codex.

What this does not show

  • One agent (Codex) and one model at low effort. Claude Code and OpenCode were not part of this run.
  • 60 paired tasks per repository bound cost well, but answer quality only to about ±10 points.
  • This measures one agent using Pairon’s project memory. It does not measure teams.

The team calculator on the homepage is a simulation. It is labelled as one. It is not this benchmark. How it works describes the product the measurement belongs to.

Request a spot

Pairon is letting teams in a few at a time. A request is not an account, and it does not connect a computer.