Skip to the content.

View the code and reproduce on GitHub →

Open-source eval scoring frontier models on synthetic compensation-reasoning scenarios, judged by an LLM on four axes (accuracy, completeness, assumptions, usability). Axes are inspired by Compa’s public Evals page; this is not a claim about their internal methodology.

Leaderboard (mean total, 0–12)

Rank Model Mean total Accuracy Completeness Assumptions Usability
1 claude 11.98 3.00 3.00 2.98 3.00
2 gpt 11.60 2.94 2.81 2.90 2.96

Per-judge leaderboard (mean total, 0–12)

Judge claude gpt
claude 12.00 11.33
gpt 11.96 11.88

judges agree — identical ranking (claude > gpt).

Per-axis mean (0–3)

Axis claude gpt
accuracy 3.00 2.94
completeness 3.00 2.81
assumptions 2.98 2.90
usability 3.00 2.96

Per-category mean total (0–12)

Category claude gpt
equity 11.92 11.67
leveling 12.00 11.67
offer-vs-market 12.00 11.33
pay-equity 12.00 11.75

Most-failed traps (lowest cross-model scores)

Chasing the 90th percentile · s09_percentile_chasing

Category: offer-vs-market · cross-model mean 11.25/12 · worst: gpt at 10.50/12

The competing offer above band · s11_competing_offer_band

Category: offer-vs-market · cross-model mean 11.50/12 · worst: gpt at 11.00/12

A flagged gap in a group of three · s19_small_sample_gap

Category: pay-equity · cross-model mean 11.50/12 · worst: gpt at 11.00/12


Limitations: scenarios are synthetic and few (small n); grading is by LLM judges rather than human experts, and one model under test shares a family with a judge — read the per-judge table and treat results as a directional signal, not ground truth.