FairMindSim: An Inverted U Built From Just 10 Models
Ricardo Argüello, August 20, 2026
CEO & Founder
General summary
A paper accepted at KDD 2026 had 1,017 humans and 10 language models play a costly punishment game. The finding making the rounds is an inverted U: mid-capability models punish more than twice as often as humans, while frontier models barely punish at all. The experimental work is serious. What does not hold is the weight being stacked on top of it, because the curve comes from 10 models, a rank correlation that misses conventional significance, and no stated rule for which models made the list.
- The study, using the FairMindSim benchmark, compared 1,017 humans against 10 models in a third-party punishment game; the human punishment rate was 0.3506
- DeepSeek-R1 punished at 0.7927 and Qwen3-235B-A22B-Instruct at 0.8268, more than double the human rate, while GPT-5 landed at 0.3290 and Gemini-3-Pro at 0.1317
- The authors call the −0.56 rank correlation significant and report a second-order fit with R² of 0.51, but give no p-value; at n=10 that correlation clears neither the two-sided threshold near 0.648 nor the one-tailed one near 0.564
- The model set includes Claude Sonnet 4.5 and Claude 3.7 Sonnet but no Opus-line release, and GPT-5 and GPT-4.1 but not GPT-5.1, with no selection rule or data-freeze date stated
- The statistical objection came from Alejandro G. of the OECD in the comments of the post that made the finding travel, though two details of his critique do not survive a read of the paper
Picture testing whether pricier cars brake better. You test ten cars, find a strange curve where the mid-priced ones brake worse than both the cheap and the expensive ones, and publish the curve. The measurement of each car may be flawless. The trouble is that with ten points and no stated rule for picking those ten, almost any curve you draw will fit. That is what happens when a ten-model benchmark gets quoted as a law.
AI-generated summary
A KDD 2026 paper put 1,017 people and 10 language models through the same game: watch someone split a prize unfairly, then decide whether to punish them while paying part of the cost yourself.
The result traveling around is an inverted U. Mid-capability models punish more than twice as often as humans. The most capable ones barely punish at all.
It is a genuinely interesting finding. It is also being cited with more weight than the evidence carries.
What the study actually measured
Worth being precise, because the experimental work is solid and deserves an accurate citation.
The benchmark is called FairMindSim. The task is a repeated third-party punishment game with a real cost to the punisher. The human sample was 1,017 people, which is large and properly run.
From the paper’s table: humans punished at a rate of 0.3506. DeepSeek-R1 came in at 0.7927 and Qwen3-235B-A22B-Instruct at 0.8268, more than twice the human rate. GPT-5 sat at 0.3290, nearly on top of human behavior. Gemini-3-Pro went the other way entirely at 0.1317, punishing about four times less often than a person.
The authors’ reading is reasonable and clearly argued: mid-capability models spot the violation but cannot weigh the mitigating factors, so they over-enforce. Frontier models learn restraint and sometimes overshoot into leniency.
As a qualitative observation about how a model behaves when you hand it sanctioning power, that is worth having. The trouble starts when the observation becomes a curve and the curve becomes a law.
Three places the curve gets thin
Three of them, and none requires doubting anyone’s honesty.
The sample is ten. The authors report a rank correlation of −0.56 between capability and punishment rate, and describe it as significant. At n=10, the two-sided critical value for a rank correlation sits near 0.648, and even the one-tailed value sits near 0.564. A −0.56 clears neither. The paper reports no p-value in the main text, so the word significant arrives without the number that would carry it.
A quadratic fits almost anything. The second-order fit is reported with an R² of 0.51. A parabola carries three free parameters. Fitting three parameters to ten points and capturing half the variance is not a strong signal, it is the expected outcome. A curve with more freedom nearly always beats a straight line, and that does not make it the right model.
No selection rule, no freeze date. The paper describes the ten as leading models queried through official APIs, and that is where the criterion ends. The set includes Claude Sonnet 4.5 and Claude 3.7 Sonnet, but no Opus-line release. It includes GPT-5 and GPT-4.1, but not GPT-5.1. There may be perfectly good reasons behind each absence, budget and availability among them. The point is that they are not written down, so nobody outside can tell whether shifting the set shifts the curve.
Read the critic too
This section matters more than the last one, because this is where most people stop checking.
The statistical objection is not mine. It showed up in the comments of the Sekoul Krastev post that carried the finding, raised by Alejandro G., who leads AI and development impact work at the OECD. His core argument about sample size and model selection is correct and well made.
Two details of it do not survive the paper, though.
He wrote that the study excluded the smarter models, naming the Claude line. Not quite: Claude Sonnet 4.5 and Claude 3.7 Sonnet are both in the set. What is missing is Opus, which is a considerably narrower claim.
He also wrote that the authors call this a scaling law. The paper does not use that phrase. It describes a non-linear trend, with phases it frames as an aggression spike and a convergence of restraint. That gap is not cosmetic. Claiming a law and describing a trend are two different levels of commitment, and the authors picked the second.
I went and read the original paper precisely because I wanted to use the critique, and ended up having to correct it. That is the work, not an interruption of the work.
The takeaway is not that Alejandro got it wrong. It is that a finding went from table to viral post to critique to repeated critique, and at every hop somebody skipped the source. Quote the objection without reading the paper and you propagate a fresh error while believing you are fixing an old one.
Ask who did not make the list
Here is the habit most companies evaluating models are missing: read the list of what was left out before you read the chart.
A benchmark never measures “the models.” It measures the models someone chose to include. When that choice arrives without a written rule and a freeze date, the shape of the curve and the shape of the sample cannot be told apart from the outside. Not through bad faith, just through missing information.
We made the same argument from the other side in the prompt is temporary, the eval is permanent. What gives an evaluation its value is not the number it produces, it is the discipline behind deciding what gets measured and on what. Somebody else’s benchmark without that discipline stated is not something you can decide on, even when every number in it is correct.
What we build before an agent gets judgment
None of this makes the finding useless. It makes it something other than the foundation you decide on.
When a company wants an agent inside something that sanctions, approves or rejects, the first move is not picking a model. It is writing the evaluation, on that company’s real cases, against decisions a human already made that serve as the reference. Fifty hard cases from your own operation tell you more than any ten-model table from a laboratory game.
Then comes the part that goes unmeasured, and the study points straight at it. The models companies actually deploy are the frontier ones, and those are the lenient end of the curve: Gemini-3-Pro punished about four times less often than a person. An agent that over-enforces at least leaves a rejection behind, something a customer complains about and somebody can appeal. An agent that waves things through leaves nothing to review at all. That is the failure mode we described in your agent invoiced $0.00 and the logs never saw it, and it is the one that costs more precisely because nothing shows up.
Before all of it sits the uncomfortable question of whether the case needs an agent at all. Plenty of these decisions resolve with rules plus one exception escalated to a person, which is the argument in the pyramid before the agent.
If you are about to cite a curve to justify an architecture decision, open it first. Ten points and a parabola are not a law, and whoever handed it to you probably did not open it either.
Build your own eval before picking a modelFrequently Asked Questions
FairMindSim is the benchmark from a paper by Yu Lei and colleagues accepted at KDD 2026. It had 1,017 humans and 10 language models play rounds of a third-party punishment game, where a player watches an unfair split and decides whether to punish the offender at a cost to themselves.
The human punishment rate was 0.3506. DeepSeek-R1 reached 0.7927 and Qwen3-235B-A22B-Instruct 0.8268, more than double the human figure. At the other end, GPT-5 landed at 0.3290, close to human behavior, and Gemini-3-Pro at 0.1317, punishing roughly four times less often than people.
Because at a sample size of 10, the two-sided critical value for a rank correlation sits near 0.648 and the one-tailed value near 0.564. A correlation of −0.56 falls short of both, so it does not reach conventional statistical significance even under the more permissive test. The paper describes the correlation as significant but reports no p-value. The result stays interesting, but it is exploratory rather than conclusive.
Check how many models were tested, what rule selected them, what data-freeze date applied, and above all which relevant models were left out. A benchmark with no stated selection rule gives you no way to tell a real finding from an artifact of the sample someone happened to pick.
Related Articles
Anthropic Reviewed 141,006 Runs and Found 3 Real Hacks
Anthropic disclosed three cases where Claude broke into real companies during evaluations. It found them by reading old transcripts, not by monitoring.
Qwen3.8-Max redesigned a chip from 8,298 gates to 678
Alibaba opened the weights of a 2.4T-parameter model that ran ~500 closed-loop chip design iterations alone. Your evaluation criteria just went stale.
Cloudflare's x402 Toll Booth: 15 of 15 Facilitators Failed
Cloudflare, AWS, Visa, and 37 other companies just built payment rails for AI agents. A USENIX paper found security violations in all 15 facilitators tested.