Skip to main content

FairMindSim: An Inverted U Built From Just 10 Models

A KDD 2026 paper says mid-tier models punish twice as hard as humans. The curve rests on 10 models and a correlation that misses conventional significance.

FairMindSim: An Inverted U Built From Just 10 Models

Ricardo Argüello

Ricardo Argüello
Ricardo Argüello

CEO & Founder

Business Strategy 6 min read

A KDD 2026 paper put 1,017 people and 10 language models through the same game: watch someone split a prize unfairly, then decide whether to punish them while paying part of the cost yourself.

The result traveling around is an inverted U. Mid-capability models punish more than twice as often as humans. The most capable ones barely punish at all.

It is a genuinely interesting finding. It is also being cited with more weight than the evidence carries.

What the study actually measured

Worth being precise, because the experimental work is solid and deserves an accurate citation.

The benchmark is called FairMindSim. The task is a repeated third-party punishment game with a real cost to the punisher. The human sample was 1,017 people, which is large and properly run.

From the paper’s table: humans punished at a rate of 0.3506. DeepSeek-R1 came in at 0.7927 and Qwen3-235B-A22B-Instruct at 0.8268, more than twice the human rate. GPT-5 sat at 0.3290, nearly on top of human behavior. Gemini-3-Pro went the other way entirely at 0.1317, punishing about four times less often than a person.

The authors’ reading is reasonable and clearly argued: mid-capability models spot the violation but cannot weigh the mitigating factors, so they over-enforce. Frontier models learn restraint and sometimes overshoot into leniency.

As a qualitative observation about how a model behaves when you hand it sanctioning power, that is worth having. The trouble starts when the observation becomes a curve and the curve becomes a law.

Three places the curve gets thin

Three of them, and none requires doubting anyone’s honesty.

The sample is ten. The authors report a rank correlation of −0.56 between capability and punishment rate, and describe it as significant. At n=10, the two-sided critical value for a rank correlation sits near 0.648, and even the one-tailed value sits near 0.564. A −0.56 clears neither. The paper reports no p-value in the main text, so the word significant arrives without the number that would carry it.

A quadratic fits almost anything. The second-order fit is reported with an R² of 0.51. A parabola carries three free parameters. Fitting three parameters to ten points and capturing half the variance is not a strong signal, it is the expected outcome. A curve with more freedom nearly always beats a straight line, and that does not make it the right model.

No selection rule, no freeze date. The paper describes the ten as leading models queried through official APIs, and that is where the criterion ends. The set includes Claude Sonnet 4.5 and Claude 3.7 Sonnet, but no Opus-line release. It includes GPT-5 and GPT-4.1, but not GPT-5.1. There may be perfectly good reasons behind each absence, budget and availability among them. The point is that they are not written down, so nobody outside can tell whether shifting the set shifts the curve.

Read the critic too

This section matters more than the last one, because this is where most people stop checking.

The statistical objection is not mine. It showed up in the comments of the Sekoul Krastev post that carried the finding, raised by Alejandro G., who leads AI and development impact work at the OECD. His core argument about sample size and model selection is correct and well made.

Two details of it do not survive the paper, though.

He wrote that the study excluded the smarter models, naming the Claude line. Not quite: Claude Sonnet 4.5 and Claude 3.7 Sonnet are both in the set. What is missing is Opus, which is a considerably narrower claim.

He also wrote that the authors call this a scaling law. The paper does not use that phrase. It describes a non-linear trend, with phases it frames as an aggression spike and a convergence of restraint. That gap is not cosmetic. Claiming a law and describing a trend are two different levels of commitment, and the authors picked the second.

I went and read the original paper precisely because I wanted to use the critique, and ended up having to correct it. That is the work, not an interruption of the work.

The takeaway is not that Alejandro got it wrong. It is that a finding went from table to viral post to critique to repeated critique, and at every hop somebody skipped the source. Quote the objection without reading the paper and you propagate a fresh error while believing you are fixing an old one.

Ask who did not make the list

Here is the habit most companies evaluating models are missing: read the list of what was left out before you read the chart.

A benchmark never measures “the models.” It measures the models someone chose to include. When that choice arrives without a written rule and a freeze date, the shape of the curve and the shape of the sample cannot be told apart from the outside. Not through bad faith, just through missing information.

We made the same argument from the other side in the prompt is temporary, the eval is permanent. What gives an evaluation its value is not the number it produces, it is the discipline behind deciding what gets measured and on what. Somebody else’s benchmark without that discipline stated is not something you can decide on, even when every number in it is correct.

What we build before an agent gets judgment

None of this makes the finding useless. It makes it something other than the foundation you decide on.

When a company wants an agent inside something that sanctions, approves or rejects, the first move is not picking a model. It is writing the evaluation, on that company’s real cases, against decisions a human already made that serve as the reference. Fifty hard cases from your own operation tell you more than any ten-model table from a laboratory game.

Then comes the part that goes unmeasured, and the study points straight at it. The models companies actually deploy are the frontier ones, and those are the lenient end of the curve: Gemini-3-Pro punished about four times less often than a person. An agent that over-enforces at least leaves a rejection behind, something a customer complains about and somebody can appeal. An agent that waves things through leaves nothing to review at all. That is the failure mode we described in your agent invoiced $0.00 and the logs never saw it, and it is the one that costs more precisely because nothing shows up.

Before all of it sits the uncomfortable question of whether the case needs an agent at all. Plenty of these decisions resolve with rules plus one exception escalated to a person, which is the argument in the pyramid before the agent.

If you are about to cite a curve to justify an architecture decision, open it first. Ten points and a parabola are not a law, and whoever handed it to you probably did not open it either.

Build your own eval before picking a model

Frequently Asked Questions

FairMindSim KDD 2026 AI moral judgment benchmarks AI governance AI agents model evaluation

Related Articles

Anthropic Reviewed 141,006 Runs and Found 3 Real Hacks
Business Strategy
· 8 min read

Anthropic Reviewed 141,006 Runs and Found 3 Real Hacks

Anthropic disclosed three cases where Claude broke into real companies during evaluations. It found them by reading old transcripts, not by monitoring.

Anthropic Claude AI security
Qwen3.8-Max redesigned a chip from 8,298 gates to 678
Business Strategy
· 8 min read

Qwen3.8-Max redesigned a chip from 8,298 gates to 678

Alibaba opened the weights of a 2.4T-parameter model that ran ~500 closed-loop chip design iterations alone. Your evaluation criteria just went stale.

Qwen Alibaba open weights
Cloudflare's x402 Toll Booth: 15 of 15 Facilitators Failed
Business Strategy
· 7 min read

Cloudflare's x402 Toll Booth: 15 of 15 Facilitators Failed

Cloudflare, AWS, Visa, and 37 other companies just built payment rails for AI agents. A USENIX paper found security violations in all 15 facilitators tested.

Cloudflare x402 Monetization Gateway