Everyday Apparatus
Societyopenalex3 min read1 month ago

The Machine Got Kinder, and No One Showed It How

Everyone assumed an AI's morals had to be taught by hand. A study of 75 systems found that human-like judgment simply grows as the model gets bigger.

A read of Scaling laws for moral machine judgement in large language models · openalex

Guided read0:00 / 5:39

Scaling law

The empirical finding that a model's performance on a task follows a predictable curve as its size (parameters) or compute increases.

Power law

A relationship where one quantity changes as a fixed exponent of another — here, moral alignment improves by a consistent but small amount each time model size multiplies.

Moral Machine

An MIT online survey that collected millions of human responses to trolley-problem-style life-or-death dilemmas, used here as the human preference baseline.

Extended reasoning models

AI models that generate intermediate thinking steps before producing a final answer, rather than responding immediately — sometimes called chain-of-thought models.

What it’s not claiming · The paper does not claim that merely scaling up language models will eventually achieve perfect or universally aligned moral reasoning without additional alignment interventions.

`★ Insight ─────────────────────────────────────` The strongest tells here are structural, not lexical: two "isn't-X-it's-Y" negative parallelisms and one "it turned out" hedge. The hushed "quiet/almost/honest" register the brief warns about is already absent, so this is a light pass — break the parallelisms, drop the hedge, ration "simply." `─────────────────────────────────────────────────`

Picture a runaway trolley. Five people are tied to the track ahead, one person stands on a side track, and your hand rests on a lever that would switch it. Most people pull. And if you change the people, a doctor against a thief, a child against someone very old, five lives against one, most of us lean the same way, even when we'd argue about the edges. Our answers scatter, but they scatter around a center. The strange question is the machine's — what it would do, and whether it would land anywhere near the rest of us.

For years the working assumption, inside the labs and out, was that this part would be hard. You can make an artificial intelligence sharper just by building it bigger, giving it more of the internal settings it tunes while learning, and it writes better, codes better, reasons better. But values felt like a different kind of thing. Morality, the thinking went, doesn't just show up. You have to put it there: teach the model with careful human feedback, hire people to judge its answers, build whole teams whose job is to bolt a conscience onto a mind that wouldn't grow one on its own. A bigger model would be more capable. Nobody expected it to be more humane.

Then a group of researchers decided to measure it. They took 75 of these systems, from small ones to giants thousands of times larger, and put each through ten thousand versions of the trolley choice, drawn from a vast survey of what actual people around the world said they would do. For every model they asked one question: how far do its choices sit from the human average? Then they lined that distance up against size.

The result was a clean, bending curve. Make a model bigger, and it drifts reliably closer to the human consensus, and it does so precisely — on the same kind of predictable track that governs how these systems get better at language or math. Nobody taught them to. Moral alignment behaves less like a value you install and more like a skill that accumulates as the thing grows.

There was one wrinkle. When a model is made to think out loud, working through the problem step by step before it answers, it lands closer to us too. But that help goes mostly to the small ones. The giants barely move. Careful reasoning is a crutch that matters most exactly where size is scarce, like telling a first-year law student to slow down and reason through the case. The student improves a great deal. The seasoned judge was already doing it.

Here is what the study cannot tell you: whether any of this is morality at all. Every scenario was the same hypothetical, steeped in Western assumptions, drawn from the same internet these models read while training. The researchers admit they can't rule out that the big models simply recognize the genre and echo what people have already written about it. The thing being measured is resemblance to an average of human preferences, not goodness. So the pattern is real, and a little eerie: bigger models look more like us. Whether looking like us is the same as being good is a question the machine cannot answer, and neither, yet, can we.

Where this sits

Open question

Does the power‑law relationship between model size and alignment with human moral judgments that the study uncovers also hold for a wider range of ethical decisions, real‑world applications, and culturally diverse populations beyond the Western‑centric Moral Machine trolley‑problem scenarios?

Next readThe AI That Reasons in Private Can Be Taught to Show Its Work