How LLMs decide what to say, and where AI bias comes from

Why LLMs Say What They Say

An AI model doesn’t have opinions. It has geometry.

Training data carves a landscape. Your prompt picks the starting point. Then one possible answer rolls downhill, following the routes the model learned to treat as likely.

Drop a prompt Where this map lies →
Choose a prompt to begin
DROP A PROMPT04 STARTING COORDINATES
01 · The landscape

The terrain is compressed writing

Imagine every recurring phrase, argument, association, omission, and style in the model’s training material pressed into one surface. Where many compatible continuations pile up, the terrain sinks into a basin. Where few fit, it rises into a ridge.

This is not a literal map inside the model. It is a way to picture a harder idea: the model builds an answer one small piece at a time, each piece chosen to fit the words already written. So the deepest basin is not the model’s settled opinion, and reaching it is not the model reasoning toward a conclusion. It is simply the answer that training made easiest to fall into and hardest to climb back out of.

  • Mainstream viewMany familiar paths converge here.
  • Expert nicheSpecialized language opens a narrower route.
  • ContrarianA coherent minority frame has its own valley.
  • FringeA shallow pocket is reachable but unstable.
  • Institutional defaultOfficial phrasing forms a dependable channel.
  • Long-tail surpriseA rare continuation can still connect.
Amber and blue-green probability terrain annotated with six answer basins Mainstream viewExpert nicheContrarianFringeInstitutional defaultLong-tail surprise
Height represents how difficult a continuation is to reach in this simplified model. Lower is more probable, not more true.
02 · Generation

Generation is rolling downhill

The same learned terrain can produce different answers because the prompt places the first ball at different coordinates. Watch two phrasings of the same subject.

Step 1 · place the prompt

A broad question starts near broad language

“Is social media safe for teens?” drops near familiar public discussion. The words in the prompt constrain which continuations fit before generation even starts.

Static terrain showing the broad prompt rolling toward Mainstream view
Broad prompt → a common basin.
Step 2 · settle

Local probability compounds

Each piece the model writes changes the slope under the next one. A slight early lean can compound into a coherent answer, because every later word has to fit the path already taken.

Step 3 · change the coordinates

Ask for a steelman

Now the prompt requests the strongest case against the subject. The terrain has not changed. The starting position and initial nudge have.

Static terrain showing a steelman prompt settling in a different answer basin
Steelman prompt → a different reachable basin.
Step 4 · compare

Same model, different answer

You didn’t change the model. You changed the coordinates.

03 · Accountability

Three things people call “bias”

The word bias hides three different interventions, three different responsible groups, and three different ideas of a fix. Conflating them makes every argument about AI less precise.

Separate the source of the slope before arguing about whether it should be changed.
KindWhat causes itWho is accountableWhat “fixing it” meansTerrain metaphor
Layer 1Corpus bias Patterns and absences in the writing the model learned from. Repetition digs; silence leaves high ground. Diffuse: authors, publishers, platforms, collectors, and the builders who choose the sampling mix. Change the material: add missing perspectives, rebalance sources, repair errors, and accept that every corpus boundary is a choice. Natural landscape. History deposited the layers before this model existed.
Layer 2Curation bias Selection, filtering, deduplication, labeling, and exclusion decide which parts of the available corpus become training material. Dataset builders and model developers who set the inclusion rules and apply them. Change the selection process: expose provenance, audit criteria, test omissions, and revise filters whose side effects exceed their purpose. Selective erosion. Some deposits remain; others are cut away before training.
Layer 3Alignment bias Post-training objectives, preference judgments, rules, and product constraints reward some answer shapes and discourage others. The actors who design the objective, supply judgments, set deployment rules, and approve the resulting behavior. Change the earthworks: revise objectives and evaluations, inspect tradeoffs, document the intended behavior, and govern who gets to choose it. Civil engineering. Basins are deepened, ridges raised, and channels redirected on purpose.
04 · Post-training

The RLHF earthmovers

Here, “RLHF” (reinforcement learning from human feedback) stands in for the broader post-training step, where a finished model is rewarded and penalized on sample answers until it behaves the way its builders intend. Toggle it and watch a scripted demonstration edit three parts of the field.

3.25 → 4.28Mainstream deepens

More nearby paths are captured by the largest basin.

1.58 → 0.48Fringe nearly flattens

A once-reachable pocket becomes much easier to escape.

σ [4.2, 0.7]Long tail becomes a channel

A long, narrow route now feeds toward the main region.

Move the earth

The answer moves even after it settled

Morph the terrain under the resting ball. It re-settles into the new field, and a ball still rolling simply follows the slope as it shifts. Post-training changes what is easy for a model to say, not only what gets blocked at the very end.

Static terrain after post-training deepens the mainstream basin and flattens the fringe basin
Post-training earthworks: the same prompt now descends through a changed field.
The governance question

Who audits the earthworks?

Labs reshape models after training, and those edits are editorial decisions by private actors. Some are plainly useful; all create tradeoffs. The durable question is not whether shaping occurs, but who names the goal, sees the side effects, and can contest the result.

06 · Limits

Where the map lies

A good metaphor makes one structure visible by hiding another. Keep the terrain, but keep its warning label attached.