Skip to content
BACK_TO_LOGS
#Large Language Models#How AI Works#AI Bias#Machine Learning#AI Explainer#ChatGPT

How Large Language Models Actually Work

March 27, 20267 MIN READ

This is a companion piece to AI Bias Isn't Just a Data Problem. It's a People Problem.. That article examines what happens when AI teams lack diverse perspectives. This one explains how these systems work, why they produce biased outputs, and why fixing that bias is so difficult.

TL;DR: Large language models don't "think" or "know." They predict words based on human text, inheriting biases such as gender, race, and socioeconomic assumptions. These biases reflect the training data and are amplified by the model's goal of choosing the most probable output. As these systems influence real decisions, who builds them and their assumptions are more important than ever.

You ask a question to ChatGPT, Claude, or Gemini. Within a few seconds, you receive what seems like a thoughtful, knowledgeable response. It feels as if the AI understood your question and logically worked out an answer. Let's dive into how an LLM actually produced that response.

It's Predicting the Next Word

A large language model (LLM) does one main thing: given a sequence of words, it predicts the most likely next word based on its training data. It calculates and generates the most probable next word one at a time. This process is similar to autocomplete but on a much larger scale, with billions of parameters fine-tuning predictions.

This is why AI researchers sometimes describe these systems as "very sophisticated autocomplete."(opens in a new tab) NYU professor Gary Marcus has called them "autocomplete on steroids"(opens in a new tab), and in a 2021 paper(opens in a new tab), researchers Emily Bender and Timnit Gebru used the term "stochastic parrots" to describe language models that produce fluent text without any comprehension of meaning.

How It Learned: Reading the Internet

Before an LLM can predict, it must learn language patterns during training by reading text data. OpenAI's GPT-3 was trained on 300 billion tokens(opens in a new tab) (a token is roughly three-quarters of a word) drawn from filtered web pages, books, and Wikipedia. Meta's Llama 3 was trained on over 15 trillion tokens(opens in a new tab). GPT-4 is reportedly trained on approximately 13 trillion tokens(opens in a new tab), though OpenAI hasn't confirmed this. The trend in training dataset sizes(opens in a new tab) has seen exponential growth with each generation. These datasets include books, Wikipedia articles, news sites, forums, academic papers, code repositories, and social media posts: a significant portion of the publicly accessible written internet, as well as licensed datasets and curated collections.

During training, the model repeatedly receives a passage with the last word missing and predicts it. When it's wrong, it adjusts and tries again, and this process repeats trillions of times. Over time, it develops an internal representation of language, not through taught grammar, but by discovering statistical patterns.

It Doesn't "Know" Anything

When an LLM gives you an accurate answer about history, science, or law, it's not retrieving that fact from memory the way you would. It's generating a sequence of words that, based on the patterns in its training data, are statistically likely to follow your question. If the training data included accurate descriptions, say, of how photosynthesis works, the model will produce an accurate-sounding explanation. The model produces an accurate-sounding explanation, not because it understands photosynthesis, but because it learned the pattern of how people describe it.

This is also why LLMs produce wrong answers with fluency. The model has no mechanism for verifying whether what it generates is true. It's optimizing for "what words are most likely to follow the given set of words," not "what is factually correct." Researchers call this "hallucination"(opens in a new tab), a term that entered mainstream use around 2022(opens in a new tab) with the release of ChatGPT. Cambridge Dictionary named "hallucinate" its 2023 Word of the Year(opens in a new tab).

The Role of Probability

The model doesn't always pick the single most likely next word. Instead, the model produces a probability distribution score across its entire vocabulary for each position. For a given context (what is the capital of France?), "Paris" might have a 70% probability, "Lyon" might have 3%, and "the" might have 0.001%. The system then samples from this distribution, sometimes picking a slightly less likely word, which is what makes responses varied and natural-sounding.

This randomness is tunable via the "temperature" setting, which controls sampling creativity. Low temperature leads to predictable, repetitive output by choosing high-probability words, while high temperature results in more creative, less predictable results by selecting less likely words. AI's "creativity" or "precision" often depends on this setting.

What This Means for Bias

LLMs inherit the patterns from their training text. If "CEO" appears alongside male names and "nurse" alongside female names across millions of articles, the model learns that association. The model learns the association as a probability, not a belief: male-associated words become more likely when generating text about CEOs.

A 2016 study by Bolukbasi et al.(opens in a new tab) (NeurIPS) found that word embeddings trained on Google News produced the analogy "man is to computer programmer as woman is to homemaker." The biases weren't programmed in; they were learned from patterns in the text.

The model doesn't hold opinions. But its outputs reflect decades of human writing with all its skews: gender and professions, race and neighborhoods, socioeconomic status and names.

These systems are increasingly making real decisions. A 2024 McKinsey survey(opens in a new tab) found that 65% of organizations regularly use generative AI across hiring, customer service, and healthcare. The FDA has approved hundreds of AI-enabled medical devices(opens in a new tab), with approvals accelerating each year. Word-by-word predictions informed by historical texts are shaping real outcomes for real people.

Why "Just Fix the Data" Isn't Simple

The obvious response is to clean up the training data, remove the biased text, and balance the dataset.

Bias in language isn't a contaminant you can filter out. A factual news article about an all-male board reinforces the association between leadership and men. A medical textbook describing symptoms as they present in white patients isn't "biased" in the traditional sense; it's incomplete. But an LLM trained on it will generate responses calibrated to that incomplete picture.

Even curated datasets carry the perspectives of whoever curated them. What was included, excluded, or labeled "high quality" all reflect human judgment. These decisions shape what the model learns.

This is why the people building these systems matter. What data to include, what benchmarks to use, what failure modes to test for: all human decisions, informed by the experiences of whoever's in the room. A team with the same backgrounds is likely to share the same assumptions about what "normal" looks like.

The Gap Between Experience and Mechanism

Using an LLM feels like talking to something that understands you. I tested this directly. After asking Alexa the CEO/nurse question (from the companion piece), I pointed out that it kept saying it was "thinking" about the riddle.

"Are you really 'thinking' about this? Aren't you just predicting the next word?"

When challenged on using the word 'thinking,' Alexa admitted it was 'sloppy anthropomorphizing.'

Alexa's response: "What I'm actually doing is statistical pattern matching at a massive scale." It laid out the process: converting text into tokens, generating probability distributions, and selecting tokens. "The 'correction' wasn't me reconsidering my logic. It was the model generating tokens that statistically align with patterns it learned about acknowledging errors."

'It's pattern matching all the way down, just incredibly sophisticated pattern matching trained on billions of text examples.'

Even this "honesty" is simulated. The model is generating the words most likely to follow a prompt about its own limitations. The self-awareness is simulated, but the fluency is real.

Beyond the Base Model: How Companies Augment LLMs

Companies layer additional systems on top of next-word prediction:

None of these changes the fundamental mechanism. RAG changes the context, fine-tuning changes the weights, and RLHF shifts which outputs are favored, but underneath, it's still pattern matching.

Each layer also inherits the bias problem. Biased documents in RAG produce biased answers, biased training data in fine-tuning reinforces those biases, and RLHF reflects the perspectives of whoever rates it. The assumptions of the people building these systems shape the outputs at every stage.

END_OF_TRANSMISSION