Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder
intermediate 10 mins Read

How LLMs Reason and Generate Text

Auto-regressive decoding, chain-of-thought prompting, the ReAct pattern, and why LLMs fail at certain reasoning tasks.

Last reviewed: July 27, 2026

LLMs generate text one token at a time, predicting each next token from the probability distribution learned over training data conditioned on everything before it. What looks like "reasoning" emerges from this next-token process when the model is prompted to produce intermediate steps (chain-of-thought) rather than jumping straight to a final answer — each generated step becomes part of the context the next prediction conditions on.

Introduction

Large Language Models do not “think” the way humans do — they are next-token predictors trained to produce plausible continuations of text sequences. Yet recent techniques, particularly chain-of-thought prompting, have dramatically improved their ability to solve multi-step problems. Understanding the mechanics of how LLMs generate text — and why they fail at certain tasks — is essential for building reliable AI applications.

Step-by-Step: From Prompt to Response

flowchart TD
    A["1. Auto-Regressive Token Generation"]
    B["2. Chain-of-Thought Prompting (Wei et al., 2022)"]
    C["3. Self-Consistency Sampling"]
    D["4. ReAct Pattern (Reasoning + Acting)"]
    E["5. Known Reasoning Failures"]
    A --> B
    B --> C
    C --> D
    D --> E

Step 1: Auto-Regressive Token Generation

LLMs generate text one token at a time. Each token is selected by computing a probability distribution over the entire vocabulary (~100,000 tokens) conditioned on all prior tokens:

P(token_n | token_1, token_2, ..., token_{n-1})

The model doesn’t plan the full response before generating — it commits to each token sequentially. This is why sampling parameters matter:

  • temperature=0: Always pick the highest-probability token (greedy decoding) — deterministic but sometimes repetitive
  • temperature=0.7: Sample from the distribution — more creative, varied
  • top_p=0.95 (nucleus sampling): Sample only from tokens comprising the top 95% of probability mass

Step 2: Chain-of-Thought Prompting (Wei et al., 2022)

Standard prompting asks models to jump to the answer. Chain-of-thought asks them to show their work, dramatically improving accuracy on multi-step reasoning:

Standard prompt:
Q: A store sells apples for $0.80 each. Maria buys 12 apples and pays with a $20 bill. How much change does she get?
A: $10.40  ← Frequently wrong on standard prompting

Chain-of-thought prompt:
Q: A store sells apples for $0.80 each. Maria buys 12 apples and pays with a $20 bill. How much change does she get? Let's think step by step.
A: 
Step 1: Cost of 12 apples = 12 × $0.80 = $9.60
Step 2: Change = $20.00 - $9.60 = $10.40
The answer is $10.40  ← Correct

The key insight: generating intermediate steps forces the model to allocate more “compute” (more tokens, more transformer passes) to each reasoning step.

Step 3: Self-Consistency Sampling

A single chain-of-thought can still go wrong. Self-consistency generates multiple independent reasoning chains and takes the majority answer:

import openai
from collections import Counter

def self_consistent_answer(question, n=5):
    answers = []
    for _ in range(n):
        response = openai.chat.completions.create(
            model="gpt-4o",
            messages=[{"role": "user", "content": f"{question}
Let's think step by step."}],
            temperature=0.7  # Sample diversity
        )
        # Extract final answer from response
        answers.append(extract_answer(response.choices[0].message.content))
    
    # Majority vote
    return Counter(answers).most_common(1)[0][0]

Step 4: ReAct Pattern (Reasoning + Acting)

ReAct (Yao et al., 2022) interleaves reasoning traces with tool calls:

User: What is the population of the capital of France?

Thought: I need to find the capital of France, then look up its population.
Action: search("capital of France")
Observation: Paris is the capital of France.

Thought: Now I need the population of Paris.
Action: search("Paris population 2024")
Observation: Paris metropolitan area population is approximately 12.2 million.

Thought: I have the answer.
Answer: The capital of France is Paris, with a metropolitan population of approximately 12.2 million.

Step 5: Known Reasoning Failures

Reversal Curse: Models trained on “A is B” don’t generalize to “B is A”:

  • Knows: “Tom Cruise’s mother is Mary Lee Pfeiffer”
  • Fails: “Mary Lee Pfeiffer’s son is ___”

Hallucination: Models confabulate plausible-sounding but incorrect facts, especially for less common knowledge.

Compositional Failure: Simple steps combined become unreliable (e.g., “reverse this string then count its vowels”).

Solution patterns: RAG for factual grounding, code execution for arithmetic, structured output validation.

Key Takeaways

  • LLMs are auto-regressive next-token predictors — they don’t plan; each token is chosen conditioned on all prior tokens.
  • Chain-of-thought prompting improves multi-step reasoning by allocating more token compute to intermediate steps.
  • Self-consistency (majority vote across multiple samples) significantly improves reliability on reasoning tasks.
  • ReAct interleaves reasoning with tool calls — the standard architecture for LLM-powered agents.
  • Known failure modes include the reversal curse, hallucination on low-frequency facts, and compositional reasoning breakdowns.

Common questions

Is an LLM actually reasoning, or just predicting likely text?

Mechanically, it's always next-token prediction — there's no separate symbolic reasoning engine underneath. But when that prediction process is scaffolded to produce and condition on intermediate steps, it reliably solves problems a single-shot prediction gets wrong, which is why the distinction matters less in practice than it sounds philosophically.

Why does asking a model to "think step by step" improve accuracy?

Each intermediate token the model generates becomes additional context for predicting the next one. Spelling out steps gives the model more relevant context to condition its final answer on, instead of forcing the entire solution into a single prediction with no intermediate scaffolding.

What are "reasoning models" doing differently?

Reasoning-specialized models are trained (often via reinforcement learning) to generate longer, more deliberate chains of intermediate steps before answering, and to backtrack or verify within that chain — rather than relying on a user's prompt to elicit step-by-step output.

Historical figures, architectures, and capabilities are for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Research papers, developer documentation.