You Don't Retrieve Memories. You Generate Them.

Curt Jaimungal Curt Jaimungal Jun 01, 2026

Audio Brief

Show transcript
This episode covers the revolutionary thesis that the human brain and large language models perform the exact same core computation by constantly predicting the next token. There are three key takeaways. First, human memory is a dynamic generative capacity rather than a static database. Second, the brain operates on a continuous feedback loop of autoregressive cognition. Third, understanding this prediction loop can optimize how humans learn skills and make decisions. Instead of retrieving saved files, human memory creatively reconstructs information in real time based on current prompts. Each immediate prediction synthesizes the entire past context to project future actions, much like anticipating the next step in a dance. This autoregressive loop means that practicing sequential skills is about mastering the transitions where one action naturally triggers the next. By viewing human cognition through the lens of artificial intelligence, we unlock a unified theory of language, thought, and potential.

Episode Overview

  • Explores a provocative thesis: that the human brain and large language models (LLMs) perform the exact same core computation by constantly predicting the "next token."
  • Challenges the traditional view of human memory, shifting it from a static database of stored facts to a dynamic, generative capacity.
  • Connects computational AI concepts with cognitive science to offer a unified, elegant theory of language, thought, and human action.

Key Concepts

  • Autoregressive Cognition: The brain operates on a continuous feedback loop. It generates a "next token" (a word, a step, or an action), perceives that output, feeds it back into the cognitive loop, and uses it to determine the subsequent step.
  • The "Pregnant Present": Even though the brain predicts only the immediate next token, this single prediction is "pregnant" with meaning. To make it, the brain must synthesize the entire past context and project a future trajectory, much like anticipating the next step while dancing.
  • Memory as Potentiality: Humans do not store literal recordings of past events. Instead, like an LLM, the brain stores the capacity to generate responses in real-time when prompted, allowing for an infinite variety of outputs from finite neural pathways.

Quotes

  • At 0:03 - "What we are doing in our minds, in our brains, is the same kind of computation. We are simply guessing the next token." - Framing the central thesis that bridges biological cognition and artificial intelligence.
  • At 0:33 - "In the process of guessing that very next token, they're taking into consideration all the past and also the likely future. It has to project this kind of trajectory." - Explaining the concept of the "pregnant present" and how short-term prediction facilitates long-range planning.
  • At 1:58 - "What we actually have stored is just the capacity to autoregressively generate." - Redefining the mechanics of human memory as a generative process rather than a retrieval system.

Takeaways

  • Reframe memory as a real-time generation: Avoid viewing memory recall as retrieving a saved file; instead, recognize it as a creative reconstruction triggered by a present "prompt."
  • Apply the prediction loop to skill learning: When practicing sequential activities like speaking, writing, or physical movement, focus on the flow of the transition from one moment to the next, as each action naturally generates the prompt for the next.
  • Use the LLM framework to understand cognitive limits: Leverage the analogy of token limits and prompt engineering to better understand how focus, context, and immediate environments shape human decision-making.