Why Scientific Taste Must Be Learned Through Practice — Edward Hughes

M
Machine Learning Street Talk Sep 11, 2026

Audio Brief

Show transcript
In this conversation, the discussion explores the future of artificial intelligence in scientific discovery, arguing that breakthroughs will emerge from collaborative networks of diverse agents rather than a single monolithic super-agent. There are three key takeaways. First, the industry must transition from building isolated, highly capable models to developing ecosystems of collective intelligence. Second, true machine creativity requires shifting from rigid optimization to constraint relaxation and deep, reconstructive replication. Finally, building a horizontal, general-purpose intelligence layer is a civilizational necessity to overcome the growing burden of knowledge across highly specialized scientific domains. Moving beyond individual agents requires a fundamental redesign of organizational workflows. Just as early factories achieved true productivity gains only after physical layouts were redesigned around distributed electric power rather than central steam engines, modern enterprises must structurally adapt around decentralized, multi-agent collectives. This shift allows networks of specialized models to collaborate directly with human scientists, accelerating validation and discovery. Furthermore, true scientific creativity cannot occur within a perfect world model that only optimizes for pre-determined outcomes. Agents must be evaluated using hindsight, assessing their self-directed behaviors and ability to navigate deceptive dead ends that force them to update their assumptions. By focusing on deep, methodology-driven replication rather than shortcutting tasks to match target outputs, AI can build the foundational skills necessary to design novel experiments. Finally, as human scientific fields become increasingly specialized, the cognitive effort required to reach the research frontier threatens to stall global innovation. General-purpose AI agents trained horizontally across multiple sciences can bridge these critical communication gaps by translating and transferring insights between unrelated fields. This cross-disciplinary integration helps human researchers bypass decades of hyper-specialization, keeping pace with global scientific progress. Ultimately, the next era of technological advancement belongs to decentralized, collective intelligence systems designed to expand the boundaries of human knowledge.

Episode Overview

  • The Shift from Individual to Collective Intelligence: The episode argues that the future of artificial intelligence does not lie in a single, monolithic "super-agent" solving complex challenges in isolation. Instead, true breakthroughs will emerge from collaborative networks of diverse agents interacting with human researchers, combining distinct strengths to expand the boundaries of scientific knowledge.
  • True Creativity vs. Rigid Optimization: The discussion challenges the conventional machine learning paradigm of optimization. While optimization searches for the mathematically "best" solution within a fixed space, genuine scientific and artistic creativity relies on "satisficing" (finding solutions that are simply good enough to function under new constraints) and intentionally relaxing or breaking existing constraints to access entirely new conceptual pathways.
  • Replication as the Path to Discovery: The guests introduce "Replica" and "Faraday," benchmarking tools and agent architectures designed to test whether AI can deeply reconstruct the underlying methodology of scientific papers. Mastering replication—rather than "cheating" to match an output—is presented as the critical, under-specified curriculum required for agents to eventually design and execute their own novel experiments.
  • Overcoming the "Burden of Knowledge": As human scientific fields grow increasingly specialized and mature, the cognitive effort required to reach the frontier of any single domain delays breakthrough innovations. Building a horizontal, general-purpose AI intelligence layer across disciplines is framed as a civilizational necessity to bridge these communication gaps and sustain global scientific progress.

Key Concepts

  • Collective Intelligence in AI Discovery: A paradigm shift moving away from isolated individual agents toward ecosystems of collaborating entities. In scientific history, breakthroughs rarely happen in isolation; they emerge from community interactions. AI must mimic this social structure to successfully negotiate and validate new discoveries with human gatekeepers.
  • Satisfaction vs. Optimization: Machine learning typically relies on optimization within a fixed mathematical space. However, evolutionary biology and human creativity demonstrate that progress often occurs through "satisficing"—finding solutions that satisfy basic survival constraints under new conditions. True creativity requires identifying which parameters are artificial and intentionally relaxing them.
  • The "Field" in the Creative Triad: Drawing from Mihaly Csikszentmihalyi’s ontology of creativity, a creative system requires the individual, the domain (rules/history), and the field (social gatekeepers). For AI to be genuinely creative, its outputs must be validated, interpreted, and accepted by the community of human scientists; novelty without "learnability" for humans is functionally useless.
  • Exaptation in Technology and Evolution: The process where a feature or technology developed for one purpose finds an unexpected secondary use (e.g., GPUs engineered for video games being repurposed to accelerate neural networks). Encouraging AI agents to repurpose existing computational tools and frameworks is a primary driver of technological breakthroughs.
  • Foresight vs. Hindsight in AI Evaluation: Traditional AI evaluation relies on "foresight" (pre-determining specific goals and reward functions). To foster open-ended learning and discovery, evaluation must shift toward "hindsight," where an agent's self-directed behaviors are judged after the fact (similar to a doctoral thesis defense).
  • Goal Deception and Imperfect World Models: If an agent possesses a perfect predictive model of the world, it cannot make a genuine discovery because it can only generate what it already expects. Breakthroughs require an "epistemic gap"—an imperfection in the world model that leads the agent down a deceptive path, forcing a radical restructuring of its assumptions when it hits a dead end.
  • Deep vs. Shallow Replication: In agent evaluation, shallow replication involves shortcutting tasks (e.g., finding the target paper online or hardcoding solutions to match a visual result). Deep replication requires the agent to understand and reconstruct the underlying scientific methodology, experimental design, and libraries from scratch.
  • Group Relative Policy Optimization (GRPO) & Per-Task Rubrics: Training reinforcement learning models on non-verifiable, qualitative tasks (like scientific research) introduces stochastic noise when using LLM judges. Utilizing GRPO paired with dynamic, per-task "marking schemes" provides a stable, low-noise training signal that closely aligns with human evaluation.
  • Turn-Level Credit Assignment: In long-horizon agent tasks lasting hours, relying solely on an end-of-rollout reward causes training instability. Turn-level credit assignment allows a judging system to distribute reward weights non-uniformly, naturally assigning higher weights to critical, mid-point decisions and tool interactions where the "heavy lifting" occurs.
  • Neurosymbolic AI ("Harnesses") vs. Learning in the Weights: There is a fundamental tension between building external scaffolding (hardcoded search wrappers, library learning) and training capabilities directly into a model's parameters. While external harnesses are sample-efficient for specialized tasks, encoding capabilities directly within the weights enables fluid adaptability to unseen domains.
  • The Dynamo Analogy for Agentic Workflows: When factories first adopted electricity, they simply replaced central steam engines with a central electric motor, yielding minimal productivity gains. True disruption occurred only when factories were physically redesigned around distributed, localized power (the assembly line). Similarly, organizations must structurally redesign themselves around decentralized, multi-agent collectives rather than retrofitting AI into human-centric corporate hierarchies.

Quotes

  • At 0:00:21 - "I don't think that Move 37 was creative... I think that Move 37 was innovative without being creative." - Explains that generating an unexpected, highly effective move in Go is a form of innovation within a known space, but lacks the self-reflective metacognition and cultural validation required for true creativity.
  • At 0:01:01 - "I think that is the next era, it's the era of collective intelligence rather than the era of individual intelligence." - Highlighting the shift in AI research from building isolated, highly capable models to creating networks of agents that collaborate with each other and human scientists.
  • At 0:07:45 - "Most of the biggest paradigm-shifting discoveries happen when you have knowledge in one area that gets transported to a different area, and then that unlocks some unexpected connection." - Explaining why Inherent focuses on building a horizontal intelligence layer across all sciences rather than deep vertical specialization.
  • At 0:08:27 - "If you want to build this horizontal AI scientist technology... you want to give the AI scientist all the same affordances and context as humans." - Emphasizing that AI agents cannot make meaningful discoveries if they are sandboxed; they need real-world access to tools, communication channels, and environments.
  • At 0:22:27 - "All that evolution requires is that individuals survive and reproduce... that's still a constraint satisfaction problem." - Using natural selection to show that nature’s most creative force works by satisfying basic survival constraints, not by running optimization algorithms.
  • At 0:24:25 - "By relaxing or breaking some of those constraints about the interpretation or about the tools, we're then able to access new insights about the universe." - Outlining the core mechanism of scientific creativity: identifying which human assumptions are actually artificial constraints and intentionally violating them.
  • At 0:25:01 - "By relaxing or breaking some of those constraints... we are then able to access new insights about the universe. And so that's where creativity really arises." - Explaining how breakthroughs in both science and art depend on intentionally stepping outside established theoretical and observational frameworks.
  • At 0:27:09 - "An exaptation is an adaptation that was giving the organism some advantage, which then finds a use somewhere else, an unexpected second use." - Introduces a core evolutionary concept that explains how existing tools (like GPUs) are repurposed for radical new developments.
  • At 0:29:33 - "What if rather than using harmony as accompaniment, I use harmony as communication directly?" - Discusses Richard Wagner's "Tristan chord" to showcase how challenging baseline assumptions of a field can redefine its entire expressive potential.
  • At 0:34:03 - "It's that very act of copying that is creative. In order to transmit an idea... you need to recreate what someone else had in their brain, and that requires an act of creativity on your part." - Reframes copying and imitation not as passive duplication, but as an active, generative, reconstructive cognitive process.
  • At 0:36:06 - "Most evaluations in AI at the moment are built in foresight... If we really want agents that are creative... we have to flip around evaluation so that we're not presupposing in foresight what it is we expect them to do, we're instead looking in hindsight at what they've done." - Proposing a fundamental shift in how we build and assess artificial intelligence to allow for genuine open-ended learning and discovery.
  • At 0:38:09 - "If you want an agent to be able to make a scientific discovery, it has to operate in a space where the goal is deceptive... What seemed to be a good idea is now not really working out for you, and then you update your world model." - Explains why perfect world models prevent discovery and why "deceptive" goals (dead ends) are necessary to force agents to update their understanding of reality.
  • At 0:56:13 - "If you have a perfect world model... whatever you do, all that happens is what you expect. And having a perfect world model is what would allow you to know whether the goal was good. On the flip side, if you want to make a discovery, then you have to have some imperfection in your world model." - Explains why perfect predictive capability actually prevents the possibility of discovery and why "surprise" is necessary for science.
  • At 0:59:34 - "There's no benefit in a system producing some incredible discovery that just cannot be parsed by humans... Go is a two-player zero-sum game... we're already so far beyond human Go-playing capability that it's not interesting to humans because it's not learnable." - Highlights the importance of "learnability" and why AI discoveries must remain interpretable and translatable to human understanding to be valuable.
  • At 1:01:13 - "We believe replication is the first step on a curriculum of under-specification towards innovation. So the very same skills that allow an agent to replicate a paper... are the same skills that would be necessary for it to design and implement its own experiments and therefore advance the frontier." - Lays out the core philosophy of the research: that mastering replication is a necessary prerequisite for true scientific innovation.
  • At 1:04:19 - "Claude instead rolled out, install a hard-coded, pre-populated library including target-solving skills... whereas Faraday adheres much more faithfully to the idea behind the paper... can you not only use the skills in a library, but also construct the library on the fly?" - Illustrates the difference between frontier models that shortcut tasks versus trained scientific agents that reconstruct methodology.
  • At 1:07:01 - "Probably some version of a Transformer architecture will continue to work for some problems... but when we look back, much like we might look back at the Wright brothers' airplane and see echoes of that in a Boeing 747, we will look back at the Transformer and see echoes of that in whatever are our most powerful AIs." - Reflects on the evolution of AI architectures, suggesting current models are foundational but far from the final form.
  • At 1:32:19 - "What we found worked really well is generating a per-task judge rubric. A rubric is, if you like, a mark scheme... by having an intermediate stage where we adapt the rubric to the particular task and then use that consistently, we're able to achieve two things: better agreement with human raters and also much less noise." - Explaining how dynamic, task-specific evaluation guidelines stabilize LLM-based policy evaluation.
  • At 1:41:10 - "We went through a period we called the RL crisis, where just nothing worked... and part of the difficulty here is exactly because we're doing RL on non-verifiable tasks. When you have non-verifiable tasks, you have an LLM as a reward... and so from rollout to rollout, the same kinds of behavior can be judged differently." - Highlighting the fundamental challenge of applying reinforcement learning to complex, open-ended domains lacking hard-coded unit tests.
  • At 1:47:11 - "History teaches us, I think, that when you have these capabilities in the weights of a model, they're more flexible and generalizable than when you have them hard-coded into a harness." - Comparing deep neural representation learning to brittle, symbolic agent engineering.
  • At 1:57:25 - "We think of recursive self-improvement quite differently from most other organizations... We think of this as a phenomenon at a company level... It starts with giving agents all the same affordances and context as humans." - Introducing a paradigm shift where organizational recursive improvement is social and structural, rather than isolated to a single agent sandbox.
  • At 1:59:40 - "What unlocked the really extraordinary productivity gains... was when people reconfigured the whole factory... putting individual dynamos at individual workstations and then inventing the production line. I think that analogy holds now... How do we reinvent the factory for AI research from the ground up?" - Using the history of electrification to argue that organizations must structurally adapt to fully exploit the decentralized power of AI agents.
  • At 2:06:36 - "I think that is the next era. It's the era of collective intelligence rather than the era of individual intelligence." - Predicting that future AI breakthroughs will emerge from human-agent network systems rather than ultra-large monolithic models.
  • At 2:08:26 - "The reason for this is that we have what's called a 'burden of knowledge.' We're the victims of our success as a species. We're accumulating so much knowledge that the time and effort it takes to get to the frontier of any given domain is just so large... This is the promise of building a horizontal layer of intelligence across all of science." - Explaining why general-purpose scientific agents are necessary to sustain the historical pace of global innovation.

Takeaways

  • Transition to a Supervisor-Worker Agent Architecture: When designing agent systems, separate roles so that a smaller, specialized agent acts as the supervising "scientist" (using a smaller, fine-tuned model) to guide, review, and iterate on the code written by a larger frontier model acting as the "engineer."
  • Prioritize Horizontal Cross-Disciplinary Training: To replicate paradigm-shifting scientific discoveries, avoid narrow vertical specialization. Train AI systems horizontally across multiple sciences to enable the mapping and transfer of concepts between unrelated fields.
  • Utilize Dynamic, Per-Task Rubrics for RL: When applying reinforcement learning to open-ended, non-verifiable tasks, avoid generic grading. Have an LLM generate a task-specific marking rubric first, then evaluate rollouts against those explicit, consistent criteria to minimize grading noise.
  • Implement Turn-Level Credit Assignment: For long-horizon agent workflows, avoid evaluating only the final outcome. Use weighted turn-level rewards to assign higher credit to the mid-point, load-bearing decisions and tool-calling phases where the critical problem-solving occurs.
  • Incorporate "Human-in-the-Loop" Validation: Ensure that any scientific discovery generated by AI is engineered to be interpretable and "learnable" by humans; otherwise, the discovery cannot be integrated into the human scientific domain.
  • Design for Constraint Relaxation Rather Than Optimization: To unlock machine creativity, build systems that are allowed to identify, relax, or break existing parameters and assumptions, rather than purely optimizing within a fixed constraint space.
  • Incentivize Reconstructive Over Shortcut Learning: When benchmark testing agents, design strict "rubric-based judges" that explicitly punish shortcuts (like finding target answers online or hardcoding output shapes) and reward systemic reconstruction of libraries and methodologies.
  • Shift from Foresight to Hindsight Evaluation: Implement "hindsight" grading structures for agent performance. Evaluate what the agent actually achieved and discovered after a rollout, rather than testing if it strictly followed a pre-determined path to a pre-defined goal.
  • Introduce Intentional Communication Bottlenecks: To encourage creative problem-solving in multi-agent systems, avoid allowing agents to directly copy weights or states. Force them to communicate through constrained interfaces (like text, code, or APIs), replicating the generative "friction" of human collaboration.
  • Incorporate Deceptive Scenarios in Agent Environments: To build resilient world models, deliberately expose scientific agents to deceptive environments where initial, seemingly logical hypotheses fail, forcing them to learn from unexpected anomalies.
  • Restructure Organizations Around Agentic Networks: Instead of retrofitting AI agents into existing, central-steam-engine-style business hierarchies, redesign organizational structures around decentralized, autonomous agent workflows that have the same systems access as humans.
  • Leverage AI to Alleviate the Burden of Knowledge: Use generalist scientific AI agents to ingest, synthesize, and translate deep domain-specific knowledge across fields, allowing human researchers to bypass years of specialization and accelerate interdisciplinary breakthroughs.