scieee AI-readable full text Open interactive document viewer

Thinking Machines: Foundations of General Intelligence and World Models

Ananta, Nair; Erin, Austin; Jason, Watson; Farnoush, Banaei-Kashani

Abstract

Despite significant advances in artificial intelligence, contemporary systems remain limited in their capacity to integrate fast adaptive behavior with structured abstract reasoning, capabilities that in humans emerge through continuous coordination across brain systems. While modern models achieve superhuman proficiency in narrow domains, they continue to falter at compositional reasoning, causal abstraction, and embodied understanding. In this first part of a two-part investigation, we analyze the foundational mechanisms underlying general intelligence through three interdependent dimensions: energy-efficient computation, embodied goal-directed control, and multi-timescale memory consolidation. Drawing on principles from neuroscience and systems theory, we argue that intelligence emerges not from monolithic optimization, but from the dynamic communication amongst processes operating across multiple temporal and representational hierarchies. We introduce a dual-loop computational framework, in which a fast inner loop supports real-time embodied adaptation, while a slower outer loop integrates memory, abstraction, and long-horizon planning. The architecture promotes energy-efficient through sparse, recurrent dynamics, maintains goal-directed behavior through predictive feedback and utility optimization, and constructs internal world models unifying perception, action, and inference. Together, these mechanisms outline a coherent foundation for scalable, adaptive, and energy-efficient general intelligence. The second part of this series extends these principles into a neuro-inspired framework for metacognitive control and self-reflective learning.

Full text

1 Thinking Machines: Foundations of General Intelligence and World Models Ananta Nair1, Erin E Austin2, Jason M Watson2, Farnoush Banaei-Kashani2 1 Dell Technologies 2 University of Colorado Denver Abstract Despite significant advances in artificial intelligence, contemporary systems remain limited in their capacity to integrate fast adaptive behavior with structured abstract reasoning, capabilities that in humans emerge through continuous coordination across brain systems. While modern models achieve superhuman proficiency in narrow domains, they continue to falter at compositional reasoning, causal abstraction, and embodied understanding. In this first part of a two-part investigation, we analyze the foundational mechanisms underlying general intelligence through three interdependent dimensions: energy-efficient computation, embodied goal-directed control, and multi-timescale memory consolidation. Drawing on principles from neuroscience and systems theory, we argue that intelligence emerges not from monolithic optimization, but from the dynamic communication amongst processes operating across multiple temporal and representational hierarchies. We introduce a dual-loop computational framework, in which a fast inner loop supports real-time embodied adaptation, while a slower outer loop integrates memory, abstraction, and long-horizon planning. The architecture promotes energy-efficient through sparse, recurrent dynamics, maintains goal-directed behavior through predictive feedback and utility optimization, and constructs internal world models unifying perception, action, and inference. Together, these mechanisms outline a coherent foundation for scalable, adaptive, and energy-efficient general intelligence. The second part of this series extends these principles into a neuro-inspired framework for metacognitive control and self-reflective learning. Keywords: Artificial General Intelligence, Energy Efficiency, Embodied Cognition, Dual-System Architectures, Predictive World Models, Hierarchical Memory, Metacognitive Control, Adaptive Learning, Causal Abstraction, Neuro-Inspired Computation 1. Introduction The recurring cycle of promise and disillusionment has long defined the nonlinear trajectory of Artificial Intelligence (AI). In recent years, monumental progress has driven a belief of machines as tools built for thought. However, the great divide between humans and machine is not measured in thought or logic, but rather what separates true intelligence is that human understanding emerges not from the ability to solely generate answers, but from the instinct to build knowledge by asking the right questions. Human intelligence is inherently inquisitive, efficient, adaptable, compositional, and constantly building upon a foundation to refine and reshape understanding through exploration, exploitation, thinking, and reflection, all qualities that continue to remain elusive even to the most advanced AI models. Despite surface-level 2 successes, the inner workings of these systems lack the epistemic drives and inductive biases that support systematic generalizable learning in humans. Modern AI systems have undergone a remarkable transformation over the past decade. The advent of the transformer architectures (Vaswani et al., 2017) has driven rapid progress and enabled successful large language models (LLMs) such as GPT-4, LLAMA, Deepseek, and Gemini (Achiam et al., 2023; Touvron et al., 2023; Liu eta al., 2024; Team et al., 2023). These systems display impressive generative fluency, and are capable of knowledge synthesization, natural language and code generation, mathematical reasoning, and cross-modal integration. Multimodal and Agentic foundations further extend these capabilities to multiple data processing streams and tool calling whereas embodiment efforts have supported robotic alignment (Driess et al., 2023; Brohan et al., 2023, Ke et al., 2025). These advancements all suggest a growing convergence between foundation models, embodiment, and multimedia processing, pushing AI toward more capable interactive systems (Kim et al., 2024). However, alongside these achievements, substantial challenges remain. When pushed into increasingly complex domains, models persistently struggle with tasks of higher order cognition. With model capabilities often measured through benchmark performance, these tests are increasingly demonstrating to be unreliable indicators of genuine understanding or adaptability. Even when models exceed human-level performance on narrow tasks, they often lack key ingredients of cognition: they operate without explicit goals, fail to infer underlying cause and effect or represent causal structure, and do not seek to iteratively minimize uncertainty to formulate stable and generalizable explanations. This entails models to further lack a principled mechanism for integrating information across short and long-term timescales and restricting their ability to consolidate experience and support continual learning. These limitations produce systems whose generalization is narrow and are largely interpolation-bound, with little capacity to transfer knowledge beyond patterns of seen data. Architectural inefficiencies further compound the problem, as heavy reliance on energy-intensive training regimes is required to learn, relearn, and adapt to changing information, resulting in costly yet rigid systems. This stands in stark contrast to the brain, which achieves parallel, data-efficient, and scalable computation while running on roughly 20 watts of power (Kováč, 2010). Together, the above challenges point not merely to issues of scale, data, or optimization, but to a fundamental misalignment between current AI architectures and the principles of biological intelligence that lead to higher order cognition. Rather than continuing to optimize existing models, in this paper, we argue that meaningful progress in AI will require rethinking of principles. The aim of this work is not to discard the successes of deep learning, but rather to extend it, by augmenting current architectures with neuro-inspired mechanisms and integrating well-established theories into a theoretical framework that fosters flexibility, adaptability, and efficiency. Admittingly, although the brain may not serve as a perfect blueprint (Engle, 2002, Noyes 2001, Kanki 2018) for intelligent adaption, it does remain the most complete and adaptable model of flexible cognition, offering essential insights for architectural design. Thus, we propose a two-part investigation into the computational foundations of intelligence. Part I, the present work, examines how intelligence emerges from system-level communications across interacting brain processes, and how these organizing principles can guide the design of extending deep learning architectures with neuro-inspired principles. Specifically, we identify three bottlenecks that constrain 3 current AI systems: the lack of energy-efficient computation, limited embodiment and goal-directed control, and the absence of multi-timescale memory consolidation. Part II (Nair et al., 2025, in preparation) builds directly on this foundation, formalizing these principles into a unified mathematical and computational framework. It integrates Wilson–Cowan population dynamics, Hopfield attractors, Kalman filtering, graph-structured dependencies with spectral operators, Lyapunov-stable dynamics, and goaloriented policy refinement. The resulting dual-loop system couples’ real-world interaction with offline consolidation under shared constraints derived from the Free Energy Principle, offering a coherent theoretical account of efficient computation, adaptive embodiment, and hierarchical memory. 1.1 The Hidden Costs of Modern AI: Bottlenecks in Energy and Architecture Many behaviors that remain profoundly challenging for artificial systems are performed effortlessly by humans, including young children (Gandhi et al., 2021; Stojnić et al., 2023). Yet from a computational perspective, these seemingly simple abilities are deceptively complex. The transition from fast, intuitive pattern recognition (System I) to slower, deliberate reasoning (System II) hinging on capacities such as compositionality, causal reasoning, and the integration of information in a global workspace (Kahneman 2011; Tversky & Kahneman 1981; De Martino et al., 2006; Baars 2005). Yet, inspired by the challenge of hard-coding System II in machines, the Cyc project attempted a large-scale effort to encode common-sense knowledge in machine-readable form. Over 35 years, it amassed 1.5 million concepts, and 25 million assertions, but its rule-based curated design limited scalability, as static symbols could not adapt to novel complexities (Lenat & Guha, 1989; (Lenat et al., 1990; Panton et al., 2006). On the contrary, human cognition is not governed by a fixed ruleset but by a dynamic, self-organizing network of hierarchical concepts that can be recombined, restructured, and repurposed in novel ways to allow for both fast and slow thinking. Humans identify relevant features, infer latent structures, and generalize from sparse, noisy data with remarkable efficiency; often to tasks never previously encountered. These abilities arise from flexible abstract representations and rich inductive biases, as well prior experience that together enable compositional goal-directed learning (Herd et al., 2021; Russin et al., 2020; O’Reilly et al., 2010, 2014, 2020; O’Reilly & Munakata, 2000; Wilson & Izmalov 2020). Abstractions, provide a mechanism for transforming experience into structured forms, supporting analogical transfer, systematic generalization, and the flexible recombination of knowledge across tasks (Lake et al., 2017). Through the construction of hierarchical world models built from distributed representations, lower-level sensory features are progressively integrated into higher-order concepts in prefrontal and temporal regions. Hippocampal–cortical interactions consolidate experiences into cortical structures, capturing statistical regularities, and supporting both memory and generalization. Prefrontal systems adaptably encode relational and compositional structures, modulated by predictive learning, attention, and goals, to allow reward circuits to be reused and recombined across contexts. These circuits disentangle salient features to capture invariant relationships, while supporting counterfactual reasoning and top-down control, enabling the brain to move beyond correlation and toward structured causal reasoning and flexible generalization. These multi-scale processes transform concrete sensory input into schemas, categories, and rules that serve as the foundations for complex problem solving. 4 Inductive biases, on the other hand, act as constraints on the system, defining an operable space of hypotheses testing, and ensuring learning remains tractable both through data-efficiency and environmental structured alignment (Wilson & Izmailov, 2020). In biological and artificial systems, such biases manifest in architectural choices and priors, such as assumptions of symbolic rules or compositional structures. However, in the brain, biases are shaped by evolutionary constraints on neural circuitry, such as hierarchical organization, recurrent connectivity, and plasticity that favor certain patterns of representation and generalization. Together, these mechanisms allow learners to transcend rote pattern recognition and construct models of the world that are adaptable and rooted in causal structure. Current state-of-the-art AI systems, especially LLMs, present a contrasting picture: they are data-hungry, narrow in scope, and brittle under distributional shift. They often fail to internalize goals or causal structures, overwrite prior knowledge during retraining, struggle to transfer skills across domains, and remain unable to showcase novelty in outcome response. The use of evaluation benchmarks, serving as a standardization for comparison, intelligence, and creation of thinking machines, further becomes complicated as benchmarks saturate soon after release and performance gains increasingly reflect task specific optimization (Perlitz et al., 2023). Commercial incentives exacerbate this trend, and iterative model updates can further degrade previously acquired capabilities. As shown in Chen & Zou (2024), GPT-3.5 and GPT-4 exhibited significant declines in code generation, mathematical reasoning, and sensitive question answering performance within three months of user data. As models are tuned for benchmark gains and often fall failure to catastrophic forgetting with new user inputs, there exists a stark disparity between taskspecific optimization and genuine cognitive flexibility (Lake et al., 2017; Wang et al., 2025; Lapuschkin et al. 2019; Geirhos et al., 2020; Ying et al., 2025, Liu et al., 2025). Despite larger model sizes, curated datasets, and enhanced capabilities, AI systems also falter on benchmarks of increasing subject-matter complexity. On GPQA, a test of multi-subject reasoning, models surpass non-expert humans but lag behind domain experts (Rein et al., 2024), whereas on Frontiers Math, focused on advanced mathematical problem-solving, the gap is even greater (Glazer et al., 2024). Furthermore, though these tests probe intuitive understanding within a domain, they neglect the novelty often seen in subject matter commentary, namely the shift away from merely extrapolating data to generating genuine discoveries that both challenge and build upon existing knowledge. While some argue that emergent capabilities within models counter these claims, evidence suggests such phenomena often reflect benchmark design rather than genuine cognitive advances. Schaeffer et al. (2023) showed that nonlinear or discontinuous metric tasks, such as multiple-choice grading, tend to produce apparent emergence, whereas linear or continuous metrics cause these abilities to vanish. On BigBench, emergent behaviors appeared in only 5 of 39 benchmarks. These restrictions are further compounded by extreme architectural energy inefficiencies. Transformer models rely on self-attention mechanisms that scales quadratically with sequence length, resulting in a rapid increase in memory and computation as input size grows. Unsurprisingly, energy requirements and power consumption have risen accordingly. The original Transformer required roughly 4,500 watts during training, whereas models such as PaLM and LLaMA-3.1-405B consumed 2.6 million and 25.3 million watts respectively, a ~600and ~5,000-fold increase (Maslej et al., 2025). Carbon emissions consequently have also risen in parallel, GPT-3 training emitted approximately 588 tons of CO₂, GPT-4 exceeded 5,000 tons, and LLaMA-3.1-405B approached 9,000 tons, all orders of magnitude above the ~18 tons emitted annually 5 by the average American. Thus, we believe, with this trajectory, AI performance gains, and the eventual scaling of AI will come at increasingly disproportionate environmental and computational costs. Just as human cognition supports efficient learning and robust generalization across novel tasks and domains (O’Reilly et al., 2014; Pezzulo et al., 2018; Juechems & Summerfield, 2019), achieving similar versatility in AI will require a shift away from monolithic, task-specific training paradigms, and toward architectures that are multimodal, hierarchical, goal-driven, and multi-timescale dependent. Without such a shift, progress toward artificial general intelligence (AGI) will likely remain constrained. On AGI benchmarks such as ConceptARC, designed to test human-like abstract reasoning and analogical problem solving, GPT-4 achieved 69% accuracy, far below the human average of 95% (Moskvichev & Mitchell, 2023). These tasks, readily learned by toddlers further showcase the disparity between children and machines. Humans solve such tasks quickly by leveraging tractable hypothesis testing on reward-driven and causal circuits, whereas LLMs rely heavily on ridged statistical correlations learned from large-scale data. Though these techniques generally afford deep networks the ability to build models of combinatorial high dimensional spaces, pattern-match, and brute force tractable search solutions for complex tasks, it does remains a brittle approach when faced with novel compositional challenges of flexible cognition as well as limitations on time and resources (Schaeffer et al., 2023; Goyal & Benjio 2022; Goldblum et al, 2023). 1.2 The Boundaries of Modern AI: Compositional Generalization, Embodiment, and Goal Alignment Transformer models combine self-attention and backpropagation to support compositional generalization. In-context learning (ICL) exemplifies this by conditioning models to move beyond mere memorization by demonstration examples at inference time. This enables models to solve complex problems without gradient updates, task-specific fine-tuning, or additional training data, capabilities that have been interpreted as evidence of emergent compositional generalization (Brown et al., 2020; Lake & Baroni, 2023, Dong et al., 2022; Elhage et al., 2021). Yet, growing evidence suggests these generalizations are often superficial and frequently a sophisticated form of pattern matching. Golchin et al. (2024) demonstrate that ICL disproportionately surface memorized training data, with model performance on few-shot tasks heavily influenced by the retrieval of memorized patterns. Furthermore, improvements from zero-shot to manyshot might reflect varying degrees of data leakage, an effect that’s amplified when demonstrations lack explicit labels as models rely on surface-level cues such as input format and distributional alignment. Performance persists even when labels are randomized, suggesting that models use shallow heuristics rather than learning structure. Min et al. (2020) similarly report and reinforce this view; replacing correct labels with random one’s results in minimal performance degradation, implying that format regularities, not semantics, often drive success. This reliance on surface cues contrasts sharply with human generalization, where dependance lies on building world models from correct conceptual associations, and constructing a deterministic understanding of causal or probabilistic dynamics rather than surface level regularities. Recent work reveals that even models with high predictive accuracy fail to internalize basic physical principles. Vafa et al. (2025) show that foundation models do not infer Newtonian dynamics, even when fine-tuned, and instead fall back on task-specific heuristics. Shojaee et al. (2025) comparably finds that large reasoning models perform well on moderate tasks but collapse on more complex variants geared at executive function, with reasoning effort 6 paradoxically decreasing as difficulty increases. Even when given explicit algorithms, models often fail to execute them, revealing a gap between simulating structured reasoning and possessing it. These limitations are especially pronounced under systematic generalization and distributional shift. Out of distribution (ODD) learning requires systematic generalization beyond training data and is a hallmark of cognitive development in humans (Hupkes et al., 2020). However, due to architectural and learning limitations, neural networks often struggle to extrapolate and instead interpolate between training and test data (De Silva et al., 2023). On benchmarks such as SCAN, COGS, and CFQ, models achieve near-perfect in-distribution accuracy but fail on novel compositions, with performance remaining highly sensitive to training conditions and random seed (Kim & Linzen 2020; Lake 2017; Rafailov et al., 2023). Even in domains where AI models demonstrate superhuman performance such as Chess, Go, Atari, Doom, Hide & Seek, and StarCraft (Silver et al.,2017; 2018; Schrittwieser et al., 2020; Vinyals et al., 2019; Ye et al, 2021, Baker et al., 2019) closer inspection reveals deep inefficiencies. These systems train on orders of magnitude more experience than human experts, consuming hundreds of millions of training examples where humans require only thousands (Lake et al., 2017). Likewise, reinforcement learning agents can master complex policies through pattern guided search yet fail to infer latent goals. The data-intensive character of current AI systems underscores a broader point, these models are not engaging in genuine learning, but rather, they construct structured representations for optimized search that convert mapping from sparse samples to dense distributions. Lake et al. (2017) emphasized the dramatic imbalance between human and machine. For instance, AlphaGo accumulating experience from over 100 million games, which far exceeding the lifetime exposure of any professional Go player, likely ~50,000 games. In Atari, while human experts practiced for ~2 hours, agents trained on 200 million frames per game, equivalent to 924 hours of play, or over 500 times more experience. Although a credible direct comparison is difficult between the two groups, as human integrate experience across many tasks and contexts, there is a consensus that human learning is far more sample-efficient, requiring orders of magnitude fewer examples and experiences, holistically shaping a world model of reusable representations. Such disparities raise concerns about the sample efficiency and training data limitations, as well as the above stated energy demands of current AI systems. Humans, including young children, leverage structured compositional priors and causal abstractions to flexibly recombine concepts. This ability enhanced by the capability to infer latent goals aids in building world models of reuseable parts. For instance, a toddler can infer the rules and objectives of a new game from minimal exposure, but agents often misidentify which features are essential to the task objective. These same toddlers also understand within a few instances the repercussions of losing a life in a game or the levels objective, but even though AI models can solve computationally intractable domains, they often fail to infer latent goals or abstract rules even after mastering environment-specific performance. This lack of flexibility highlights a core deficiency in current learning systems, i.e., mastery without understanding. In procedurally generated reinforcement environments, this manifests as agent’s overfitting training correlations and utilizing high-volume experience to aid the converge on effective policies rather than learning the underlying structure for rewards (Koch et al., 2021; Di Langosco et al., 2022). This inherently poses significant challenges when scaling to tasks constrained on the availability, diversity, and efficiency of training data and computational resources, like many tasks in the real-world. Phenomenon such as goal 7 misgeneralization and objective robustness are particularly evident in open-ended or procedurally generated environments where superficial behavioral success masks a deeper failure to represent the task's true structure and objectives. In Koch et al. (2021) and Di Langosco et al. (2022), agents were trained to traverse a maze and collect a coin reliably, however during testing when the coins position was randomized, agents ignored the coin and proceeded directly to the maze’s endpoint. This decoupling of task learning from intent is a failure rarely observed in human learners. In other tasks where agents collected keys with the objective to open chests, the agents mastered the task for training environments that contained twice as many chests as keys. Yet in testing, where keys outnumbered chests, agents prioritized collecting all available keys, often before opening chests. In both these cases, agents seemed to overfit to training-time statistical regularities, and misinterpreting correlation (maze navigation and key availability) as causation (reward acquisition). We believe, these findings reveal a fundamental gap in current machine intelligence: even when agents are proficient at task execution, they often misidentify which features of the environment are essential to task objectives. At the core of this fragility lies the absence of an immersive experience or world model, that serve as a lived-in representation that binds perception, reward, and causal structure into coherent understanding. Without structured representations of causal relationships and their associated rewardpunishment contingencies, agents are confined to heuristic search. Thus, current approaches fail to scale with complexity and make solving complex problems particularly difficult. In scenarios of task switching, reinforcement learning agents succeed through contextual modulation and meta-learning adjustments of policies or latent state representations in response to changing reward contingencies. LLMs, similarly display a parallel illusion of flexibility, and task execution proficiency often masks a shallow grasp of underlying understand. Task switching, that utilizes techniques such as in-context learning serves as a control signal that reconfigures activation patterns across attention heads without changing weights, leading to adaptations that are ephemeral. Other methods such as retrieval augmented generation and finetuning enable selective activation of specialized sub-modules that implement on-demand cognitive restructuring. However, in both RL agents and LLMs, traces of cognitive restructuring, flexible hub reconfiguration, and plasticity appear only in nascent form, remaining architecturally prescribed rather than emergent, with weight-based updates lacking the layered self-organizing temporal dynamics that shape cortical learning. A large limitation arises from the lack of embodied grounding, reward prescription, and consolidation dynamics. Uncertainty control remains symbolic rather than predictive, and restructuring is statistical rather than causal. Nonetheless, architectures that couple retrieval, tool use, and continual optimization hint at the first steps toward persistent, goal-conditioned adaptation, an early precursor of neuromorphic autonomy (Wang et al., 2023; Park et al.,2023). In humans, embodiment grounds learning in direct interaction with cause and effect, physical constrains, and tangible experience enabling concepts to be built from sensory-motor modalities and structured feedback rather than isolated abstract symbol manipulation. Such interactions internalize the dynamics of the environment and transform problems that are computationally intractable into efficiently solvable tasks through structured predictive modes and physical constraints. In computational terms, when a system can model its own dynamics, even intractable problems become tractable through abstraction, constraint satisfaction, structured inference, and continued feedback. In humans, compositional and causal generalization emerges from embodied learning in concert with goal-specific learning, metacognition, and the strategic use of constraints to guide search. Knowledge acquisition is not governed solely by policy- 8 driven reinforcement of state-action transitions focused on reward attainment, but by a complex mesocortical dopaminergic prediction–error system that signals the prediction and gating of reward, unexpected outcomes, absence of anticipated reward, as well as punishment (Mollick et al., 2020). These signal shape not only immediate learning but also the subsequent restructuring of experience, driving decision policies and cognitive strategies across tasks. Fronto-striatal loops that evaluate outcomes also reconfigure task sets, thus enabling flexible rule switching, belief updating under uncertainty, and reorganization of internal models in light of new evidence. In this manner, local episodes of success and failure can scale into global adjustments of how tasks should be represented, attempted, solved, and adapted to novel contexts (Cole et al., 2013; Rigotti et al., 2013). Neurological cognitive restructuring is a hierarchical phenomenon that unfolds primarily through rapid functional retuning rather than wholesale rewiring. The expression of experience through moment-tomoment changes is demonstrated through reconfiguration of fronto-parietal networks or flexible hub, that shift representational geometry within prefrontal and orbitofrontal neural populations and neuromodulatory adjustments that tune learning rates with the value assignments. These fast processes update task sets, beliefs, and control policies on the fly for real-world interaction, and are housed in the inner loop (discussed below). For extended timescales, housed in the outer loop (also discussed below), these longer-lasting structural modifications such as synaptic weight changes, dendritic growth, and systems-level reorganization, only emerge with consolidation, reconsolidation, or exposure of information at different levels of processing, leading to the slow but stabilized creation of new cognitive regimes. During task execution and learning across multiple timescales, restructuring manifests as coordinated neural changes in network topology balancing flexibility and selection (Bassett et. al., 2011). Task switching and novel rule acquisition recruit fronto-parietal circuits that transiently reconfigure whole-brain coupling, with node flexibility predicting subsequent learning. At the representational level, mixed-selectivity neurons in prefrontal cortex enable high-dimensional recoding, effectively rotating population codes to implement new rules without changing anatomical pathways. Studies also show, the revision of beliefs engages a distributed update network spanning dorsal Anterior Cingulate Cortex (dACC)/dorsal medial Prefrontal Cortex (dmPFC), anterior insula, lateral PFC, and parietal cortex, where precision-weighted prediction errors adjust learning rates and arbitrate between competing models (Friston et al., 2010; Huber et al., 2015; O’Reilly 2010; Euston et al., 2012). The orbitofrontal cortex provides a map of latent task states, enabling identical stimuli to be reinterpreted as contingencies shift, such that when memories are reactivated, reconsolidation integrates new evidence into durable circuit changes. Within the default mode network (DMN), situated in the medial prefrontal and posterior cingulate hubs, supporting self-referential models and high-level priors, this restructuring appears as a dynamic shift in segregation, integration, and valuation coupling (Raichle 2015). Rewarding experiences strengthen DMN coupling with mesolimbic circuitry, and the experience of reward enhances task-dependent connectivity between DMN and ventral striatum, and the rise of individual differences (Dobryakova & Smith 2022). Evidence of these interactions suggest positive outcomes retuning self-referential priors through DMN– reward coupling and vmPFC–hippocampal communication consolidates goal-consistent schemas (Van Kesteren et. al., 2010). Lastly, motivational signals further retune restructuring by reallocating control rather than rewiring circuitry. Dorsal anterior cingulate integrates expected value, effort cost, and control 9 demand to set proactive versus reactive modes, while dopaminergic gating in corticostriatal loops regulates working memory updating and policy flexibility. Together, all these findings suggest that immersive experiences reshape cognition, causing dynamic reconfiguration in valuation and memory systems, with the shifting weighting of internal models, yielding durable functional change. Embodiment further grounds these mechanisms in direct interaction with physical cause and effect, enabling concepts to be constructed from sensory–motor modalities and structured feedback rather than abstract symbol manipulation. Such interactions embed environmental dynamics into predictive models, transforming problems that are computationally intractable in the abstract (NP) into tractable ones (P) through abstraction, constraint satisfaction, and structured inference. Compositional and causal generalization emerges from this interplay of embodied learning, goal-specific reinforcement, metacognition, and strategic constraints. Over time, embodiment, reward learning, and cognitive restructuring converge to ground concepts in lived experience. Dopaminergic prediction–error signals reinforce adaptive policies while suppressing ineffective ones, and prefrontal systems restructure and generalize across contexts. Across tasks, information, and rewards, restructuring unfolds as a hierarchy, from computational retuning to representational recoding and network reconfiguration, together consolidating new modes of thought. This synergy allows humans to progress from grasping an object, to abstracting principles of tool use, and eventually flexibly reconfiguring individual learnings into collective strategies for novel environments. Flexible cognition thus emerges not from isolated symbolic routines, but from the continuous cycle of embodied interaction, reward-guided updating, and restructuring of representations across tasks. 1.3 The Architecture of Learning: Dual Loop Shared Workspaces and Hierarchical Memory The absence of a shared neural representational workspace constrains current AI systems, locking them into narrow domains of competence and rendering them brittle when extrapolating beyond familiar distributions (Marcus, 2018). In contrast, the brain deploys high-dimensional population codes to flexibly encode experience (Stringer et al., 2019; Gao & Ganguli, 2015), projecting these representations onto lowdimensional manifolds that compress structure while preserving generative capacity. This dimensionality reduction not only stabilizes neural dynamics but also provides a substrate for abstraction and generalizable thought. Within the brain, the global workspace, is a distributed network that integrates compact representations across dispersed modules, supporting the sharing, recombination, and reuse for reasoning, thinking, and planning. Deep networks, with their massive parallel feature extractors, can approximate aspects of System I intuition, while explicit reasoning modules in symbolic or neuro-symbolic systems mimic aspects of System II, enabling context sensitive stepwise deliberation (LeCun et al., 2015). Human cognition, in contrast, integrates these processes seamlessly as subsystems that emerge from a single underlying optimization framework; fast, context-sensitive priors shape deliberate cognition, while structured inference refines and reshapes those priors over time. This bidirectional exchange allows flexible adaptation to novel situations, a capacity absent in current AI systems, which overfit to familiar patterns and falter under distributional shifts. Consistent with global workspace theory, the brains shared workspace emerges from prefrontal- 16 consolidating knowledge. Through connectivity between the loops at different levels of the hierarchy, the system can iterate over information across multiple levels of consolidation, scale, and complexity. This organization enables experience to progress from rapid adaptation to refinement and, ultimately, to the formation of long-term abstract durable knowledge. Learned attractors in the outer loop guide efficient exploration and decision-making in the inner loop, while predictive error signals determine which states are elevated for global dissemination and integration. The system thus minimizes energy under both local and global constraints, integrating associative memory, goal-directed learning, and uncertainty estimation. Over iterative cycles, this coupling begins to support schema formation and hierarchical control, analogous to hippocampal–neocortical interactions in biological consolidation, preserving both efficiency and coherence, and achieving the balance between parallel and serial search that biological cognition exemplifies. Consistent with the Global Workspace Theory, in which neural populations broadcast information across specialized thalamo-cortical circuits, our framework, defines the workspace as the mediating interface that links outer-loop reflection with inner-loop action.The thalamic gating mechanisms that enable broadcast in the brain, are here driven by prediction, attention, and an overarching computational uncertainty objective to consolidate controls and determine which inner-loop states are elevated for global long-term dissemination and storage. Over iterative cycles, this coordination yields increasingly refined abstractions and schemas, much like the biological integration of hippocampal rapid learning with slower neocortical consolidation. As outer-loop knowledge is reintroduced into the real-world through the inner loop, it establishes a continuous cycle of refinement and expansion. As the outer loop shapes more effective real-world search and decision-making strategies, optimizing energy and time efficiency, the inner loop gathers sensory input for consolidation and updating through embodiment and reward. Over successive iterations, this bidirectional exchange creates increasingly sophisticated abstractions and knowledge clusters within the outer loop, scaffolding more flexible and powerful forms of cognition. In this manner, the same foundational mechanisms that support simple learning apply to scaling behavior to more complex forms. 3. Conclusion In this work, we established the conceptual foundations of our framework, proposing that intelligence arises not from monolithic optimization but from dynamic communication among processes operating across multiple temporal and representational scales. We outline the neural principles underlying a dual-loop architecture, in which a fast inner loop enables real-time embodied adaptation, while a slower outer loop supports abstraction and long-horizon planning. This discussion considers how intelligence emerges in the brain and how these same principles can extend the capabilities of machine intelligence, addressing key limitations in current architectures. The three challenges highlighted here are mitigated through energyefficient computation achieved via sparse, recurrent dynamics; goal-directed learning maintained through predictive feedback and utility optimization; and multi-timescale consolidation that unifies fast real-time embodied adaptation with slower consolidation and long-horizon planning. By framing these mechanisms as interdependent processes operating across distinct timescales, we identify a minimal set of organizing principles through which cognition can remain both scalable and generalizable. In Part II, we formalize these principles within a theoretical, neuro-inspired dynamical systems framework that couples distributed experts through shared energy constraints, demonstrating how structured interactions among heterogeneous 17 subsystems, each pursuing complementary objectives, can yield stable, adaptive, and compositional behavior. Taken together, these two works advance a unified account of how efficient computation, embodied control, and hierarchical memory interact to produce the flexible, integrated intelligence characteristic of both biological and artificial systems. References Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., ... & McGrew, B. (2023). Gpt-4 technical report. arXiv preprint arXiv:2303.08774 Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., ... & Ballas, N. (2023). Selfsupervised learning from images with a joint-embedding predictive architecture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 15619-15629). Baars, B. J. (1993). A cognitive theory of consciousness. Cambridge University Press. Baars, B. J. (1997). In the theater of consciousness: The workspace of the mind. Oxford University Press, USA. Baars, B. J. (2002). The conscious access hypothesis: origins and recent evidence. Trends in cognitive sciences, 6(1), 47-52. Baars, B. J. (2005). Global workspace theory of consciousness: toward a cognitive neuroscience of human experience. Progress in brain research, 150, 45-53. Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., & Mordatch, I. (2019, September). Emergent tool use from multi-agent autocurricula. In International conference on learning representations. Bassett, D. S., Wymbs, N. F., Porter, M. A., Mucha, P. J., Carlson, J. M., & Grafton, S. T. (2011). Dynamic reconfiguration of human brain networks during learning. Proceedings of the National Academy of Sciences, 108(18), 7641-7646. Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., ... & Zitkovich, B. (2023). Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818. Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33, 18771901. Chen, L., Zaharia, M., & Zou, J. (2024). How is ChatGPT’s behavior changing over time?. Harvard Data Science Review, 6(2). 18 Cole, M. W., Reynolds, J. R., Power, J. D., Repovs, G., Anticevic, A., & Braver, T. S. (2013). Multi-task connectivity reveals flexible hubs for adaptive task control. Nature neuroscience, 16(9), 1348-1355. Cushman, F., & Morris, A. (2015). Habitual control of goal selection in humans. Proceedings of the National Academy of Sciences, 112(45), 13817-13822. De Havas, J. A., Parimal, S., Soon, C. S., & Chee, M. W. (2012). Sleep deprivation reduces default mode network connectivity and anti-correlation during rest and task performance. Neuroimage, 59(2), 17451751. De Martino, B., Kumaran, D., Seymour, B., & Dolan, R. J. (2006). Frames, biases, and rational decisionmaking in the human brain. science, 313(5787), 684-687. Di Langosco, L. L., Koch, J., Sharkey, L. D., Pfau, J., & Krueger, D. (2022, June). Goal misgeneralization in deep reinforcement learning. In International Conference on Machine Learning (pp. 12004-12019). PMLR. Dobryakova, E., & Smith, D. V. (2022). Reward enhances connectivity between the ventral striatum and the default mode network. NeuroImage, 258, 119398. Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., ... & Sui, Z. (2022). A survey on in-context learning. arXiv preprint arXiv:2301.00234. Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Wahid, A., ... & Florence, P. (2023). Palme: An embodied multimodal language model. Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., ... & Olah, C. (2021). A mathematical framework for transformer circuits. Transformer Circuits Thread, 1(1), 12. Engle, R. W. (2002). Working memory capacity as executive attention. Current directions in psychological science, 11(1), 19-23. Euston, D. R., Gruber, A. J., & McNaughton, B. L. (2012). The role of medial prefrontal cortex in memory and decision making. Neuron, 76(6), 1057-1070. Friston, K. (2010). The free-energy principle: a unified brain theory?. Nature reviews neuroscience, 11(2), 127-138. Gandhi, K., Stojnic, G., Lake, B. M., & Dillon, M. R. (2021). Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of others. Advances in neural information processing systems, 34, 9963-9976 Gao, P., & Ganguli, S. (2015). On simplicity and complexity in the brave new world of large-scale neuroscience. Current opinion in neurobiology, 32, 148-155. 19 Geirhos, R., Jacobsen, J. H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., & Wichmann, F. A. (2020). Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11), 665-673. Gladstone, A., Nanduru, G., Islam, M. M., Han, P., Ha, H., Chadha, A., ... & Iqbal, T. (2025). EnergyBased Transformers are Scalable Learners and Thinkers. arXiv preprint arXiv:2507.02092. Glazer, E., Erdil, E., Besiroglu, T., Chicharro, D., Chen, E., Gunning, A., ... & Wildon, M. (2024). Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai. arXiv preprint arXiv:2411.04872. Golchin, S., Surdeanu, M., Bethard, S., Blanco, E., & Riloff, E. (2024). Memorization in in-context learning. arXiv preprint arXiv:2408.11546. Goldblum, M., Finzi, M., Rowan, K., & Wilson, A. G. (2023). The no free lunch theorem, kolmogorov complexity, and the role of inductive biases in machine learning. arXiv preprint arXiv:2304.05366. Goyal, A., & Bengio, Y. (2022). Inductive biases for deep learning of higher-level cognition. Proceedings of the Royal Society A, 478(2266), 20210068. Hafner, D., Pasukonis, J., Ba, J., & Lillicrap, T. (2023). Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104. Herd, S., Krueger, K., Nair, A., Mollick, J., & O’Reilly, R. (2021). Neural mechanisms of human decision-making. Cognitive, Affective, & Behavioral Neuroscience, 21(1), 35-57. Horovitz, S. G., Braun, A. R., Carr, W. S., Picchioni, D., Balkin, T. J., Fukunaga, M., & Duyn, J. H. (2009). Decoupling of the brain's default mode network during deep sleep. Proceedings of the National Academy of Sciences, 106(27), 11376-11381. Huber, R. E., Klucharev, V., & Rieskamp, J. (2015). Neural correlates of informational cascades: brain mechanisms of social influence on belief updating. Social Cognitive and Affective Neuroscience, 10(4), 589-597. Hupkes, D., Dankers, V., Mul, M., & Bruni, E. (2020). Compositionality decomposed: How do neural networks generalise?. Journal of Artificial Intelligence Research, 67, 757-795. Juechems, K., & Summerfield, C. (2019). Where does value come from?. Trends in cognitive sciences, 23(10), 836-850. Kahneman, D. (2011). Fast and slow thinking. Allen Lane and Penguin Books, New York. Kanki, B. G. (2018). Cognitive functions and human error. In Space safety and human performance (pp. 17-52). Butterworth-Heinemann. 20 Ke, Z., Jiao, F., Ming, Y., Nguyen, X. P., Xu, A., Long, D. X., ... & Joty, S. (2025). A survey of frontiers in llm reasoning: Inference scaling, learning to reason, and agentic systems. arXiv preprint arXiv:2504.09037. Kim, N., & Linzen, T. (2020). COGS: A compositional generalization challenge based on semantic interpretation. arXiv preprint arXiv:2010.05465. Kim, Y., Kim, D., Choi, J., Park, J., Oh, N., & Park, D. (2024). A survey on integration of large language models with intelligent robots. Intelligent Service Robotics, 17(5), 1091-1107. Koch, J., Langosco, L., Pfau, J., Le, J., & Sharkey, L. (2021). Objective robustness in deep reinforcement learning. arXiv preprint arXiv:2105.14111, 2. Kováč, L. (2010). The 20 W sleep‐walkers. EMBO reports, 11(1), 2-2. Lake, B. M., & Baroni, M. (2023). Human-like systematic generalization through a meta-learning neural network. Nature, 623(7985), 115-121. Lake, B. M., Ullman, T. D., Tenenbaum, J. B., & Gershman, S. J. (2017). Building machines that learn and think like people. Behavioral and brain sciences, 40, e253. Lapuschkin, S., Wäldchen, S., Binder, A., Montavon, G., Samek, W., & Müller, K. R. (2019). Unmasking Clever Hans predictors and assessing what machines really learn. Nature communications, 10(1), 1096. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. nature, 521(7553), 436-444. Lee, S. W., Shimojo, S., & O’doherty, J. P. (2014). Neural computations underlying arbitration between model-based and model-free learning. Neuron, 81(3), 687-699. Lenat, D. B., Guha, R. V., Pittman, K., Pratt, D., & Shepherd, M. (1990). Cyc: toward programs with common sense. Communications of the ACM, 33(8), 30-49. Lenat, D. B., & Guha, R. V. (1989). Building large knowledge-based systems; representation and inference in the Cyc project. Addison-Wesley Longman Publishing Co., Inc. Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., ... & Piao, Y. (2024). Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 Liu, B., Li, X., Zhang, J., Wang, J., He, T., Hong, S., ... & Wu, C. (2025). Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems. arXiv preprint arXiv:2504.01990. Marcus, G. (2018). Deep learning: A critical appraisal. arXiv preprint arXiv:1801.00631. 21 Maslej, N., Fattorini, L., Perrault, R., Gil, Y., ....& Oak, S. (2025). Artificial Intelligence Index Report 2025. Stanford Institute for Human-Centered Artificial Intelligence. arXiv preprint arXiv:2504.07139 McClelland, J. L., McNaughton, B. L., & O'Reilly, R. C. (1995). Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory. Psychological review, 102(3), 419. Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., & Zettlemoyer, L. (2022). Rethinking the role of demonstrations: What makes in-context learning work?. arXiv preprint arXiv:2202.12837. Mollick, J. A., Hazy, T. E., Krueger, K. A., Nair, A., Mackie, P., Herd, S. A., & O'Reilly, R. C. (2020). A systems-neuroscience model of phasic dopamine. Psychological review, 127(6), 972. Moskvichev, A., Odouard, V. V., & Mitchell, M. (2023). The conceptarc benchmark: Evaluating understanding and generalization in the arc domain. arXiv preprint arXiv:2305.07141. Nair, A. (2021). A Mathematical Approach to Constraining Neural Abstraction and the Mechanisms Needed to Scale to Higher-Order Cognition. arXiv preprint arXiv:2108.05494. Nair, A., Austin, E. E., Watson, J. M., & Banaei-Kashani, F. (2025). Thinking Machines II: A dual-system framework for metacognitive control and learning [Manuscript in preparation]. Nair, A., & Banaei-Kashani, F. (2022). Bridging the gap between artificial intelligence and artificial general intelligence: A ten commandment framework for human-like intelligence. arXiv preprint arXiv:2210.09366. Noyes, J. (2001). Human error. People in Control: Human factors in control room design, 3-16. O’Reilly, R. C. (2010). The what and how of prefrontal cortical organization. Trends in neurosciences, 33(8), 355-361. O’Reilly, R. C. (2020). Unraveling the mysteries of motivation. Trends in cognitive sciences, 24(6),425434. O'Reilly, R. C., Hazy, T. E., Mollick, J., Mackie, P., & Herd, S. (2014). Goal-driven cognition in the brain: a computational framework. arXiv preprint arXiv:1404.7591. O’Reilly, R. C., Herd, S. A., & Pauli, W. M. (2010). Computational models of cognitive control. Current opinion in neurobiology, 20(2), 257-261. O'Reilly, R. C., & Munakata, Y. (2000). Computational explorations in cognitive neuroscience: Understanding the mind by simulating the brain. MIT press. 22 O’Reilly, R. C., Nair, A., Russin, J. L., & Herd, S. A. (2020). How sequential interactive processing within frontostriatal loops supports a continuum of habitual to controlled processing. Frontiers in Psychology, 11, 380. Panton, K., Matuszek, C., Lenat, D., Schneider, D., Witbrock, M., Siegel, N., & Shepard, B. (2006). Common sense reasoning–from Cyc to intelligent assistant. In Ambient Intelligence in Everyday Life: Foreword by Emile Aarts (pp. 1-31). Berlin, Heidelberg: Springer Berlin Heidelberg. Park, J. S., O'Brien, J., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023, October). Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology (pp. 1-22). Perlitz, Y., Bandel, E., Gera, A., Arviv, O., Ein-Dor, L., Shnarch, E., ... & Choshen, L. (2023). Efficient benchmarking of language models. arXiv preprint arXiv:2308.11696. Pessoa, L. (2008). On the relationship between emotion and cognition. Nature reviews neuroscience, 9(2), 148-158. Pezzulo, G., Rigoli, F., & Friston, K. J. (2018). Hierarchical active inference: a theory of motivated control. Trends in cognitive sciences, 22(4), 294-306. Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 53728-53741. Raichle, M. E. (2015). The brain's default mode network. Annual review of neuroscience, 38(1), 433-447. Rasch, B., & Born, J. (2013). About sleep's role in memory. Physiological reviews. Rigotti, M., Barak, O., Warden, M. R., Wang, X. J., Daw, N. D., Miller, E. K., & Fusi, S. (2013). The importance of mixed selectivity in complex cognitive tasks. Nature, 497(7451), 585-590. Russin, J., O’Reilly, R. C., & Bengio, Y. (2020). Deep learning needs a prefrontal cortex. Work Bridging AI Cogn Sci, 107(603-616), 1. Sämann, P. G., Tully, C., Spoormaker, V. I., Wetter, T. C., Holsboer, F., Wehrle, R., & Czisch, M. (2010). Increased sleep pressure reduces resting state functional connectivity. Magnetic Resonance Materials in Physics, Biology and Medicine, 23(5), 375-389 Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are emergent abilities of large language models a mirage?. Advances in Neural Information Processing Systems, 36, 55565-55581. Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., ... & Silver, D. (2020). Mastering atari, go, chess and shogi by planning with a learned model. Nature, 588(7839), 604-609. 23 Shani, C., Soffer, L., Jurafsky, D., LeCun, Y., & Shwartz-Ziv, R. (2025). From tokens to thoughts: How LLMs and humans trade compression for meaning. arXiv preprint arXiv:2505.17117 Shojaee, P., Mirzadeh, I., Alizadeh, K., Horton, M., Bengio, S., & Farajtabar, M. (2025). The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity. arXiv preprint arXiv:2506.06941. Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., ... & Hassabis, D. (2018). A general reinforcement learning algorithm that masters chess, shogi, and Go through selfplay. Science, 362(6419), 1140-1144. Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., ... & Hassabis, D. (2017). Mastering the game of go without human knowledge. nature, 550(7676), 354-359. Stanford Institute for Human-Centered Artificial Intelligence. (2024). Artificial Intelligence Index Report 2024. Stanford University. https://hai.stanford.edu/ai-index/2024-ai-index-report Stringer, C., Pachitariu, M., Steinmetz, N., Carandini, M., & Harris, K. D. (2019). High-dimensional geometry of population responses in visual cortex. Nature, 571(7765), 361-365. Stojnić, G., Gandhi, K., Yasuda, S., Lake, B. M., & Dillon, M. R. (2023). Commonsense psychology in human infants and machines. Cognition, 235, 105406. Team, G., Anil, R., Borgeaud, S., Alayrac, J. B., Yu, J., Soricut, R., ... & Blanco, L. (2023). Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805. Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M. A., Lacroix, T., ... & Lample, G. (2023). Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971. Tversky, A., & Kahneman, D. (1981). The framing of decisions and the psychology of choice. science, 211(4481), 453-458. Vafa, K., Chang, P. G., Rambachan, A., & Mullainathan, S. (2025). What has a foundation model found? using inductive bias to probe for world models. arXiv preprint arXiv:2507.06952. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30 Van Kesteren, M. T., Fernández, G., Norris, D. G., & Hermans, E. J. (2010). Persistent schema-dependent hippocampal-neocortical connectivity during memory encoding and postencoding rest in humans. Proceedings of the National Academy of Sciences, 107(16), 7550-7555. Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., ... & Silver, D. (2019). Grandmaster level in StarCraft II using multi-agent reinforcement learning. nature, 575(7782), 24 350-354. Wang, W., Jiang, G., Linzen, T., & Lake, B. M. (2025). Rapid Word Learning Through Meta In-Context Learning. arXiv preprint arXiv:2502.14791. Wang, Z., Cai, S., Chen, G., Liu, A., Ma, X., & Liang, Y. (2023). Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents. arXiv preprint arXiv:2302.01560. Wilson, A. G., & Izmailov, P. (2020). Bayesian deep learning and a probabilistic perspective of generalization. Advances in neural information processing systems, 33, 4697-4708. Ying, L., Collins, K. M., Wong, L., Sucholutsky, I., Liu, R., Weller, A., ... & Tenenbaum, J. B. (2025). On Benchmarking Human-Like Intelligence in Machines. arXiv preprint arXiv:2502.20502. Zhou, G., Pan, H., LeCun, Y., & Pinto, L. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning 2024. URL https://arxiv. org/abs/2411, 4983.