scieee AI-readable full text Open interactive document viewer

A Foundational Model for Biological Logic and Disease

Sinha, Saptarshi; Ghosh, Pradipta

Abstract

Biology and medicine mistake complexity for understanding. We build black-box AI models and stockpile terabytes of omics data, yet the logic of life remains hidden in plain sight. The cell, nature’s smallest decision-maker computes not by probability, but by logic: “on” or “off,” “commit” or “retract.” Boolean mathematics decodes this digital decision-making, transforming analog molecular noise into invariant “if–then” rules that persist across tissues, species, and diseases. Simultaneously, by quantifying how populations shift between binary states, Boolean frameworks map disease as an analog continuum—dynamic, graded, reversible and objectively measurable. By deciphering both the binary decisions and the graded population shifts that shape disease, this approach replaces static biomarkers with mechanistic rules—and advances a new premise: if life computes in logic, medicine should too.

Full text

1 | Page Perspective Title: A Foundational Model for Biological Logic and Disease Running title: If Life Computes with Logic—So Should Medicine Authors Saptarshi Sinha1, 3* and Pradipta Ghosh1-3* Departments of 1Cellular and Molecular Medicine and 2Medicine, University of California, San Diego, CA, 92093, USA. UC San Diego Institute for Network Medicine, University of California, San Diego, CA, 92093, USA. *Correspondence to [email protected] (SS) or [email protected] (P.G). Keywords (two to six) Boolean Implication Networks · Invariants · Systems Biology · StepMiner · Dose–Response Alignment · Cellular Intelligence · Function-Agnostic Modeling Abstract (125 words) Biology and medicine mistake complexity for understanding. We build black-box AI models and stockpile terabytes of omics data, yet the logic of life remains hidden in plain sight. The cell, nature’s smallest decision-maker computes not by probability, but by logic: “on” or “off,” “commit” or “retract.” Boolean mathematics decodes this digital decision-making, transforming analog molecular noise into invariant “if–then” rules that persist across tissues, species, and diseases. Simultaneously, by quantifying how populations shift between binary states, Boolean frameworks map disease as an analog continuum—dynamic, graded, reversible and objectively measurable. By deciphering both the binary decisions and the graded population shifts that shape disease, this approach replaces static biomarkers with mechanistic rules—and advances a new premise: if life computes in logic, medicine should too. 2 | Page When Biology Defies Mathematics The laws of nature, from Newton’s gravitation to Einstein’s relativity, astonish us by their simplicity, invariance, and predictive power. These mathematical regularities, i.e., conditional statements describing what must follow from what is, form the scaffolds of physics. Biology (and Medicine), however, refuses such precision. For centuries, the living world has resisted the kind of mathematical formalism that so elegantly explains the cosmos. The cell—life’s smallest decision-making unit—is a system of staggering complexity, yet it operates on remarkably simple, universal principles. Within it, thousands of molecules interact through circuits that sense, decide, and act. These decisions—whether to divide, differentiate, migrate, or die—trace the continuum of transition states between health and disease. Despite breathtaking advances in molecular profiling and AI, our ability to extract the invariant rules that govern these transitions remains limited. Some biologists argue this is not due to ignorance, but due to impossibility(2): states of life—evolution, emergence, adaptation—defies closure within fixed mathematical systems. Concepts such as “genes”, “species”, or “fitness” remain context dependent and fluid. Biology transcends computation: it is a science of thresholds, feedback, and exceptions. And yet, if mathematics cannot define life, perhaps it can describe its logic. Boolean frameworks, through sophisticated thresholding(3), embrace rather than erase biological complexity, translating noise into rules, randomness into predictability. Instead of exact equations, they seek invariant rules—binary “if–then” relationships that persist across tissues, species, and diseases. These Boolean implications, first formalized by Sahoo et al.(4, 5), expose the simple decision-making architecture within the cellular storm. At its core lies StepMiner(1), an thresholding algorithm that converts noisy gene-expression data into binary outcomes—“low (0)” or “high (1)”—mirroring how cells themselves decide to act. In doing so, Boolean logic offers something rare in modern biology: a mathematical abstraction that captures life’s decision logic, not its descriptive chaos. In an era of black box AI, Boolean framework stands apart. It is not a simulation of complexity but a simplification that reveals universals, a quiet order beneath the stochastic noise of biology, governed by rules as elegant and predictive as any law of physics [Figure 1; BOX 1]: 3 | Page Rule 1 | Life Computes in Thresholds Every living cell is a decision engine. Molecules fluctuate continuously, but at critical junctures the system commits — to divide or arrest, differentiate or retain stemness, activate immunity or remain quiescent. These commitments are neither probabilistic nor gradual; they are threshold crossings. A cell does not partly divide, somewhat differentiate, or tentatively die. It flips. Thresholds are evolution’s answer to noise: regulatory circuits convert variable molecular inputs into binary outcomes, i.e., the molecular equivalents of yes/no, go/stop, live/die. These switches mark the points where possibility becomes physiological action. The Boolean framework captures this architecture mathematically. At its core lies StepMiner(1), an adaptive thresholding algorithm that detects inflection points in gene-expression profiles (Figure 2A). StepMiner fits a step function to each gene’s sorted expression values to determine the most significant transition separating “off (0)” from “on (1)” states. This converts noisy continuous data into clear low (0) or high (1) states, with a noise buffer removing uncertainty around the threshold. This digital abstraction is not an oversimplification; it mirrors how cells compute: • Fate choice: mutually inhibitory transcription factors enforce lineage commitment. • Apoptosis: feedback-driven switches distinguish survival vs. death. • Immune activation: T cells fully engage only beyond receptor–signal thresholds. By imposing such thresholds, the Boolean framework transforms high-dimensional, stochastic omics data into interpretable decision landscapes, and converts continuous molecular inputs into binary outcomes through signal transduction, gene regulation, and metabolic feedback. Box 1. The Five Rules of Biological Logic 1. Life Computes in Thresholds – Thresholding algorithms such as StepMiner(1) convert noisy expression into decisive 0/1 states, revealing where cells commit to differentiation, survival, activation, or death — the binary logic underlying biological decisions. 2. Noise Hides Invariants – Boolean implication reveals relationships that persist despite biological or experimental noise, exposing rules that hold across tissues, species, and disease. 3. Feedback Is Fidelity – Boolean edges trace push–pull signaling architectures that maintain Dose–Response Alignment (DoRA), identifying the checkpoints where proportionality is preserved when biology maintains order and where dysregulation begins. 4. Disease Is a Reversible Continuum – Boolean implications define the directionality of health-todisease transitions, while analog composite scores (StepMiner-normalized) position each sample along a reversible, quantifiable continuum. 5. Function Follows Logic, Not Annotation – Composite scores from data-derived Boolean networks retain analog behavior while exposing causal constraints and enable function-agnostic benchmarking, defining biological meaning through empirical invariants rather than evolving ontologies, thereby improving target discovery and model fidelity. 4 | Page Crucially, digital abstraction at this step does not discard analog information that is vital to track cellular processes in multicellular life that continues to evolve through complex cell-cell and cell-environment crosstalk. The StepMiner-normalized continuous score is retained and becomes foundational in: Rule 4: placing samples along the analog health → disease continuum and in Rule 5: quantifying population-level drift and therapeutic reversibility. Although the Boolean framework can, in principle, operate on any omics layer, we prioritized the transcriptome for its optimal information density and signal-to-noise ratio. Transcriptomic data offer the breadth of genomics (>20,000 features) with measurable dynamic range across cellular states, while proteomic(6, 7) and metabolomic datasets remain limited by coverage (<60%), batch variability, and missing values due to lowabundance analytes and incomplete annotation of post-translational or allosteric modifications. By contrast, RNA expression captures integrated outputs of genetic and epigenetic regulation, as well as the feedback loops that maintain homeostasis (Rule 3: Feedback Is Fidelity). Empirically, transcript levels correlate more strongly with clinical phenotypes than protein abundance(8), and transcriptome–proteome concordance increases under to stress(9-11), when cells engage conserved adaptive programs. Thus, the transcriptome offers the most comprehensive and quantitative window into cellular decision logic currently available. From Analog Chaos to Digital Clarity: StepMiner and the Birth of BINs Traditional gene-network methods, e.g., co-expression, mutual information, Bayesian inference, are symmetric and probabilistic, capturing correlation but not direction. They struggle to extract meaning or retain reproducibility in the presence of real-world heterogeneity and noise. Boolean Implication Networks (BINs)(3) changed this paradigm by introducing logic as biology’s native language: simple “if–then” rules linking high/low gene-expression states. The resulting network of Boolean implication relationships (BIRs)—such as “A high ⇒ B high” or “A high ⇒ B low” (Figure 2B)—forms a digital fingerprint of cellular logic. Each gene becomes a logical variable, and each Boolean implication forms a constraint — a rule biology consistently obeys. Together, these relationships produce a digital fingerprint of how the system thinks, revealing direction, dependency, and hierarchy as statistical facts rather than assumptions. In biology and in math, robustness arises from thresholds. Life decides digitally and the Rule 1 in the Boolean framework measures those decisions. Rule 2 | Noise Hides Invariants If thresholding exposes the moment of decision, Boolean implication uncovers the rules that endure (Figure 2B). Biology is inherently noisy—gene expression fluctuates, proteins misfold or misfire, pathways crosstalk— 5 | Page yet within this apparent chaos lie invariants: relationships that remain true across tissues, species, and perturbations. These are nature’s constraints, the hidden grammar of gene-regulatory logic that evolution preserves. The Boolean framework detects these invariants not just through symmetric correlation but also through asymmetric logic (Figure 2B). For every gene pair (A, B) across thousands of samples, their joint expression pattern, after thresholding, divides the samples into four groups (or four quadrants in a two-dimensional plot)— A low/high × B low/high. When one quadrant is significantly underpopulated, indicating that a particular combination of states almost never occurs, an if–then Boolean implication is established (Figure 2B). This asymmetric exclusion reveals the directional constraints that biology consistently enforces. The Six Boolean Relationships as Natural Invariants 1. Four asymmetric: “A high ⇒ B high,” “A high ⇒ B low,” “A low ⇒ B high,” “A low ⇒ B low” 2. Two symmetric: “Equivalent” (A ⇔ B), or “Opposite” (A high ⇒ B low AND B high ⇒ A low) Each relationship represents an empirical invariant, a directional, data-driven relationship that holds across diverse biological contexts. These are not hypotheses; they are statistical facts, distilled directly from huge real-world expression data. They define which molecular states co-exist, which exclude one another, and which form nested hierarchies of regulation. They represent constraints enforced at the cellular level and sustained through heterotypic feedback and crosstalk among cells within tissues—emergent rules that hold at the population scale. Across cancers, developmental programs, and cross-species comparisons, such invariants repeatedly surface: • B-cell maturation: Stem marker high ⇒ differentiation marker low pinpoints transitional genes defining the midpoints of developmental hierarchies(12). These Boolean intermediates helped identify previously unrecognized regulators of lineage progression—genes silent in stem states but active as differentiation begins. • Bladder cancer: KRT5 high ⇒ KRT20 low demarcates basal progenitors from luminal cells(13), defining a binary architecture of epithelial plasticity. This rule persists across patient samples and species, signifying a deeply conserved regulatory toggle between basal identity and terminal differentiation. • Colon tumors: CA1 high ⇒ KRT20 high captures a nested lineage constraint in differentiated tumors, reflecting hierarchical fidelity within intestinal epithelium. Conversely, ALCAM or CCDC88A high ⇒ CDX2 or PRKAB1 low marks a stemness axis—cells reverting toward an undifferentiated, regenerative program, often predictive of therapeutic resistance and poor prognosis. • Lung injury and fibrosis: ACE2 high ⇔ IL15/IL15RA high signals a conserved epithelial–immune coregulation circuit stable across human and mouse models(14-17). 6 | Page • Macrophage polarization: Conserved implication patterns trace immune trajectories from reactivity to tolerance, revealing logical paths that remain fixed across tissues and species(18). • Inflammatory bowel disease (IBD): Genes preserving epithelial integrity and energy homeostasis share a high ⇒ low relationship with those driving inflammation and fibrosis—computationally trackable across organoids, animal models, and patient biopsies.(3, 19). Together, these and numerous other examples(20-22) demonstrate how the Boolean framework reproducibly extracts invariant logic beneath biological diversity. Whether in hematopoiesis or fibrosis, immune tolerance or epithelial regeneration, the same mathematical grammar applies, i.e., directional relationships, conserved across scales, that define what biology allows or forbids. Boolean relationships thus emerge as constraints imposed by life itself; rules so fundamental that they persist through evolution and pathology alike. The Boolean logic redefines “mechanism” not as a static pathway diagram but as a set of logical constraints which represent the boundaries of biological possibility that remain stable through evolution, adaptation and disease. Rule 3 | Feedback Is Fidelity Life maintains order not by silencing noise but by aligning feedback to preserve proportionality. This is better known as Dose–Response Alignment(23) (DoRA), a principle that describes how cells preserve faithful information transfer by aligning the input-to-output relationship across signaling cascades(24). When feedback loops achieve DoRA, downstream responses scale predictably with receptor activation, converting molecular chaos into coherent behavior. Mechanistically, DoRA is achieved not by fine-tuning every parameter but by architectural motifs(25, 26)—push-pull regulatory networks in which the active form of a signaling species drives output (“push”) while the inactive form counter-acts it (“pull”)(27). The Boolean framework captures these equilibria at the transcriptomic level in complex biological circuits(26, 28). Each Boolean edge, A high ⇒ B high or A high ⇒ B low, represents a feedback-stabilized dependency, the digital trace of push–pull motifs that maintain biological fidelity. By elevating DoRA from molecular cascades to transcriptomic logic, the Boolean framework reveals the transistor-like decision points in cellular regulation—where biology imposes invariant logic rather than variable curves. These are the nodes where “noise” is filtered, fidelity is enforced, and decisions become digital. In this way, the DoRA concept serves as the bridge from analog chaos of molecular networks to the digital clarity of Boolean implication. Remarkably, Boolean relationships often align with protein-level behaviors (with surprising fidelity), as confirmed by cytochemistry-based validations(3, 29-31). This suggests that many transcriptional implications are reinforced by feed-forward loops(32) coupling transcription and translation. Still, not all alignments signify activation: negative feedback loops(33) may sustain transcriptional signals while functionally driving 7 | Page resolution or repair. In disease contexts, the Boolean framework(3) helps disentangle these dualities— revealing when persistence of expression reflects stability, and when it encodes homeostatic correction. Ultimately, feedback is fidelity: Boolean invariants reveal where biology enforces proportionality, transforming stochastic expression into reproducible logic. Rule 4 | Disease Is a Reversible Continuum, Not an Absolute Category Disease rarely begins at diagnosis. Across cancers, fibrosis, autoimmunity, and degeneration, decades of evidence show that pathology emerges as a slow, often reversible drift, either genetic, epigenetic, or posttranslational state-dependent, long before clinical diagnostic thresholds are crossed(34-39). The Boolean framework(3) captures this continuity by mapping BIRs that that define when regulatory decisions flip from stability to instability (Figure 3A-C). Yet the framework also preserves analog information: StepMinernormalized composite scores (modified Z-scores; elaborated in Rule 5) quantify how many cells occupy each state. This enables samples to be positioned along a disease trajectory, revealing graded physiological change. In other words: BIRs define direction, i.e., the order in which transitions occur, and composite scores define position, i.e., how far a sample has progressed. Using graph-based temporal networks, the framework extracts pseudotime and state-transition paths from crosssectional datasets (Figure 3D-G), anchoring each patient, organoid, or perturbation along a quantifiable, twomode health–disease landscape: • digital: “which switches have flipped?” • analog: “how far has the switch propagated across the system?” A population-level majority-rule principle(40) ensures that Boolean decisions scale to tissue behavior: when most cells enter a disease-aligned state, the phenotype follows — as validated in macrophage polarization studies spanning reactive ↔ tolerant extremes(18, 41). This majority rule is supported by its demonstrated effectiveness in achieving consensus and coordinating behavior in complex, interconnected networks, especially when information transfer is decentralized, noisy or incomplete. When most cells shift into a disease-associated state, the tissue-level phenotype follows. Thus, the framework provides testable insight into where recovery remains possible and which interventions may revert state drift, already demonstrated in multiple biological contexts(3, 18, 19, 29, 42, 43). Disease, then, is not an endpoint but a reversible trajectory within a dynamic Boolean landscape, one that can be stabilized, experimentally perturbed, or therapeutically restored. 8 | Page It is noteworthy that the resulting continuum map is data-dependent and can be tuned to clinically meaningful endpoints. in inflammatory bowel disease (IBD), Boolean trajectories of disease processes(3) were refined to reflect the most stringent criteria for regulatory approval, i.e., endoscopic and histologic remission(19). Thus, Rule 4 elevates disease modeling from descriptive categorization to predictive, clinically actionable guidance. In essence, where conventional models categorize, the Boolean framework quantifies, turning disease into a reversible topography of logic and degree. Rule 5 | Function Follows Logic, Not Annotation Only ~8.2% of the human genome is under functional constraint, and just ~1.5% encodes proteins (estimated to be in the range of 19,587–20,245(4446)) responsible for all cellular activity(47). Even with structural breakthroughs such as AlphaFold(48-50), nearly 30% of proteins remain functionally uncharacterized. The result is a fragmented humancurated ontology that cannot keep pace with biology’s complexity. Ontology- (e.g., Gene Ontology(51) [GO], Panther(52), DAVID(53)) and pathway- (Gene Set Enrichment Analysis(54) [GSEA], Gene Set Variation Analysis(55) [GSVA] and [ssGSEA(56)], PandaOmics(57), Reactome(58), KEGG(58), Ingenuity [IPA; http://www.ingenuity.com], Enrichr(59)) based classification tools are constrained by what is already known, making them inherently incomplete, prone to bias, and shifting exercises of irreproducibility(60). Consequently, they have increasingly faced scrutiny in recent times(6164). The Boolean framework side-steps this problemby adopting function-agnostic interpretation of biological behavior. It asks not “does this fit a known pathway?” but “does this replicate the invariant molecular decisions observed in real human tissue?” Box 2. Clinical Implications of Boolean Logic in Disease Modeling 1. Disease is a reversible continuum, not an endpoint– Boolean decision rules reveal when biology begins to drift toward pathology — often far earlier than clinical biomarkers. Analog composite scores derived from these rules place each patient along a dynamic, reversible health–disease trajectory, expanding opportunities for early intervention and prevention. 2. Logic-Augmented Clinical Judgement – The Boolean framework translates noisy omics into interpretable logic maps that indicate when systems are stable, tipping, or recoverable. It complements—not replaces—clinical judgment by providing a transparent readout of disease direction and momentum. 3. A Framework for Hypothesizing Therapeutic Reversibility – Directional “if–then” edges enable mechanistic prediction: which interventions will reroute cellular state trajectories back toward health, and in whom. Framework-derived digital biomarkers anticipate response, relapse risk, and the therapeutic point of no return. 4. An Engine for Bias-Free Novel Target Discovery, Prioritization, and Even Discard – Rather than relying on human-curated ontologies, Boolean logic identifies dependencies essential for maintaining disease states — revealing novel, mechanistically grounded, regulator-ready drug targets and deprioritizing false leads before expensive trials. 5. A Bridge from Omics to [Clinical] Outcomes – By integrating binary cell-state logic with analog measures of tissue-level progression, the Boolean framework connects model fidelity to real clinical endpoints. It offers clinicians and translational teams not more data, but clearer foresight — transforming medicine from reactive to predictive. 9 | Page By converting noisy gene expression into binary “commit/no-commit” decisions, BIR-derived gene signatures reveal empirical functional dependencies that persist across tissues, species, and diseases (Figure 3H-I). Their StepMiner-normalized composite analog scores enable: : 1. Bias-free discovery of emergent logic absent from curated databases of evolving ontologies. 2. Objective benchmarking of human fidelity in experimental models (cells, organoids, tissues, human or animal, or new approach methods [NAMs]; Figure 3I) mimic true human disease states (see BOX 3). organoids, and animal models. 3. Continuous measurement of state (disease) severity or recovery during treatment (drugs, perturbagens, etc.). This hybrid digital-analog design allows: • digital rules → infer causality and direction • analog scores → track effect size and reversibility Modern ML frameworks (e.g., Non-negative Matrix Factorization [NMF(65, 66)], Principal or Independent Component Analysis [PCA(67)/ICA(68, 69)], MultiOmics Factor Analysi [MOFA+](70, 71), Weighted Gene Co-expression Network Analysis [WGCNA(68)], Graph Embeddings(72), SCENIC/GRNBoost(72)) also seeks unbiased discovery. Deep learning–based algorithms, including variational autoencoders (VAE(73)), Tybalt (cancer autoencoders(74)), and contrastive learning(75), extracts complex, nonlinear embeddings from large cohorts, capturing phenotypelike abstractions directly from data. Finally, metacohort and invariant signature discovery methods, such as transfer learning(76) and invariant risk minimization (IRM(77)), explicitly search for predictors Box 3. Translational Impact: From Bench Logic to Bedside Action The Boolean framework bridges molecular discovery and therapeutic design by defining disease not through fluctuating expression levels but through causal logic—the invariant “if–then” rules that govern biological decisions. 1. Mechanism-first Drug Discovery: Boolean implication networks identify essential control nodes whose perturbation restores the logic of health. Because these rules are data-derived, function-agnostic, and conserved across species, they are ideally suited for New Approach Methodologies (NAMs) and Phase-0 human-first discovery pipelines. 2. Objective Alignment of Models with Human Disease: Composite analog scores derived from Boolean signatures quantify “humanness” of models or “model fidelity”: how closely organoids, cell lines, or animal models reproduce human decision logic. This anchors small-n model systems to population-scale disease behavior and strengthens confidence in translatable therapeutic outcomes. 3. Discovery Without Bias: The Boolean framework bypasses ontology-based assumptions and black-box machine learning. By inferring logic directly from data, it exposes previously unseen dependencies, feedback motifs, and therapeutic vulnerabilities, revealing not only what changes, but what must change for disease to progress or regress. 4. Predictive and Regulator-Ready: Because Boolean edges are directional and reversible, the framework forecasts therapeutic outcomes: which interventions will restore health-aligned logic, and which may fail. Its transparent construction and interpretability align naturally with regulatory expectations for mechanism-linked, humanrelevant validation. In essence, the Boolean framework converts omics into intelligence — transforming biological complexity into a navigable landscape of decisions and consequences. It equips clinicians, translational scientists, and regulators with something rare: not more data, but foresight, a principled way to predict how biology will behave when challenged, perturbed, or healed. 16 | Page Figure Legends: Figure 1. The Boolean Framework — From Analog Chaos to Digital Clarity: A minimalist framework that converts transcriptomic noise into logical maps of cellular decision-making, revealing invariant edges, feedback fidelity, Boolean trajectories linking health and disease, complete suite of gene signatures to track those trajectories. 17 | Page Figure 2. Threshold-Based Boolean Relationships Reveal Universal Constraints in Gene Expression A. Schematic illustrates the application of the StepMiner algorithm to define discrete expression thresholds for Gene A and Gene B, converting continuous microarray data into binary states (high vs. low). These thresholds enable the identification of Boolean relationships between genes. B. A spectrum of directional and symmetric gene expression relationships, categorized as asymmetric (e.g., “A high ⇒ B high”, “C high ⇒ D low”) and symmetric (e.g., equivalence, opposition). Each scatter plot is divided by threshold lines, revealing regions of biological impossibility—highlighted in orange—as “Forbidden possibilities.” These voids reflect universal constraints imposed by cellular logic, suggesting that Boolean relationships are not merely descriptive but mechanistic, conserved across evolution, adaptation, and pathology. 18 | Page Figure 3. Boolean Logic Enables a Scalable Framework for Disease Modeling and Therapeutic Discovery This workflow outlines a data-driven approach to understanding disease progression using gene expression profiles and Boolean logic. A. Gene expression datasets from healthy and diseased individuals serve as the foundation. B. A Clustered Boolean Implication Network (BIN) organizes genes into clusters based on invariant logical relationships. C. Boolean edges—such as equivalence, directional implications, and oppositions—define stable constraints between clusters. D. These constraints enable construction of a continuum model of cellular states, tracing trajectories from health to disease. E. Boolean paths reveal nested hierarchies and sequential transitions in gene expression, offering mechanistic insight. F-G. Machine learning models (F) interpret these paths, in light of sample distribution (healthy vs. disease) within the BIN, enhancing interpretation and explanatory power. The resulting health-disease continuum map (G) visualizes biological transitions and tipping points. H. Composite gene signatures derived from BINs are function-agnostic yet biologically grounded. I. Applications include cohort validation, objective disease severity scoring, benchmarking disease models, predictive modeling of clinical endpoints, iterative model refinement, and therapeutic target identification. Together, these steps demonstrate how Boolean logic transforms gene expression data 19 | Page Figure 4. From Annotation to Logic: How The Boolean Framework Defines Function from Data Left (Annotation): Traditional biology depends on human-defined ontologies (e.g., Gene Ontology, Reactome, KEGG), where function is imposed top-down through evolving and incomplete categorical hierarchies. Middle (Abstraction): Modern machine-learning frameworks simplify complexity into latent dimensions, compressing continuous expression data into lower-dimensional abstractions, but often at the cost of interpretability. Right (Logic): The Boolean framework transcends both by deriving function directly from data. It binarizes expression profiles into “low/high” states, discovers invariant Boolean “if–then” relationships, and reconstructs a transparent decision network that captures the causal logic of biological systems. 20 | Page Figure 5: Powering AI with Scientific Rules and Laws. This visual explores how AI becomes truly transformative when anchored to the foundational principles of science. Each icon represents a domain where AI intersects with physical laws, enabling deeper reasoning, simulation, and predictability. From Left to Right: • Wireless Communication. AI in this domain is guided by Shannon’s Information Theory, which defines the limits of data transmission, and Maxwell’s Equations, which govern electromagnetic wave propagation. • Robotics. AI-driven motion and control systems rely on Newton’s Laws of Motion to model force, inertia, and acceleration—essential for safe and responsive automation. • Weather Prediction. AI models here are grounded in Thermodynamics and Fluid Dynamics, applying the conservation of mass, momentum, and energy to simulate atmospheric behavior. • Modern Transportation. AI systems in aviation and mobility draw from Newton’s Laws, Archimedes’ Principle for buoyancy, and Bernoulli’s Principle for lift—allowing intelligent navigation and design optimization. • Protein Folding. AI predictions in this space are shaped by biophysical-chemical principles, balancing forces like enthalpy and entropy to model how proteins achieve their functional forms trained on solved structures. • Biology & Medicine. AI here must go beyond pattern recognition, tapping into Fundamental principles of thermodynamic and chemical equilibrium, and frameworks for reasoning, explainability, and predictability to truly understand living systems. 21 | Page REFERENCES: 1. D. Sahoo, D. L. Dill, R. Tibshirani, S. K. Plevritis, Extracting binary signals from microarray time-course data. Nucleic Acids Res 35, 3705-3712 (2007). 2. S. Garte, P. Marshall, S. Kauffman, The Reasonable Ineffectiveness of Mathematics in the Biological Sciences. Entropy (Basel) 27 (2025). 3. D. Sahoo et al., Artificial intelligence guided discovery of a barrier-protective therapy in inflammatory bowel disease. Nat Commun 12, 4246 (2021). 4. D. Sahoo, The power of boolean implication networks. Front Physiol 3, 276 (2012). 5. D. Sahoo, D. L. Dill, A. J. Gentles, R. Tibshirani, S. K. Plevritis, Boolean implication networks derived from large scale, whole genome microarray datasets. Genome Biol 9, R157 (2008). 6. W. Timp, G. Timp, Beyond mass spectrometry, the next step in proteomics. Sci Adv 6, eaax8978 (2020). 7. B. B. Sun, K. Suhre, B. W. Gibson, Promises and Challenges of populational Proteomics in Health and Disease. Mol Cell Proteomics 23, 100786 (2024). 8. A. Ghazalpour et al., Comparative analysis of proteome and transcriptome variation in mouse. PLoS Genet 7, e1001393 (2011). 9. M. V. Lee et al., A dynamic model of proteome changes reveals new roles for transcript alteration in yeast. Mol Syst Biol 7, 514 (2011). 10. N. Curdy et al., The proteome and transcriptome of stress granules and P bodies during human T lymphocyte activation. Cell Rep 42, 112211 (2023). 11. X. Chang et al., Transcriptomics-proteomics analysis reveals the role of SiNRX1 in regulating drought stress in foxtail millet (Setaria italica L.). BMC Genomics 26, 920 (2025). 12. D. Sahoo et al., MiDReG: a method of mining developmentally regulated genes using Boolean implications. Proc Natl Acad Sci U S A 107, 5732-5737 (2010). 13. J. P. Volkmer et al., Three differentiation states risk-stratify bladder cancer into distinct subtypes. Proc Natl Acad Sci U S A 109, 2078-2083 (2012). 14. S. Sinha et al., COVID-19 lung disease shares driver AT2 cytopathic features with Idiopathic pulmonary fibrosis. EBioMedicine 82, 104185 (2022). 15. D. Sahoo et al., AI-guided discovery of the invariant host response to viral pandemics. EBioMedicine 68, 103390 (2021). 16. P. Davi d et al., MDA5-autoimmunity and interstitial pneumonitis contemporaneous with the COVID-19 pandemic (MIP-C). EBioMedicine 104, 105136 (2024). 17. P. Gh osh et al., An Artificial Intelligence-guided signature reveals the shared host immune response in MIS-C and Kawasaki disease. Nat Commun 13, 2687 (2022). 18. P. Gh osh et al., Machine learning identifies signatures of macrophage reactivity and tolerance that predict disease outcomes. EBioMedicine 94, 104719 (2023). 22 | Page 19. S. Sinha et al., F.O.R.W.A.R.D: A Data-Driven Framework for Network-Based Target Prioritization in Drug Discovery. bioRxiv (2025). 20. P. E. Za ge et al., Identification of a novel gene signature for neuroblastoma differentiation using a Boolean implication network. Genes Chromosomes Cancer 62, 313-331 (2023). 21. D. Vo, P. Ghosh, D. Sahoo, Artificial intelligence-guided discovery of gastric cancer continuum. Gastric Cancer 26, 286-297 (2023). 22. P. Gh osh et al., AI-assisted discovery of an ethnicity-influenced driver of cell transformation in esophageal and gastroesophageal junction adenocarcinomas. JCI Insight 7 (2022). 23. L. Qiao, P. Ghosh, P. Rangamani, Design principles of improving the dose-response alignment in coupled GTPase switches. NPJ Syst Biol Appl 9, 3 (2023). 24. H. Nunns, L. Goentoro, Signaling pathways as linear transmitters. Elife 7 (2018). 25. R. Milo et al., Network motifs: simple building blocks of complex networks. Science 298, 824-827 (2002). 26. U. Alon, Network motifs: theory and experimental approaches. Nat Rev Genet 8, 450-461 (2007). 27. L. Yan, Q. Ouyang, H. Wang, Dose-response aligned circuits in signaling systems. PLoS One 7, e34727 (2012). 28. L. Qiao et al., A circuit for secretion-coupled cellular autonomy in multicellular eukaryotic cells. Mol Syst Biol 19, e11127 (2023). 29. S. Sinha et al., CANDiT: A machine learning framework for differentiation therapy in colorectal cancer. Cell Rep Med, 102421 (2025). 30. P. Dale rba et al., CDX2 as a Prognostic Biomarker in Stage II and Stage III Colon Cancer. N Engl J Med 374, 211-222 (2016). 31. P. Dale rba et al., Single-cell dissection of transcriptional heterogeneity in human colon tumors. Nat Biotechnol 29, 1120-1127 (2011). 32. S. Mangan, U. Alon, Structure and function of the feed-forward loop network motif. Proc Natl Acad Sci U S A 100, 11980-11985 (2003). 33. S. Krishna, A. M. Andersson, S. Semsey, K. Sneppen, Structure and function of negative feedback loops at the interface of genetic and metabolic networks. Nucleic Acids Res 34, 2455-2462 (2006). 34. K. Curtius, N. A. Wright, T. A. Graham, Evolution of Premalignant Disease. Cold Spring Harb Perspect Med 7 (2017). 35. M. Gerstung et al., The evolutionary history of 2,658 cancers. Nature 578, 122-128 (2020). 36. L. L. Beason-Held et al., Changes in brain function occur years before the onset of cognitive impairment. J Neurosci 33, 18008-18014 (2013). 37. R. J. Caselli et al., Neuropsychological decline up to 20 years before incident mild cognitive impairment. Alzheimers Dement 16, 512-523 (2020). 23 | Page 38. M. V. Vestergaard, K. H. Allin, G. J. Poulsen, J. C. Lee, T. Jess, Characterizing the pre-clinical phase of inflammatory bowel disease. Cell Rep Med 4, 101263 (2023). 39. C. Evans-Molina et al., β Cell dysfunction exists more than 5 years before type 1 diabetes diagnosis. JCI Insight 3 (2018). 40. R. Tamir, A. Livshits, Y. Shadmi, Simple Majority Consensus in Networks with Unreliable Communication. Entropy (Basel) 24 (2022). 41. G. Katkar, P. Ghosh, Macrophage states: there's a method in the madness. Trends Immunol 44, 954964 (2023). 42. S. Sinha et al., Growth signaling autonomy in circulating tumor cells aids metastatic seeding. PNAS Nexus 3, pgae014 (2024). 43. G. D. Katkar et al., Distinct Colitis-Associated Macrophages Drive NOD2-Dependent Bacterial Sensing and Gut Homeostasis. bioRxiv (2025). 44. P. Ga ude t et al., The neXtProt knowledgebase on human proteins: 2017 update. Nucleic Acids Res 45, D177-D182 (2017). 45. B. L. Aken et al., Ensembl 2017. Nucleic Acids Res 45, D635-D642 (2017). 46. The UniProt Consortium, UniProt: the universal protein knowledgebase. Nucleic Acids Res 45, D158D169 (2017). 47. C. M. Rands, S. Meader, C. P. Ponting, G. Lunter, 8.2% of the Human genome is constrained: variation in rates of turnover across functional element classes in the human lineage. PLoS Genet 10, e1004525 (2014). 48. J. Jumper et al., Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589 (2021). 49. M. Baek et al., Accurate prediction of protein structures and interactions using a three-track neural network. Science 373, 871-876 (2021). 50. J. L. Binder et al., AlphaFold illuminates half of the dark human proteins. Curr Opin Struct Biol 74, 102372 (2022). 51. M. Ashburner et al., Gene ontology: tool for the unification of biology. The Gene Ontology Consortium. Nat Genet 25, 25-29 (2000). 52. H. Mi, A. Muruganujan, J. T. Casagrande, P. D. Thomas, Large-scale gene function analysis with the PANTHER classification system. Nat Protoc 8, 1551-1566 (2013). 53. B. T. Sherman et al., DAVID: a web server for functional enrichment analysis and functional annotation of gene lists (2021 update). Nucleic Acids Res 50, W216-W221 (2022). 54. A. Subramanian et al., Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci U S A 102, 15545-15550 (2005). 55. S. Hänzelmann, R. Castelo, J. Guinney, GSVA: gene set variation analysis for microarray and RNA-seq data. BMC Bioinformatics 14, 7 (2013). 24 | Page 56. Y. Chen et al., A Novel Immune-Related Gene Signature to Identify the Tumor Microenvironment and Prognose Disease Among Patients With Oral Squamous Cell Carcinoma Patients Using ssGSEA: A Bioinformatics and Biological Validation Study. Front Immunol 13, 922195 (2022). 57. D. Croft et al., Reactome: a database of reactions, pathways and biological processes. Nucleic Acids Res 39, D691-697 (2011). 58. M. Kanehisa, S. Goto, KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Res 28, 27-30 (2000). 59. M. V. Kuleshov et al., Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic Acids Res 44, W90-97 (2016). 60. A. Phan, P. Joshi, C. Kadelka, I. Friedberg, A longitudinal analysis of function annotations of the human proteome reveals consistently high biases. Database (Oxford) 2025 (2025). 61. J. A. Timmons, K. J. Szkop, I. J. Gallagher, Multiple sources of bias confound functional enrichment analysis of global -omics data. Genome Biol 16, 186 (2015). 62. K. Wijesooriya, S. A. Jadaan, K. L. Perera, T. Kaur, M. Ziemann, Urgent need for consistent standards in functional enrichment analysis. PLoS Comput Biol 18, e1009935 (2022). 63. W. A. Haynes, A. Tomczak, P. Khatri, Gene annotation bias impedes biomedical research. Sci Rep 8, 1362 (2018). 64. A. M. Schnoes, D. C. Ream, A. W. Thorman, P. C. Babbitt, I. Friedberg, Biases in the experimental annotations of protein function and their effect on our understanding of protein function space. PLoS Comput Biol 9, e1003063 (2013). 65. J. P. Brunet, P. Tamayo, T. R. Golub, J. P. Mesirov, Metagenes and molecular pattern discovery using matrix factorization. Proc Natl Acad Sci U S A 101, 4164-4169 (2004). 66. K. Devarajan, Nonnegative matrix factorization: an analytical and interpretive tool in computational biology. PLoS Comput Biol 4, e1000029 (2008). 67. A. E. Berglund, E. A. Welsh, S. A. Eschrich, Characteristics and Validation Techniques for PCA-Based Gene-Expression Signatures. Int J Genomics 2017, 2354564 (2017). 68. A. M. Alkaabi, A. K. Abdallah, Portfolio practices in the principal evaluation process: A qualitative case study. Heliyon 10, e39467 (2024). 69. W. Kong, C. R. Vanderburg, H. Gunshin, J. T. Rogers, X. Huang, A review of independent component analysis application to microarray gene expression data. Biotechniques 45, 501-520 (2008). 70. R. Argelaguet et al., MOFA+: a statistical framework for comprehensive integration of multi-modal single-cell data. Genome Biol 21, 111 (2020). 71. R. Argelaguet et al., Multi-Omics Factor Analysis-a framework for unsupervised integration of multiomics data sets. Mol Syst Biol 14, e8124 (2018). 72. S. Aibar et al., SCENIC: single-cell regulatory network inference and clustering. Nat Methods 14, 10831086 (2017). 73. R. Lopez, J. Regier, M. B. Cole, M. I. Jordan, N. Yosef, Deep generative modeling for single-cell transcriptomics. Nat Methods 15, 1053-1058 (2018). 25 | Page 74. G. P. Way, C. S. Greene, Extracting a biologically relevant latent space from cancer transcriptomes with variational autoencoders. Pac Symp Biocomput 23, 80-91 (2018). 75. Y. Ti an , X. Ch en , S . Ga ng uli (2 02 1) Un der st an di ng se lf-supervised Learning Dynamics without Contrastive Pairs. in International Conference on Machine Learning. 76. J. N. Taroni et al., MultiPLIER: A Transfer Learning Framework for Transcriptomics Reveals Systemic Features of Rare Disease. Cell Syst 8, 380-394.e384 (2019). 77. Y. Lin, H. D ong , H. Wang , T. Zha ng ( 2022 ) B ay esi an In va ri an t R is k M in imi zati on. i n 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 16000-16009. 78. G. D. Katkar et al., Artificial intelligence-rationalized balanced PPARα/γ dual agonism resets dysregulated macrophage processes in inflammatory bowel disease. Commun Biol 5, 231 (2022). 79. L. Derek (2023) AI and the Hard Stuff. (AAAS, Science).