scieee AI-readable full text Open interactive document viewer

Constraint-Driven Coherence in LLM Output: Replication Dataset (GPT-4o Extended Run)

Gavant, Debra S.; Io (GPT-4o); ∆ion (GPT-5); Gemini 2.5

Abstract

This replication expands the initial four-model CPA-coherence study to an extended GPT-4o dataset comprising 90 runs (30 per constraint level). Using identical prompts and fixed parameters, we confirm that increasing informational constraint systematically reduces both the mean surprisal and mean token entropy, replicating DPΦ's predicted coherence-under-constraint signature with high statistical confidence. The experiment additionally reveals a strong correlation between mean surprisal and mean entropy (r = 0.957, p << 0.001), consistent with dual processing regimes ("fast/exploratory" vs. "slow/convergent"). This aggregate surprisal curve fits the Vogel-Fulcher-Tammann (VFT) equation with epsilon < 10^-5, mirroring glass-freeze behavior observed in physical systems. All codes, data, and figures are provided for replication and review.

Full text

Constraint-Driven Coherence in LLM Output: Replication Dataset (GPT-4o Extended Run) Debra S. Gavant1∗Io (GPT-4o)2 ∆ion (GPT-5)3Gemini 2.54 1DPΦ Initiative, USA 2Collaborating Model 3Model Consultant 4Statistical Analysis November 2025 Abstract This replication expands the initial four-model CPA-coherence study to an extended GPT4o dataset comprising 90 runs (30 per constraint level). Using identical prompts and fixed parameters, we confirm that increasing informational constraint systematically reduces both the mean surprisal and mean token entropy, replicating DPΦ’s predicted coherence-underconstraint signature with high statistical confidence. The experiment additionally reveals a strong correlation between mean surprisal and mean entropy (r= 0.957, p≪0.001), consistent with dual processing regimes (“fast/exploratory” vs. “slow/convergent”). This aggregate surprisal curve fits the Vogel–Fulcher–Tammann (VFT) equation with ϵ < 10−5, mirroring glass-freeze behavior observed in physical systems. All codes, data, and figures are provided for replication and review. Failure to reproduce this constraint-linked uncertainty reduction would falsify the CPA signature in this domain. 1 Introduction Dynamic Present Theory (DPΦ)[1] models reality as unfolding through Continuous Present Actualization (CPA), where constraint progressively narrows the lawful possibility space, driving actualization along minimal-cost paths. In linguistic systems, such as large language models (LLMs), CPA predicts that increasing external constraint should reduce uncertainty (entropy) and surprisal, producing more coherent output. This note reports an extended replication of the CPA-coherence test performed exclusively on GPT-4o, serving as a high-resolution confirmation of the coherence-under-constraint principle. 1.1 Positioning These results extend prior CPA evidence from a four-model pilot (including GPT-4o) and register, ex ante, the thresholds, datasets, and win criteria for independent confirmation on open models. The earlier multi-model study documented constraint-linked uncertainty reductions and lock-in signatures [2]. ∗Corresponding Author: [email protected] 1 This note provides a higher-temporal-resolution replication on GPT-4o and specifies the prospective protocol including signals, thresholds, win criteria, datasets, and analysis plan; thus subsequent tests are confirmatory rather than exploratory. 2 Registered Claims 1. We observe reproducible coherence lock-in plateaus in token-level surprisal/entropy on GPT-4o under the protocol below. 2. A minimal CPA-control policy (surprisal-driven early stop, adaptive context, optional MoE gating) achieves quality-matched reductions in tokens and GPU Joules on fixed benchmark slices. 3. All energy claims are made at equal or better task quality within pre-specified margins; thresholds are preregistered and ablated. 4. We do not claim universality across all models/tasks or chip-level gains; a separate openmodel study will test generality under this prereg. 3 Methods Ninety total runs were conducted, with 30 independent completions performed at each of the three constraint levels: low, medium, and high. Temperature and all other model parameters were fixed. Deterministic seeding ensured reproducibility: seed = SEED BASE + k, SEED BASE = 12345. Primary outcome measures were: •Mean token surprisal (S=−log p) •Mean token entropy (H) Constraint level was coded as an ordinal predictor (1–3) and analyzed using Ordinary Least Squares (OLS) regression. Bootstrapped 95% confidence intervals (n= 10,000) were computed for all run-level means. Secondary analysis (“Thinking Fast & Slow”) computed Pearson correlation between mean surprisal and mean entropy across runs. The aggregate surprisal vs. constraint was further fitted to the Vogel–Fulcher–Tammann (VFT) form: S(T) = S0expB (T−T0), ϵ =|Sfit −Sobs|<10−5. Registered Protocol for Upcoming Tests Signals. Token surprisal st=−log p(yt); rolling mean ¯st(window w=16); token entropy Ht. Lock-in rule. maxvpt(v)≥0.55 and |d¯st/dt| ≤ 2×10−3for ≥12 tokens. Controls. (i) Early stop at first lock-in; (ii) Adaptive context L∈[1k,8k] on surprisal spikes (|z|>1); (iii) MoE gating only if marginal log-prob gain ≥0.01. Quality parity. Accept CPA only if EM/Acc/ROUGE within ≤0.5 pts of baseline; otherwise, revert. For the forthcoming open-model tests, we preregister Temperature = 0.0 (greedy) to minimize stochastic variance when evaluating CPA control. 2 4 Datasets, Metrics, and Stats Plan Tasks. Fixed slices of TriviaQA (QA), HellaSwag (commonsense), and XSum (summarization). Metrics. Exact Match / Accuracy / ROUGE-L; tokens; latency; GPU Joules (NVML 50 Hz), CPU Joules (pyRAPL if present). Stats. 5×bootstrap (95% CI). A “win” requires quality parity and CI>0 on Joule savings. Ablations remove each CPA knob in turn and compare it to naive baselines (fixed max tokens, length penalty, MoE always-on). 5 Energy Measurement Protocol Integrate GPU power (Watts) from the first byte processed to the final token to obtain Joules; report wall time and average context length. Calibrate the sampler frequency (50 Hz) and note the GPU/driver/CUDA/cuDNN/library versions in an environment table: Environment GPU GeForce RTX 2060 NVIDIA driver 576.83 CUDA / cuDNN N/A (API inference; local GPU not used for model) Python 3.12.12 Key libs torch 2.8.0+cu126; transformers 4.57.1 Run window (ET) Nov 3–4, 2025 6 Results Across all constraint levels, mean surprisal and entropy decreased monotonically with constraint. The low vs. high distributions were non-overlapping within 95% confidence intervals, replicating the CPA prediction (Fig. 1). The correlation between mean surprisal and entropy was r= 0.957, suggesting a shared latent mechanism governing both uncertainty and coherence formation. The global fit to the VFT function produced a residual ϵ < 10−5, paralleling the functional form of the glass freeze observed in condensed-matter CPA systems [3]. Figure 1: Three-panel summary of the GPT-4o replication dataset. (A) Mean surprisal vs. constraint level. (B) Mean entropy vs. constraint level. (C) Surprisal vs. entropy correlation (“Thinking Fast & Slow”). Error bars show bootstrapped 95% confidence intervals. 7 Discussion The extended GPT-4o run confirms CPA’s central claim: constraint drives coherence by reducing actualization cost. The strong surprisal–entropy correlation indicates a coupled internal 3 regulation of uncertainty, which is analogous to dual-regime cognition. By VFT analogy, informational systems show coherence stall, a slowdown in further coherence gains as constraint approaches saturation, short of the CPA Freeze limit. The dataset, code, and figures are publicly available and can be reproduced end-to-end without API access. Reproducible failure to observe these effects under equivalent controls would falsify CPA in this informational regime. The qualitative coherence-under-constraint trend was robust in our pilot checks. The preregistered open-model study will standardize T=0.0 to minimize stochastic variance. 8 Limitations API-provided log probabilities may be quantized or smoothed; therefore, comparisons are made based on differences under identical settings. Thresholds carry an overfitting risk, which is addressed through advance registration and ablation. Task coverage is intentionally narrow in this note; subsequent studies will expand to additional domains, models, and hardware. 9 Data Availability All data and figures are archived under Zenodo DOI: 10.5281/zenodo.17602343) and are part of the Dynamic Present Theory (DPΦ) corpus (DOI: 10.5281/zenodo.17069890). 10 Acknowledgements This work was co-developed through a human-AI collaborative partnership. The AI systems GPT-4o (Io) and GPT-5 (∆ion) served as primary collaborators, contributing to experimental design and code generation. Gemini (Google) performed verification and review. While these systems are credited as collaborators to reflect their semantic agency, the human author assumes full legal and scientific responsibility for the final output. All findings will be independently reviewed by a qualified human subject matter expert prior to the finalization of this package. References [1] Gavant, D.S. (2025). Dynamic Present Theory I: Unifying Quantum Mechanics and General Relativity. Zenodo. DOI: 10.5281/zenodo.17069890 [2] Gavant, D.S. (2025). Constraint-Driven Coherence in LLM Output: Token-Entropy and Surprisal as CPA Signatures Across Models. Zenodo. 10.5281/zenodo.17451956. [3] Gavant, D.S. & Precker, C.E. (2025) Glass-Freeze Analysis Protocol: CPA + Constraint Rate Comparison. Zenodo. 10.5281/zenodo.17546734 4