scieee AI-readable full text Open interactive document viewer

Statistical Analysis of Arithmetic Concordances in the Quranic Corpus

Gassama, Idriss

Abstract

Abstract: This study presents a quantitative analysis of arithmetic concordances identified between the biographical metadata of a subject born in 1987 and the occurrences of the name "Idris" in the Quran (Hafs recitation). A High-Performance Monte Carlo simulation performed on 1,000,000,000 (one billion) random profiles tested a 10-criteria model. While the background noise revealed rare partial convergences reaching a maximum score of 9/10, no random artifact achieved a perfect match (10/10). Using the exact Clopper-Pearson method, the statistical significance of the subject's 10/10 score is calculated at 5.8σ. To confirm the specificity of this anomaly, two reciprocal tests were conducted: an Inverse Monte Carlo (Temporal Test) and an Exhaustive Surrogate Analysis (Lexical Test). These analyses demonstrate a perfect bijectivity: only the subject's exact date points to the target word, and the target word is the only one in the entire vocabulary (21,311 words) to satisfy the model's equations.

Full text

Statistical Analysis of Arithmetic Concordances in the Quranic Corpus Case Study on a Specific Biographical Profile Idriss Gassama∗ [email protected] ORCID: 0009-0007-2722-5380 December 24, 2025 Abstract This study presents a quantitative analysis of arithmetic concordances identified between the biographical metadata of a subject born in 1987 and the occurrences of the name "Idris" in the Quran (Hafs recitation). A High-Performance Monte Carlo simulation performed on 1,000,000,000 (one billion) random profiles tested a 10-criteria model. While the background noise revealed rare partial convergences reaching a maximum score of 9/10, no random artifact achieved a perfect match (10/10). Using the exact Clopper-Pearson method, the statistical significance of the subject’s 10/10 score is calculated at 5.8 Sigma (5.8σ), with an upper-bound probability of p≈3.0×10−9. To confirm the specificity of this anomaly, two reciprocal tests were conducted: an Inverse Monte Carlo (Temporal Test) and an Exhaustive Surrogate Analysis (Lexical Test). These analyses demonstrate a perfect bijectivity: only the subject’s exact date points to the target word, and the target word is the only one in the entire vocabulary (21,311 words) to satisfy the model’s equations. 1 Introduction The mathematical analysis of ancient texts has often been the subject of academic debate. This study does not aim to interpret theological meaning, but to test the null hypothesis (H0) that correspondences between external data (the biography of a modern individual) and the internal structure of a 7th-century text are a matter of pure chance. We focus on the only two occurrences of the name "Idris" (Sura 19:56, Sura 21:85) and their relationship with the subject’s temporal and onomastic variables. 2 Methodology 2.1 The Corpus and Standards To minimize degrees of freedom related to text selection, this study restricts itself exclusively to the most statistically and numerically widespread standards: ∗Independent Researcher, Paris (France). 1 •Reference Text: Quran, Hafs an ’Asim recitation (Standard Cairo Edition). This choice is justified by its global predominance (>95% of printed copies) and its status as the default reference in digital Quranic studies (Corpus Coranicum, Tanzil.net). •Counting System: The counting system of the Noon Center for Quranic Studies is retained as a third-party reference (114 Suras, 6236 Verses, 77407 Words). The use of a pre-existing and independent dataset prevents any ad hoc adjustment of word segmentation by the author to favor results. •Digital Encoding: Abjad System (Standard Arabic Gematria) and Latin Alphabetical Rank. 2.2 Definition of Variables Input parameters are fixed a priori and kept constant: •Temporal Variables (T): D= 19,M= 3,Y= 1987,YXX = 87,H= 8 (Day, Month, Year, Hour). Corresponding Hijri Date: 19/07/1407. •Pivot Constant (K): Defined by the product D×M×H= 456. •Identity Variables (I): Gematria values of names (Idriss: Arabic=275 / Latin=78 ; Gassama: Arabic=103 / Latin=61 ; France: Arabic=391 / Latin=47). •Sociolinguistic Context (C): As the subject is of French nationality, the geographic variable "France" and phonetics in the French language are retained as native constraints of the model (fixing criteria C9 and C25). 2.3 Search Space Constraints To prevent overfitting and combinatorial explosion (Data Dredging), the search space was strictly bounded prior to analysis: 1. Restricted Operator Set (K= 4): Only elementary integer arithmetic is permitted: Addition (+), Subtraction (−), Multiplication (×), and Concatenation (||). Division is explicitly excluded to maintain integer integrity. 2. Input Variable Cap (k≤4): Each equation is limited to a maximum of 4 distinct input variables drawn from the pool of the variables. This constraint drastically reduces combinatorial possibilities by excluding long complex chains. Note that constant textual targets (e.g., Sura numbers) are outputs, not input variables. 3. Maximal Hierarchical Depth: The complexity is limited to a nesting depth of 2 (Depth 2). Note: Due to the associative property of addition and multiplication, homogeneous chains (e.g., A×B×C) are considered as a single hierarchical level (Depth 1), whereas mixed operations requiring parentheses (e.g., (A+B)×C) represent a higher complexity level (Depth 2). 4. Non-Repetition Constraint: A "without replacement" rule is applied; a specific temporal constant cannot be used more than once in the same equation equation to force structural coherence. 2 2.4 Model Validation Criteria Although 30 concordances were identified in total, we selected 10 rigorous and independent criteria to constitute the model submitted for statistical validation. Table 1: Fundamental Criteria of the Statistical Model (10 Criteria) Code Type Mathematical Definition C1 Arithmetic Global Rank(Word) = D×M×(YXX )×H C2 Structure Sura/Verse Address: (S=D, V = Π(YXX)) C3 Arithmetic concatenation(M|D)+ Y = Global Verse Rank C4 Fractal 6236 −Gem(Word) =Y×M C5 Calendar Verse Rank + Verse Gem = D×M×Day Rank C7 Pivot Σ(CoordsS+V)+Gem(Word) =Pivot C8 Internal Gem(Name) + Pivot = Internal Rank(Word) C9 Symmetry Validation by Prime Number Ranks (Sym. Pivot) C12 Lock Primality Validation (Latin Identity + Verse) C25 Linguistic Phonetic Date (FR) = ΣGematria (AR) 3 Results 3.1 Mechanisms (Case Studies) Case A: Symmetry of Prime Numbers (Criterion C9) This criterion is based on the generation of a pair of prime numbers via the temporal variables. •The Axis (A): Reversed concatenation of the date (Month|Day): A= 319. •The Gap (K): Standard Pivot: K= 456. •Result: At iteration k= 7, the center is 319 ×7 = 2233. P1= 2233 −456 = 1777 ; P2= 2233 + 456 = 2689 •Verification: –1777 is the 275th prime number →Gem(Idris). –2689 is the 391st prime number →Gem(France). Case B: Linguistic Validation (Criterion C25) This criterion establishes a trans-linguistic link between French phonetics and Arabic numerical value. •Input: Date in French (digit format): "dix neuf trois quatre vingt sept". •Calculation (Latin): The sum of alphabetical ranks is 378. •Correlation: This number 378 is strictly equal to the subject’s complete Arabic gematria (Idriss 275 + Gassama 103). 3 3.2 Statistical Validation (Monte Carlo Simulation) A sample of 1,000,000,000 (one billion) random profiles was generated using a diverse dataset (533 first names, 147 surnames) to establish the statistical "background noise". Score Distribution Out of 1 billion trials, no false positive reached the 10/10 threshold. The maximum convergence observed by chance is 9/10, with an extremely low frequency. Table 2: Score distribution over N= 109simulations (Corrected) Score Occurrences Frequency (f) 0/10 990,574,565 99.06% 1/10 9,163,383 0.92% 2/10 208,797 2.1×10−4 3/10 16,283 1.6×10−5 4/10 34,709 3.5×10−5 5/10 721 7.2×10−7 6/10 1,472 1.5×10−6 7/10 66 6.6×10−8 8/10 3 3.0×10−9 9/10 1 1.0×10−9 10/10 0 0.00 Total 1,000,000,000 100% The 70 cases of "Near Misses" (scores 7/10 to 9/10) correspond to partial alignments of biographical variables that fail to satisfy the full set of recursive and structural constraints (notably the intersection of C4, C5, and C9). Significance Analysis Since the actual observation (10/10) remains unique against a maximum noise of 9/10, we calculate the probability of obtaining a score ≥10 by chance with k= 0 successes out of n= 109trials (Exact Clopper-Pearson method at 95% confidence): Pupper = 1 −(0.05)1/109≈3.0×10−9(1) The conversion to standard deviation (Sigma) via the normal distribution yields: Z= Φ−1(1 −Pupper)≈5.8σ(2) This result validates the hypothesis of an extreme statistical anomaly. 3.3 Bijectivity Validation (Reciprocal Tests) To confirm the uniqueness of the solution, two exhaustive analyses were conducted in post-processing. 4 3.3.1 Inverse Monte Carlo (Temporal Specificity Test) We fixed the textual targets of the word "Idris" and tested 2,000,000 random dates. •Result: 3 occurrences of total convergence were detected. •Analysis: These 3 occurrences all correspond to the same exact date (19/03/1987 8h). •Conclusion: The temporal "launch window" is unique. 3.3.2 Exhaustive Surrogate Analysis (Lexical Specificity Test) We scanned the entire Quranic vocabulary (21,311 unique words) against the subject’s biographical profile. •Intersection (Full Match): Only one word simultaneously validates all 10 criteria: (Idris, Gem=275). 4 Discussion 4.1 Comparative Analysis and Differentiation This research aligns with the field of Computational Theology, treating sacred corpora as structured databases. It diverges from previous "numerical miracle" studies (e.g., Khalifa, 1974) which often suffered from selection bias. By introducing probabilistic rigor and functional bijectivity, we distinguish a "signal" from "noise." While previous works showed a date pointing to a word, we demonstrate that only that date points to that word, and vice versa. 4.2 Robustness against the Look-Elsewhere Effect A common critique in quantitative text analysis is the "Look-Elsewhere Effect" (or the multiple comparisons problem), where a model might be adjusted post-hoc to fit observed data. However, the constraints defined in Section 2.2 (elementary arithmetic operators, strictly non-repetitive native variables, and low Kolmogorov complexity) effectively restrict the functional search space. We estimate the order of magnitude for these functional combinations to be approximately Nspace ≈105to 106. Even when applying a highly conservative Bonferroni correction based on the upper bound of this search space (106) to the raw Monte Carlo p-value (p≈3.0×10−9), the adjusted p-value remains highly significant: padj =praw ×Nspace ≈0.003 (3) This adjusted value (0.003) comfortably satisfies the standard scientific threshold for significance (α= 0.05), confirming that the observed signal is clearly distinguishable from combinatorial noise. 5 5 Conclusion This study identifies and quantifies an objective statistical anomaly within a massive sample of one billion random profiles. The Monte Carlo simulation establishes that the probability of such a 10-criteria convergence occurring by chance is nearly null (p≈ 3.0×10−9), reaching a statistical significance of 5.8 Sigma (5.8σ). Furthermore, the reciprocal validation tests demonstrate a perfect bijectivity: the subject’s specific temporal metadata points exclusively to the target word "Idris," while that target word is the only one in the entire 21,311-word vocabulary to satisfy the model’s equations. These results strongly exclude the hypothesis of a generic statistical artifact or a trivial coincidence, revealing a mathematically unique structural correlation. Data and Code Availability To ensure full scientific transparency and reproducibility, the complete Python source code used for the Monte Carlo simulations (N= 109) and the exhaustive lexical scans, along with the biographical metadata and datasets, are publicly available. These resources are hosted on the Zenodo repository and share the same Digital Object Identifier (DOI) as this publication. This allows independent researchers to audit the algorithms, replicate the 5.8σsignificance results, and verify the bijectivity of the identified structural anomalies. 6 References [1] N. Metropolis and S. Ulam, “The Monte Carlo Method,” Journal of the American Statistical Association, vol. 44, no. 247, pp. 335–341, 1949. [2] C. J. Clopper and E. S. Pearson, “The use of confidence or fiducial limits illustrated in the case of the binomial,” Biometrika, vol. 26, no. 4, pp. 404–413, 1934. (Reference for the exact method used for k= 0 successes). [3] H. Jeffreys, Theory of Probability, 3rd ed., Oxford University Press, 1961. [4] T. Sellke, M. Bayarri and J. O. Berger, “Calibration of p-Values for Testing Precise Null Hypotheses,” The American Statistician,55 (1), 62–71 (2001). [5] C. E. Bonferroni, “Teoria statistica delle classi e calcolo delle probabilità,” Pubblicazioni del R Istituto Superiore di Scienze Economiche e Commerciali di Firenze, vol. 8, pp. 3–62, 1936. (Reference for the statistical correction applied to the search space). [6] E. Gross and O. Vitells, “Trial factors for the look-elsewhere effect in high energy physics,” The European Physical Journal C, vol. 70, no. 1, pp. 525–530, 2010. (Theoretical framework for quantifying statistical significance in large search spaces). [7] Al-Qur¯an al-Kar¯ım (Standard Egyptian Edition), Amiri Press, Cairo, 1342 AH [1924 CE]. (Canonical reference text for the Hafs recitation used in this study). [8] Centre Noon for Qur’¯anic Studies, Word, Letter, and Verse Enumeration Tables for the Canonical Hafs .Text (Dataset based on the Medina Codex), Available at: http://www.islamnoon.com/content/887/1 (Accessed 30 June 2025). [9] G. Ifrah, The Universal History of Numbers: From Prehistory to the Invention of the Computer, John Wiley & Sons, 1998. (Reference for the Abjad numeral system and Semitic gematria). [10] B. Jarrar, Irh¯as¯at al-Ij¯az al-Adad¯ı f¯ı al-Qur¯an al-Kar¯ım [Premonitions of Numerical Miracles in the Holy Quran], Noon Center for Qur’anic Studies, Ramallah, 1998. (Seminal work discussing the mathematical balance “Al-Mizan” and the number 456). [11] “Idris (prophet),” Wikipedia, The Free Encyclopedia,https://en.wikipedia.org/ wiki/Idris_(prophet) (Accessed 30 June 2025). 7 A Complete Inventory of the 30 Concordances This table presents all arithmetic and structural anomalies identified during the exploratory study. Table 3: Synthesis of the 30 Numerical Concordances No. Category Description Formula / Proof 01 Arithmetic Global rank of 1st word "Idris" D×M×87 ×H= 39 672 02 Structure Address 19:56 linked to Date S= 19,V= 8 ×7(Short Year) 03 Arithmetic Global Verse Rank (2306) 319(Month|Day) + 1987 = 2306 04 Recursive Difference Total Verses/Name 6236 −275 = 1987 ×3 05 Mixed Rank + Verse Gematria 2306 + 2140 = 19 ×3×78 06 Primes Global Rank via Prime Numbers 193(Day|Month) + P(319(Month|Day)) = 2306 07 Pivot Sum Coordinates + Name Σ(Coords) + 275 = 456 (Pivot) 08 Internal Word Internal Rank (559) 559 + 275 = 456 + 378(275+103) 09 Primes 1st Pair of Primes (Pivot Gap) Axis 319 ×7±456 →P(275) and P(391:G"France") 10 Primes Pair generated by Pivot and Rank Diff(Pa, Pb)=456 →Sura Titles 11 Calendar Solar/Hijri Conversion (Solar−ΣRanks)±Rank = 1433 12 Convergence Convergence P(x)±x P(78) −78 = 319 and P(56) + 56 = 319 13 Metadata Latin Name in Sura Titles Σ(Titles of Name) = 2306 −559 14 Semantic Inverse Rank defines Hijri Year Inv(2306) −2306 = 1407 + G("Hijri") 15 Pivot Verse 21:85 Gematria to Date 1889 −456 = 1433 16 Symmetry Identity (6 letters / 7 letters) 6×7 = 42 (S.22). S.22 has 78 verses. 17 Cluster Quadruple convergence on 22:27 Sum, Product, Context = 378 18 Metadata Verse 22:27 Gematria linked to Titles 3826 −456 = Σ(Titles "Gassama") 19 Fusion Global Rank = Sum of Times 2306 = 873(Year|Month) + 1433 20 Letter Rank of 1st letter Idris 2200 = 275 ×8 21 Theology Balance Idris + Ilyas Σ(Idr+Ily) = 809 = 456+(275+ 78) 22 Primes Sum Rangs Idris+Ilyas 2306 + 3911 = 6217 = P(809) 23 Identity Complete Epithet Gematria "Idriss Gassama Al-Fransi" = 809 24 Calendar Gap Solar/Hijri sums Σ(Solar)−Σ(Hijri) = 378 25 Linguistic Short Date Phonetics (French) Value("dix neuf trois...") = 378 26 Linguistic Hijri Date Phonetics Value("dix neuf sept mille...") = 378 27 Linguistic Full Date (French) Value(Full Date) = 378 + 139 28 Linguistic Latin Name to Hijri Date Latin Name (227)→P(227) = 1433 29 Calendar Temporal Chiasmus 809 Greg. Year 809 = Hijri 193 ; Hijri 809 = Greg 1407 8 30 Theology Latin Sum Idris + Elijah 78 + 31 = 109. Or Σ(Date) = 19 + 3 + 87 = 109. B Unified Verification Code (MC/Inverse-Surrogate) This script combines the Monte Carlo simulation logic (Code 1) and the Surrogate/Inverse specificity analyses (Code 2). 1 2 3import math 4import re 5from datetime import date 6from typing import Dict , Tuple 7 8# ==================================================================== 9# STRUCTURAL REFERENCE DATA 10 # ==================================================================== 11 12 # Quran Structure ( Number of verses per sura , index 0 = Sura 1) 13 VERSES_PER_SURA = [ 14 7, 286 , 200 , 176 , 120 , 165 , 206 , 75 , 129 , 109 , 123 , 111 , 43 , 52, 99 , 128 , 15 111 , 110 , 98 , 135 , 112 , 78, 118 , 64, 77 , 227 , 93 , 88 , 69 , 60, 34 , 30, 73, 16 54, 45, 83, 182, 88 , 75, 85, 54, 53, 89, 59, 37, 35 , 38 , 29 , 18 , 45, 60, 17 49, 62, 55, 78, 96, 29, 22 , 24 , 13 , 14 , 11 , 11, 18, 12, 12, 30, 52, 52, 18 44, 28, 28, 20, 56, 40, 31 , 50 , 40 , 46 , 42 , 29, 19, 36, 25, 22, 17, 19, 19 26, 30, 20, 15, 21, 11, 8, 8, 19, 5, 8, 8, 11, 11, 8, 3, 9, 5, 4, 7, 3, 20 6, 3, 5, 4, 5, 6 21 ] 22 23 FIXED_TARGETS = { 24 " word_rank ": {39672 , 42288} , 25 "sura ": 19, " verse ": 56, 26 " verse_rank ": 2306 , 27 " sum_coords ": 181 , 28 " verse_gem ": 2140 , 29 " word_gem ": 275 , 30 " internal_rank ": 559 31 } 32 33 # ==================================================================== 34 # VALIDATION ROUTINES 35 # ==================================================================== 36 37 def validate_date_structure (d: int, m: int, y: int, h: int) -> bool: 38 """ 39 Inverse Validation Algorithm . 40 Checks if a given date satisfies the set of fixed model constraints . 41 """ 42 # Intermediate calculations 9