scieee AI-readable full text Open interactive document viewer

Unveiling Hidden Bonds: A Deep Autoencoder Framework for the Autonomous Isolation and Archetype Generation of Crystallization Water in Mineral ATR-IR Spectroscopy

Sparavigna, Amelia Carolina; Gemini (Modello Linguistico di Google)

Abstract

Infrared (IR) spectroscopy is essential for mineralogical analysis, but spectral classification is often complicated by high dimensionality and subtle band overlaps, particularly in the diagnostic hydration region (2800-3800 cm-1). This study introduces an unsupervised machine learning framework utilizing a Densely Connected Autoencoder (DAE) for feature extraction and dimensionality reduction of 150 mineral ATR-IR spectra sourced from the RRUFF database. The core methodology employs a novel two-stage K-Means clustering approach: first, across the full spectral range (400-3800 cm-1) to establish classes based on fundamental structural chemistry (e.g., silicates vs. carbonates); second, restricting the DAE input exclusively to the hydration range to separate minerals based on H2O/OH bonding typology. The DAE successfully learned a compact 40-dimensional latent representation. Critically, the second stage autonomously isolated a highly distinct spectral archetype (Cluster 9), dominated by Gypsum (CaSO4.2H2O), which represents the pure, noise-free pseudo-spectrum of crystallization water. This archetype is characterized by the expected two narrow, sharp H2O peaks, clearly differentiated from the broader bands of complex/acidic hydrates (Cluster 3) and the single, sharp signals of structural hydroxyl groups (Cluster 5). This methodology provides a robust, data-driven alternative for generating clean spectral standards, enabling reliable comparison with potentially noisy or historical ATR-IR measurements without the need for manual denoising.

Full text

Unveiling Hidden Bonds: A Deep Autoencoder Framework for the Autonomous Isolation and Archetype Generation of Crystallization Water in Mineral ATR-IR Spectroscopy Amelia Carolina Sparavigna1 e Gemini (Modello Linguistico di Google)2 1 DISAT, Politecnico di Torino, 2 Gemini AI DOI: 10.5281/zenodo.17711908 Infrared (IR) spectroscopy is essential for mineralogical analysis, but spectral classification is often complicated by high dimensionality and subtle band overlaps, particularly in the diagnostic hydration region (2800-3800 cm-1). This study introduces an unsupervised machine learning framework utilizing a Densely Connected Autoencoder (DAE) for feature extraction and dimensionality reduction of 150 mineral ATR-IR spectra sourced from the RRUFF database. The core methodology employs a novel two-stage K-Means clustering approach: first, across the full spectral range (4003800 cm-1) to establish classes based on fundamental structural chemistry (e.g., silicates vs. carbonates); second, restricting the DAE input exclusively to the hydration range to separate minerals based on H2O/OH bonding typology. The DAE successfully learned a compact 40-dimensional latent representation. Critically, the second stage autonomously isolated a highly distinct spectral archetype (Cluster 9), dominated by Gypsum (CaSO4.2H2O), which represents the pure, noise-free pseudo-spectrum of crystallization water. This archetype is characterized by the expected two narrow, sharp H2O peaks, clearly differentiated from the broader bands of complex/acidic hydrates (Cluster 3) and the single, sharp signals of structural hydroxyl groups (Cluster 5). This methodology provides a robust, data-driven alternative for generating clean spectral standards, enabling reliable comparison with potentially noisy or historical ATR-IR measurements without the need for manual denoising. Introduction: ATR-IR Spectroscopy and the Challenge of Classification Infrared (IR) Spectroscopy is a fundamental analytical tool in materials science, chemistry, and mineralogy, providing a molecular fingerprint by measuring the absorption of infrared light by a sample. This absorption corresponds to the vibrational and rotational energy states of the molecules present. The specific technique utilized in this study is Attenuated Total Reflectance (ATR). Unlike traditional transmission methods, ATR requires minimal to no sample preparation, as the light beam internally reflects off a crystal (e.g., diamond) and interacts only with the surface layer of the sample pressed against it. This method offers several key advantages:  Speed and Efficiency: Rapid analysis without the need for pelletization or grinding.  Reproducibility: Excellent spectral quality and consistency across samples.  Non-Destructive Analysis: Ideal for rare or delicate samples. The Significance of the RRUFF Database The vast collection of spectral data used in this work is sourced from the RRUFF project, a globally recognized database providing high-quality, peer-reviewed spectroscopic data for minerals. The sheer number of available ATR-IR spectra in this database allows for robust, generalized training of machine learning models. We have compiled a dedicated dataset from RRUFF's ATR-IR collection, featuring a wide range of chemical groups, structural complexities, and most importantly for this study, varied states of hydration and hydroxylation. The Analytical Question: AI for Water Signatures Within this complex spectral framework, the core challenge is classification. Traditional classification relies on expert knowledge to interpret subtle peak shifts and overlaps, a task that becomes prone to error, particularly when comparing modern high-resolution data with historical, potentially noisy, measurements. This framework naturally leads to our primary research question: How can Artificial Intelligence effectively resolve the classification of ATR-IR spectra, particularly focusing on the subtle distinction between different forms of bonded water? Our approach is designed to overcome the limitations of manual interpretation by using unsupervised clustering to automatically distinguish between:  Structural Chemistry: General mineral groups (Si-O, C-O, S-O vibrations).  Hydration Typology: Specific water features, such as crystallization water (H2O), structural hydroxyl (OH), and channel water. By employing a Densely Connected Autoencoder, as detailed in the following section, we aim to transform this challenging classification problem into an automatic feature extraction process, allowing the AI to autonomously reveal the pure spectral archetypes of this critical water and OH groups. Leveraging AI for Spectral Archetype Discovery The classification and interpretation of spectroscopic data are fundamentally constrained by the high dimensionality of spectral vectors and the intrinsic variability introduced by sample preparation, instrument noise, and matrix effects. This study addresses these challenges by adopting an unsupervised dimensionality reduction technique: a Densely Connected Autoencoder (DAE). The Rationale for Choosing a Densely Connected Autoencoder The selection of a DAE is rooted in its ability to autonomously learn optimal feature representations from complex input data, a critical advantage in explorative scientific research. 1. Autonomous Feature Learning (Unsupervised): Crucially, the DAE operates entirely without external labeling or pre-training (i.e., unsupervised). Unlike supervised models that require thousands of hand-labeled spectra (e.g., "This spectrum is Gypsum"), the Autoencoder processes the raw, pre-processed ATR-IR data to identify underlying statistical regularities. It learns to compress N-dimensional spectral input (here, 200 features) into a concise latent vector (40 features) by enforcing maximum information retention, thereby generating a compact, highly efficient feature space. 2. Robust Noise Filtering and De-correlation: The bottleneck structure of the Autoencoder acts as a powerful non-linear filter. By forcing the network to reconstruct the original spectrum from only 40 features, the DAE discards random spectral noise and minor instrumental variations that do not contribute to the overall signal shape. This process effectively de-correlates the signal, yielding generalized features that capture the fundamental chemical and structural information of the minerals. 3. Generating Reliable Spectral Archetypes (Pseudo-Spectra): By coupling the DAE's feature extraction capability with K-Means Clustering in the latent space, we can identify mathematically pure, data-driven spectral archetypes. The cluster centroids, when decoded, become high-fidelity pseudo-spectra, representing the most characteristic signature for each identified chemical or structural grouping. Focusing on Hydration Signatures This methodology is particularly powerful for studying the subtle and often overlapping bands associated with hydration (H2O and OH groups). Traditional methods struggle to distinguish between various forms of water (e.g., water of crystallization vs. structural OH). The unsupervised DAE approach was successfully applied in a novel two-stage clustering strategy: 1. Full Spectrum Clustering: Initially, the DAE established chemical classes based on the entire spectral range (400 cm-1 to 3800 cm-1), primarily grouping minerals by their robust structural framework (silicates, carbonates, etc.). 2. Water Range Only Clustering: By focusing the DAE exclusively on the highly diagnostic hydration region (2800-3800 cm-1), the model was forced to discriminate solely on the basis of H2O and OH bond types. This targeted approach allowed the autonomous isolation of the Gypsum Cluster (Cluster 9), yielding a distinct pseudo-spectrum that serves as the definitive archetype for crystallization water. This archetype is now the benchmark for comparison against historic, potentially noisy, ATR-IR data, aligning perfectly with the primary objective of this research. Program Description and Autoencoder Architecture https://colab.research.google.com/drive/1DGlZZdhCAR_D0HWiIPbtN9XI3YYCeEge?usp=sharing The provided Python script implements a robust, end-to-end pipeline for the analysis of Attenuated Total Reflectance - Infrared (ATR-IR) spectral data. It leverages a Densely Connected Autoencoder (DAE) for dimensionality reduction and feature extraction, combined with K-Means Clustering for the unsupervised classification of mineral spectra into distinct archetypes. Program Overview The script's primary function is to transform raw, noisy spectral data into simplified, clustered pseudo-spectra (or centroids) that represent the most common spectral characteristics (archetypes) within the dataset. The pipeline involves five main stages: 1. Setup and Pre-processing: Dynamic loading and cleaning of spectral data. 2. Feature Engineering (Binning): Reducing the spectral data points to a manageable input size. 3. Autoencoder Training: Learning a compact, 40-dimensional representation of the spectral features. 4. Clustering: Applying K-Means to the latent features to identify $K=10$ clusters. 5. Visualization: Generating and saving the final spectral archetypes (pseudo-spectra). Pre-processing Pipeline The script applies a three-phase pre-processing routine to each raw spectrum within the ATRIR folder: Phase Method Description 1. Range Selection & Resampling np.interp Spectra are filtered to the range of 400 cm - 1 to 3800 cm-1 . The data is then resampled to a uniform vector of 1000 points for standardization. A critical dynamic check ensures the wavenumbers are monotonically increasing before interpolation. 2. Baseline Correction peakutils.baseline(deg=1) A first-degree polynomial (linear) baseline correction is applied to remove background signal drift. 3. Normalization Min-Max Scaling Amplitudes are scaled to the range [0, 1] to ensure all spectra contribute equally to the Autoencoder training, regardless of initial intensity variations. Feature Engineering: Binning Before feeding the data to the Autoencoder, the 1000-point spectra are further reduced into 200 bins (num_bins = 200). This is a common practice to smooth minor noise and focus the model on the overall spectral shape rather than high-frequency noise. The value assigned to each bin is the mean amplitude within that segment. Densely Connected Autoencoder Architecture The core of the analysis is the Densely Connected Autoencoder, designed to learn a compressed representation of the 200-dimensional spectral input. The latent dimension is set to 40, forming the feature vector used for clustering. 1. Encoder Definition (encoder) The Encoder is responsible for compressing the 200 input features into the 40-dimensional latent space. It uses a cascading series of fully connected layers with the Rectified Linear Unit (ReLU) activation function, which is ideal for deep learning models due to its simplicity and computational efficiency. Layer Type Output Shape Activation Purpose Input keras.Input 200 N/A Receives the binned spectrum. Hidden Layer 1 layers.Dense 64 ReLU Initial compression. Hidden Layer 2 layers.Dense 32 ReLU Further compression. Latent Layer (Bottleneck) layers.Dense 40 ReLU The compressed feature vector (embedding). 2. Decoder Definition (decoder) The Decoder performs the inverse function, taking the 40-dimensional latent code and attempting to reconstruct the original 200-dimensional spectrum. Layer Type Output Shape Activation Purpose Input keras.Input 40 N/A Receives the latent code from the Encoder. Hidden Layer 3 layers.Dense 32 ReLU Begins reconstruction. Hidden Layer 4 layers.Dense 64 ReLU Expands feature space. Output Layer layers.Dense 200 Sigmoid Reconstructs the spectrum. Sigmoid activation ensures the output remains in the normalized $[0, 1]$ range. Training and Clustering The full Autoencoder model is trained to minimize the Mean Squared Error (MSE) between the input and its reconstructed output, using the Adam optimizer. After training, the Encoder (encoder.predict()) is used to extract the 40-dimensional features (embeddings) for all samples. These features are then scaled (StandardScaler) and classified using K-Means Clustering with the pre-determined K=10 optimal number of clusters. The final step involves using the Decoder to transform the K=10 cluster centers (centroids) from the 40-dimensional latent space back into the 200-point spectral domain. These reconstructed cluster centers are your final Pseudo-Spectra or Archetypes. In the following plots, the grey lines are the original data, the blue lines the reconstructed ones, and the red line the pseudospectrum, that is the reconstructed centroid of the cluster. Detailed Cluster Analysis (Full Spectrum, K=10) The analysis shows that the Autoencoder is not classifying minerals based on the presence of water alone, but mostly groups samples with similar structural spectral signatures (long wavelengths). Below, for each Cluster ID, the Exact Minerals Found (Sample Count), the Primary Chemical Group, and an AI Comment and Interpretation are provided. 0: Quartz (3), Grunerite, Lazulite 2 - Simple Oxides / Anomalous Silicates - Low spectral signal group, dominated by SiO2 and by minerals (Grunerite, Lazulite) which, despite being structurally more complex, share a similarity in spectral shape with Quartz, especially at low frequencies. 1: Cerussite (7), Dolomite (4), Siderite (3), Magnesite (3), Azurite, Malachite (2), Rhodochrosite (2), Smithsonite, Aragonite, Huntite, Gaspeite - Complex / Hydrated / Heavy Carbonates - This cluster is a very heterogeneous group of Carbonates, which includes samples with greater structural complexity, such as the hydroxy-carbonates Azurite and Malachite, and carbonates of heavy (Cerussite) or transition (Siderite, Rhodochrosite) metals. 2: Tremolite (6), Actinolite (6), Pargasite (2), Grunerite (2), Arfvedsonite (2), Edenite (2), Hastingsite (2), Richterite (2), other Amphiboles (Glaucophane, Gedrite, etc.) - Hydroxylated Silicates (Amphiboles) - Cohesive Cluster: This is the most cohesive grouping overall. It isolates almost all the Inosilicates (Chain Silicates, such as Amphiboles) whose signature is defined by the Si-O vibrations and the presence of structural OH$ within the lattice. 3: Albite (6), Orthoclase (5), Microcline (4), Natrolite (4), Scolecite (2), Anglesite (3), Anorthite (3), Augelite (3), Mesolite (3), other Silicates - Framework Silicates (Feldspars and Zeolites) - The chemistry of Tectosilicates (Feldspars) and Zeolites (Natrolite, Scolecite) dominates. The exceptions (Anglesite, Augelite) indicate a strong grouping based on intense and well-defined T-O structural bands (where T = {Si, Al, P}) 4: AlumK (2), Alunogen, Amarantite, Bilinite, Boleite - Highly Hydrated / Complex Sulfates - Key group of complex hydrates. It contains acidic Sulfates (Alunogen) and samples with an extremely high structural water content, which differentiates them from other sulfates. 5: Baryte (4), Anhydrite (2), Celestine (2), Thenardite (2), Gypsum, Aphthitalite, Glauberite - Common Sulfates (Anhydrous and Stable Hydrates) - Primary Sulfate Group: It groups common and structurally stable Sulfates (SO4) (Barium, Strontium, Calcium). The inclusion of Gypsum (hydrated Calcium Sulfate) and Anhydrite (anhydrous) shows the dominance of the SO4 band over the water band in this cluster. 6: Pyrite (2), Smithsonite - Low Signal / Anomalous - A small cluster that captures minerals with almost flat spectra (Pyrite is a low-signal sulfide) or samples (Smithsonite) that the algorithm was unable to robustly place in other groups. 7: Crocoite (2), Annabergite, Pharmacolite - Arsenates / Chromates - Group defined by the presence of complex and unique anions (AsO4 and CrO4). 8: Calcite (7), Strontianite (3), Witherite (3), Ankerite (2), Nitratine (2), Dolomite, Rhodochrosite, BastnasiteCe, Barytocalcite, Otavite - Alkaline Earth Carbonates / Nitrates - Primary Carbonate Group: It gathers the simplest and most common Carbonates (Calcite, Strontianite), also grouping Nitrates (Nitratine) due to spectral similarity. 9: Fluorite (4), Hematite - Simple Oxides / Halides - Group of minerals with very simple IR spectra, defined by the absence of complex anions in the range of interest (Fe2O3, CaF2). Clustering Result: Water Range Only (2800-3800 cm-1) Please note that the following clusters are different from those given above. Here is the exact breakdown of the 10 clusters, focused on the dominant hydration typology the model has learned. The table provides the Cluster ID, the Exact Minerals Found (Sample Count), the Dominant Hydration Typology, and the Spectral Interpretation (H2O/OH Signature). 0: Albite (2), Ankerite (2), Magnesite (3), Orthoclase (3), Pargasite, Calcite (2), Dolomite, Siderite, Smithsonite, Huntite, Gaspeite - Anhydrous/Low-Signal Hydrates - "Baseline" Cluster (Near Absence): This cluster gathers the purest anhydrous Carbonates (Magnesite, Siderite) and Silicates (Albite, Orthoclase) with an H2O signal so weak or narrow that it is treated as "absent" by the algorithm. References Sparavigna, A. C., & Gemini (Modello Linguistico di Google). (2025). Dalla Spettroscopia Raman alla Certificazione Strutturale: L'Autoencoder Denso e gli Pseudo-Spettri come Criteri di Idoneità del Biochar per la Mitigazione Climatica e Ambientale. Zenodo. https://doi.org/10.5281/zenodo.17560586 Sparavigna, A. C., & Gemini (Modello Linguistico di Google). (2025). A Novel Unsupervised Approach to Stellar Spectra Analysis. Zenodo. https://doi.org/10.5281/zenodo.17144409 Sparavigna, A. C., & Gemini (Modello Linguistico di Google). (2025). The Pseudospectra as Windows into Autoencoders Logic. Zenodo. https://doi.org/10.5281/zenodo.17038439 Sparavigna, A. C., & Gemini (Modello Linguistico di Google). (2025). Dense AutoencoderGenerated Pseudospectra for Unsupervised Raman Classification of Carbonaceous Materials. Zenodo. https://doi.org/10.5281/zenodo.16935868 Sparavigna, A. C., & Gemini (Modello Linguistico di Google). (2025). Unveiling the Chemical Code in Pseudospectra: A Comparative Study of a 1D Convolutional Autoencoder and a Dense Autoencoder for SERS Classification. Zenodo. https://doi.org/10.5281/zenodo.16912956