scieee AI-readable full text Open interactive document viewer

Physics-Informed CSI Prediction for Mobile LEO-ISL Links: An ISAC-Enhanced Transformer Approach

Wang, Houtianfu; Akan, Ozgur B.

Abstract

This paper presents a Channel State Information (CSI) prediction method for mobile Low Earth Orbit Non-Terrestrial Network (LEO-NTN) links. High Doppler shifts and transient occlusions in these environments render traditional estimation methods outdated by the time they are applied. Our approach integrates physical channel models, specifically knife-edge diffraction (KED) and radar cross-section (RCS) based scattering, into a spatio-temporal Transformer architecture. We enforce consistency between predicted CSI and Integrated Sensing and Communications (ISAC) measurements of range and velocity through a differentiable output constraint. The model achieves $2.33{\pm}0.47$\,ms inference latency with standard parameters ($T{=}25$, $d{=}256$) while maintaining sensing overhead below $10\%$ through the use of narrow-band chirps. Evaluated on 120,000 samples, the proposed Physics-Informed Spatio-temporal Transformer (PIST) reduces normalized mean square error (NMSE) by $8.1\%$ compared to baseline methods without sensing integration ($p{<}0.01$), enabling more robust beamforming and adaptive modulation in high-Doppler LEO links. Performance improvements are consistent across X-band, Ka-band, and THz frequency ranges. We provide comprehensive evaluation metrics including latency measurements, statistical analysis, and implementation details to ensure reproducibility.

Full text

1 Physics-Informed CSI Prediction for Mobile LEO-ISL Links: An ISAC-Enhanced Transformer Approach Houtianfu Wang, Student Member, IEEE, Ozgur B. Akan, Fellow, IEEE Abstract—This paper presents a Channel State Information (CSI) prediction method for mobile Low Earth Orbit NonTerrestrial Network (LEO-NTN) links. High Doppler shifts and transient occlusions in these environments render traditional estimation methods outdated by the time they are applied. Our approach integrates physical channel models, specifically knifeedge diffraction (KED) and radar cross-section (RCS) based scattering, into a spatio-temporal Transformer architecture. We enforce consistency between predicted CSI and Integrated Sensing and Communications (ISAC) measurements of range and velocity through a differentiable output constraint. The model achieves 2.33±0.47 ms inference latency with standard parameters (T=25,d=256) while maintaining sensing overhead below 10% through the use of narrow-band chirps. Evaluated on 120,000 samples, the proposed Physics-Informed Spatio-temporal Transformer (PIST) reduces normalized mean square error (NMSE) by 8.1% compared to baseline methods without sensing integration (p<0.01), enabling more robust beamforming and adaptive modulation in high-Doppler LEO links. Performance improvements are consistent across X-band, Ka-band, and THz frequency ranges. We provide comprehensive evaluation metrics including latency measurements, statistical analysis, and implementation details to ensure reproducibility. Index Terms—LEO satellite communications, inter-satellite link (ISL), non-terrestrial networks (NTN), mobile Doppler channels, integrated sensing and communications (ISAC), CSI prediction, physics-informed learning, Transformer network I. INTRODUCTION SINCE the dawn of the Space Age, the number of satellites deployed in Low Earth Orbit (LEO) has increased rapidly [1], [2]. Fig. 1contrasts a sparse orbital scene in 1965 with a 2015 environment dense with active satellites, spent rocket stages, and fragmentation debris. This growth reflects both rapid advances in space technology and the escalating challenge of orbital sustainability. The recent deployment of largescale satellite constellations—“mega-constellations”—by national space agencies and commercial entities has further accelerated this trend [3], [4]. While these networks promise significant improvements in global connectivity (enhanced capacity, reduced latency, and higher throughput), they also exacerbate orbital congestion by adding active satellites alongside defunct spacecraft, abandoned rocket bodies, and collisiongenerated debris. Recurring near-miss incidents with debris The authors are with Internet of Everything Group, Department of Engineering, University of Cambridge, CB3 0FA Cambridge, UK. Ozgur B. Akan is also with the Center for neXt-generation Communications (CXC), Department of Electrical and Electronics Engineering, Koc¸ University, 34450 Istanbul, Turkey (email:[email protected]) Fig. 1. Illustration of orbital debris growth over time. fragments underscore the growing risk of the Kessler Syndrome [5], [6]. The consequences of this orbital crowding extend beyond physical collision risks to fundamentally alter electromagnetic propagation conditions: floating debris induces complex scattering phenomena—creating non-line-of-sight (NLOS) propagation paths, generating multipath fading on inter-satellite links (ISLs), and causing abrupt signal disruptions through transient occlusion effects. These disturbances significantly affect channel state information (CSI) accuracy—the foundation for adaptive modulation, precision beamforming, and interference mitigation—leading to reduced spectral efficiency and service availability [7], [8]. Crucially, in mega-constellations, even a tail-risk of debris-scatter-induced NLOS events as low as 10−7can erode the six-nines (99.9999%) availability targets demanded by cutting-edge services. Given these propagation impairments, the core challenge lies in predicting a channel governed by the complex and nonlinear interplay of orbital mechanics and electromagnetic wave propagation. Traditional estimation techniques, which 2 often rely on simplified linear assumptions, prove increasingly inadequate in this highly dynamic environment [9], [10]. Specifically, in high-Doppler LEO links, the coherence time scales as Tc=O(1/fD)(e.g., Tc≈κc/fD), shrinking as carrier frequency and relative velocity grow. Here, κc∈[0.2,0.5] is an empirical constant reflecting the chosen coherence-time definition and fading model. To address these challenges, this paper proposes a PhysicsInformed Spatio-temporal Transformer (PIST) framework that integrates historical orbital data with real-time ISAC sensing (Fig. 2). The key contributions are: •Output-differentiable ISAC–channel consistency. A light head gθmaps the predicted CSI b Ht+1 to (ˆr, ˆv)and enforces a differentiable term Lcons(ˆr, ˆv, b Ht+1;rISAC, vISAC), yielding ∇θLcons = 0 and closing a physical loop at the output (beyond feature concatenation). •Channel-layer physics injection. KED and RCS-scaled glints are injected at the channel layer with SGP4 geometry and shown to reduce to 3D-GBSM in limiting regimes (no edges, σ→0, low speed). •Mobile-NTN practicality. Under T=25, d=256 the predictor runs at 2.33 ms with ∼1.6M params, FP32 ≈6.4MB (FP16 ≈3.2MB), and sub-10% sensing overhead. •Validated gains. Over 100 held-out batches we obtain 8.1% NMSE reduction vs. no-sensing baseline (p<0.01, paired t-test). As a preventive, future-facing study, we validate via standards-aligned, physics-consistent simulations rather than one-to-one present-day measurements. When the aggregate control latency ∆(acquisition+feedback+inference+scheduling) occupies a fixed fraction of the coherence time Tc=O(1/fD), estimation-only CSI becomes stale; in high-Doppler LEO this often occurs once ∆≳αTcwith α∈[0.2,0.5]. We restore timeliness within the same loop by enforcing an output-differentiable ISAC–channel consistency on b Ht+1. Under T=25, d=256, measured inference is 2.33 ms with sub-10% sensing overhead, consistent with mobile-NTN real-time constraints. The remainder of this paper is organized as follows. Sec. III details data curation, where raw Two-Line Element (TLE) datasets are converted and classified into active satellites and debris. Sec. IV introduces a channel model that accounts for both line-of-sight and debris-induced multipath effects. Sec. V presents the PIST network, which uniquely fuses geometric feature extraction with Transformer self-attention guided by electromagnetic constraints. Extensive simulations in Sec. VI demonstrate PIST’s superiority over conventional methods, particularly under severe scattering conditions and varying debris densities. We conclude in Sec. VII with a discussion of implications for next-generation satellite network design and future research in space-domain awareness. II. RELATED WORK We organize the landscape along key axes to motivate and position the proposed Physics-Informed Spatio-temporal Prediction Framework LEO Satellite ISAC System Real-time data Historical data CSI Prediction Offline Training Online Prediction Fig. 2. Overview of the proposed PIST framework integrating historical data and real-time sensing. Transformer (PIST) framework. We first establish the necessity of prediction over estimation for high-Doppler LEO links. We then analyze the limitations of mainstream data-driven approaches. Finally, we contrast two paradigms for injecting external knowledge— sensing-aided communications and physics-informed learning—to identify the specific research gap that PIST addresses. A. Estimation/Mapping vs. Prediction Channel estimation and mapping target instantaneous states when end-to-end latency is well below the coherence time Tc. In contrast, channel prediction, which extrapolates h(t+∆), becomes necessary when the aggregate delay ∆(feedback + inference + scheduling) exceeds a modest fraction of Tc(e.g., ∆≳0.2–0.3Tc). In high-Doppler regimes, Tcscales roughly as O(1/fD)(for example, Tc≈κc/fDwith κc∈[0.2,0.5] depending on the model), making LEO particularly delaysensitive; see Sec. V and Table Vfor our measured inference latency. We therefore target one-step, short-horizon prediction with strictly causal side information, rather than static mapping. B. Sensing-Aided Communications (Terrestrial →NTN) Terrestrial works fuse vision/LiDAR/radar or GNSS/location to infer beams, blockage, or link states; representative studies include camera-aided beam and blockage prediction and real-world beam tracking [11], [12], LiDAR-aided beam and blockage prediction [13], [14], [15], mmWave radar-aided beam prediction with on-device learning [16], and geolocation-aided beam/power decisions [17]. These cues are largely semantic or geometric and the learned targets are mappings or link states rather than short-term CSI dynamics. For NTN/LEO, the most causal and lightweight observables are the range and radial velocity of dominant reflectors or occluders. An ISAC radar-like branch can deliver these kinematics with modest overhead, providing a direct and lightweight physical linkage to Doppler and blockage dynamics [18]. We leverage (b R, bv)at the channel layer to regularize prediction, which avoids heavy perception 3 stacks and better matches LEO real-time constraints. This stance aligns with recent TVT advances on mmWave ISAC for connected autonomous vehicles, where multibeam sensing is exploited for precise object localization under mobility [19]. Such results corroborate the value of lightweight sensing for real-time control loops in high-dynamics settings. C. Data-Driven, Graph, and Spatio-Temporal Approaches Early works use LSTM/GRU to learn temporal correlation in terrestrial and satellite links; for example, [10], and [30] mitigates channel aging in massive-MIMO LEO systems. Transformers avoid recurrence and support parallel training for CSI feedback or prediction [20], [21]. To capture spatial correlations, GNNs have also been applied. SAMS-GNN [22] employs multiscale graph convolutions for multi-band CSI; subsequent spatio-temporal GNNs add temporal edges. However, these “black-box” approaches share common limitations. They can overfit scarce orbital data and ignore governing physics, limiting robustness under debris-induced NLOS and fast Doppler. These risks amplify under distribution shift (for example, orbital regimes, blockage statistics, or sensor latency changes), where purely data-driven models often degrade unless regularized by physics. Furthermore, graph-based architectures typically lack explicit channel-layer physics and output-level coupling to ISAC observables, making interpretability and causal generalization challenging in LEO. We therefore include a graph-recurrent hybrid (GRN) as a strong supervised baseline that captures spatial-temporal correlations while remaining comparable in task definition and latency. D. Physics-Informed Learning for Wireless PIML surveys and tutorials indicate that injecting physics improves plausibility and generalization in wireless channel learning [23], while adaptive PINNs aim to scale constraintbased training via transfer or meta-learning [24]. In parallel, physics-informed generative models couple geometric channels with modern generators to synthesize physically consistent samples [25]. Another line (for example, ReVeal) incorporates Maxwell-consistent constraints for spectrum cartography [31]. In contrast, our constraints operate at the channel layer via KED (knife-edge diffraction) and RCS (radar crosssection) priors and are enforced on model outputs, making the physics signal explicitly output-differentiable and directly tied to CSI dynamics (see Sec. V-C). Most prior PIML works regularize fields or geometry via PDE/residuals or structural priors; few provide an output-level constraint specialized to LEO diffraction and scattering. E. NTN/LEO Blockage and Scattering NTN/LEO research often emphasizes standardization aspects and large-scale channel models [26], [27], geometrybased stochastic models and 3D GBSM for LEO [28], or coverage under shadowing and distance-dependent blockage [29]. F. Summary and Positioning Scope of baselines. DRL methods primarily optimize sequential resource or beam scheduling policies rather than onestep numerical CSI regression; direct score-wise comparison is thus task-mismatched. We include a graph-recurrent hybrid (GRN) to represent graph-plus-temporal predictors and keep the supervised prediction scope comparable across models. Table Isummarizes the key families along task, whether physics is injected at the channel layer, whether the constraint is output-differentiable, and NTN readiness. Compared with sensing-aided mapping (vision, LiDAR, radar, GNSS), we rely on ISAC-measured range and radial velocity as a direct and lightweight physical linkage to CSI, rather than semantic occupancy. Compared with generic PIML, we encode KED/RCS at the channel layer and translate them into an output-differentiable training signal. Finally, unlike offline or heavy pipelines, our model sustains millisecond-level latency (see Table V) compatible with high-Doppler LEO links. Strong long-sequence predictors such as temporal convolutional networks (TCN) and state-space models (SSM/S4) target similar horizons but entail different inductive biases and input interfaces. We restrict numerical baselines to supervised temporal/graph models under identical inputs, leaving TCN/SSM exploration to future work. In summary, while prior works characterize blockage statistically or through geometry-based fading [26], [27], [28], [29], explicit, differentiable modeling that links observable kinematics (R, v)to short-horizon CSI remains under-explored. We fill this gap by injecting ISAC-measured (b R, bv)and coupling them with KED/RCS priors inside the predictor, yielding an outputlevel, differentiable constraint tailored to LEO dynamics. III. DATA ACQUISITION AND PREPROCESSING A. Satellite Data Collection and Classification We retrieve raw orbital data from Space-Track, a repository of Two-Line Element (TLE) sets for thousands of LEO objects. Each TLE provides classical orbital elements and ballistic coefficients at a reference epoch. Objects are classified as: •Transmitters/Receivers (TX/RX): Identified by standardized International Designators in their TLEs. •Debris: Lacking operational designators and exhibiting higher area-to-mass ratios. This ensures that only active links and representative debris trajectories are retained for further processing. B. Orbital Propagation and Coordinate Conversion Classified TLEs are propagated via the SGP4 algorithm [32], accounting for Earth’s gravitational harmonics, atmospheric drag, and solar radiation pressure. The resulting position and velocity time series are transformed into a common Earth-Centered Inertial (ECI) frame at a unified epoch, yielding time-indexed ECI trajectories for all objects. 4 TABLE I POSITIONING AMONG RELATED FAMILIES. Work family Task Sensing Phys@Channel Output-diff NTN Vision/LiDAR/Radar/GNSS-aided [11], [12], [13], [14], [16], [17] Mapping Cam/LiDAR/Radar/GNSS – – △ Transformer CSI [20], [21] Pred./fb None – – △ SAMS-GNN [22] Pred. None – – △ PIML/PINN surveys [23], [24] Learn – PDE/Loss – △ Physics-informed generative [25] Synth – Struct. implicit/no △ NTN/LEO channel modeling [26], [27], [28], [29] Modeling – Geom./GBSM – ✓ This work (PIST) Pred. ISAC (R, v)KED/RCS ✓ ✓ Phys@Channel: physics injected at channel layer (KED=knife-edge diffraction; RCS=radar cross-section; Geom./GBSM=geometry-based stochastic modeling; PDE/Loss=PDE priors in loss; Struct.=structural priors). Output-diff: outputdifferentiable consistency. NTN: ✓ready for Non-Terrestrial Network/Low Earth Orbital scenario; △partially applicable (terrestrial/generic, no NTN-specific or real-time validation); — not applicable/not reported. PIST latency ∼2–3 ms; see Table V. C. Spatial Filtering To maintain a compact, unobstructed scenario, we apply: 1) Regional constraint: Retain only objects whose ECI coordinates lie within predefined longitudinal/latitudinal bounds. 2) Line-of-sight check: Exclude any TX–RX or debris interaction occluded by Earth’s curvature. This spatial filtering guarantees that all relevant objects colocate in an unobstructed sector, streamlining subsequent channel modeling. D. ISAC-Enabled Debris Sensing Building on recent work in THz-based debris detection [18], our ISAC subsystem uses a similar dual-frequency waveform architecture to achieve both high data throughput and robust debris sensing in crowded LEO. A high-frequency carrier (Ku/Ka or potentially THz band) is dedicated to communications, providing large bandwidth and narrow beams for high data rates. In parallel, a lower-frequency radar signal (Lband or S-band chirp) is used for debris sensing, as its longer wavelength provides deeper penetration. This division of labor exploits the strengths of each band: the short wavelength of the high-frequency link maximizes throughput, while the long wavelength of the radar enhances environmental awareness. Sensing echoes are processed every Tsense, which is aligned (greater than or equal to) with the model inference latency, to obtain noisy range and radial-velocity estimates, ˆ Rk(t)=Rk(t)+wR,k(t),ˆvk(t) = vk(t)+wv,k(t), with wR,k ∼ N(0, σ2 R)and wv,k ∼ N(0, σ2 v). These ISACderived features are fed into the PIST network as additional inputs (see Sec. V-A) and are also used to define the sensingconsistency loss term in training (Sec. V-C). Importantly, PIST itself is trained exclusively on historical orbital priors derived from TLE datasets, and the ISAC sensing branch serves solely to capture rare, transient deviations—such as debris that shift off their predicted trajectories after collision events. A practical choice for the sensing signal is an L-band Linear Frequency Modulated (LFM) chirp. We adopt out-ofband narrow chirps with a short duty cycle, whose matchedfiltering gain affords sub-meter range resolution. Let Bsbe the sensing bandwidth, PRF the repetition rate, and ρthe duty factor. The average spectral occupancy is then ρBs, and the average power consumed is ρPpeak. In our framework, this lowfrequency L/S-band radar operates in an entirely separate, nonoverlapping frequency band from the high-frequency Ku/Ka communication link, ensuring mutual isolation via simple frequency-division multiplexing (FDM). With µs-level chirps (pulse-compressed) and a sub-10% duty factor (ρ<0.1), operating at ms-level frame periodicity, the added resource load is negligible to Ku/Ka payload scheduling. Furthermore, any sensing-induced latency is inherently absorbed by the predictive nature of the framework. This FDM approach avoids the need for additional time slots or code-division schemes, allowing continuous, simultaneous sensing and communication. In this paper, “debris” refers to catalog-grade objects (typically >10 cm) that are individually trackable; sub-10 cm populations are absorbed into the background term n0(t)in (9). The simulated 3–5 equivalent large reflectors represent dominant scatterers within the TX–RX Fresnel corridor as a stress-test scenario, not a global density estimate. IV. CHANNEL MODELING METHODOLOGY Building on the preprocessed orbital data, which forms the geometric backbone of our framework, the time-varying channel characteristics in LEO satellite-debris environments are modeled through a composite response comprising the line-ofsight (LOS) component and debris-induced non-line-of-sight (NLOS) components, as illustrated in Fig. 3. Extending the theoretical framework of inter-satellite link (ISL) geometric modeling in [33], we incorporate critical enhancements for debris-rich orbital environments, including debris size and material characteristics, detailed knife-edge diffraction modeling, debris classification, and refined noise modeling. This work adopts a single-polarization channel model, a standard methodology found in the majority of the LEO CSI prediction literature [30], [18]. This approach is justified as practical LEO systems commonly use circularly polarized (CP) antennas to mitigate polarization mismatch losses, rendering the explicit modeling of a time-varying polarization state less critical. Our model therefore concentrates on the channel parameters that change most rapidly—gain, phase, and delay—which aligns with the current research consensus. Consequently, the overall 5 Space Debris Low Earth Orbit LEO Satellite LOS Path NLOS Path Fig. 3. Illustration of LEO satellite communication channel with space debris. The direct LOS path (orange solid line) and NLOS scattered paths (green and purple dashed lines) contribute to the composite channel response h(t). Space debris creates multipath propagation effects that significantly impact CSI estimation. channel complex gain consists of the direct LOS term, discrete scattering contributions from large debris, and background noise from mediumand small-sized debris. This unified model captures essential propagation phenomena—including free-space path-loss decay, dynamic shadowing due to debris occlusion, and Doppler frequency shifts arising from the relative motion of the satellite transmitter (TX), satellite receiver (RX), and scattering debris. In our simulation, the channel is evaluated at discrete time instances with a fixed sampling interval ∆t, enabling the tracking of fast time variations typical in LEO scenarios. The following subsections retain the original structure while integrating these enhancements. A. LOS Path Model The LOS component includes free-space propagation and knife-edge diffraction (KED) induced by large obstacles such as individual debris bodies. •Free-space term: Let R=|rRX(t)−rTX(t)|be the instantaneous distance, λthe wavelength, and Gt, Grthe transmitter and receiver antenna gains. The free-space amplitude is AFS =pGtGr λ 4πR.(1) •Knife-edge diffraction (KED) applicability and model: To conservatively account for occlusions by compact debris, we adopt the ITU-R P.526 single-edge diffraction model to compute the shadow factor [34]. The effective edge is tied to the projected half-size along the LoS, heff ≈D/2. Let rF=pλd1d2/(d1+d2)denote the first-Fresnel radius; a commonly used heuristic is that for D≳0.5rFthe single-edge KED well approximates UTD/dual-edge loss, whereas for D≪rFdiffraction is negligible and the link is LoS-dominant. Under this setting, define ν=heff s2(d1+d2) λd1d2 ,(2) with heff =D/2. For ν > −0.7, the diffraction loss is Ldiff (ν)≈6.9+20 log10p(ν−0.1)2+1+ν−0.1dB, (3) and the corresponding field gain is F(ν) = 10−Ldiff (ν)/20.(4) •LOS complex envelope: Incorporating free-space delay τLOS =R/c and wavenumber k= 2π/λ, the LOS term becomes hLOS(t)=AFSe−j2πfcτLOS F(ν) =pGtGr λ 4πRe−j2πfcτLOS F(ν).(5) When a large debris object blocks the LOS path, set heff = D/2in the above expressions to compute the corresponding diffraction effect. B. NLOS Scattering Model The NLOS component accounts for scattering by debris, subdivided into discrete paths from large debris and background noise from medium/small debris. Debris are classified by size relative to wavelength: Scattering Regime Size Condition RCS Approximation Rayleigh D/λ < 0.1σ∝D6/λ4 Mie 0.1≤D/λ ≤10 Oscillatory behavior Optical (Geometric) D/λ > 10 σ≈π(D/2)2 TABLE II SCATTERING REGIMES AS A FUNCTION OF DEBRIS SIZE D. Consider the X-band study case at f= 10 GHz (hence λ=c/f ≈0.03 m). Based on the standard scattering-regime boundaries—Rayleigh/Mie at D/λ = 0.1and Mie/Optical at D/λ = 10—we define: Threshold1:D λ= 0.1⇒D= 0.1λ= 0.003 m Threshold2:D λ= 10 ⇒D= 10λ= 0.30 m Debris with D < Threshold1= 0.003 m lie in the deep Rayleigh regime and are neglected. Debris with D≥ Threshold2= 0.30 m enter the optical regime and generate discrete scattering paths. Debris in the intermediate range 0.003 ≤D < 0.30 m produce continuous volume scattering, contributing to the background-noise term. 1) Debris Size Classification (Large Debris): Debris larger than Dthresh = 0.30mare categorized into four classes based on physical dimensions for further analysis:              L1 : 0.3< D ≤0.5 m, L2 : 0.5< D ≤1.0 m, L3 : 1.0< D ≤5.0 m, L4 : D > 5.0 m. 6 For X-band frequencies (at 10 GHz), each class is assigned typical RCS values derived from the NASA’s Standard Breakup Model (SEM) [35]: RCS units and conversion. All logarithms are base-10 and SI units are used (σin m2,λin m). We define [σ]dBsm = 10 log10 σ, hσ λ2idB = 10 log10 σ−20 log10 λ, hence [σ]dBsm =hσ λ2idB + 20 log10 λ. Class D/λ Range [σ/λ2]dB [σ]dBsm L1 10 – 16.7 ≈11.1 −19.4 L2 16.7 – 33.3 ≈14.6 −15.9 L3 33.3 – 166.7 ≈18.6 −11.9 L4 ≥166.7 ≈20.5 −10.0 TABLE III RCS VALUES FOR DEBRIS SIZE CATEGORIES AT X-BAND (10 GHZ) The calculation process exemplified for L1 debris: RCSL1 = [RCS/λ2]dB + 20 log10(λ) ≈11.1 dB + (−30.5 dB) =−19.4 dBsm These dBsm values are converted to linear units (σdebris in m2) for channel modeling: σdebris = 10RCSdBsm /10 (6) 2) Discrete Scattering Paths: Each large debris iyields a bistatic scattering path given by the radar equation: hi(t) = pGtGr λ√σdebris,i (4π)dt,idr,i e−jk(dt,i+dr,i),(7) where dt,i and dr,i are distances from transmitter to debris and debris to receiver, respectively. 3) Forward-Scattering Enhancement: For incidencescattering angles θ < 5◦, a forward-scattering enhancement factor is applied [36]: E(θ) = 1 + Kmax exp −1 2θ σθ2,(8) with Kmax ∝D/λ and σθ=λ/D. The amplitude of each path is multiplied by pE(θ). C. Composite Channel Response Combining LOS and all NLOS components, the overall channel impulse response is h(t) = hLOS(t) + X i:Di≥0.30m pE(θi)hi(t)+n0(t),(9) where n0(t)∼ CN(0, N0)denotes background noise. Expanding, we have h(t) = pGtGr λ 4πR e−jkRF(ν) + N X i=1 pE(θi)pGtGr λ√σi 4πdt,idr,i e−jk(dt,i+dr,i) +n0(t). (10) This enhanced channel model rigorously captures diffraction by large debris, size-specific scattering via accurate RCS quantification, forward-scattering peaks, and background noise from small debris, providing a robust foundation for CSI estimation in debris-rich LEO environments. D. Model Sanity Checks 1) S1 (LOS limit).: When no catalog-grade debris is present (e.g., D < 0.3m) or the diffraction parameter ν≤−0.7, the field factor F(ν)→1and h(t)reduces to the pure free-space term in (5), i.e., LOS with FSPL and delay only. 2) S2 (Knife-edge consistency).: For ν > −0.7, the shadowing factor follows ITU-R P.526, hence the diffraction loss Ldiff (ν)and field gain F(ν)are consistent with a standard engineering model. 3) S3 (RCS scaling).: The scattering term uses σdebris = 10RCSdBsm /10 with a forward-scattering enhancement E(θ) that scales as Kmax ∝D/λ and σθ=λ/D, matching Rayleigh/Mie/optical regimes and yielding plausible amplitude boosts near θ≈0◦. E. Connection to 3D - GBSM Under the narrowband, stationary–scatterer limit, the forward model in (10) reduces to a discrete specular sum h(t, τ) = L X ℓ=1 aℓej2πfD,ℓtδ(τ−τℓ),(11) where LOS and debris–induced paths are indexed by ℓ, with amplitudes aℓ∝E(θℓ) rℓpRCSℓ,(12) and (τℓ, fD,ℓ)determined by geometry. Consequently, the power–delay and Doppler spectra coincide with those of 3D–GBSM, Sh(τ) = X ℓ|aℓ|2δ(τ−τℓ), Sh(fD) = X ℓ|aℓ|2δ(fD−fD,ℓ), (13) i.e., the forward–scattering factor E(θ)re-weights nearforward angles but does not change the spectral support {(τℓ, fD,ℓ)}. Hence, under these assumptions, our debrisaware prior is statistically consistent with 3D–GBSM. V. DL BASED CSI PREDICTION Building upon the comprehensive channel modeling methodology established in the previous section, we now explore advanced approaches for CSI prediction in LEO satellite communication systems. The accurate modeling of timevarying channel characteristics—including LOS components, debris-induced multipath effects, and Doppler shifts—provides essential physical insights that can be leveraged for predictive modeling. As LEO constellations operate in increasingly congested orbital environments, traditional prediction methods face significant challenges in capturing the complex spatiotemporal dependencies of satellite channels [37]. This section 7 introduces a novel physically-informed deep learning framework that integrates the geometric and electromagnetic properties of satellite communications with advanced neural network architectures to achieve superior CSI prediction performance under dynamic orbital conditions. A. Overview of the PIST Architecture The proposed PIST architecture combines historical CSI, orbital kinematics, and ISAC measurements (range/velocity) through a Transformer encoder for low-latency CSI prediction. Inputs concatenate Fgeometric features with {ˆ Rk,ˆvk}Nd k=1; the backbone comprises a lightweight Conv stem, multi-layer Transformer encoder, and hierarchical temporal/target aggregation (Fig. 4). Training minimizes NMSE with an outputdifferentiable sensing-consistency term (Sec. V-C). B. Detailed Model Components A lightweight 1D-Conv stem extracts local spatio-temporal patterns; a multi-layer Transformer encoder models long-range dependencies; and a hierarchical temporal/target aggregation pools salient steps and disambiguates multi-target kinematics. The encoder ingests concatenated orbital features and ISAC-derived range/velocity, exposing debris-aware cues to the prediction head. Two fully-connected layers output real and imaginary CSI components. This compact stack preserves millisecond-level inference while retaining debris-aware expressiveness. C. Sensing Consistency Loss To ensure full differentiability with respect to network parameters, we define a sensing-consistency penalty based on debris kinematics predicted by a lightweight regression head that shares the PIST backbone. Let (ˆrk,ˆvk)=gθb Ht+1be the predicted range/velocity for occluder k, and let (rISAC k, vISAC k) be the co-temporal ISAC measurements. With Ndoccluders, a numerically stable, scale-aware form is Lsense =1 Nd Nd X k=1 "ˆrk−rISAC k max(|rISAC k|, εr)2 +ˆvk−vISAC k max(|vISAC k|, εv)2#,(14) where εr, εvare small positive constants (e.g., εr=1 m, εv=0.1m/s) to avoid division by zero and prevent undue amplification at very small magnitudes. The total loss is L=LNMSE +αLsense.(15) 1) Output-differentiable consistency: Let b Ht+1 = fθ(Xt−T+1:t)and (ˆr, ˆv)=gθ(b Ht+1). We define an ISAC–channel term Lcons(ˆr, ˆv, b Ht+1;rISAC, vISAC)and set L=LNMSE +αLcons. Because (ˆr, ˆv, b Ht+1)depend on θ, ∂L ∂θ =∂LNMSE ∂θ +α∂Lcons ∂(ˆr, ˆv, b Ht+1)·∂(ˆr, ˆv, b Ht+1) ∂θ = 0, so the constraint backpropagates to fθ, gθ, i.e., it is outputdifferentiable. 2) Penalty-Method Interpretation: Let the consistency constraints be ϕ1(hθ(x), x)=ˆr−rISAC and ϕ2(hθ(x), x)=ˆv−vISAC. The overall objective min θLNMSE(θ)+α 2 X i=1 Eh(ϕi(hθ(x), x))2i leads to the gradient ∇θL=∇θLNMSE(θ)+α 2 X i=1 E[2 ϕi(hθ(x), x)∇θϕi(hθ(x), x)] , so predictions are explicitly pulled toward the sensingconsistent manifold. D. Complex Domain Processing Strategy In terms of the complex domain processing strategy, a logarithmic amplitude-phase decoupling method is proposed: by taking the logarithm of the amplitude and compressing the dynamic range, together with the phase cyclic normalization operation, the complex CSI is mapped into a continuous differentiable space. Simultaneously, a two-channel supervisory mechanism is designed where the main loss function calculates the mean-square error in the normalized complex space, while the auxiliary loss function restricts the phase continuity. This dual supervision effectively overcomes the defects of traditional real-valued networks in destroying the integrity of the complex domain. E. Innovations and Advantages Compared with the prior art, this architecture, depicted in Fig. 4, realizes the deep integration of a data-driven model and physical a priori knowledge through the synergy of geometric semantic feature extraction and spatio-temporal attention fusion. This physics-informed feature engineering provides a robust interpretive foundation for the model, which is further enhanced by its capability for real-time sensor fusion. While the model is initially trained on historical CSI data generated from preprocessed orbital trajectories and channel simulations, its architecture inherently supports real-time sensor fusion. The Transformer encoder’s attention mechanism dynamically weights ISAC-provided debris parameters alongside historical patterns, enabling sub-second adaptation to emerging congestion scenarios. This dual-timescale processing—where convolutional layers capture persistent geometric relationships and attention heads prioritize sudden environmental changes—allows continuous refinement of CSI predictions without retraining. This design paradigm retains the ability of deep learning to capture complex channel dynamics, but also ensures that the prediction results conform to the physical reality by embedding the electromagnetic propagation mechanism, providing a new theoretical framework for the channel prediction problem of satellite communication systems. 8 Fig. 4. Architecture of the Physics-Informed Spatio-temporal Transformer (PIST) network. The model consists of a feature extraction module using stacked 1D convolutional layers, a Transformer encoder for spatio-temporal dependency modeling, and an attention-based feature aggregation module leading to the final prediction head. F. Complexity Analysis The computational complexity of the proposed PIST network can be evaluated by analyzing the number of floatingpoint operations required for each prediction. For a model with hidden dimension d, sequence length T, and input dimension Cin (given by the feature width Fplus 2Ndsensing channels; Cin=48 in our setup), each prediction requires approximately OCin ×T×d+T2×d+T×d2operations. The most computationally intensive components are the selfattention mechanisms that scale quadratically with sequence length (T2). For a typical configuration with T=25, Nd=5, and d=256, the PIST network contains approximately 1.6M parameters. By comparison, an LSTM-RNN with equivalent hidden dimensions would require about 0.82M parameters, roughly half the size of the PIST model. This difference stems from the PIST architecture’s more elaborate structure combining convolutional feature extraction, multi-head attention mechanisms, and specialized aggregation modules designed to capture orbital dynamics. While the PIST model demands more computational resources, experimental results demonstrate that it achieves significantly lower NMSE (0.53) compared to LSTM-RNN (1.10) approaches. This substantial improvement in prediction accuracy justifies the additional computational cost. Modern satellite processors have sufficient computational power to support real-time CSI prediction with low latencies, making the PIST model viable for practical deployment. For an encoder layer with hidden size d, sequence length T, and input channels Cin, the multiply–accumulate operations (MACs) are approximately CinTd+T2d+Td2. Under our settings (e.g., d=256,T=25,Cin=48), this is ≈2.1M MACs per layer; a 4-layer stack totals ≈8.4M MACs (∼8–17 MFLOPs depending on MAC-to-FLOP counting), which matches the measured 2.33 ms inference latency in Table Vand Fig. 8. Future optimization strategies include model quantization (reducing model size by up to 75%), knowledge distillation, parameter pruning, and replacing quadratic attention mechanisms with linear alternatives. Additionally, adaptive prediction scheduling based on orbital environment conditions can further reduce the average computational load by approximately 60%. These optimizations will be explored in future work to enhance the efficiency of PIST while maintaining its superior prediction accuracy in congested orbital environments. VI. SIMULATION AND RESULT ANALYSIS This section reports CSI prediction results for LEO ISL links. Model. Enhanced Transformer with 48-dim input; 4 encoder layers (2 heads), width 256, dropout 0.2. Data. 120,000 samples pairing ECI position/velocity with CSI. Physics-informed features: distance, radial velocity, azimuth, elevation, Doppler, relative 3D coordinates (8-d). Sliding windows form sequences that predict the next-step CSI (MMSE-estimated at 15 dB SNR). Stratified split: 70%/15%/15%. Training. AdamW (lr 10−3,wd10−2), batch 32, 100 epochs, OneCycleLR, mixed precision. Early stopping on validation NMSE (patience = 10) with best-epoch checkpoint; results are averaged over seeds {0,1,2}under the same grid (batch/lr/wd) across models. Train/validation NMSE decreased smoothly and plateaued under OneCycleLR across all seeds. 9 Fig. 5. Comparison of CSI prediction performance for PIST, Transformer, RNN, and GRN methods under three different frequency bands (X band, Ka band, THz band) when channel debris number is five. Fig. 6. Comparison of CSI prediction performance for PIST, Transformer, RNN, and GRN methods in the X band (10 GHz) with varying channel debris densities (3, 4, and 5). Metric. Normalized mean-square error (NMSE): NMSE =E  1 N N X n=1 ˆ h(t+n∆t)−h(t+n∆t) 2 2 ∥h(t+n∆t)∥2 2  . Baselines. We compare PIST with a standard Transformer, an LSTM, and a graph-enhanced recurrent network (GRN). For fairness, all first use only historical inputs; the sensing contributions (feature inputs and loss) are quantified via the ablation in Sec. VI-E. A. Comparison of Different Communication Frequencies Across X/Ka/THz (Fig. 5) with a fixed debris count (5, 0.5 m), PIST attains the lowest NMSE. Gains are largest at X/Ka where dynamics are milder; at THz, all models degrade due to higher Doppler and D/λ, yet PIST remains best. Physics-aware features and the output-differentiable consistency improve robustness to Doppler spikes and transient occlusions compared with purely data-driven baselines and the heavier GRN. Note that NMSE>1arises from highDoppler, transient-occlusion dynamics, underscoring the need for physics priors and output-level ISAC consistency. B. Comparison of Different Channel Density Fig. 6(X-band) and Table IV summarize the effect of debris density (3/4/5 at 0.5 m). As density rises, multipath becomes TABLE IV NMSE RESULTS UNDER DIFFERENT CHANNEL DENSITY AND COMMUNICATION FREQUENCIES Debris Freq (Hz) PIST Transformer LSTM GRN 5 1e10 0.5323 1.0216 1.1054 1.0592 3e10 0.5364 1.1778 1.1256 1.0564 1e12 0.5497 1.1069 1.3005 1.0269 4 1e10 0.7986 1.0438 1.0779 1.0622 3e10 0.7817 1.0221 1.1265 1.0990 1e12 0.6376 1.1036 1.1607 1.0378 3 1e10 0.8584 1.0386 0.9966 1.0720 3e10 0.6727 1.1325 1.1267 1.1221 1e12 0.5883 1.0699 1.0944 1.0721 richer but more non-stationary: LSTM and GRN degrade, Transformer holds up moderately, while PIST improves relative to all baselines. The sensing-aided, physics-regularized encoder leverages additional specular/near-specular paths as informative constraints rather than nuisance. Physically, the improvement tightens instantaneous channelquality estimates (CQI), enabling safer use of higher-order modulation and coding and reducing beam misalignment events under fast LEO dynamics. The net effect is higher spectral efficiency and more reliable links at millisecond-level inference cost. For quantified gains, see Sec. VI-E (Table VI). C. Comparison of Different Debris Size We vary equivalent diameters {0.5,1,5,7}m at 10/30 GHz and 1THz with debris counts 3and 5(Fig. 7). At 10 GHz, larger objects yield richer, more stable reflections and lower NMSE. At 30 GHz and 1THz, growing D/λ and Doppler induce faster phase jitter and higher-order scattering, raising NMSE for all models; PIST remains best but the margin narrows. These trends align with (i) ∆f∝fleading to larger phase jitter at higher f, (ii) increased scattering complexity with D/λ, and (iii) shorter coherence time Tcat higher fthat diminishes the value of historical samples. In highfrequency scenarios (≥30 GHz), more refined priors (e.g., dynamic Doppler compensation, partitioned scattering models) or stronger long-sequence encoders can further help. D. Performance Comparison: Accuracy vs. Latency To assess the practical viability of the proposed PIST model, Fig. 8presents a critical trade-off analysis between prediction accuracy (NMSE) and computational cost. A more detailed breakdown of the latency, including its standard deviation over 100 test batches, is provided in Table V. This analysis highlights not only the average speed but also the performance consistency of each model. Fig. 8and Table Vshow the latency–accuracy trade-off. The PIST-Baseline (no sensing aided) attains the NMSE of 0.5323 at millisecond-level latency (2.33 ms). Enabling the sensing-consistency loss (Aux-Loss) further reduces NMSE to 0.4889 under the same backbone (Table VI), while maintaining millisecond-level latency. This highlights a practical accuracy–latency sweet spot for onboard inference in LEO links.