scieee AI-readable full text Open interactive document viewer

Robust NLoS Localization in 5G mmWave Networks: Data-based Methods and Performance

6G-MUSICAL

Abstract

Ensuring smooth mobility management while employing directional beamformed transmissions in 5G millimeterwave networks calls for robust and accurate user equipment (UE) localization and tracking. In this article, we develop neural network-based positioning models with time- and frequencydomain channel state information (CSI) data in harsh non-line-ofsight (NLoS) conditions. We propose a novel frequency-domain feature extraction, which combines relative phase differences and received powers across resource blocks, and offers robust performance and reliability. Additionally, we exploit the multipath components and propose an aggregate time-domain feature combining time-of-flight, angle-of-arrival and received path-wise powers. Importantly, the temporal correlations are also harnessed in the form of sequence processing neural networks, which prove to be of particular benefit for vehicular UEs. Realistic numerical evaluations in large-scale line-of-sight (LoS)-obstructed urban environment with moving vehicles are provided, building on full ray-tracing based propagation modeling. The results show the robustness of the proposed CSI features in terms of positioning accuracy, and that the proposed models reliably localize UEs even in the absence of a LoS path, clearly outperforming the stateof-the-art with similar or even reduced processing complexity. The proposed sequence-based neural network model is capable of tracking the UE position, speed and heading simultaneously despite the strong uncertainties in the CSI measurements. Finally, it is shown that differences between the training and online inference environments can be efficiently addressed and alleviated through transfer learning.

Full text

IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, VOL. XX, NO. XX, MM YYYY 1 Robust NLoS Localization in 5G mmWave Networks: Data-based Methods and Performance Roman Klus ,Member, IEEE, Jukka Talvitie , Julia Equi , Gabor Fodor , Johan Torsner, and Mikko Valkama , Fellow, IEEE Abstract—Ensuring smooth mobility management while employing directional beamformed transmissions in 5G millimeterwave networks calls for robust and accurate user equipment ( UE ) localization and tracking. In this article, we develop neural network-based positioning models with timeand frequencydomain channel state information ( CSI ) data in harsh non-line-ofsight ( NLoS ) conditions. We propose a novel frequency-domain feature extraction, which combines relative phase differences and received powers across resource blocks, and offers robust performance and reliability. Additionally, we exploit the multipath components and propose an aggregate time-domain feature combining time-of-flight, angle-of-arrival and received path-wise powers. Importantly, the temporal correlations are also harnessed in the form of sequence processing neural networks, which prove to be of particular benefit for vehicular UEs. Realistic numerical evaluations in large-scale line-of-sight ( LoS )-obstructed urban environment with moving vehicles are provided, building on full ray-tracing based propagation modeling. The results show the robustness of the proposed CSI features in terms of positioning accuracy, and that the proposed models reliably localize UE s even in the absence of a LoS path, clearly outperforming the stateof-the-art with similar or even reduced processing complexity. The proposed sequence-based neural network model is capable of tracking the UE position, speed and heading simultaneously despite the strong uncertainties in the CSI measurements. Finally, it is shown that differences between the training and online inference environments can be efficiently addressed and alleviated through transfer learning. Index Terms—5G New Radio, channel state information, deep learning, non-line-of-sight, positioning, tracking, vehicular systems I. INTRODUCTION EXPANDING to the millimeter-wave (mmWave) frequencies allows to harness large channel bandwidths in the fifth generation ( 5G ) New Radio ( NR ) mobile communication networks, which improves the network capacity, peak data rates, and latency characteristics compared to legacy systems [1], [2]. In such mmWave networks, beamforming active antenna arrays are a critical technology, allowing directional transmission and reception capabilities, thus improving the link budget while mitigating co-channel interference and providing the basis for angle-based cellular positioning. In general, accurate real-time knowledge of the locations of the network user equipment ( UE ) is critical to ensure Copyright (c) 20xx IEEE. Personal use of this material is permitted. However, permission to use this material for any other purposes must be obtained from the IEEE by sending a request to [email protected]. Limited subset of early-stage results presented at IEEE SPAWC 2022 [14]. R. Klus, J. Talvitie, and M. Valkama are with Tampere University, Finland. J. Equi and J. Torsner are with Ericsson Research, Helsinki, Finland. G. Fodor is with Ericsson Research and with KTH, Stockholm, Sweden. This work was supported by the Academy of Finland (grants #319994, #323244, #328214, #338224, and #357730). G. Fodor and J. Equi were supported by the EU project 6G-MUSICAL, project ID:101139176. Data and codes openly available at https://doi.org/10.5281/zenodo.12204893. smooth and seamless mobility management, efficient handover management, and improved reliability of the radio link, while also allowing for location-based services ( LBS ) [3]–[6]. Baseline UE localization builds commonly on Global Navigation Satellite System ( GNSS ) based approaches. However, terrestrial positioning utilizing 5G and other signals of opportunity is of increasing importance, and is also the main technical scope of this article. This is well motivated, as the availability of GNSS is known to be compromised not only indoors but also in outdoor urban areas [4], [7], while the large channel bandwidths and directional antenna systems deployed in 5G allow for accurate timeand angle-based measurements. There is generally a wide selection of positioning methods available in the literature [4], covering both model-based Bayesian filtering approaches [8]–[13] as well as data-driven machine learning ( ML ) based methods [14]–[31]. Majority of the works focus on the line-of-sight ( LoS ) scenarios, where the channel includes a direct propagation path from the transmitter ( TX ) to the receiver ( RX ), and thus the UE location can be estimated geometrically by utilizing pseudo-ranges and/or angular information based on radio measurements. Good examples of such LoS -oriented 5G positioning works include [11], [12], building commonly on Bayesian filtering methods such as different variants of extended Kalman filter ( EKF ) and particle filter ( PF ). In case the LoS path is unavailable, geometry-based Bayesian filtering models still exist, e.g., [9], [10], [13]. Such methods are, however, typically limited to single-bounce scenarios with a single path per scattering point, and can thus easily become unreliable in realistic scattering environments while being also computationally heavy and complex [32], [33]. Therefore, UE positioning and tracking in non-line-of-sight ( NLoS ) or LoS -ambiguous scenarios call for novel solutions and models, capable of ensuring fast and reliable operation in realistic scattering environments with feasible realtime computational complexity. This is the main technical focus of this article, with a specific emphasis on vehicular systems in challenging urban environments, where NLoS scenarios commonly occur with realistic network deployments [34]. In this article, building on our initial work in [14], we propose and describe efficient ML based models for NLoS or LoS -obstructed positioning that offer robust performance, low operational complexity, good generalization properties, and wide architectural options. We harness the temporal correlation of the channel features in vehicular systems and focus on sequence processing neural network ( NN ) methods as the fundamental ML engine. We propose two alternative feature sets capable of describing the radio channel and further NN - based positioning, namely frequency-domain and time-domain This article has been accepted for publication in IEEE Transactions on Vehicular Technology. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/TVT.2024.3456958 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, VOL. XX, NO. XX, MM YYYY 2 channel state information ( CSI ) features. Additionally, we include the vehicle speed and heading as additional model outputs to enable efficient tracking using a single model. ML -based positioning has been addressed in the recent literature, e.g. in [14]–[18], [20]–[31]. To this end, received signal strength ( RSS ) measurements were adopted in [19]–[21], in indoor IEEE 802.11 network ( Wi-Fi ) fingerprinting context, while the corresponding 5G deployment was considered in [35]. The work in [24] proposed a CSI-based fingerprinting model that utilizes the amplitude response of the channel. Due to the classifier-like NN , the proposed system is, however, limited to small deployments. The authors of [15] developed an NN -based feature extractor with Wi-Fi CSI measurements considering the amplitude response only, followed by a k -Nearest Neighbors ( k-NN ) positioning algorithm. Similar approaches building on channel amplitude response measurements were considered also in [25] and [26]. In [27], a Wi-Fi positioning approach utilizing the channel phase response as a feature was proposed. The method is, however, not suitable for large-scale scenarios due to the classifier-based NN , while the considered phase slope calculation is also subject to ambiguities. The work in [22], in turn, presented a 5G positioning system, however, being limited to the LoS scenario while utilizing only the beam-specific reference signal received power ( RSRP ) values as the features. In [23], a positioning system that generates probability maps using an NN model was described with a feature representation that transforms the frequency-domain uplink ( UL ) CSI data to a delay-domain. The probability maps enable efficient sensor fusion, yet their scale directly affects the complexity of the underlying NN model. Furthermore, [29] described a paradigm to produce a high-accuracy 5G localization dataset, building on channel frequency response measurements. Different hybrid solutions also exist in the literature, either in terms of aiding Bayesian filtering solutions through ML methods or fusing measurements from various different sensors [16]. To this end, [17] proposed a series of recurrent neural network ( RNN ) models to replace the EKF and thus enable implicit and data-driven learning. Furthermore, a sensor fusion approach using reinforcement learning-assisted particle filtering is described in [10]. Neural network based 5G fingerprinting and GNSS data fusion were, in turn, considered in [18]. Recently, in [36], an NN model classifying different propagation paths from time-domain CSI in the form of path parameters was paired with geometry-based positioning algorithm considering LoS and single-bounce NLoS paths. ML -based localization with propagation time measurements has also gained interest in recent works [14], [30], [33]. In [30], a bidirectional RNN is employed to track the UE based on the time-of-arrival ( ToA ) measurements from multiple nodes, while the authors of [33] utilize a convolutional neural network ( CNN ) to estimate the accurate ToA from the raw channel impulse responses in LoS / NLoS scenarios. However, only LoS measurements are used for the actual localization. Finally, angular information at either the gNodeB ( gNB ) or the UE can also be used for ML -based localization, as shown in [31]. The main related works and their relevant aspects are summarized in Table I. Importantly, the NLoS positioning under rich and realistic scattering is not explicitly addressed, in particular in the context of beamforming 5G mmWave networks. Thus, complementary to the existing literature, this article focuses on ML -based reliable network localization using timeand frequency-domain CSI data with emphasis on challenging NLoS scenarios, while noting also various relevant uncertainty aspects. The application focus is on vehicular systems in urban environments with 5G mmWave deployments and CSI features that can be obtained through 3GPP standardized UL and/or downlink ( DL ) measurements and corresponding signaling. The contributions and novelty compared to the existing ML -based positioning literature can be stated and summarized as follows: • We introduce, derive, and evaluate efficient frequencydomain CSI features in the form of sparse power and phase measurements, and their combinations, for ML - based positioning models and compare their robustness with the ones introduced previously in the literature; • We also introduce and evaluate alternative time-domain path-wise CSI features and demonstrate their effectiveness in wireless positioning scenarios while finding the relevant and best-performing feature combinations; • We develop a novel hybrid NN processing model in terms of instantaneous and sequence data processing for simultaneous UE location, velocity, and heading tracking using the above channel-based features; • We evaluate the performance of different features and processing models in a large-scale, realistic, urban scenario with full ray-tracing-based channel measurements under harsh NLoS conditions in the context of 28 GHz mmWave 5G network, while also considering realistic training and measurement uncertainties; • We show that the proposed time-domain and frequencydomain features outperform the benchmark solutions, especially when combined with the sequence processing ML model, in terms of the positioning accuracy and complexity – in particular, in the challenging multi-bounce scattering environments, where the proposed positioning approach achieves comparable accuracy in both LoS and NLoS conditions, regardless of the number of bounces; • We also address the important practical issue of environment or gNB deployment differences between the training phase and the actual online inference phase, through transfer learning, and show that specializing to the current deployment is feasible; • Finally, we address the complexity of the developed methods, in comparison to the prior-art, while also openly share the data and codes for research reproducibility and transparency. For clarity it is stated that selected time-domain features were initially considered in [14], however, the adopted NN models were lacking the advanced sequence processing capabilities. Additionally, no frequency-domain features were considered, while the transfer learning aspects were also fully neglected. The rest of this article is organized as follows. Section II introduces the network measurements applicable for positioning and their acquisition with 3GPP compatible reference signals and measurement procedures. Section III introduces the proposed frequency-domain and time-domain CSI data features and the This article has been accepted for publication in IEEE Transactions on Vehicular Technology. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/TVT.2024.3456958 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, VOL. XX, NO. XX, MM YYYY 3 TABLE I SUMMARY OF RELATED WORKS Reference Technology Features ML model Tracking Uncertainty Vehicular NLoS Feat. Label Envir. Systems [14] 5G mmWave AoA and ToF DNN, LSTM ✓ ✓ ✓ ×✓ ✓ [15] Wi-Fi channel amplitude response CNN × × × × × ✓* [18] 5G mmWave RSRP DNN ×✓× × × × [22] 5G mmWave RSRP DNN × × × × × ×** [23] Wi-Fi channel frequency response DNN ×✓× × × ✓ [24] Wi-Fi channel amplitude response DNN ✓× × × × × [25] Wi-Fi channel amplitude response CNN × × × × × ×* [26] Wi-Fi channel amplitude response ML ✓× × × × ×* [27] Wi-Fi channel phase response DNN ✓× × × × × [29] 5G mmWave channel frequency response CNN ×✓× × × × [30] Unspecified ToA RNN ✓× × × × × [31] BLE AoA CNN ×✓× × × ✓*** [33] LTE channel impulse response CNN ×✓× × × ✓* [36] mmWave path-wise CSI DNN ×✓× × ✓ ✓** This Work 5G mmWave timeand frequency-domain CSI DNN, LSTM ✓ ✓ ✓ ✓ ✓ ✓ * Some NLoS samples available in evaluation; ** Removing the detected NLoS samples before positioning; *** RX signal subject to Rayleigh fading. related pre-processing. Additionally, the NN processing models and architectures are described incorporating both instantaneous and sequence models. Section IV describes the considered urban vehicular positioning scenario and evaluation environment, together with the practical measurement or data uncertainties. Additionally, the obtained numerical results are presented and analyzed, while also considering the important aspect of specializing to the prevailing gNB deployment through transfer learning and addressing the processing complexity in terms of parameter counts. Finally, Section V concludes the work. II. POSITIONING MEASUREMENTS AND DATA ACQUISITION This section introduces the signals and standardized measurements, available in 5G NR , to extract positioning data. Specific focus is on synchronization signal block ( SSB ) based measurements in DL, in terms of frequency-domain data, while in timedomain we harness multi-round trip time ( MRTT ) and UL angleof-arrival ( AoA ) based multipath measurements utilizing UL sounding reference signal ( SRS ) and DL positioning reference signal ( PRS ). The relevant measurements and data acquisition methods are illustrated conceptually in Fig. 1, while being described in detail below. For clarity, we state that the frequencyand time-domain measurements are alternative approaches to obtain positioning features and data. A. Signals and Measurements for 5G NR Positioning Measuring the received signal strength is one common approach utilized in wireless positioning. In 5G NR , we distinguish RSRP , reference signal received quality ( RSRQ ), and received signal strength indicator ( RSSI ), including their beamspecific and resource-specific alternatives. Their acquisition is defined in [37] building on different reference signals, such as UL SRS and DL synchronization signal ( SS ) and PRS . The corresponding measurements are called UL - SRS - RSRP , SS - RSRP and DL - PRS - RSRP , respectively. Signal strength measurements are vital for numerous network functions, such as mobility management, and thus regularly collected. Compared to signal strength-based measurements, propagation delay or time-of-flight ( ToF )-based ranging benefits from large transmission bandwidths while being less sensitive to channel effects, such as reflections, diffractions, and scattering. To relax clock synchronization requirements between TX and RX , the 5G NR standard supports MRTT measurements [38] where the gNB measures the round trip time, denoted as ”gNB Rx–Tx time difference”, based on PRS transmission in DL and SRS transmission in UL . In addition, the UE measures the time between receiving the PRS and sending the SRS , denoted as the ”UE Rx–Tx time difference”, which is reported to the gNB [37] in order to solve the channel-dependent propagation delay. Alternatively, the ToF can be estimated indirectly at the gNB via time-difference-of-arrival ( TDoA ) and the related positioning calculations, as defined in [38]. Obtaining the ToF directly at the UE is currently not explicitly standardized. Beamformed radio access provides inherent support for angle estimation and corresponding angle-based positioning schemes. In the current 5G NR standard, angle estimation is directly specified only for the gNB -side angular information either via the uplink angle-of-arrival ( UL-AoA ) or the downlink angleof-departure ( DL-AoD ) [38]. The UL-AoA is defined as the estimated azimuth and vertical angles of a UE , observed at a gNB [37], based on UL SRS . The exact angle estimation method used in the UL-AoA is not specified, which allows for performance optimization. On the other hand, the estimation of the gNB angle at the UE using the DL-AoD is practically restricted to the use of spatial power measurements, i.e., DL - PRS - RSRP , which limits the achievable angle estimation accuracy [38]. Importantly, estimating the ranges and angles also for paths beyond the LoS component is feasible [13], offering added value to the positioning task [9]. Since the current 5G NR standard does not specify accurate estimation of path ranges or angles at the UE side, for a DL -based positioning we consider observing frequency-domain CSI measurements based on SS s transmissions by the gNB s. This is beneficial since the SS s are periodically transmitted and thus systematically available. Additionally, for an UL -based positioning method, we consider observing time-domain multipath measurements, including path-wise angles and propagation delays, directly at the gNB s. For obtaining the path-wise angles and propagation delays in practice, it is possible to exploit the above-discussed 5G NR specified UL-AoA and MRTT methods, respectively. This article has been accepted for publication in IEEE Transactions on Vehicular Technology. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/TVT.2024.3456958 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, VOL. XX, NO. XX, MM YYYY 4 DL SSB/PRS UL SRS gNodeB UL AoA gNB Rx-Tx td LMF MDT bRSRP IDs location gNodeB UL AoA gNB Rx-Tx td UE bRSRP CSI IDs UE Rx-Tx td DL SSB/PRS *idle mode measurements only Multi-RTT UL-AoA Frequency Domain Time Domain Feature Pre-processing NN Model x y h v Feature Pre-processing NN Model x y h v Beam-CSI UL-AoA MRTT PATH-RP Fig. 1. Illustration of the network data acquisition scheme with two gNB s, a UE , and a network localization entity represented by the Location Management Function ( LMF ). Different UL and DL reference signals form the physical basis for obtaining positioning measurements and data. Additionally, the baseline neural processing chains from features to UE location, heading and velocity are highlighted, for both timeand frequency-domain feature scenarios. B. Frequency-Domain Channel Measurements through Beamformed DL SSs The frequency-domain CSI relates to the frequency response ( FR ) of the effective channel between the gNB and the UE . Considering orthogonal frequency-division multiplexing ( OFDM ) based transmission, the antenna-element-wise FR at subcarrier n, denoted as H(n)∈CNRX×NTX , can be written as [9], [13] H(n) = K−1 X k=0 hke−j2πnτkFs NaRX(θAOA,k)aH TX(θAOD,k),(1) where K is the number of paths, while hk , τk , θAOA,k and θAOD,k are the complex path coefficient, ToF , AoA and angle-ofdeparture ( AoD ) for the kth path, respectively. Furthermore, NTX and NRX are the numbers of transmit and receive antennas in respective order, Fs is the sampling frequency, and N is the OFDM Fast Fourier Transform ( FFT ) size. Finally, aTX(·)∈CNTX and aRX(·)∈CNRX are the steering vectors, which define the phases per antenna element with respect to the array center, for given AoD and AoA . Considering further the analog phased-arrays in mmWave systems, the TX and RX apply beamforming weights  bTX ∈CNTX and  bRX ∈CNRX , respectively. The corresponding effective beamformed channel at subcarrier n , considered as the frequency-domain CSI, can then be expressed as H(n) =  bH RXH(n) bTX.(2) In practice, besides noise and interference, the CSI estimation can suffer from inaccuracies [28] due to radio frequency ( RF ) impairments, clock and frequency offsets between the gNB and the UE , and imperfect timing advance information. Furthermore, due to signaling overhead, CSI is often reported per blocks of subcarriers, which reduces the CSI resolution in frequency. In this paper, we consider obtaining the frequency-domain CSI via 5G NR SSs, transmitted periodically in DL by all gNBs. C. Time-Domain Multipath Measurements through MRTT and UL-AoA In time-domain, the radio propagation channel can be modeled as a composition of individual propagation paths with path-specific propagation delay, power gain, phase shift, AoD , and AoA , together with additional distortion and interference, among other channel effects. The antenna-element-wise channel impulse response H(τ)∈CNRX×NTX can be written as a function of propagation delay τas [34], [36] H(τ) = K−1 X k=0 hkaRX(θAOA,k)aH TX(θAOD,k)δ(τ−τk)(3) where δ(·) is a Dirac delta function (i.e., a unit impulse). Similar to the frequency-domain representation, while again assuming analog phased-arrays, the effective beamformed channel impulse response can be written as H(τ) =  bH RXH(τ) bTX.(4) In this paper, the measured time-domain CSI includes the path delays τk , the path powers |hk|2 , and the gNB side path angles θAOA,k . For the path delays τk and path powers |hk|2 , the estimation procedure is assumed to exploit 5G NR MRTT measurements [38], as discussed in Section II-A . Furthermore, for estimating the gNB side path angles θAOA,k , it is also possible to utilize UL-AoA measurements [38] based on UL SRS transmissions, as noted in Section II-A. The overall data acquisition concept is illustrated in Fig. 1, high-lighting the different considered measurements. In general, within the current 5G NR standard, the LMF is responsible for the localization and related signaling management while the positioning calculations can be carrier out either at the UE or the network side. Moreover, reporting the MRTT and UL-AoA measurements between the gNB and the LMF is supported by the so-called NR Positioning Protocol A [39]. Different alternative ways to arrange for labelled training data include crowdsourcing, crowdsensing, as well as utilization of synthetic data. These are discussed further in Section IV.G. III. PROPOSED METHODS This section describes and introduces the novel approach of utilizing frequency-domain CSI with relative phase, while also addressing the time-domain CSI data pre-processing. In addition, the proposed architectures, hyperparameters, and training algorithm of the proposed NN models are presented. This article has been accepted for publication in IEEE Transactions on Vehicular Technology. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/TVT.2024.3456958 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, VOL. XX, NO. XX, MM YYYY 5 Finally, important system-level implementation alternatives and aspects are discussed. A. Frequency-Domain CSI Data Preprocessing 1) Proposed Relative Phase Approach: 5G mmWave networks operate at high carrier frequencies, at and beyond 24 GHz, with wavelengths approaching the millimeter-scale. In mobile scenarios, utilizing absolute phase responses is highly impractical, as movement of a few millimeters in distance results in a full rotation of the phase. Furthermore, as is wellknown, the frequencies and wavelengths relate through λ=δs=c/f, (5) where δs is the propagation distance between two points with equal phase, λ is the signal wavelength, c is the speed of light, and fis the signal frequency. In this article, we consider obtaining the frequency-domain CSI in the resolution of 12 subcarriers, which refers to a bandwidth of one resource block ( RB ) in 5G NR , denoted as ∆fRB . The CSI is interpreted at the 6th subcarrier of each RB , and thus the corresponding subcarrier index for the mth RB observation is given as nRB,m= 6 + m∆fRB with m= 0, ..., M −1 , where M is the number of RB s. Moreover, we propose to take advantage of differential phase measurements between neighboring RB s. Based on (1) - (2) , and when considering an individual propagation path, the phase difference between subcarriers is equal across the spectrum and completely determined by the ToF through the complex exponential term. Possible phase rotations due the other terms in (1) are constant over all subcarriers, and thus do not induce phase difference between the subcarriers. Thus, for an individual path with ToF of τ0 , the phase difference between two consecutive resource blocks can be given as ∆ϕ= 2πτ0∆fRB .(6) While the above expression builds on a single propagation path, we utilize this approach in this work also in case of realistic multipath propagation. As elaborated further below, the differential phase approach allows to mitigate the effect of phase periodicity, and thus extract relevant features for the proposed NN-based positioning. To this end, the linkage between the relative phase and a specific propagation distance is unambiguous only when the relative phase is within one phase cycle ( ∆ϕ < 2π ). A distance dφ , which inflicts the full 2π cycle of the relative phase between two neighboring RBs, can be solved based on (6) as dφ=c ∆fRB (7) by denoting τφ∆fRB = 1 , where τφ=dφ/c is the corresponding ToF resulting in a full phase cycle. By using the relative phase difference ∆ϕ , instead of an absolute phase, as the frequency-domain feature, the positioning performance can be significantly improved, as shown in Section IV. Although the distance ambiguity issue still remains with the phase difference recurrence at every dφ meters, it is greatly improved compared to the recurrence level with an absolute phase at every δs meters, as f≫∆fRB. 2) Frequency-Domain CSI Features and Visualization: To provide a short illustration, we consider a single representative user path along an urban environment, as shown in Fig. 2a (for further details of the environment, refer to Section IV). Then, Fig. 2b and Fig. 2c demonstrate the utilized frequencydomain CSI feature representations along the path, including the proposed features and the features from the related literature. Specifically, the raw, complex channel response, further referred to as complex frequency response ( FR-Complex ), is depicted in Fig. 2b, top, and Fig. 2b, center, which show the real and imaginary parts of FR-Complex for 10 consecutive resource blocks along the path. The feature is obtained from (2) as real(H(nRB,m)) and imag(H(nRB,m)) for m= 1, ..., 10. Furthermore, Fig. 2b, bottom, depicts the channel power response, denoted as the power-domain frequency response ( FR-Power ). Such channel feature is utilized, e.g., in [24], [25], and can be expressed via (2) as 10log10(|H(nRB,m)|2) . In 5G NR , the FR-Power corresponds to a RB -wise RSRP measurement, defined as the average power of the resource elements carrying the reference symbols. As an input feature, we also re-scale FR-Power to the normalized range of [0,1]. Then, the top graph of Fig. 2c visualizes the proposed relative phase difference as the frequency-domain feature, further denoted as the relative phase frequency response ( FR-Phase ). Specifically, building on the discussion in Section III-A 1, the FR-Phase can be obtained and expressed following (2) as ∆ϕ(m) = arg(H(nRB,m)) −arg(H(nRB,m−1)) (8) for RB indices m= 1, ..., M −1 . The dependency between the signal path lengths and the relative phases, especially in LoS regions, is clearly visible in the figure. Moreover, it can be seen that the feature magnitude is recurring with a path propagation distance at every dφmeters, as derived in (7). 3) Proposed Combined Feature: To utilize the maximum information enclosed in the measured channel responses, we further propose the so-called power and relative phase frequency response ( FR-Power/Phase ) approach as the ultimate frequencydomain feature. This approach combines the FR-Phase and FR-Power by transforming the FR-Phase to the complex unit circle, with subsequent element-wise multiplication with the re-scaled FR-Power. This is expressed as ¯ PRB(m)exp(j∆ϕ(m)) (9) where ¯ PRB(m) refers to the re-scaled normalized power for RB indices m= 0, ..., M −1 . Furthermore, since the FR-Phase has one element less than FR-Power , we extend the FR-Phase array with an additional element for m= 0 by defining ∆ϕ(0) = 0 . The real and imaginary components of the proposed FR-Power/Phase feature set are visualized in Fig. 2c, center and bottom, respectively. The proposed feature allows to accommodate the advantages of both received power and relative phases in a single complex feature vector, while relaxing the distance ambiguity of the relative phase feature. B. Time-Domain CSI Data Preprocessing The time-domain CSI data utilized for localization includes propagation delays τk , powers |hk|2 , and gNB side path angles This article has been accepted for publication in IEEE Transactions on Vehicular Technology. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/TVT.2024.3456958 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, VOL. XX, NO. XX, MM YYYY 6 100 150 200 250 300 x [m] 0 50 100 150 200 250 300 350 y [m] 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 Normlaized channel gain [-] (a) 0 100 200 300 400 500 600 -0.5 0 0.5 Norm. value [-] FR-Complex (real) 0 100 200 300 400 500 600 -0.5 0 0.5 Norm. value [-] FR-Complex (imag) 0 100 200 300 400 500 600 Distance travelled [m] 0 0.5 1 Norm. gain [dB] FR-Power (b) 0 100 200 300 400 500 600 -2 0 2 Diff. phase [rad] FR-Phase 0 100 200 300 400 500 600 -1 0 1 Norm. value [-] FR-Power/Phase (real) 0 100 200 300 400 500 600 Distance travelled [m] -1 0 1 Norm. value [-] FR-Power/Phase (imag) (c) Fig. 2. An example UE track in urban environment is shown in (a) where the gNB location is depicted with the red rectangle while the track color represents the mean normalized power across the resource blocks. The frequency-domain features FR-Complex and FR-Power as well as the proposed FR-Phase and FR-Power/Phase are visualized in (b) and (c), respectively, along the UE track shown in (a). In (b) and (c), different colors represent the ten different RB allocations within the full SSB transmission bandwidth. θAOA,k for the observed LoS and NLoS paths. To this end, the measured path-wise propagation delays τk , referred to as path-wise time-of-flight ( path-ToF ), are transformed to propagation distances by multiplying them with the speed of light. The path-wise AoA s, θAOA,k , are obtained at the gNB side, and transformed to directions in Cartesian coordinates in the preprocessing, to omit the zero-crossing problem with cyclic angular data. This results in a robust AoA feature, called path-wise angle-of-arrival ( path-AoA ) in the following, which is less susceptible to angular deviation and related uncertainties. The path-wise received powers ( path-RP s) |hk|2 are expressed in decibels (dBm) to overcome the extremely low feature magnitudes in linear scale. Such feature is the timedomain equivalent of the frequency-domain FR-Power , which accumulates all paths into the same observed frequency-domain measurement. Similar to the propagation delay feature, the path power feature includes information on the path propagation distance, but most importantly, it also provides information on the number and type of channel interactions, such as reflections, diffraction, or scattering, within the radio path. In general, different combinations of the time-domain CSI data can be adopted. The aggregated path-ToF+AoA and pathToF+RP+AoA features, proposed in this work, are the most powerful ones, as shown through the numerical results. C. NN Model Architectures and Hyperparameters Among the various alternative data-aided approaches, we restrict ourselves to NN models in this work, which currently dominate the ML area due to their performance, scalability, generalization properties, and dynamic architecture options [40]. 1) Activation Function: In this work, we utilize the Gaussian error linear unit ( GELU ) [41] as the non-linear activation function. Its main advantages over traditional rectified linear unit ( ReLU ) include resistance to a “dying ReLU” problem [42], differentiability at all values while having also been shown to offer improved performance already in a number of applications such as natural language processing [41]. It can be defined as GELU(x) = xΦ(x) [43], [44] where Φ(x) is the cumulative distribution function of the standard normal distribution. The function can also be approximated for faster processing as GELU(x) = x 2tanh r2 π(x+Cx3)!,(10) where C= 0.044715 . Compared to ReLU , the higher complexity of GELU is compensated by the faster convergence of the model, as well as the corresponding improved positioning performance, based on our complementary experiments. 2) Utilized NN Architectures: As the functional NN layers, we utilize in this work both densely connected layers and long-short term memory ( LSTM ) [45] layers, the later being used only in the sequence-based implementation of models. Importantly, the LSTM layer is a recurrent-based layer capable of preserving long-term and short-term trends within the data. The architecture of the densely-connected model is depicted in Fig. 3a. It consists of 5 densely connected layers with GELU activation functions and a single densely connected layer with linear activation and 2 neurons as the output, estimating the UE position. The architecture of the sequence processing capable model is, in turn, shown in Fig. 3b. It consists of 5 densely connected layers after the input, with a single LSTM layer connected in parallel with the 5 th dense layer. The concatenated output of these layers is then fed to an LSTM layer with 5 neurons at the output with linear activation. Specifically, the densely connected layers serve as instantaneous feature extractors, while the intermediate LSTM layer learns the temporal features. Due to the considered parallel architecture, the last functional layer has access to both instantaneous and temporal features. The resulting output is then divided into a positioning output with 2 variables, a velocity output with a single variable, and a heading output with an additional tanh(·) activation and 2variables. 3) Data Structures, Normalization and Training: In general, the input dimensions vary based on the selected features This article has been accepted for publication in IEEE Transactions on Vehicular Technology. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/TVT.2024.3456958 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, VOL. XX, NO. XX, MM YYYY 7 I N P U T 5 1 2 G E L U 5 1 2 G E L U 5 1 2 G E L U 5 1 2 G E L U 5 1 2 G E L U 2 x y (a) v I N P U T 1 0 2 4 G E L U 5 1 2 G E L U 5 1 2 G E L U 5 1 2 G E L U 1 2 8 G E L U 5 1 2 G E L U 5 x, y T G H Hx Hy (b) Fig. 3. Architecture and hyperparameters of the densely connected NN , in (a), and of the sequence processing NN , in (b). Each layer is specified by the number of neurons and an activation function. and deployment scenario. For frequency domain features, and when considering the evaluation scenario described in Section IV containing 3gNB s, 16 beams, and 10 RB s, the input size is either 480 for FR-Power and FR-Phase or 960 for FR-Power/Phase and FR-Complex . When considering the timedomain CSI in the same scenario, the individual path-ToF , path-AoA , and path-RP features have each an input size of 15 , the combined path-ToF+AoA and path-RP+AoA features have 30 inputs, and finally the aggregated path-ToF+RP+AoA feature has an input size of 45 . Some of the features are also normalized prior to the training, as demonstrated already along Fig. 2b and Fig. 2c. Specifically, all power-related quantities as well as path-wise ToF measurements, when first converted to pseudoranges, are all normalized between [0,1] within the overall sets of available measurements. Finally, all angle and phase quantities are, by design, within [−π, π] . We emphasize that each different feature scenario and combination corresponds to an individual NN , trained and deployed on its own. The vast set of numerical results, provided in Section IV, provides the corresponding mutual performance comparisons. All considered NN models are trained using the Adam optimizer [46] with learning rates of 0.001 for the first 200 epochs, and then an early-stopping mechanism based on validation performance for additional 500 epochs, while iteratively reducing the learning rate to 0.0005 and 0.0001 after each stop. The lowered learning rates ensure a fine-tuned performance with a small number of epochs. The mean squared error (MSE) loss was selected for each output, and for the sequence-based NN model, the loss weights were selected as 0.8 , 0.1 , and 0.1 for positioning loss, velocity loss, and heading loss, respectively. Furthermore, stemming from the deployment area of around 550 ×370 m 2 (see Fig. 4), the position labels are reduced by a factor of 300 to accelerate the training. D. System-Level Implementation Alternatives and Aspects In general, there are alternative ways to organize and implement the use of the CSI measurements and data for NN training and actual online inference processing for localization. These are discussed below, in relation to the proposed methods and the data acquisition visualized in Fig. 1, while noting also the important role of UE radio resource control (RRC) state. To this end, the time-domain CSI data, i.e., the MRTT -based ToF measurements and the SRS-based UL-AoA measurements, are by definition obtained at the network side. Thus, in this case, it is natural to also perform both the model training as well as the localization inference processing at the network side. Consequently, there is no need for additional signaling or feedback, and all training data from different UE s is inherently gathered together for training the model. Importantly, since MRTT and UL-AoA require scheduled SRS and PRS transmissions, time-domain measurements are only available in the connected mode when it comes to the UE RRC state. Frequency-domain RSRP and other CSI measurements are collected from periodic and always available SSB transmissions at the UE side, thus enabling utilization of efficient data crowdsourcing methods. Despite a possible technical capability to perform training at the UE , assuming individually trained models at different UE s can be considered unrealistic. Therefore, UE s are expected to periodically share such measurement data with the network for NN training, for example, through Minimization of Drive Testing ( MDT ) messaging in the form of raw measurements and location tags, or alternatively as locally pre-trained models following the principle of federated learning (FL). Interestingly, unlike with MRTT and UL-AoA , the DL frequency-domain CSI and RSRP measurements can be collected and obtained also in the RRC idle mode as part of standard mobility management procedures. This can be considered a great asset enabling continuous data collection and localization with very low power consumption. Furthermore, assuming a pre-shared model from the network for the final inference phase, the UE can perform localization independently without supplementary signaling with the gNB – allowing thus UE localization and tracking also in the idle mode. IV. EVALUATION ENVIRONMENT AND RESULTS A. Evaluation Scenario and Assumptions The evaluation environment builds on ray-tracing-based channel measurements utilizing Wireless Insite®software [47]. We employ the map-based METIS Madrid grid [48], recognized as the relevant urban scenario by 3GPP in 5G NR specifications [34]. The Madrid grid layout introduces generally a rich radio propagation environment with different street widths and open areas, empowering generalization and scalability. The simulated urban scenario illustrated in Fig. 4 contains three 5G mmWave gNB s operating at 28 GHz, such that clear NLoS regions also exist along the streets. Each gNB is equipped with a uniform cylindrical antenna array with 4 elevated layers, each with 16 antenna elements placed at 5 m height. The beam configuration includes 16 beams with uniformly separated azimuth angles and a common down-tilted elevation angle fixed at 10 deg. The AoD and ToF measurements are obtained based on the corresponding characteristics of the radio propagation This article has been accepted for publication in IEEE Transactions on Vehicular Technology. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/TVT.2024.3456958 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, VOL. XX, NO. XX, MM YYYY 8 0 100 200 300 400 500 x [m] 0 50 100 150 200 250 300 350 y [m] LOS NLOS gNB Fig. 4. Illustration of the METIS Madrid map-based deployment and evaluation scenario with 40 simulated UE tracks while distinguishing the LoS and NLoS regions. Detailed example paths on one crossing are shown on the right. path with the highest received power, building on the signals and measurement procedures described in Section II. The obtained AoD and ToF measurements are exposed to substantial measurement errors, as discussed further in Section IV-B . The beam-wise frequency-domain CSI measurements are obtained from SSB transmissions as an average of received subcarrier powers per RB, with 120 kHz subcarrier spacing. Measurements with path-loss higher than 160 dB are not considered, while in general the environment shown in Fig. 4 possesses large areas and street segments with severe multi-bounce phenomena. The combined time-domain and frequency-domain dataset consists of 40 vehicle-like user tracks, where the UE collects measurements at 100 ms intervals. The UE locations are initialized with random locations along the streets, and the UE s move within the area by considering an equal probability to advance in any direction at intersections. The UE velocity varies between 20 km/h and 60 km/h depending on the present street and possible proximity of intersections while when approaching an intersection, the UE decelerates at 3 m/s 2 until reaching a fixed velocity of 20 km/h for smooth turning. After the turn, the UE accelerates at 2 m/s 2 until reaching a street-specific speed limit. The speed limit is generally defined as 40 km/h, apart from the top horizontal street which has the speed limit of 20 km/h (see Fig. 4) and the wider street below the pedestrian street having a limit of 60 km/h. The exact UE trajectories and associated measurement locations are different for each simulated user track. As this work is heavily focused on NLoS positioning, Fig. 4 visualizes the simulated tracks with the LoS / NLoS indication at each sampled location. We note that in order to efficiently track moving UE s with varying velocities through the sequence processing models, a sufficiently rich training dataset is needed with representative velocity statistics. The available 40 user tracks are distributed into 32 UE traces for training, 4 for validation, and the remaining 4 for the actual testing. The validation and testing paths are carefully selected, to avoid any area-specific bias in the evaluation. Furthermore, as the work focuses on the NLoS positioning performance, we validated the consistency of the LoS / NLoS split across the datasets. The distribution of the samples in the individual datasets based on the number of LoS gNB s is consistent with approx. 35% NLoS samples, 60% of samples having a single LoS gNB , and only 5% samples having 2gNB s in LoS . The distribution suggests that the traditional model-based solutions, such as trilateration, are not applicable in the considered scenario. In total, there are 25 181 samples in the dataset. B. Network Data Uncertainties In this work, we take into account the important practical aspect of uncertainties in the measurements and thereon in the corresponding features. To this end, the frequency-domain CSI is impaired in its FR-Complex representation with complex additive white Gaussian noise ( AWGN ) samples with magnitude equal to 30% of the corresponding channel estimate’s root mean square ( RMS ) magnitude. Such represents large practical measurement uncertainties. The other related features such as the FR-Power/Phase are impaired correspondingly, through the transformations from the impaired FR-Complex to amplitude/power and phase domains. To impair the time-domain features, we impose impairments separately to path-ToF , path-RP and path-AoA quantities. The path-ToF feature uncertainty is an AWGN with standard deviation ( std ) equal to 10 m. We consider the constant uncertainty scale regardless of the ToF magnitude, as the measurement errors are mostly resulting from hardware inaccuracies and timing offsets in the UE s and gNB s. Furthermore, we impair the path-RP feature with an AWGN with 2 dB std , which corresponds to the maximum impairment of ±6 dB range with 99.7% certainty, defined by 3GPP as the required absolute measurement accuracy for SS - RSRP [49]. The path-AoA is, in turn, impaired with discretized accuracy of 22.5◦ ( 360◦/16 beams), rather than with a randomized value, to incorporate the gNB limitations in accurately determining the AoA. Finally, as reviewed in the Introduction, a large majority of the state-of-the-art works, such as [15], [18], [21], [23]–[25], [28], [29], [35], utilize channel amplitude or power response, or even the integrated received power, as the positioning feature. Hence, in the following, the results with FR-Power feature represent essentially the state-of-the-art reference approach when it comes to the frequency-domain features. In the timedomain feature case, the use of the individual dominant path features has been considered in [5], [8], [30]–[32], [36], thus serving as the main reference schemes. Additionally, the stateof-the-art schemes build commonly on snap-shot NNs without harnessing the temporal correlation. For research reproducibility, data and codes are openly available at https://doi.org/10.5281/zenodo.12204893. C. Numerical Results with Dense Snap-Shot NNs We next provide and analyze the results obtained with dense NN based ML models while considering both the frequencydomain and time-domain features as well as the impacts of the feature density or granularity in the two considered domains. To establish an understanding on the baseline or reference performance, we start with the results under perfect measurements (no uncertainties), while then show also the performance under practical measurement uncertainties. 1) Results with Frequency-Domain Features: First, we analyze and compare the different frequency-domain CSI features introduced in Section III-A and their positioning capabilities with a densely connected snap-shot NN with 5 hidden layers. We thus split all the user tracks into individual samples and compare the performance without considering the temporal dependencies within sequences or additional uncertainties, to focus on the quality of the features themselves. This article has been accepted for publication in IEEE Transactions on Vehicular Technology. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/TVT.2024.3456958 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, VOL. XX, NO. XX, MM YYYY 9 0 10 20 30 40 Positioning error [m] FR-Complex FR-Power FR-Phase FR-Power+FR-Phase FR-Power/Phase (a) 0 5 10 15 Positioning error [m] path-ToF path-RP path-AoA path-ToF+AoA path-RP+AoA path-ToF+RP+AoA (b) Fig. 5. Distributions of positioning errors on the testing data when evaluating (a) different frequency-domain features, and (b) different time-domain features, with dense snap-shot NN and with no measurement uncertainties. Fig. 5a visualizes the distributions of the positioning errors on the testing dataset for each feature. Each boxplot marks the median (center) as well as the first and third quartiles ( 25th and 75th percentiles) encapsulated in the box, while the whiskers mark the values of 5th and 95th percentiles. The results show that the proposed FR-Power/Phase feature representation enables the most efficient training in terms of positioning error and that considering the FR-Complex features as the input provides the poorest performance. The FR-Power and FR-Phase features achieve comparable median performance, but in terms of outliers, FR-Power performs better. The 95th percentiles, referring essentially to the presence of outliers, of FR-Complex and FR-Phase are significantly higher than those of the remaining methods, as shown quantitatively in Table II. Furthermore, the feature combination denoted as FR-Power + FR-Phase represents the simple concatenation of the corresponding individual features. The numerical results show that the positioning performance is improved when compared to the individual features, but the novel FR-Power/Phase feature – utilizing the same, yet pre-processed inputs – provides superior performance. The table high-lights in bold the best performance numbers in different cases. We next further investigate the impact of the feature representation by considering the LoS and NLoS data separately, with the results being shown in Table II. We can observe that the proposed FR-Power/Phase feature representation achieves the lowest positioning errors by a considerable margin, when compared to the other methods ( 3.35 m and 4.70 m mean positioning error in LoS / NLoS , respectively) addressed earlier in the literature. By comparing the performance in LoS and NLoS scenarios, we can observe some increase in the error in NLoS , however, the exact impact is clearly feature-dependent. Furthermore, when considering the 95th percentiles of the error distributions, we can observe that the errors related to the FR-Complex and FR-Phase features are drastically increased, in both LoS and NLoS scenarios, while FR-Power + FR-Phase and FR-Power/Phase features sustain a relatively stable performance across the majority of the testing samples. Overall, the obtained results clearly show and demonstrate that utilizing the novel FR-Power/Phase feature offers the best performance by a large margin, clearly outperforming the earlier state-of-the-art in the field of frequency-domain features. Thus, in the further frequency-domain feature related evaluations, we consider only the FR-Power/Phase feature representation. 2) Results with Time-Domain Features: Next, we evaluate the positioning capabilities and performance when utilizing the different time-domain features ( path-ToF , path-RP , and path-AoA ) as the input data. We also evaluate the combination of the features, while the model can consider up to 5 dominant multipath components above the 160 dB path-loss threshold. Fig. 5b visualizes the achieved positioning results, showing that the proposed combinations of path-ToF and AoA or path-ToF , RP and AoA are the two best performing aggregate features. The results also suggest that the path-RP measurement provides less relevant information to the model than the path-ToF , which the model can directly interpret as normalized pseudo-range measurement. This can be seen by comparing the individual features (path-ToF vs. path-RP), as well as the cases where they are combined with path-AoA. The impacts of the features as well as the standalone performance in LoS / NLoS are summarized in Table III, while also highlighting the best-performing features in each scenario. The table shows that the combination of all features (pathToF+RP+AoA) together with path-ToF+AoA offer the best results across all statistics. The corresponding performance of path-RP+AoA lags already behind. When evaluating the individual features, the path-AoA provides high-accuracy positioning capabilities with less than 2 m median positioning error in NLoS , as it can effectively capture the propagation patterns within the given deployment. The path-RP and path-ToF provide, in turn, significantly poorer performance as individual features, especially when considering the higher percentile errors. These results thus clearly prove the value of the directional measurements. Additionally, when compared to the results presented in Table II, relative performance improvement can be observed, which we credit to stronger interpretability of time-domain measurements as model inputs compared to the frequency-domain CSI features. Notably, meter-scale positioning accuracy can be reached through the time-domain features also in NLoS. 3) Impact of Feature Granularity: Next, we assess and compare the performance of the snap-shot NN model while varying the granularity or sparsity of the input measurements. We again separate the testing dataset into LoS and NLoS parts, and first evaluate the frequency-domain data as resource block-level ( RB-level ) features, and their mean values across the RB s as the bandwidth-level ( BW-level ) features. The RB-level features are obtained perRB , which contains 12 subcarriers with a subcarrier This article has been accepted for publication in IEEE Transactions on Vehicular Technology. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/TVT.2024.3456958 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ IEEE TRANSACTIONS ON VEHICULAR TECHNOLOGY, VOL. XX, NO. XX, MM YYYY 16 [27] X. Wang, L. Gao, and S. Mao, “CSI Phase Fingerprinting for Indoor Localization With a Deep Learning Approach,” IEEE Internet Things J., vol. 3, no. 6, pp. 1113–1123, 2016. [28] P. Ferrand, A. Decurninge, and M. Guillaud, “DNN-based localization from channel estimates: Feature design and experimental results,” in Proc. IEEE GLOBECOM, 2020, pp. 1–6. [29] K. Gao, H. Wang, H. Lv, and W. Liu, “Toward 5G NR high-precision indoor positioning via channel frequency response: A new paradigm and dataset generation method,” IEEE J. Sel. Areas Commun., 2022. [30] D. Lynch, L. Ho, M. MacDonald, and M. O’Neill, “Localisation in wireless networks using deep bidirectional recurrent neural networks,” in Proc. Int. Joint Conf. Neural Networks (IJCNN), 2020, pp. 1–8. [31] Z. HajiAkhondi-Meybodi, M. Salimibeni, A. Mohammadi, and K. N. Plataniotis, “Bluetooth low energy and CNN-based angle of arrival localization in presence of Rayleigh fading,” in Proc. IEEE ICASSP, 2021, pp. 7913–7917. [32] Y. Xie, L. Zhou, Y. Zhang, H. Huan, and Z. Zhang, “Simultaneous localization of scatterers and target user based on indoor prior information in NLOS environments,” IEEE Trans. Veh. Technol., vol. 71, no. 11, pp. 11 729–11 740, 2022. [33] T. Feigl, E. Eberlein, S. Kram, and C. Mutschler, “Robust ToA-estimation using convolutional neural networks on randomized channel models,” in Proc. IPIN, 2021, pp. 1–8. [34] 3GPP, “Study on channel model for frequencies from 0.5 to 100 GHz,” 3GPP, Tech. Rep. 38.901, 3 2022, version 17.0.0. [35] M. M. Butt, A. Rao, and D. Yoon, “RF fingerprinting and deep learning assisted UE positioning in 5G,” in Proc. IEEE VTC-Spring, 2020, pp. 1–7. [36] Y. Chen, J. Palacios, N. Gonz ´ alez-Prelcic, T. Shimizu, and H. Lu, “Joint initial access and localization in millimeter wave vehicular networks: a hybrid model/data driven approach,” in Proc. IEEE SAM, 2022, pp. 355–359. [37] 3GPP, “Physical layer measurements,” 3GPP, Tech. Rep. 38.215, 1 2021, version 16.4.0. [38] 3GPP, “Stage 2 functional specification of User Equipment (UE) positioning in NG-RAN,” 3GPP, Tech. Rep. 38.305, 12 2021, version 16.7.0. [39] 3GPP, “NG Radio Access Network (NG-RAN); Stage 2 functional specification of User Equipment (UE) positioning in NG-RAN,” 3GPP, Tech. Rep. 38.305, 9 2022, version 16.8.0. [40] I. Goodfellow, Y. Bengio, and A. Courville, Deep learning. MIT press, 2016. [41] D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” arXiv preprint arXiv:1606.08415, 2016. [42] L. Lu, Y. Shin, Y. Su, and G. E. Karniadakis, “Dying relu and initialization: Theory and numerical examples,” arXiv preprint arXiv:1903.06733, 2019. [43] Y. Wang et al., “Transformer-based acoustic modeling for hybrid speech recognition,” in Proc. IEEE ICASSP, 2020, pp. 6874–6878. [44] J. Xiao, X. Fu, A. Liu, F. Wu, and Z.-J. Zha, “Image De-raining Transformer,” IEEE Trans. Pattern Anal. Mach. Intell., 2022. [45] S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997. [46] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980, 2014. [47] Remcom. Wireless InSite - 3D Wireless Prediction Software. Accessed: Jan 27, 2021). [Online]. Available: https://www.remcom.com/ wireless-insite-em-propagation-software [48] A. Rauch et al., “Fast algorithm for radio propagation modeling in realistic 3-D urban environment,” Advances in Radio Science, vol. 13, pp. 169–173, 11 2015. [49] 3GPP, “NR; Requirements for support of radio resource management,” 3GPP, Tech. Rep. 38.133, 9 2022, version 16.13.0. [50] Y. Assayag, H. Oliveira, E. Souto, R. Barreto, and R. Pazzi, “Indoor positioning system using synthetic training and data fusion,” IEEE Access, vol. 9, pp. 115 687–115 699, 2021. [51] A. Capponi et al., “A survey on mobile crowdsensing systems: Challenges, solutions, and opportunities,” IEEE Communications Surveys & Tutorials, vol. 21, no. 3, pp. 2419–2465, 2019. [52] “Synthetic data,” IEEE Standards Association, Mar 2023. [Online]. Available: https://standards.ieee.org/industry-connections/synthetic-data/ [53] A. Castellani, S. Schmitt, and S. Squartini, “Real-world anomaly detection by using digital twin systems and weakly supervised learning,” IEEE Trans. Ind. Informat., vol. 17, no. 7, pp. 4733–4742, 2020. BIOGRAPHIES Roman Klus is a Doctoral Researcher at Tampere University (TAU), Finland. He received his Ing. degree (Czech equivalent of master’s) in the field of Electronics and Communications from Brno University of Technology in 2019. His research focuses on modern machine learning approaches in 5G and beyond networks, especially utilizing neural network structures in mobility management and positioning. He is a Graduate Student Member of IEEE. Jukka Talvitie (S’09, M’17) received the M.Sc. and D.Sc. degrees from the Tampere University of Technology, Finland, in 2008 and 2016, respectively. He is currently a University Lecturer with the Unit of Electrical Engineering, Tampere University (TAU), Finland. His research interests include signal processing for wireless communications, radio-based positioning and sensing, radio link waveform design, and radio system design, particularly concerning 5G and beyond mobile technologies. Julia Equi received her Ph.D. degree in electrical engineering from Telecom Paris-Tech, France, in 2014. From 2015 to 2016, she was with the Communication Systems Division at Link ¨ oping University, Sweden. Currently, she is a senior researcher with Ericsson Research, Jorvas, Finland. G ´ abor Fodor (Senior Member, IEEE) received the Ph.D. degree in electrical engineering from the Budapest University of Technology and Economics in 1998, the Docent degree from the KTH Royal Institute of Technology, Sweden in 2017, and the D.Sc. degree from the Hungarian Academy of Sciences (Doctor of MTA) in 2019. He is currently a Master Researcher with Ericsson Research and an Adjunct Professor with the KTH Royal Institute of Technology. He is currently serving as an Editor for IEEE WIRELESS COMMUNICATIONS and as an Associate Editor in Chief for IEEE COMMUNICATIONS MAGAZINE. Johan Torsner is currently a Research Manager at Ericsson Research, leading Ericsson’s research activities in Finland. He has held several positions within research and development and has been involved in concept development and standardization of wireless systems from 3G to 6G. He holds an MSc in wireless communication from the Royal Institute of Technology, Stockholm. Mikko Valkama [S’00, M’01, SM’15, F’22] received his D.Sc. (Tech.) degree (with honors) from Tampere University of Technology, Finland, in 2001. In 2003, he was with the Communications Systems and Signal Processing Institute at SDSU, San Diego, CA, as a visiting research fellow. Currently, He is a Full Professor and the Head of the Unit of Electrical Engineering at the newly formed Tampere University, Finland. His general research interests include radio communications, radio localization, and radio-based sensing, with particular emphasis on 5G and 6G mobile radio networks. This article has been accepted for publication in IEEE Transactions on Vehicular Technology. This is the author's version which has not been fully edited and content may change prior to final publication. Citation information: DOI 10.1109/TVT.2024.3456958 This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/