scieee AI-readable full text Open interactive document viewer

Edge Intelligence for Human Activity Recognition Using Resource-Constrained Sensors

Alexia Moreno; Datar Sean; Yu-En Yen; Yu-Ting Huang

Abstract

Abstract: Human activity recognition (HAR) systems are evolving beyond traditional high-bandwidth sensors toward resource-efficient alternatives that can operate on edge devices while preserving user privacy. This survey examines HAR approaches using low-resource sensing modalities including inertial measurement units (IMUs), structural vibrations, acoustic emissions, capacitive sensing, and pressure sensors. We analyze signal processing pipelines for these modalities, feature extraction techniques optimized for limited computational resources, and lightweight machine learning models suitable for microcontroller deployment. The survey covers data augmentation strategies for small datasets, transfer learning approaches from simulated to real environments, and sensor fusion techniques that combine multiple low-cost sensors. Additionally, we discuss energy harvesting for self-powered operation, real-time constraints in edge computing, and the trade-offs between model complexity and recognition accuracy in resource-constrained settings.

Full text

Edge Intelligence for Human Activity Recognition Using Resource-Constrained Sensors Alexia Moreno [email protected] University of Michigan Yu-Ting Huang [email protected] University of Michigan Yu-En Yen [email protected] University of Michigan Datar Sean [email protected] University of Michigan Abstract Human activity recognition (HAR) systems are evolving beyond traditional high-bandwidth sensors toward resourceefficient alternatives that can operate on edge devices while preserving user privacy. This survey examines HAR approaches using low-resource sensing modalities including inertial measurement units (IMUs), structural vibrations, acoustic emissions, capacitive sensing, and pressure sensors. We analyze signal processing pipelines for these modalities, feature extraction techniques optimized for limited computational resources, and lightweight machine learning models suitable for microcontroller deployment. The survey covers data augmentation strategies for small datasets, transfer learning approaches from simulated to real environments, and sensor fusion techniques that combine multiple lowcost sensors. Additionally, we discuss energy harvesting for self-powered operation, real-time constraints in edge computing, and the trade-offs between model complexity and recognition accuracy in resource-constrained settings. 1 Introduction Modern sensing systems increasingly rely on continuous data collection from wearable and embedded devices to support applications such as health monitoring, activity tracking, and ambient intelligence. While cloud-based machine learning has traditionally powered these services, sending high-frequency sensor streams to remote servers introduces fundamental limitations: wireless transmission consumes significantly more energy than local computation, intermittent connectivity prevents reliable real-time inference, and streaming raw sensor data raises privacy and security concerns. As these systems proliferate into everyday environments, the assumption of abundant computation, memory, and power no longer holds. These constraints motivate a decisive shift toward edge computing, where intelligence moves closer to the sensor and computation occurs directly on low-power microcontrollers. Unlike smartphones or cloud servers, such devices often operate with less than a few hundred kilobytes of RAM, milliwatt-level power budgets, and no operating system support. Yet, they are expected to perform always-on sensing, real-time inference, and multimodal processing without sacrificing accuracy or user experience. This growing mismatch between algorithmic complexity and hardware capability defines the central challenge addressed by this work. This paper investigates how sensing modalities, signal processing strategies, and lightweight fusion architectures can be designed to make human activity recognition feasible on severely resource-constrained edge platforms. By systematically examining the interplay between sensing choices, computational overhead, and deployment costs, we aim to bridge the gap between what state-of-the-art HAR models require and what low-power embedded devices can sustain in practice. 2 Sensing Modalities Sensors for HAR constantly improve due to Moore’s Law, where the number of transistors can be doubled every two years within a same-sized integrated circuit. The improvements that have been noticed pertain to both low-level statistical analysis and large-scale semantic designations. There is important variance in how signal processing is accomplished based on the semantics that sensors support. Some signals are left unfiltered, providing insightful data that may be drastically inaccurate due to noise, while others are filtered to create clean analog or even digital signals. Analog signals are transmitted over cheaper hardware, and provide more continuous data that, while somewhat noise-susceptible, gives insight into activities that are highly detailed. Digital signals are noise-immune and can reliably be transmitted over larger distances, but have the drawback of disregarding data in transitional states that make it difficult to perform fine-grained, personalized data gathering for deep learning. Inertial Measurement Units (IMUs) are extremely useful in HAR. They are used primarily in automated manufacturing facilities, navigation for cars by estimating momentum, and robotics that specialize in service, where estimating inertia 1 is critical for the safety of people surrounding the machines. Given that IMUs are primarily meant to measure velocity, object orientation, and gravity, there is a wide range of human activities that provide IMUs with accurate data. This includes walking behavior to evaluate a person’s mood, and elevation changes by people at laborious work, like tower builders, through the gravity measurement, transmitting safety concerns to other electromechanical devices in use. They are designed similarly to gyroscopes, allowing for machine learning on the basis of nuances with how people orient their bodies under different circumstances. Thus, the sensing principle of inertial tracking directly correlates to recognizing complex motion patterns and body postures in real-time. [1, 39] Structural vibrations are a unique source for HAR in that they can observe the behavior of masses of people with the tradeoff of needing extra manual guidance, like a mapping of an area. They aren’t very useful when attached to an environment of an individual, because there is no consistent trend in waveforms other than the potential for a dangerous platform that shakes or makes a worker uncomfortable about stomping their feet. For groups of people, concerns over riots in stadiums can be observed successfully, but the battery to each node must be manually supplied on a daily basis, making these sensors cost inefficient and not low-resource, meaning much research is needed. Fortunately, these nodes today are very efficient at identifying shape deformations, giving insight into what kind of group HAR (stampeding, sports game jostling, riots) are detrimental and under a specified duration. Consequently, structural vibration sensing links induced mechanical waves to collective crowd behaviors, enabling non-line-of-sight activity monitoring. [ 13 , 20 ] Acoustic emissions are typically used to detect stress in materials, but due to their low-resource maintenance are actively used for HAR to determine low-level statistics. They can analyze ultrasonic waves that detail changes in a person’s physiology to determine health conditions that regular vibrations would disrupt. They rely on the piezoelectric effect, which is somewhat cheap, and they can be placed in large quantities as their signals can easily be processed together to form image creation. A limitation with these sensors is that they must be placed in contact with a region of interest. They cannot receive nor send signals through the air as this is not a reliable medium. Consequently, systems other than an individual’s body conditions cannot be monitored, making HAR deployment very difficult unless different devices can benefit from learning body conditions, which could be analyzing people running, jumping, or carrying a heavy weight at work. This allows for the derivation of internal physiological states from surfacelevel acoustic propagations, bridging the gap between material stress and human health monitoring. [18, 30] Capacitive sensing is a cost-efficient means of monitoring the electromechanical devices that interplay with workplace HAR. However, since they measure electrostatic field displacement, the internal materials must be as large as possible for accuracy, and are therefore not scalable. Fortunately, for the purpose of edge computing, size is less of a concern as a multiplicity of these sensors can be placed far away from one another. Also, they are temperature resilient, so they can last longer and perform under conditions that provide more data. This principle links electrostatic field perturbations directly to human proximity and touch interactions, facilitating device-free activity recognition. [32, 33] Pressure sensors are currently effective for HAR in health, and not working environments. The advantages are that they are very miniature due to MEMS technology, within a limited region can accurately detect blood pressure changes, and do not disrupt the areas of analysis. They can also be used with fiber optics, being placed on microscopic fiber to detect thermal expansion, which may eventually become useful data for HAR but right now is too removed from semantic consideration. This direct measurement of force application is fundamental for distinguishing weight-bearing activities and fine-grained contact interactions in healthcare settings. [41, 53] 3 Signal Processing and Feature Extraction Signal processing serves as the foundational stage in Human Activity Recognition (HAR) on edge devices, ensuring that raw sensor streams are transformed into compact, meaningful representations suitable for lightweight machine learning models. While cloud-assisted HAR platforms also employ preprocessing pipelines, they generally operate with far fewer limitations on compute, memory, and communication bandwidth—enabling them to support more complex models and higher data throughput [ 24 ]. In contrast, edge deployments running on microcontrollers must aggressively reduce noise, dimensionality, and computational overhead. As a result, efficient preprocessing and feature engineering remain indispensable for achieving real-time inference under stringent hardware constraints. 3.1 Preprocessing and Noise Reduction Wearable and structural sensors frequently capture motionindependent disturbances originating from environmental vibrations, quantization artifacts, sensor drift, and electromagnetic interference. To mitigate these effects, low-pass and band-pass filtering are commonly applied to IMU, acoustic, and vibration signals to isolate human motion frequencies typically within 0.3–15 Hz [ 28 ]. For instance, normal walking often exhibits dominant spectral components around 1–2 Hz, while running typically shifts these peaks toward 2 2.5–5 Hz, illustrating why restricting the passband to human locomotion ranges is effective. Wavelet-based denoising further enhances transient motion analysis by preserving local time–frequency characteristics while suppressing highfrequency noise [ 29 ]. Proper windowing strategies such as sliding windows with Hamming or Hann functions ensure temporal continuity while balancing latency and recognition accuracy [4]. 3.2 Feature Extraction for Edge Inference Feature extraction converts preprocessed signals into lowdimensional descriptors tailored for lightweight models. Classical approaches categorize features into the following domains: Time-domain features capture intensity and temporal variability using metrics such as mean, variance, root-meansquare (RMS), zero-crossing rate, and signal magnitude area [ 34 ]. These features are computationally inexpensive and thus ideal for microcontroller-based HAR pipelines. Frequency-domain features are obtained through the Fast Fourier Transform (FFT) to reveal dominant frequencies, spectral entropy, and power spectral density, which correlate strongly with activity periodicity such as gait cycles [ 49 ]. Frequency analysis is particularly beneficial for distinguishing activities with similar amplitude patterns but distinct rhythmic signatures. Time-frequency features derived from Short-Time Fourier Transform (STFT) and wavelet transforms provide joint temporal and spectral localization, enabling the analysis of nonstationary signals such as transitions between walking, sitting, and lifting [ 15 ]. These hybrid representations preserve motion dynamics while avoiding the memory demands of deep temporal encoders. Recent studies emphasize that careful feature design can outperform raw end-to-end pipelines on constrained devices by significantly reducing inference costs without sacrificing accuracy [ 42 ][ 37 ]. This highlights an important distinction in edge HAR: while deep learning dominates cloudscale recognition, handcrafted or hybrid features remain a competitive and often superior choice for microcontroller execution[9, 14]. 3.3 Sensor Fusion Sensor fusion integrates heterogeneous sensing modalities to improve robustness, reduce ambiguity, and compensate for limitations inherent in individual sensors[ 44 ]. Since lowcost edge systems often rely on minimally capable devices, such as IMUs, capacitive nodes, structural vibration sensors, and pressure arrays, fusion plays a critical role in extracting reliable activity semantics with minimal hardware overhead [10]. 3.3.1 Fusion Architectures. Sensor fusion occurs at three primary hierarchical layers: Data-level fusion combines raw or minimally processed sensor streams (e.g., accelerometer + gyroscope) before feature extraction. Although this preserves maximum information, synchronization complexity and increased sampling bandwidth limit its suitability for ultra-low-power edge devices [48]. Feature-level fusion merges modality-specific descriptors, such as vibration-based periodic features with IMU-derived pose signatures into a unified embedding [ 38 ]. This level of fusion offers an optimal balance between computation efficiency and recognition accuracy, making it the dominant strategy for microcontroller-based HAR. Decision-level fusion aggregates predictions from independent classifiers through weighted voting, Bayesian updates, or ensemble learning [ 25 ]. This approach excels when modalities operate asynchronously or exhibit different noise characteristics; however, it often introduces additional latency and memory overhead. 3.4 Lightweight Fusion for Edge Devices To adhere to resource constraints, recent studies adopt lightweight probabilistic fusion, complementary filters, and Kalmanfilter-based inference for pose reconstruction and gait tracking [ 7 ]. Attention-based multimodal fusion has emerged as a promising direction, selectively amplifying salient features from each modality while discarding redundant information [ 27 ]. These strategies reduce computational demands and enhance interpretability, two critical factors for embedded HAR deployments operating under strict energy budgets. Furthermore, sensor fusion encourages privacy-preserving intelligence by reducing dependence on cameras and highbandwidth transmissions [ 45 ]. By combining multiple lowpower sensors, HAR systems can achieve high recognition accuracy even in ambiguous scenarios without resorting to invasive perception modalities. Collectively, advances in signal processing, feature extraction, and multimodal fusion enable HAR systems to operate robustly on edge hardware, bridging the gap between wearable sensing constraints and real-world activity recognition requirements. To illustrate these design trends, representative HAR studies spanning different modality combinations, fusion strategies, and deployment characteristics are summarized in Table 1. 4 Lightweight Machine Learning and Data Optimization for Microcontrollers The deployment of Human Activity Recognition (HAR) on resource constrained edge devices marks a transition from centralized high performance computing to decentralized 3 Table 1: Representative HAR studies with modalities, fusion levels, and deployment characteristics. Study Modality Fusion Model P L Bao & Intille (2004) [4] Multi-IMU Featurelevel Decision tree No No Lane et al. (2015) [24] IMU + audio Featurelevel SVM/GMM Yes Yes Xu et al. (2019) [48] Wearable IMU Data/feature CNN/LSTM Yes Yes Mehrang et al. (2018) [29] IMU (acc+gyro) Featurelevel SVM No No energy efficient intelligence. However, this transition encounters two fundamental bottlenecks, data scarcity and limited computational power. Unlike fields such as Computer Vision or Natural Language Processing, where abundant large scale datasets enable the training of powerful general purpose models, HAR generally lacks datasets of comparable scale and standardization to those foundational benchmarks in vision and language. Common public datasets such as UCI HAR [ 3 ], WISDM [ 21 ], and PAMAP2 [ 35 ] contain only a small number of subjects and limited activity classes, which leads to insufficient variability for training generalizable models. Sensor readings also exhibit wide variation caused by device placement differences, sampling frequencies, environmental conditions, and individual characteristics of users. Consequently, deep learning models trained on small or personalized datasets often suffer from severe overfitting. In addition, microcontrollers used in wearable sensing systems impose strict computational constraints. Typical platforms offer only tens to hundreds of kilobytes of SRAM and approximately one megabyte of Flash memory, and often lack floating point hardware support. Combined with the requirement for low power operation to enable continuous sensing, these constraints necessitate the development of models and learning strategies that are inherently efficient. 4.1 Data Augmentation Strategies for Low Resource Wearable Sensing Data augmentation is an effective approach for mitigating overfitting in low-resource settings, yet applying imageinspired augmentation directly to sensor time series can distort temporal structure or alter the physical meaning of a signal. For instance, flipping an accelerometer signal may reverse its gravitational component and fundamentally change the interpretation of the underlying motion. Consequently, effective augmentation for Human Activity Recognition must preserve both temporal relationships and physical plausibility. Crucially, the selection of these strategies must also account for deployment constraints: complex generative methods are often restricted to offline training, whereas lightweight transformations are essential for efficient, on-device continual learning. Manual heuristics remain a cornerstone of augmentation due to their interpretability and efficiency. Jittering introduces small Gaussian perturbations to mimic sensor noise, while magnitude scaling simulates variations in user movement intensity. Because of their low computational cost, these operations are particularly well-suited for on-device adaptation to new users. Time warping [ 16 ], which adjusts the temporal progression of a sequence, addresses natural differences in motion speed between individuals. Furthermore, combining time warping with random data masking [ 17 ] encourages models to learn robust features by simulating signal dropouts common in wireless wearable networks . Moving beyond manual heuristics, automated frameworks like AutoAugHAR [ 52 ] have emerged to dynamically select optimal augmentation policies, though this data-driven search often incurs a high computational cost. More advanced techniques synthesize data entirely; Generative Adversarial Networks (GANs) [ 12 ] can learn sensor distributions to create realistic synthetic signals. However, due to the heavy computational resources required, GAN-based generation is typically limited to offline training phases. A more lightweight alternative applicable to resource-constrained settings is Mixup [ 46 ], along with time series specific variants, which combine different sequences to produce intermediate samples. Although these blended examples do not necessarily correspond to actual human motions, they smooth decision boundaries and improve robustness to noise and segmentation uncertainty without the latency of deep generative models. 4.2 Simulation Based Learning and Transfer Techniques for HAR Collecting labeled sensor data is labor intensive and often impractical, especially when rare or hazardous events must be captured. Simulation based methods have therefore emerged as a promising way to reduce reliance on real world data. Virtual IMUs embedded in physics based environments, such as Unity, enable the generation of large volumes of accurately labeled synthetic data by attaching simulated sensors to digital avatars. Such data can cover a wide range of body types, motion styles, and sensor placements, and can safely represent scenarios that are difficult to record in real situations, such as falls [23]. Nevertheless, synthetic IMU signals inevitably differ from real measurements. These discrepancies arise from sensor characteristics, including bias drift, axis misalignment, and soft tissue movement, as well as human variability in gait 4 patterns, limb proportions, and movement fluency. Environmental factors, such as inconsistent sensor placement and sampling rate mismatches, further widen this gap. To bridge these sim-to-real disparities, domain adaptation techniques are essential. Approaches such as Domain-Adversarial Neural Networks (DANN) [ 47 ] and Maximum Mean Discrepancy (MMD) minimization [ 6 ] explicitly align feature distributions to ensure the model remains invariant to domain shifts. The efficacy of this strategy is well-documented; for instance, Kwon et al. [ 22 ] demonstrated that pre-training on massive synthetic IMU datasets generated from videos (IMUTube) could improve F1-scores by approximately 12% on the RealWorld dataset and consistently boost performance on standard benchmarks like OPPORTUNITY compared to training on limited real data alone. Complementing this, crosssubject transfer learning provides a practical deployment strategy. A base model trained on a large public or synthetic dataset can be fine-tuned using only a small amount of userspecific data [ 40 ], effectively reducing the computational burden while enabling rapid adaptation to new users. 4.3 Efficient Neural Architectures for Microcontroller Deployment Even with sufficient data, deploying HAR models on microcontrollers requires architectures designed specifically for extreme memory and compute constraints. Many devices offer less than 256 KB of SRAM, which makes model size and peak memory usage essential considerations. This has led to a shift toward purpose built lightweight architectures rather than scaled down versions of larger models. Modern designs reduce multiply accumulate operations and minimize activation memory. TinyHAR [ 51 ] demonstrates that depth-wise separable convolutions combined with simplified attention mechanisms can deliver competitive accuracy while operating within the constraints of microcontrollers. Further improvements are achieved by architectures like TinierHAR [ 5 ], which utilize depthwise separable convolutions to minimize computational load and temporal attention modules to identify and discard irrelevant information from the signal sequence. While convolutional architectures provide excellent efficiency, recurrent networks often remain superior in capturing long term temporal dependencies. Lightweight recurrent variants, including residual LSTM designs [ 2 ], reduce parameter counts while maintaining expressive temporal modeling capabilities. Additional studies show that hybrid convolutional recurrent structures can outperform purely convolutional or recurrent models on HAR benchmarks under microcontroller constraints [31]. In addition to supervised learning, self supervised learning has emerged as a powerful strategy for wearable sensing. Methods such as contrastive predictive coding learn robust representations from large quantities of unlabeled data collected during regular device operation [ 11 ]. Because wearable devices naturally accumulate unlabeled motion data continuously, this enables on device models to improve without additional labeling effort. Collectively, these developments in data augmentation, simulation based training, transfer learning, and lightweight neural architectures illustrate that deep learning for HAR can be successfully deployed on microcontrollers despite constraints on data availability and hardware capacity. 5 System Constraints The deployment of HAR systems on edge devices presents significant challenges that go beyond algorithmic design. Resource-constrained platforms impose hard limits on energy availability, processing latency, and model complexity, forcing designers to carefully balance performance against these constraints. This section examines three critical systemlevel considerations: energy harvesting for self-powered operation, real-time processing requirements, and the fundamental trade-offs between model complexity and recognition accuracy. 5.1 Energy Harvesting for Self-Powered Operation HAR workloads that require continuous or frequent sensing and inference place significant demands on the available power budget. To mitigate this power demand, energy harvesting can be utilized. In particular, wearable sensors and IoT devices can be self-powered through biomechanical energy harvesting using triboelectric nanogenerators (TENGs) and piezoelectric nanogenerators (PENGs) [26]. Typical power demands of low-power HAR sensors operating in active mode range from 100 𝜇 W to 1 mW. A commercial accelerometer-based HAR system (HARAC) operating at 10 Hz consumes 60.35 𝜇 W on average for sensing alone, with a per-sample energy cost of 5.5 𝜇 J and sleep mode consumption of 6 𝜇 W [ 19 ]. The feature extraction and processing pipeline contributes further power consumption, with MCU consumption during active sampling ranging from 322-480 𝜇 W depending on sensor type. This consumption can be minimized through duty-cycled operation with deep-sleep modes between samples on ultra-low-power processors. Data transmission via Bluetooth Low Energy (BLE) requires 2.72 mW average, representing a significant portion of the power budget. For a typical deployment scenario with 10 Hz sampling over 10-second windows, HARAC requires 892.12 𝜇 J total energy per cycle (603.5 𝜇 J sensing + 288.62 𝜇 J transmission), primarily due to the need to transmit 3axis accelerometer data. In contrast, the HARKE approach, which directly uses kinetic energy harvesting (KEH) voltage 5 patterns for activity recognition, requires only 184.61 𝜇 J per cycle (88.4 𝜇J sensing + 96.21 𝜇J transmission), a significant energy savings achieved by eliminating accelerometer power consumption and reducing transmitted data volume. This dramatic reduction demonstrates that mW-level energy harvesting from TENGs and PENGs reported in [ 26 ] can sustain continuous HAR workloads. This energy efficiency comes with an accuracy trade-off. HARKE achieves 80-86% recognition accuracy depending on body placement, representing a 4-15% reduction compared to traditional accelerometer-based approaches. However, this gap narrows to only 4.17% when devices are placed closer to the central part of the body [19]. 5.2 Real-Time Processing Requirements Real-time HAR processing on the edge demands that inference keep pace with incoming sensor data to maintain system stability. Recent deployments on microcontrollers demonstrate these constraints. A DeepConv LSTM model on Arduino Nano BLE achieved 21 ms inference latency with a 5-second window [ 50 ], while CNN models on STM32L476 MCUs require 0.35 MHz for neural network processing at 1 classification per second [ 43 ]. When systems cannot meet these window-by-window deadlines, the recognition becomes unreliable. Beyond inference speed, edge HAR systems must manage hardware interrupts, background tasks, and sensor I/O concurrently. This makes deterministic timing essential — algorithms cannot introduce variable latency without compromising real-time operation. Studies on RISC-V based MCUs show inference latency ranges from 9 𝜇 s to 16 ms depending on model complexity [8]. Memory constraints pose another critical challenge for real-time processing. Quantized HAR models must fit within limited flash and RAM. Standard deployments use 136.51 KB flash for model weights and 29.1 KB RAM for activations [ 50 ], although further quantization can reduce memory footprint to as little as 0.05-23.17 KB [ 8 ]. Common MCUs like ESP32 offer 320 KB SRAM and 4 MB flash [ 36 ], whereas STM32 variants provide 520 KB total memory [ 8 ]. As described in the next section, balancing model complexity against these hardware limits while maintaining accuracy represents a fundamental design trade-off in edge HAR deployments. 5.3 Model Complexity and Accuracy Trade-offs More complex models usually achieve better recognition accuracy, but deploying them on microcontrollers with limited RAM, flash, and power budgets is challenging. On one end, reducing models too aggressively risks losing the capacity to distinguish fine-grained activities. On the other end, full-sized models are impractical because they exceed memory constraints and consume too much power for resourceconstrained sensors [8, 50]. To manage this trade-off, HAR systems combine quantization, pruning, architectural search, and adaptive inference. Daghero et al. [ 8 ] used mixed-precision quantization (2-, 4-, and 8-bit) with hyperparameter optimization to generate a set of efficient CNNs for HAR. Their models span more than an order of magnitude in memory (0.05-23.17 KB), latency (9 𝜇 s-16 ms), and energy consumption (0.05-61.59 𝜇 J). Additionally, their adaptive inference technique enables a single deployed model to operate in over 20 different modes, dynamically adjusting computational complexity based on input difficulty. Zhou et al. [ 50 ] compressed a DeepConv LSTM from 513.23 KB to 136.51 KB using full integer quantization for Arduino Nano deployment, achieving 97.09% accuracy with 189.6 KB flash and 29.1 KB RAM usage—a 73% size reduction with only 1.15% accuracy loss. These cyclical trade-offs require deeper investigation to enable continued progress in HAR on resource-constrained sensing platforms. 6 Conclusion Human activity recognition on resource-constrained platforms is transitioning from traditional, data-intensive systems toward lightweight, scalable solutions that prioritize privacy, efficiency, and real-time performance. By leveraging low-cost sensing modalities, optimized signal processing pipelines, and compact machine learning architectures, HAR can now operate directly on edge devices without relying on continuous cloud connectivity. While these advances reduce computation and energy demands, challenges remain in handling heterogeneous sensor characteristics, ensuring robustness under diverse environmental conditions, and balancing accuracy with model complexity. Continued progress in sensor fusion, adaptive learning, and energy-aware system design will be critical for deploying reliable, autonomous, and privacy-preserving HAR systems in everyday environments. Ultimately, the convergence of efficient sensing, embedded intelligence, and sustainable operation positions edge-based HAR as a foundational component of next-generation ubiquitous computing and human-centered applications. References [1] Norhafizan Ahmad, Raja Ariffin Raja Ghazilla, Nazirah Khairi, and Vijayabaskar Kasi. 2013. Reviews on Various Inertial Measurement Unit (IMU) Sensor Applications. International Journal of Signal Processing Systems 1 (01 2013), 256–262. doi:10.12720/ijsps.1.2.256-262 [2] Sarab AlMuhaideb, Lama AlAbdulkarim, Deemah Mohammed AlShahrani, Hessah AlDhubaib, and Dalal Emad AlSadoun. 2024. Achieving More with Less: A Lightweight Deep Learning Solution for Advanced Human Activity Recognition (HAR). Sensors 24, 16 (8 2024), 5436. doi:10.3390/s24165436 6 [3] Davide Anguita, Alessandro Ghio, Luca Oneto, Xavier Parra, and Jorge L. Reyes-Ortiz. 2013. A Public Domain Dataset for Human Activity Recognition using Smartphones. In Proceedings of the 21st European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN). 24–26. [4] L. Bao and S. Intille. 2004. Activity recognition from accelerometer data. In Pervasive Computing. [5] Sizhen Bian, Mengxi Liu, Vitor Fortes Rey, Daniel Geißler, and Paul Lukowicz. 2025. TinierHAR: Towards Ultra-Lightweight Deep Learning Models for Efficient Human Activity Recognition on Edge Devices. arXiv preprint arXiv:2507.07949 (10 2025), 163–169. doi:10.1145/3715071. 3750410 [6] Kaixuan Chen, Dalin Zhang, Lina Yao, Bin Guo, Zhiwen Yu, and Yunhao Liu. 2021. Deep Learning for Sensor-based Human Activity Recognition: Overview, Challenges and Opportunities. ACM Computing Surveys (CSUR) 54, 4 (2021), 1–40. doi:10.1145/3447744 [7] W. Czajewski. 2023. Lightweight fusion filters for edge HAR. Sensors (2023). [8] Francesco Daghero, Alessio Burrello, Chen Xie, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, and Daniele Jahier Pagliari. 2022. Human Activity Recognition on Microcontrollers with Quantized and Adaptive Deep Neural Networks. ACM Trans. Embed. Comput. Syst. 21, 4, Article 46 (Aug. 2022), 28 pages. doi:10.1145/3542819 [9] R. Garg et al . 2022. TinyHAR: An efficient and lightweight neural network for human activity recognition on microcontrollers. In IEEE International Conference on Pervasive Computing and Communications (PerCom). [10] S. Ha. 2020. Multimodal sensor fusion for HAR. ACM IMWUT (2020). [11] Harish Haresamudram, Irfan Essa, and Thomas Plötz. 2023. Investigating Enhancements to Contrastive Predictive Coding for Human Activity Recognition. In 2023 IEEE International Conference on Pervasive Computing and Communications (PerCom). 232–241. doi:10.1109/ percom56429.2023.10099197 [12] Debapriya Hazra and Yung-Cheol Byun. 2020. SynSigGAN: Generative Adversarial Networks for Synthetic Biomedical Signal Generation. Biology 9, 12 (12 2020), 441. doi:10.3390/biology9120441 [13] Jag Humar, Ashutosh Bagchi, and Hongpo Xu. 2006. Performance of vibration-based techniques for the identification of structural damage. Structural Health Monitoring 5, 3 (8 2006), 215–241. doi:10.1177/ 1475921706067738 [14] A. Ignatov. 2017. Real-time human activity recognition using deep learning on the edge. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. 1555–1563. [15] A. Ignatov. 2018. Real-time HAR using wavelet features. Pattern Recognition (2018). [16] Brian Kenji Iwana and Seiichi Uchida. 2021. An empirical survey of data augmentation for time series classification with neural networks. PLoS ONE 16, 7 (7 2021), e0254841. doi:10.1371/journal.pone.0254841 [17] Chi Yoon Jeong, Hyung Cheol Shin, and Mooseop Kim. 2021. Sensordata augmentation for human activity recognition with time-warping and data masking. Multimedia Tools and Applications 80, 14 (3 2021), 20991–21009. doi:10.1007/s11042-021-10600-0 [18] Richard Kapur. 2016. Acoustic Emission in Orthopaedics: A State of the Art Review. Journal of Biomechanics 49 (10 2016). doi:10.1016/j. jbiomech.2016.10.038 [19] Sara Khalifa, Guohao Lan, Mahbub Hassan, Aruna Seneviratne, and Sajal K. Das. 2018. HARKE: Human Activity Recognition from Kinetic Energy Harvesting Data in Wearable Devices. IEEE Transactions on Mobile Computing 17, 6 (2018), 1353–1368. doi:10.1109/TMC.2017. 2761744 [20] Marcel Koch, Thomas Pfitzinger, Fabian Schlenke, Fabian Kohlmorgen, Roland Groll, and Hendrik Wöhrle. 2024. Recognition of Human Activities Based on Ambient Audio and Vibration Data. IEEE Access 12 (2024), 174399–174412. doi:10.1109/ACCESS.2024.3457912 [21] Jennifer R. Kwapisz, Gary M. Weiss, and Samuel A. Moore. 2011. Activity recognition using cell phone accelerometers. ACM SIGKDD Explorations Newsletter 12, 2 (3 2011), 74–82. doi:10.1145/1964897.1964918 [22] Hyeokhyen Kwon, Gregory D Abowd, and Thomas Ploetz. 2020. IMUTube: Automatic Extraction of Virtual on-body Accelerometry from Video for Human Activity Recognition. In Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (IMWUT), Vol. 4. ACM, 1–29. doi:10.1145/3411839 [23] Hyeokhyen Kwon, Gregory D. Abowd, and Thomas Plötz. 2021. Complex Deep Neural Networks from Large Scale Virtual IMU Data for Effective Human Activity Recognition Using Wearables. Sensors 21, 24 (12 2021), 8337. doi:10.3390/s21248337 [24] N. Lane et al . 2015. Can smartphones transform sensing? IEEE Pervasive Computing (2015). [25] O. Lara. 2013. A survey on mobile activity recognition. Comput. Surveys (2013). [26] Long Liu, Xinge Guo, Weixin Liu, and Chengkuo Lee. 2021. Recent Progress in the Energy Harvesting Technology—From Self-Powered Sensors to Self-Sustained IoT, and New Applications. Nanomaterials 11, 11 (2021). doi:10.3390/nano11112975 [27] J. Ma. 2022. Attention-based multimodal fusion for wearable HAR. In ACM IMWUT. [28] J. Mäntyjärvi. 2001. Recognizing human motion with multiple acceleration sensors. In IEEE Pervasive Computing. [29] S. Mehrang. 2018. Multilevel wavelet-based denoising for wearable HAR. Sensors (2018). [30] Z. Nazarchuk, Valentyn Skalskyi, and Oleh Serhiyenko. 2017. Acoustic Emission. doi:10.1007/978-3-319-49350-3 [31] Francisco Javier Ordóñez and Daniel Roggen. 2016. Deep Convolutional and LSTM Recurrent Neural Networks for Multimodal Wearable Activity Recognition. Sensors 16, 1 (2016), 115. doi:10.3390/s16010115 [32] Robert Puers. 1993. Capacitive sensors: When and how to use them. Sensors and Actuators A Physical 37-38 (6 1993), 93–105. doi:10.1016/ 0924-4247(93)80019-d [33] Jing Qin, Li-Juan Yin, Ya-Nan Hao, Shao-Long Zhong, Dong-Li Zhang, Ke Bi, Yong-Xin Zhang, Yu Zhao, and Zhi-Min Dang. 2021. Flexible and Stretchable Capacitive Sensors with Different Microstructures. Advanced Materials 33, 34 (7 2021), e2008267. doi:10.1002/adma.202008267 [34] A. Reiss and D. Stricker. 2012. Deep feature analysis for HAR. Sensors (2012). [35] Attila Reiss and Didier Stricker. 2012. Introducing a New Benchmarked Dataset for Activity Monitoring. In 2012 16th International Symposium on Wearable Computers. 108–109. doi:10.1109/ISWC.2012.13 [36] Bidyut Saha, Riya Samanta, Soumya Ghosh, and Ram Babu Roy. 2024. Towards Sustainable Personalized On-Device Human Activity Recognition with TinyML and Cloud-Enabled Auto Deployment. doi:10.48550/arXiv.2409.00093 [37] P. Sarma et al . 2020. Feature engineering for human activity recognition with inertial sensors. IEEE Sensors Journal 20, 14 (2020), 8060– 8072. [38] S. Singh. 2019. Feature-level fusion for HAR. In IEEE PerCom. [39] Isaac Skog and Peter Händel. 2006. Calibration of A MEMS inertial measurement unit. XVII IMEKO World Congress Metrology for a Sustainable Development (01 2006). [40] Elnaz Soleimani and Ehsan Nazerfard. 2020. Cross-subject transfer learning in human activity recognition systems using generative adversarial networks. Neurocomputing 426 (11 2020), 26–34. doi:10.1016/j.neucom.2020.10.056 7 [41] P. Song, Z. Ma, J. Ma, L. Yang, J. Wei, Y. Zhao, and X. Wang. 2020. Recent Progress of Miniature MEMS Pressure Sensors. Micromachines 11, 1 (2020), 56. doi:10.3390/mi11010056 [42] A. Stisen. 2015. Smart device HAR benchmarking. Pervasive and Mobile Computing (2015). [43] STMicroelectronics. 2024. Human Activity Recognition Case Study. https://www.st.com/content/st_com/en/st-edge-ai-suite/casestudies/human-activity-recognition.html. Accessed: December 2025. [44] B. Wang et al . 2019. A survey on multimodal sensor fusion for human activity recognition. Information Fusion 51 (2019), 25–43. [45] Y. Wang. 2022. Privacy-preserving multimodal sensing. IEEE IoT Journal (2022). [46] Kristoffer K Wickstrøm, Michael Kampffmeyer, and Robert Jenssen. 2022. Mixing up contrastive learning: Self-supervised representation learning for time series. Pattern Recognition Letters 155 (2022), 54–61. [47] Garrett Wilson, Janardhan Rao Doppa, and Diane J. Cook. 2020. MultiSource Deep Domain Adaptation with Weak Supervision for TimeSeries Sensor Data. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Virtual Event, CA, USA) (KDD ’20). Association for Computing Machinery, New York, NY, USA, 1768–1778. doi:10.1145/3394486.3403228 [48] C. Xu. 2019. Data-level fusion for wearable HAR. Information Fusion (2019). [49] Z. Yan. 2015. Spectral features for activity recognition. IEEE IoT Journal (2015). [50] Haotian Zhou, Xiujun Zhang, Yu Feng, Tongda Zhang, and Lijuan Xiong. 2025. Efficient human activity recognition on edge devices using DeepConv LSTM architectures. Scientific Reports 15, 1 (2025), 10633. doi:10.1038/s41598-025-98571-2 [51] Yexu Zhou, Haibin Zhao, Yiran Huang, Till Riedel, Michael Hefenbrock, and Michael Beigl. 2022. TinyHAR: A Lightweight Deep Learning Model Designed for Human Activity Recognition. In Proceedings of the 2022 ACM International Joint Conference on Pervasive and Ubiquitous Computing (ISWC ’22). 89–93. doi:10.1145/3544794.3558467 [52] Yexu Zhou, Haibin Zhao, Yiran Huang, Tobias Röddiger, Murat Kurnaz, Till Riedel, and Michael Beigl. 2024. AutoAuGHAR: Automated Data Augmentation for Sensor-based Human Activity Recognition. Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies (IMWUT) 8, 2 (5 2024), 1–27. doi:10.1145/3659589 [53] Yizheng Zhu and Anbo Wang. 2005. Miniature fiber-optic pressure sensor. IEEE Photonics Technology Letters 17, 2 (2005), 447–449. doi:10. 1109/LPT.2004.839002 8