scieee AI-readable full text Open interactive document viewer

Exploring multimedia principles and learning from errors in augmented reality and video

Candido, Vito; Cattaneo, Alberto A. P.; Petko, Dominik

Abstract

Background This study compares the effectiveness of immersive augmented reality and video instruction in procedural learning by examining two pedagogical approaches: standard demonstration-based training (DBT) and DBT enhanced with common errors. Based on cognitive theory of multimedia learning and cognitive load theory, we expected AR to reduce extraneous cognitive load (ECL) and improve learning outcomes through enhanced signaling and spatial and temporal contiguity compared to video. Aims To compare the effectiveness of immersive AR and video instruction in procedural learning, examining DBT and DBT with common errors. Sample 114 Swiss vocational students (69 male) learned a new T-shirt folding technique. Methods A 2 × 2 design was employed with two factors: technology (AR vs. video) and pedagogical approach (DBT vs. DBT with common errors). Participants completed retention and transfer tasks, with measurements including folding accuracy, success rates, completion times, knowledge test scores, and self-reported cognitive load. Results Contrary to expectations, the tested structural equation model demonstrated that AR led to higher ECL, resulting in worse performance than video this predicting success in both retention and transfer tasks. The DBT with common errors successfully improved folding accuracy, particularly when using AR. Conclusions These results suggest the limited impact of signaling and spatial and temporal contiguity when learning materials are optimized by applying coherence and segmenting principles. The tested pedagogical approach shows promise, particularly in AR applications, although further research is needed to draw definitive conclusions.

Full text

Exploring multimedia principles and learning from errors in augmented reality and video Vito Candido a,b,* , Alberto Cattaneo a , Dominik Petko b a Swiss Federal University for Vocational Education and Training, Via Besso 84, 6900, Lugano, Switzerland b University of Zurich, Kantonsschulstrasse 3, 8001, Zürich, Switzerland ARTICLE INFO Keywords: Augmented reality Cognitive theory of multimedia learning Cognitive load theory Demonstration based training ABSTRACT Background: This study compares the effectiveness of immersive augmented reality and video instruction in procedural learning by examining two pedagogical approaches: standard demonstration-based training (DBT) and DBT enhanced with common errors. Based on cognitive theory of multimedia learning and cognitive load theory, we expected AR to reduce extraneous cognitive load (ECL) and improve learning outcomes through enhanced signaling and spatial and temporal contiguity compared to video. Aims: To compare the effectiveness of immersive AR and video instruction in procedural learning, examining DBT and DBT with common errors. Sample: 114 Swiss vocational students (69 male) learned a new T-shirt folding technique. Methods: A 2 ×2 design was employed with two factors: technology (AR vs. video) and pedagogical approach (DBT vs. DBT with common errors). Participants completed retention and transfer tasks, with measurements including folding accuracy, success rates, completion times, knowledge test scores, and self-reported cognitive load. Results: Contrary to expectations, the tested structural equation model demonstrated that AR led to higher ECL, resulting in worse performance than video this predicting success in both retention and transfer tasks. The DBT with common errors successfully improved folding accuracy, particularly when using AR. Conclusions: These results suggest the limited impact of signaling and spatial and temporal contiguity when learning materials are optimized by applying coherence and segmenting principles. The tested pedagogical approach shows promise, particularly in AR applications, although further research is needed to draw definitive conclusions. 1. Introduction Over the past decade, the use of augmented reality (AR) in education has seen a significant increase (Akçayır & Akçayır, 2017; L´ opez-Belmonte et al., 2020). AR is defined as the simultaneous, geometrically aligned, and real-time combination of real and virtual objects (Azuma, 1997). Technical solutions for experiencing AR range from using smartphones to dedicated head-mounted displays (HMDs) AR glasses, such as the Microsoft HoloLens 2. Although AR is recognized as a valuable tool in professional settings for supporting procedural execution (Yin et al., 2023), its effectiveness as a learning aid, particularly for procedural learning, has not been firmly established when compared to current best practices such as video (Chang et al., 2022). Currently, video is considered one of the most effective media for teaching procedures, due to its ability to demonstrate processes through animations (Berney & B´ etrancourt, 2016; H¨ offler & Leutner, 2007). The cognitive theory of multimedia learning (CTML; Mayer, 2020) and demonstration-based training (DBT; Ashford et al., 2006; Rosen et al., 2010) provide guidance on how this medium and multimedia materials in general can be optimized to support learning. CTML offers evidence-based principles for designing instructional materials effectively, while DBT specifies the essential steps and their optimal sequence for teaching procedures. A specific DBT application demonstrated improved learning outcomes when demonstrations included common errors (van der Meij & Flacke, 2020). In contrast, we have less evidence regarding the effectiveness of AR in procedural learning. AR can display animations directly within the real-world environment, a potentially powerful feature for procedural * Corresponding author. E-mail addresses: [email protected] (V. Candido), [email protected] (A. Cattaneo), [email protected] (D. Petko). Contents lists available at ScienceDirect Learning and Instruction journal homepage: www.elsevier.com/locate/learninstruc https://doi.org/10.1016/j.learninstruc.2025.102244 Received 7 November 2024; Received in revised form 23 September 2025; Accepted 29 September 2025 Learning and Instruction 101 (2025) 102244 Available online 6 October 2025 0959-4752/© 2025 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY license ( http://creativecommons.org/licenses/by/4.0/ ). tasks. HMDs allow learners to interact hands-free with the real environment, providing high levels of instructional support through animated instructions. However, it remains unclear whether CTML principles apply equally well to AR (Mutlu-Bayraktar et al., 2019; Çeken & Tas¸kın, 2022). Additionally, no existing research has explicitly applied DBT guidelines—including common errors or not—to procedural learning in AR. Finally, AR’s high level of assistance might enable learners to correctly execute procedures without necessarily achieving deep learning about the procedure to perform. Therefore, AR might particularly benefit from DBT highlighting common errors, an approach known to foster deeper understanding during learning. Considering these gaps, our study investigates the application of CTML principles considering the technology used (AR vs. video) and additionally adding the instructional approach adopted (DBT vs. DBT with the presentation of common errors). We selected a specific T-shirt folding technique as the procedure to learn. The accuracy of the final result may strongly depend on signaling—a principle from CTML stating that learning outcomes improve when relevant instructional information is highlighted. Visualizing directly on the concrete T-shirt where it should be grasped for correct folding might benefit participants using AR. This visualization could demonstrate that AR provides additional affordances beyond the animations offered by video, thereby enhancing instructional support. On the other hand, although the procedure can be completed in a few seconds, it presents numerous potential errors that can compromise the accuracy of the final fold, offering opportunities to illustrate common mistakes. This paper helps expand our understanding of the effectiveness of CTML principles when using immersive technologies; moreover, it examines an instructional approach that could provide an effective solution for learning procedures through AR. 1.1. Cognitive load and augmented reality As the effectiveness of CTML principles is deeply linked to how information is processed during learning—a topic extensively explored by cognitive load theory (CLT)—before addressing them it is important to introduce the CLT. “CLT aims to explain how the information processing load induced by learning tasks can affect students’ ability to process new information and to construct knowledge in long-term memory” (Sweller et al., 2019, pp. 261–262). CLT identifies three types of CL: intrinsic cognitive load (ICL), extraneous cognitive load (ECL), and germane cognitive load (GCL). ICL reflects the task’s inherent complexity and comprises two components: an objective one, determined by the task’s nature and the interactivity of the different elements, and a subjective one, which varies according to the individual’s cognitive abilities and prior knowledge (Kalyuga, 2005; Sweller et al., 1998, 2011, 2019). ECL stems from the design and organization of educational materials: well-structured materials can reduce it by focusing attention on the task, while poorly designed materials can increase it by generating distraction or confusion (Chen et al., 2023). Furthermore, ICL and ECL are not independent from each other. For learning tasks involving low levels of ICL, a design that typically elicits high levels of ECL has no impact. Such an impact can instead be observed both on levels of perceived ECL and on learning outcomes when students have to perform highly interactive tasks. Finally, according to Sweller (2010), GCL corresponds to the working memory resources actually devoted to learning the essential elements of a task (that is, those that generate ICL). GCL does not constitute a third, independent load; rather, it depends on the levels of ICL and on how much ECL is imposed by the material or the task. When ECL is low and ICL does not exceed working memory capacity, the “free” resources can be used to deeply process the intrinsic elements, thereby increasing GCL, which can be defined as the processing devoted to schema construction and deeper understanding. Throughout the text, we will generically refer to the term CL whenever the scales used in the analyzed studies do not allow for distinguishing between its specific components; otherwise, we will explicitly indicate which component of CL has been examined. The effects of AR on CL levels have been partially examined in the literature. Akçayır and Akçayır’s (2017) review reported mixed results, with some studies showing increased CL levels and others showing decreased levels. Bacca-Acosta et al. (2021) found that AR can mainly decrease CL. Buchner et al. (2021a) highlighted that in this respect, the application design quality is the most influential factor: well-designed AR applications reduce CL, while poorly designed ones increase it. However, AR research presents a significant gap: most studies have focused on smartphone applications, neglecting HMDs solutions. HMDs use generates a greater sense of immersion and presence, aspects that warrant particular attention. Immersion and sense of presence are terms usually connected to the study of virtual reality. Virtual reality (VR) is defined as a computergenerated, three-dimensional environment that users can explore and interact with through specialized devices (American Psychological Association, 2018). Although the term VR is sometimes used more broadly to refer to desktop applications, it is almost always associated with the use of HMDs. While VR can employ technical solutions similar to AR, what changes is the proportion of computer-generated elements displayed: with AR, computer-generated elements are superimposed onto the real world, whereas with VR the entire environment surrounding the user is computer-generated. This complete ‘immersion’ can increase motivation but also cognitive load, potentially compromising learning outcomes (Han et al., 2021). This phenomenon manifests in both declarative and procedural tasks, where immersive solutions generally produce inferior results (Barreda-´ Angeles et al., 2021; Frederiksen et al., 2020). Given the technological similarity and relative novelty of HMD devices for untrained users, we hypothesize that HMDs-AR might produce analogous effects. Although no specific meta-analyses exist on the effect of HMDs-AR on CL, Buchner et al. (2021b) reviewed several relevant studies. Three of them reported significantly higher CL levels compared to non-HMDs technologies (Gross et al., 2018; Hochreiter et al., 2018; Kawai et al., 2010), while two observed a reduction (Tsai & Huang, 2018; Young et al., 2016). Only Tsai and Huang (2018) examined learning outcomes, while others focused solely on usability variables. Recently, [Authors] (2025) compared HMDs-AR, smartphone AR, and video for learning problem-solving strategies, applying CTML principles. The results provided moderate support for the null hypothesis concerning ECL levels across the three conditions, emphasizing the primary importance of design over the technology used. However, the video condition outperformed the AR versions in the retention task. These findings partially challenge the idea that technology itself does not influence CL levels ([Authors], 2025; Buchner et al., 2021b). As highlighted by Hochreiter et al. (2018), different technical solutions can determine different CL levels, independently from the learning materials. However, which component of CL is specifically affected by the technical solution remains unclear, particularly regarding ECL, as scales distinguishing between different components are rarely employed. 1.2. Applying cognitive theory of multimedia learning principles in AR to manage cognitive load The CTML (Mayer, 2020) outlines design principles to effectively manage CL in multimedia materials using an evidence-based approach. The principles are designed to minimize extraneous processing, optimize essential processing, and facilitate generative processing. Although theorized separately, these concepts can be traced back to ECL, ICL, and GCL. These principles include spatial contiguity, which states that learning improves when corresponding words and pictures are presented near each other, reducing ECL by minimizing the need to search for related information (Schroeder & Cenkci, 2018, g =0.65). Similarly, the temporal contiguity principle asserts that simultaneous presentation of words and pictures enhances learning by preventing cognitive overload from holding information in working memory, especially with V. Candido et al. Learning and Instruction 101 (2025) 102244 2 complex materials (Ginns, 2006, d =0.78). The signaling principle posits that adding cues to guide attention to relevant elements enhances learning, particularly in complex real-world contexts where complexity cannot be reduced (van Gog, 2021; Richter et al., 2016, d =0.35; Schneider et al., 2018, g =0.33). AR can enhance the implementation of these principles by overlaying information directly onto real-world environments, further reducing ECL. For spatial contiguity, AR can integrate text and images directly onto real objects, minimizing the need for mental mapping. In terms of temporal contiguity, AR can present corresponding words and images simultaneously in the user’s field of view, facilitating immediate action. For signaling, AR can provide precise visual cues in the real world, guiding attention to relevant elements and improving learning outcomes. In other words, we expect AR to provide a more effective implementation of spatial and temporal contiguity. Displaying animations directly overlaid on the object to be manipulated in the real world can also optimize signaling, reducing the steps required between understanding the gesture and performing it, thereby lowering ECL levels compared with video. With video, the user receives the instructions on the screen but then has to act on the T-shirt, thus worsening spatial contiguity. Similarly, this also affects temporal contiguity, because the user cannot act immediately but must first look at the T-shirt. In both cases, the use of AR applies the principles more efficiently, because it enables learning and action to occur in the same place, and the time that elapses between receiving the instructions and acting depends only on the participant, not on the application. Because these principles aim to reduce extraneous processing, a decrease in ECL levels is plausible. Apart from the application of single multimedia learning principles, AR offers technical affordances that can potentially reduce both ICL and ECL when compared to learning with traditional media or in real-world contexts (Cheng & Tsai, 2012). On one hand, AR can reduce ICL by making abstract or invisible phenomena visible and observable. A clear example is the visualization of magnetic fields: AR eliminates the need for the learners to invest cognitive resources on imagining the phenomenon, thereby allowing them to focus on understanding the underlying physical laws (Vieyra & Vieyra, 2022). In this manner, the informational enrichment optimizes cognitive load. Furthermore, AR provides simulation opportunities through which the learner can actively deepen their understanding by interacting directly with the real world, potentially increasing GCL. On the other hand, AR provides tools for managing ECL that surpass the application of signaling or spatial and temporal contiguity principles. For instance, it can be employed to reduce the complexity of the real environment by acting as a perceptual filter. This application, known as "Diminished Reality" (Cheng et al., 2022), makes it possible to inhibit stimuli and information that are irrelevant or too complex for a novice learner, thereby preventing the cognitive overload that would hinder learning (Fiorella & Mayer, 2021). This particular form of AR can be considered as a form of reverse signaling or to a direct, real-world application of the coherence principle, which would be impossible to implement without the use of this technology. Although some of these affordances are shared with other technologies, such as computer simulators, the unique added value of HMD-AR lies in its capacity to integrate these optimized learning scenarios directly into the real-world context (Cheng & Tsai, 2012). Research has shown that applying CTML principles in AR contexts can lead to improved performance, greater efficiency, and reduced CL compared to learning as usual (Küçük et al., 2016; Lai et al., 2019; Thees et al., 2020; Wang et al., 2018; Wu et al., 2017). However, most studies have compared AR to learning as usual rather than to acknowledged media that can support that kind of learning, such as video, which has already been proven effective for procedural learning. Therefore, further investigation is needed to determine whether AR offers advantages over existing multimedia approaches in applying CTML principles to manage CL and optimize learning outcomes. 1.3. Learning from errors The ability to learn from errors is crucial for learning, especially in the development of procedural skills. This approach has been extensively investigated in mathematics (Darabi et al., 2018) and has also proven its relevance to motor skills (Chien & Chen, 2018). In their review, van der Meij and Flacke (2020) identify three approaches to handling errors. The first is the error-tolerant approach, which often uses “training wheels”: the learners are presented with a simplified environment where certain advanced functions are temporarily disabled, thus reducing the cost of making mistakes. Only later, as learners gain more experience, more complex features are introduced (Carroll & Carrithers, 1984). The second is the error-induced approach, known as Error Management Training (EMT). Here, learners tackle intentionally complex tasks and receive targeted support (e.g., essential feedback and heuristics) to help them recognize and manage errors independently. This model is based on the principle that working on complex activities—with limited but precise guidance—promotes deeper learning and greater error tolerance (Keith & Frese, 2008). Finally, the third is the error-guided approach, often referred to as guided error training. In this case, learners are shown both the correct procedure and an incorrect one, enabling them to compare the two and develop a stronger understanding of the topic. The goal is to offer deeper insights by directly contrasting correct and incorrect solutions (Durkin & Rittle-Johnson, 2012). All three approaches have proven significantly effective in supporting learning compared with methods that show only the correct procedure. However, given the nature of our task, it was not possible to simplify the interaction with a real T-shirt to promote exploration (first approach). Moreover, because the procedure can be performed in about 2 s, it cannot be classified as “complex” (second approach). We therefore adopted the third approach, showing both the correct procedure and the most frequent errors side by side to encourage participants’ reflection (Rohbanfard & Proteau, 2011; van der Meij & Flacke, 2020). 1.4. Differences in learning from video and augmented reality Learning a procedure with the support of video or AR shares several aspects. Both media allow for animations, thus enhancing learning outcomes (H¨ offler et al., 2007). Additionally, both can be optimized based on the principles of CTML (Mayer, 2021; Çeken & Tas¸kın, 2024) and, consequently, can be designed to optimize levels of CL (Brame, 2016; Buchner et al., 2021b). However, AR inherently applies the principles of spatial and temporal contiguity by displaying animations and 3D models directly in the real world, thus reducing even minimal attentional shifts between observed information and real-world actions. These shifts are inevitably required when learning a procedure with video support. In the study by Thees et al. (2020), although no significant differences in learning outcomes were identified between AR and a control group, the ECL levels reported by students in the AR condition were significantly lower. Conversely, in the study by Altmeyer et al. (2020), although no significant differences in ECL emerged between AR and a control group, a small effect size favoring the AR condition was observed. Studies that directly compared AR and video suggest that AR might support significantly better learning outcomes than screen-based animations (Çeken & Tas¸kın, 2024). Nonetheless, there remains a risk of increased ECL when familiarity with the technology is still low ([Authors], 2025). In addition to the uncertainty of the reported findings, there is also a lack of studies directly comparing AR and video in procedural learning contexts where the possibility of receiving spatial information in real-time becomes crucial. 1.5. Research questions and hypotheses This study aims to investigate the effectiveness of AR designed according to the principles of CTML in supporting the learning of a motor V. Candido et al. Learning and Instruction 101 (2025) 102244 3 procedure compared to video instruction. It also examines the impact of different instructional approaches, particularly the use of common errors, on learning outcomes and cognitive load. The following research questions are addressed: RQ1: How does applying the principles of signaling and spatial and temporal contiguity in HMDs-AR and video affect the levels of ECL and learning outcomes? RQ2: Do the instructional approach, the technology used, or their interaction influence differences in folding accuracy? Folding accuracy is determined by comparing a perfectly folded T-shirt with the one folded by each participant, combining the following measures: the left and right shoulder size, the total shoulder size, and the belly size (see section 2.5.1 for details). To address RQ1, we tested a structural equation model (SEM) to examine the relationships between the variables of interest. Based on the model depicted in Fig. 1, we hypothesize that: H1.The use of HMDs-AR leads to lower levels of ECL compared to video. This hypothesis is justified by the assumption that optimized implementation of signaling, spatial and temporal contiguity in HMD-AR will decrease ECL levels (Mayer, 2020). H2.Higher levels of perceived ICL increase ECL levels. This hypothesis is plausible, as learning materials perceived as highly complex(i. e. high ICL) may amplify the salience of deficiencies in instructional design, thereby increasing ECL (Chen et al., 2023). H3.Higher levels of ECL negatively affect performance, as measured by success in the guided learning and results on the knowledge test. According to cognitive load theory, allocating cognitive resources to elements that are irrelevant to the core learning content impairs learning outcomes (Sweller, 1994). H4.The performance achieved during the learning phase predicts success in retention and transfer tasks. Meaningful learning determines good retention and transfer results (Mayer, 2020). For RQ2, we propose that: H5.The introduction of common errors improves procedure performance accuracy compared to DBT. Having the opportunity to compare correct performance with common errors allows for a deeper understanding of the procedure (van der Meij & Flacke, 2020). H6.The adoption of HMDs-AR improves folding accuracy compared to video. Being able to visualize directly on the real shirt where it should be grasped will make it easier to understand where to grip, in accordance with the spatial, temporal, and signaling principles (Mayer, 2020). H7.The HMDs-AR benefits more from an instructional approach that highlights potential mistakes that might be made. AR offers a high level of assistance, which may pose the risk of reduced reflection. Learning from errors increases reflection levels (van der Meij & Flacke, 2020). 2. Materials and methods To test our hypotheses, we employed a 2 ×2 research design. The “technology” factor included the levels “video” versus “HMDs-AR,” and the “instructional approach” factor compared “DBT-based practice supported by technology” with “DBT-based practice supported by technology enhanced with information on how to avoid common errors”. Microsoft HoloLens 2 was used for the HMDs-AR conditions, and an HP EliteBook 840 G6 displayed the videos on a 14-inch screen controlled via a USB foot pedal for playback. 2.1. Material: T-shirt folding The selected procedure was chosen for instructional, experimental and technical reasons. From an instructional perspective, the selection of this procedure was motivated by the necessity of ensuring high precision at pinch points. Indeed, accurately folding the T-shirt requires extreme precision in grasping the fabric, as variations of even a few millimeters could compromise the symmetry of the final fold (Fig. 2). Additionally, from an experimental point of view, the procedure involves manipulation in two dimensions and in three dimensions, potentially enhancing the effectiveness of augmented reality, which can directly represent depth in the real world (Fig. 3). During the procedure, beyond showing the different steps required to fold the T-shirt, participants assigned to the common error condition were also shown some of the main errors they could make, along with the incorrect final folding result (Fig. 4). The relative simplicity of the procedure also allowed students to learn it quickly, ensuring completion of the entire experimental protocol within the allocated 1-h timeframe. From a technical standpoint, the primary advantage lies in the pinch gesture used to fold the T-shirt, which was easily and accurately recognized by HoloLens 2, thereby preventing potential usability issues within the application that could have negatively impacted levels of ECL. 2.2. Developing applications according to multimedia learning principles The design of both the AR application and the video was developed Fig. 1. Hypotheses tested in the model for RQ1. V. Candido et al. Learning and Instruction 101 (2025) 102244 4 in accordance with the principles of CTML. To reduce ICL levels, the following principles were applied: the segmenting principle, by dividing the procedure into various steps, and the modality principle, by using spoken instructions instead of written ones. To minimize ECL levels, the coherence principle was applied, excluding all information not strictly necessary for learning the folding technique. Throughout the entire application, we used narrated animations to demonstrate how to perform the procedure in different steps. All information was provided contextually in accordance with the principles of spatial and temporal contiguity. After the audio message ended, a brief text field summarized only the key points. This correctly applied the redundancy principle by avoiding simultaneous presentation thereby preventing the transient effect linked to audio’s ephemeral nature. Lastly, to clearly highlight where to pinch the fabric, as well as the consequences of errors in the group with common errors, the signaling principle was utilized (Fig. 5). To foster GCL, the voice principle and immersion principle were used. We delivered verbal information with a human voice instead of a robotic voice and provided the information using a device that allowed high levels of immersion. Table 1 summarizes the principles studied, applied, and not applied. 2.2.1. Differences between AR applications and video Although the principles applied were the same for both AR and video, they also varied according to the specific characteristics of each technology. In AR, it is possible to directly visualize real-world elements, thus enhancing spatial contiguity. The signaling principle was applied to the T-shirt shown in the video, and directly on the real T-shirt in AR. The automatic detection of hand movements through hand tracking improved temporal contiguity by recognizing when an action was correctly completed and providing immediate positive audio feedback, followed by the instructions for the next step. Students who watched the video received only a single positive feedback message at the end of the procedure, which assumed that reaching that point meant all steps had been performed correctly. In the video, the viewing pace was specifically managed by inserting pauses in the footage to allow for the completion of tasks and by providing students with a pedal to control playback, facilitating action even when both hands were occupied. The animations depicting the movements needed to fold the T-shirt correctly also differed slightly between the two technologies: they were highlighted by three-dimensional animations in AR, while the video showed real hands. However, in both cases, the purpose remained explanatory and not decorative (see Fig. 6). 2.3. Demonstration based training for teaching a procedure Demonstration-based training is grounded in the theory of observational learning (Bandura, 1986), which proposes that learning from observation or a model involves four interrelated processes: attention, retention, production, and motivation. According to this approach, beyond cognitive and motivational components, it is necessary for the learner to concretely practice the procedure (Grossman et al., 2013). Fig. 2. Step 1 of folding the T-shirt: pinching at two specific points on the fabric. Fig. 3. Subsequent steps required to correctly complete the folding of the T-shirt, including moving within three dimensions to ensure proper execution. V. Candido et al. Learning and Instruction 101 (2025) 102244 5 This pedagogical approach has proven effective in numerous studies, as summarized in the meta-analysis by Ashford et al. (2006) —reporting effect sizes (δuBi) of 0.77 for measures of movement dynamics—as well as in the review by Rosen et al. (2010). The combination of DBT and CTML principles has already proven effective in supporting learning outcomes when applied to video materials (van der Meij & van der Meij, 2016). Given the known effectiveness of DBT in combination with CTML principles, this approach was adopted to support the learning of this procedure. 2.4. Participants Two power analyses were conducted. For RQ1, examining the effects through PLS-SEM, the analysis followed Kock and Hadaya’s (2018) methodology (minimum beta =0.3, inverse square root and gamma-exponential methods), requiring 87 participants. For RQ2, the power analysis for the three-factor mixed ANOVA indicated a required sample size of 48 participants (effect size =0.25). The study involved 114 participants from a Swiss commercial vocational school, aged between 15 and 35 years (M =17.9, SD =2.68, 69 male). None of the participants had prior experience with HMDs-AR. Participation was voluntary, and all participants provided informed consent before data collection. Before admitting participants to the study, we asked them to demonstrate all the techniques they could use to fold a T-shirt by providing them one. Knowing the target technique was an exclusion criterion. Anyway, none of the participants were familiar with the technique explained. Participants were randomly assigned to four experimental conditions (see Table 2). 2.5. Measures and instruments This section reports the measurements taken for T-shirt folding, the knowledge test, CL, and motivation. 2.5.1. T-shirt folding learning outcomes To assess performance, a comprehensive dataset was collected. The variables analyzed included the number of content exposures required to complete the procedure during the learning phase and, for each task: success in correctly folding the T-shirt; folding accuracy; and the time taken to complete the folding. The success of the folding was determined by verifying the use of the proposed technique. The number of content exposures was quantified by counting how many times the participants viewed the instructional content before they were able to correctly fold the T-shirt. The accuracy of the T-shirt folding was evaluated by measuring the T-shirt at predefined points, as illustrated in Fig. 7, and comparing these measurements with those of an ideally folded T-shirt prepared beforehand. The difference between the measured values and the ideal values was converted into a percentage to calculate the error percentage for each measurement using the following formula: Error percentage (Δᵢ) =(| Measured valueᵢ - Ideal valueᵢ|/Ideal valueᵢ) ×100, where i =1, 2 and 3. The total error percentage for each folding was obtained by summing the error percentages for each of the four measurements: Total error percentage =∑ᵢ =1 3 Δᵢ. Original scores were transformed by first calculating the mean (sum/4) and then subtracting it from 100 to invert the scale (100 - mean), so that higher scores indicate better performance. For example: if original error was 48.5, then 48.5/3 =16.16, and 100 - 16.16 =83.84 (final accuracy score). The reported scores thus reflect the percentage of correct responses, with 100 representing perfect performance. The accuracy value of the fold on the left side of the collar was excluded because the two pinch points on the shirt could not ensure symmetry at that location, which instead depended on how the sleeve was covered. The time taken to fold the T-shirt was calculated for each phase. During the learning phase, no time limits were imposed; the participants could spend as much time as they wished to observe the materials and fold the T-shirt. However, the researchers monitored the number of Fig. 4. Mistakes in hand placement and result in final folding. V. Candido et al. Learning and Instruction 101 (2025) 102244 6 attempts each participant needed to complete the procedure correctly. In the retention and transfer phases, times were also recorded; however, a time limit of 1 min was established for the retention task, while for each of the transfer tasks, the limit was 90 s. 2.5.2. Knowledge test For this study, we developed an ad hoc test consisting of five main questions and four follow-up questions. The first question was a dragand-drop task in which students had to reconstruct the order of the procedure steps. The four subsequent questions showed the initial positioning of the hands and asked the students if they thought the hands were positioned correctly. If students believed that the hands were not positioned correctly, they were required to choose among four response options to indicate which errors had been made. If students did not indicate any error in hand positioning, no further questions were asked. The maximum achievable score was 11. 2.5.3. Cognitive load To detect the levels of CL and differentiate its three components, we used the scale validated by Klepsch et al. (2017). This scale is concise, with eight items (two for ICL, three for GCL, and three for ECL), and is suitable for repeated administrations. It is effective in detecting ECL in multimedia environments (Skulmowski, 2022) and is appropriate for tasks grounded in multimedia learning principles (Krieglstein et al., 2023). Due to the variable reliability of Klepsch’s scale ([Authors], 2025), for ECL levels, in the first administration, we also included the ECL subscale by Krieglestein et al., 2023, consisting of five items. 2.5.4. Motivation As a control variable, we included levels of motivation using the RIMMS scale (Loorbach et al., 2014), which was developed by Keller Fig. 5. Signaling used inside the application Note. Frames of animations without signaling are on the left, with signaling applied on the right. Table 1 Principles studied and applied in the applications. Principle to reduce ICL Principle to minimize ECL Principle to foster GCL Modality * Spatial contiguity ✓Immersion principle ✓ Segmenting * Temporal contiguity ✓Voice * Pre-training ◯Signaling ✓Embodiment ◯ Redundancy * Generative activity ◯ Coherence * Personalization ◯   Image ◯ Note. ICL =Intrinsic Cognitive Load; ECL =Extraneous Cognitive Load; GCL = Germane Cognitive Load. ✓ =Principle applied and investigated; * =Principle applied but not investigated; ◯ =Principle not applied. V. Candido et al. Learning and Instruction 101 (2025) 102244 7 (1987). This scale evaluates levels of attention, relevance, confidence, and satisfaction. It is composed of 12 items, with 3 items for each of the aforementioned dimensions. 2.6. Procedure After randomly assigning participants to each condition, they were familiarized with the tools they would use. Participants in the video conditions were shown how to play and pause the video using the pedal. For conditions employing HoloLens 2, participants were shown how to wear the device comfortably, followed by device calibration to optimize the hologram display. Subsequently, all participants learned the T-shirtfolding technique entirely through their assigned technology: they first watched the full procedure—shown, in the AR condition, as a life-size holographic video precisely aligned with the real T-shirt—and then completed each step with a matching step-by-step guide delivered via Fig. 6. Differences in animation between video (left) and AR (right) Note. The text was displayed only at the end of the audio. It represented a summary of the presented content and was not redundant. Table 2 Overview of participants per condition. Condition Instrument Participants Age (mean ±SD, range) Male Participants AR-DBT+Microsoft HoloLens 2 30 17.9 ±2.21 (15–23) 14 AR-DBT Microsoft HoloLens 2 28 18.5 ±3.77 (16–35) 18 VideoDBT+ Video 28 18.1 ±2.80 (15–25) 17 VideoDBT Video 28 17.1 ±1.34 (15–20) 20 Note. AR-DBTþ=Augmented Reality with Demonstration-Based Training plus common errors; AR-DBT =Augmented Reality with Demonstration-Based Training; VIDEO-DBTþ=Video with Demonstration-Based Training plus common errors; VIDEO-DBT =Video with Demonstration-Based Training. Fig. 7. Measurements for the quality of folding Note. A =Shoulder right; B =Shoulder left; C =Shoulder total; D =Belly. The back protrusion measurement was excluded because none of the students reported this mistake. V. Candido et al. Learning and Instruction 101 (2025) 102244 8 the same technology. In the two conditions with errors, animations of typical errors and their consequences for T-shirt folding were also shown. The video for the VIDEO-DBT condition lasted 2 min, while the VIDEO-DBT +condition was 2 min and 45 s long. The AR-DBT and ARDBT +conditions did not have predefined timing, because after the initial complete animation viewing, the application followed the participants’ execution. The participants could pause the videos or restart the AR applications at any time. The researchers facilitated the use of technologies without intervening in the T-shirt folding procedure. That is, the researchers were authorized to assist the participants if they encountered issues with the technology (e.g., the video pedal or the AR tap). In the case of difficulty with T-shirt folding, the researchers were only authorized to restart the application or video. After each attempt, the participants were asked whether they deemed their execution correct. If the answer was negative, they were asked whether they wished to review the explanation again; if the response remained negative, the experiment was terminated, and the participant was asked to complete a brief questionnaire. Otherwise, they could watch the video again or restart the application without limitations until they judged the folding to be correct. If the researcher intervened in the T-shirt folding, the attempt was automatically considered a failure. After the guided phase, all the participants answered a questionnaire on CL levels, motivation, and the knowledge test. The administration of the questionnaire served a dual purpose: to evaluate the variables of interest during the guided phase and to create a pause of approximately 10 min before the retention task. Only participants who successfully completed the guided learning proceeded to the retention task, while the others were taught the folding technique by the researcher. The retention task involved folding the same shirt again without any technological assistance within a 1-min time limit. After this phase, the participants completed the CL scale again, referring to the task they had just performed. They then faced three transfer tasks with a 90-s time limit each. The transfer tasks included folding a smaller shirt, the same shirt with which the technique was learned but rotated 180◦, and a larger T-shirt in a vertical position (Fig. 8). After the transfer tasks, the participants responded to the CL scale one last time, considering the three transfer tasks. Finally, if necessary, participants were shown any errors they had made. The entire procedure is summarized in Fig. 9. 2.7. Data analysis The data were analyzed using Jamovi software (Version 2.6.25) for inferential and Bayesian statistics and ADANCO (2.4.1) for Partial Least Squares Structural Equation Modeling. A hybrid model was employed, combining both reflective measurements for latent variables (e.g., ECL) and formative measurements for emergent variables (e.g., Performance). This approach was preferred over a covariance-based model as it is more effective for small sample sizes and complex models (Hair et al., 2021). The main independent variables were the type of technology used (AR vs. video) and the instructional approach (common errors vs. standard). The dependent variables included folding accuracy, success rate at various stages of the procedure, completion time, and knowledge test results. Potential confounding variables, such as age, gender, and motivation levels, were controlled for (see Appendix A for details). Ease of use of the applications was not controlled because the researchers acted as facilitators, aiming to mitigate any difficulties participants encountered when using the application. 3. Results In this section, we present the results according to the research questions and hypotheses. To examine the first research question, we tested a SEM. For the second research question, we used a three-factor mixed ANOVA, examining both withinand between-subject effects in terms of accuracy and the interaction between technology and instructional approach. We focus on the results for accuracy in all tasks (guided, retention, and transfer tasks). The results concerning the other dependent variables are available in Appendix B. Fig. 8. Representation of the three T-shirts used in the transfer tasks Note. Transfer 1: Child’s T-shirt presented with the same orientation; Transfer 2: Same T-shirt rotated 180◦; Transfer 3: Larger polo shirt, presented rotated 90◦. V. Candido et al. Learning and Instruction 101 (2025) 102244 9 Thees, M., Kapp, S., Strzys, M. P., Beil, F., Lukowicz, P., & Kuhn, J. (2020). Effects of augmented reality on learning and cognitive load in university physics laboratory courses. Computers in Human Behavior, 108. https://doi.org/10.1016/j. chb.2020.106316 Tsai, C.-H., & Huang, J.-Y. (2018). Augmented reality display based on user behavior. Computer Standards & Interfaces, 55, 171–181. https://doi.org/10.1016/j. csi.2017.08.003 van der Meij, H., & Flacke, M.-L. (2020). A review on error-inclusive approaches to software documentation and training. Technical Communication, 67(1), 83–95. https://www.jstor.org/stable/27340099. van der Meij, H., & van der Meij, J. (2016). Demonstration-based training (DBT) in the design of a video tutorial for software training. Instructional Science, 44(6), 527–542. https://doi.org/10.1007/s11251-016-9394-9 van Gog, T. (2021). The signaling (or cueing) principle in multimedia learning. In R. E. Mayer, & L. Fiorella (Eds.), The Cambridge handbook of multimedia learning. Cambridge handbooks in psychology (pp. 221–230). Cambridge University Press. Vieyra, R., & Vieyra, C. (2022). Immersive learning experiences in Augmented Reality (AR): Visualizing and interacting with magnetic fields. In S. L. Macrine, & J. M. B. Fugate (Eds.), Movement matters: How embodied cognition informs teaching and learning (pp. 217–235). The MIT Press. Wang, M., Callaghan, V., Bernhardt, J., White, K., & Pe˜ na-Rios, A. (2018). Augmented reality in education and training: Pedagogical approaches and illustrative case studies. Journal of Ambient Intelligence and Humanized Computing, 9, 1391–1402. https://doi.org/10.1007/s12652-017-0547-8 Wang, C., Li, J., Li, H., Xia, Y., Wang, X., Xie, Y., & Wu, J. (2022). Learning from errors? The impact of erroneous example elaboration on learning outcomes of medical statistics in Chinese medical students. BMC Medical Education, 22(1), 469. https:// doi.org/10.1186/s12909-022-03460-1 Wang, Y. B., Zhang, W., & Salvendy, G. (2010). A comparative study of two hazard handling training methods for novice drivers. Traffic Injury Prevention, 11(5), 483–491. https://doi.org/10.1080/15389588.2010.489242 Wu, P.-H., Hwang, G.-J., Yang, M.-L., & Chen, C.-H. (2017). Impacts of integrating the repertory grid into an augmented reality-based learning design on students’ learning achievements, cognitive load and degree of satisfaction. Interactive Learning Environments, 26(2), 221–234. https://doi.org/10.1080/10494820.2017.1294608 Yin, Y., Zheng, P., Li, C., & Wang, L. (2023). A state-of-the-art survey on augmented Reality-assisted digital twin for futuristic human-centric industry transformation. Robotics and Computer-Integrated Manufacturing, 81. https://doi.org/10.1016/j. rcim.2022.102515 Young, K. L., Stephens, A. N., Stephan, K. L., & Stuart, G. W. (2016). In the eye of the beholder: A simulator study of the impact of google glass on driving performance. Accident Analysis & Prevention, 86, 68–75. https://doi.org/10.1016/j. aap.2015.10.010 Vito Candido, junior researcher at the Swiss Federal University for Vocational Education and Training and PhD student at the University of Zurich. Research topics: immersive technologies and cognitive psychology. Alberto Cattaneo, professor at the Swiss Federal University for Vocational Education and Training. Research topics: educational technology, instructional design, multimedia learning, teacher education. Dominik Petko, professor of Teaching and Educational Technology at the University of Zurich. Research topics: teaching methodology and educational technology. V. Candido et al. Learning and Instruction 101 (2025) 102244 16