Full text
Research Paper Recommended citation: Rachmat, A., Watterson, C., & Lundqvist, K. (2025). Comparing The Impact of Generative AI Chatbots on Students' Reflective Thinking in Learning Programming. In Kangaslampi, R., Langie, G., Järvinen, H.-M., & Nagy, B. (Eds.), SEFI 53rd Annual Conference. European Society for Engineering Education (SEFI), Tampere, Finland. DOI: 10.5281/zenodo.17631485. This Conference Paper is brought to you for open access by the 53rd Annual Conference of the European Society for Engineering Education (SEFI) at Tampere University in Tampere, Finland. This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 International License.
Comparing The Impact of Generative AI Chatbots on Students' Reflective Thinking in Learning Programming A Rachmat a, 1 , C Watterson b, K Lundqvist c, a Victoria University of Wellington, Wellington, New Zealand, 0009-0004-5811-5369 b Victoria University of Wellington, Wellington, New Zealand, 0000-0001-9471-6015 c Victoria University of Wellington, Wellington, New Zealand, 0000-0002-5514-1519 Conference Key Areas: Digital tools and AI in engineering education Keywords: Reflective Thinking, Programming Courses, Generative AI, Chatbot ABSTRACT This full paper explores how a specifically tuned chatbot can support students' reflective thinking. Research in reflective thinking is relatively uncommon in computer science education. In addition, implementing reflective thinking in teaching can be challenging for academics due to the need for student personalized support. Students also often view reflection as an additional assignment rather than a way to improve their learning. Despite these challenges, programming courses can benefit from supporting student learning through specific reflective activities to foster students' reflective thinking. Supporting students' reflective thinking in programming can be achieved by triggering the process through reflective dialogue. A generative AI-based chatbot can facilitate individual and contextual conversations. This study investigates the potential of tuning large language models to trigger reflective dialogue as a learning tool. This study examined the interaction between eleven participants and the use of generative AI learning tools in a controlled laboratory setting, with one-on-one interactions between the researcher, participant, and chatbots. An experimental research design was conducted, including a control group (ChatGPT) and an experiment group (our proposed chatbot). A mixed-methods approach was employed, combining both quantitative data from pre-test and posttest questionnaires using Kember's Reflective Thinking Scale with qualitative data from observations and interviews to assess participants' reflective thinking. The results did not show any statistically significant differences in scores on the Kember's Reflective Thinking Scale. However, thematic analysis of the qualitative data 1 Agatha Rachmat A Rachmat Agatha[email protected] .
revealed that participants in the experimental group demonstrated more reflective thinking behaviours. 1 INTRODUCTION Reflective activity is not a new topic in Engineering and Computer Science Education. In learning programming, students have to apply problem-solving skills such as reasoning, questioning, and evaluation. Effective problem-solving requires reflective thinking. Students who neglect to employ reflective thinking skills often find themselves unprepared, lacking in planning and systematic approach, and failing to consider alternative solutions (Avcı, 2022; Vivian et al., 2013). However, students struggle to understand the purpose of reflection and its relevance to the course or study. Reflection is often encouraged as a writing assignment that requires writing and linguistic skills (Chan et al., 2021; Chng, 2018). Academics also often struggle to facilitate reflective activities in their courses. Academics have to plan, prepare, organise, implement, and provide necessary feedback. These tasks are even more challenging in large classes. Moreover, not all lecturers or instructors possess the knowledge or the skills to facilitate effective reflective learning activities or to improve students' reflective thinking (Chan et al., 2021) The enormous number of available and open sources of pre-trained large language models (LLMs) supported by software developer frameworks to build chatbots has created an opportunity for personalized generative Artificial Intelligence (Gen AI) tools (HuggingFace, 2025; Langchain, 2025; Ollama, 2025). Therefore, it is possible to develop generative AI-based learning tools specifically designed to support students in engaging in reflective thinking. This study introduces a specifically tuned chatbot to students and analyses the effects on their reflective thinking. 2 BACKGROUND 2.1 Reflective Thinking Dewey developed the concept of reflective thinking as "making meaning” to gain a deeper understanding of connections to other experiences and ideas (Clarà, 2015; Dewey, 1998). This concept aligns with constructivist learning theory, where students build their knowledge in connection to their existing knowledge (Berssanette & de Francisco, 2021; Cooperstein & Kocevar-Weidinger, 2004). Kember's tools for assessing reflective thinking have been extensively used in the field (Kember et al., 2000; Leung & Kember, 2003). Kember’s research identifies four (4) levels of reflective thinking in learning (Kember, 1999; Kember et al., 2000; Kember, McKay, Sinclair, & Wong, 2008): 1. Habitual Action: Learners do routine tasks without understanding the concept or theory. Experts who have done similar things many times may be at this level. It's like second nature to experts, while learners might just follow instructions without thinking about it. 2. Understanding: Learners show some evidence of comprehension of the concept or topic. They are beginning to grasp the basics and can apply them in a practical sense.
3. Reflection: Learners are able to apply the concept to real-world situations, accompanied by personal insight, and start to question and analyse their understanding and learning process. At this level, learners demonstrate metacognition, considering their own thinking and cognitive strategies (Sargent, 2015). 4. Critical Reflection: This is the highest level of reflective thinking, where shifts in perspective on fundamental beliefs related to key concepts are rare and significant. Students are required to possess reflective thinking to guide and make connections between abstract and dynamic concepts in learning programming (Medeiros et al., 2019). The Kember Reflective Thinking Scale is utilised as an analytical tool to examine the observational and interview data in this study. 2.2 Chatbots in Programming Courses Since the rise of the popularity of Gen AI chatbots such as ChatGPT, Claude and Gemma, lecturers' and students' perspectives on utilising generative AI chatbots in teaching and learning are often divided into those who see potential benefits and those who are opposed to using gen AI chatbot in education (Denny et al., 2023; Essel et al., 2024; Lau & Guo, Aug 7, 2023; Rachmat et al., 2025) Our previous intervention study were conducted in an introductory computer programming course in the School of Engineering and Computer Science. Employing a mixed-methods research to analyse the result of the intervention of using generative AI chatbot as learning tool. The study confirmed that with proper introduction and prompt guidance, learners were more inclined to utilise chatbots without negatively impacting their learning (Rachmat et al., 2025). Additionally, it has been shown that chatbots can assist students who require support while learning programming. Moreover, there is evidence of reflective learning in their learning process while assisted by chatbots (Rachmat et al., 2025). To further analyse the interaction between student and chatbot, as well as the effect of utilising chatbot on students’ reflective thinking process during learning, we conducted a qualitative study in a more isolated environment to reduce external factors that might occur in classes room settings. 3 METHOD 3.1 Development of a Chatbot for Reflective Thinking Building upon existing literature and our previous study (Rachmat et al., 2025; White et al., 2023), we developed a Retrieval-Augmented Generation (RAG) chatbot with the Large Language Models (LLMs) framework that serves as an adaptive facilitator. By simulating the interaction between the chatbot and suggested prompts developed in our previous study, learners are guided through iterative questioning and response cycles. Our developed chatbot behaviour was modified to generate questions that align with the learning objectives of the inquiry. This facilitated a structured and incremental exploration of the subject matter, mirroring a reflective dialogue. The pre-trained LLMs were selected and explored using in-context learning (ICL). ICL enables rapid adaptation to new tasks using few-shot examples, without
requiring updates. The gemma2:2b model was selected due to its superior performance after applying ICL, as its small size makes it more accessible. The chosen pre-trained model was then integrated into the RAG with conversational memory to maintain the context of the conversation, additional knowledge of Swift programming languages was added, and an in-context learning to alter the behaviour of the chatbot. The developed chatbot was then tested with three academics before the laboratory study. The three academics were represented by: one expert in programming languages, another participant who had struggled to learn programming, and a third individual who had never learned programming. Based on the testing process, the chatbot underwent significant improvements, particularly in its ability to generate targeted questions and break down complex topics into manageable learning steps. 3.2 Research Design The study was conducted in a controlled laboratory setting with one-on-one interaction between the researcher, participant, and a chatbot. There were two chatbots used, and fourteen potential participants were randomly assigned to two groups: the control group (CG) and the experimental group (EG). CG was allocated to use the commercially available chatbot, ChatGPT, while EG was assigned to use our developed chatbot. The study employed a convenience sampling strategy, which is a non-probability sampling method. Convenience sampling entails selecting study participants who are readily available and willing to participate. Participants were recruited using flyers distributed within university campuses. The criteria included no previous mobile development learning experience and having less than a year of programming experience. Potential participants were provided with a detailed explanation of the study, indicating they were allowed to opt out at any time. The study consisted of three 1-hour sessions that took place over 3 weeks. Three sessions of learning were designed to help participants develop a simple mobile app. In each session, participants were given a task to solve based on the provided learning objectives: • Session 1: To understand the User Interface (UI) framework and create a simple UI. • Session 2: To understand Swift programming language (Data type and struct) and improve previously built UI. • Session 3: To navigate between UI. Participants were also encouraged to just follow their usual learning habits and will not be reviewed or tested based on their ability. Participants were also encouraged to ask the chatbot. 3.3 Data Collection In order to investigate partcipants’ learning experiences following the study, we employed a combination of observational method and semi-structured interview as instruments of data collection. The focus of the interview was learning experiences, struggles, and how they utilised the chatbot. Observation checklists and notes were
used as instruments during the learning process to record participants’ interactions with the chatbot, document identified problems and fill in additional information related to the learning process. The interview and observation data were evaluated using Kember’s previous research on reflective thinking levels and learning approaches (Kember, 1999; Kember et al., 2008; Leung & Kember, 2003). 4 FINDINGS AND ANALYSIS Data was gathered by observing participants during the three-week learning period and interview session. These notes were inputted into NVIVO 15 to start the coding process. The initial coding session used an open coding method, and then codes were reevaluated and further coded in the second round of the coding process. From the coding process, themes that emerged are divided into categories of learning process and reflective thinking. The themes that emerged from observational data show a learning process experienced by the participants, which are: Learning difficulties and Learning strategy. Table 3 reveals that participants in both the control group (CG) and experimental group (EG) exhibit difficulties in understanding the code and syntax. The EG group demonstrates struggles and stalls in identifying what to ask and where to begin defining the problem they are facing (2. Does not know what to ask). Table 3. Participants Learning difficulties and strategies Themes Sub-themes Control group (CG) n=7 Experiment Group (EG) n=7 Learning Difficulties 1. Trying to understand code & syntax 4 3.71 2. Does not know what to ask 1.14 2.57 3. To process new or more information 1 1.42 4. Misconception 1.57 0.71 5. To break down task 0.42 0.28 6. To understand the programming concept 0.57 0.42 Learning Strategies 1. Find more information from material given 2.85 2.71 2. Ask follow-up question (to instructor) 9.42 8.57 3. Ask the chatbot 3.28 6.14 4. Ask for the answer (to instructor) 2.71 0.57 5. Trial and error 1.86 2.14 6. Decomposition 2.85 0.86 7. Visualisation 0.14 1.14 The numbers show an average of observable instances per person
Due to their learning difficulties, the participants are more likely to request guidance from both instructors and chatbots. The CG participants were observed to more likely to seek for the answer rather than guidance, compared to the EG participants. The themes for how the chatbot is being utilised are: Get unstuck, Assist with syntax, Explanation & clarification, Hinder learning process, and Improve the learning strategy. Table 4. Chatbot’s Impact on Learning Pattern Control group (CG) n=7 Experiment Group (EG) n=7 Get unstuck 0.42 3.50 Assist with syntax 0.71 2.16 Explanation & clarification 0.28 2.71 Hindering learning process 0.85 0.28 Improve the learning strategy 0 0.33 The numbers show an average of observable instances per person Table 4 shows the impact of the chatbot on participants’ learning within three sessions. Participants within the EG were using the chatbot to Get Unstuck, which represents the functions of chatbots in providing prompts to start coding, exploring keywords that can be further examined and understood, and offering clues rather than immediate answers. Meanwhile, CG participants had fewer observable instances of asking for assistance to Get Unstuck. Moreover, more observable instances of the Hindering learning process were shown in the CG group. Examples of Hinder learning process: the chatbot provides a long response that causes participants to feel overwhelmed and unable to process the information and the chatbot provides immediate answer to solve the task given. Table 5. Participants’ Reflective Thinking Levels Based on Observation Level Control group (CG) n=7 Experiment Group (EG) n=7 Habitual Action 5.4 3.14 Understand 2.14 3 Reflection 0.28 4.4 Critical Reflection 0 0.27 The numbers show an average of observable instances per person Table 5 shows participants in CG have more observable instances that align with lower levels of Kember's Reflective Thinking Scale, and participants in EG have more observable instances with higher levels of Kember's Reflective Thinking Scale.
5 DISCUSSION Based on the participants' learning strategies, it was found that they were more likely to seek guidance from instructors because they were less comfortable using chatbots for learning, which aligns with previous studies that mentioned students have concerned towards gen AI chatbot for learning (de Kereki & Garrido, May 8, 2024; Rachmat et al., 2025). Further analysis of the interaction between participants and their assigned chatbot reveals that the EG utilised the chatbot to overcome obstacles. Specifically, it was observed that EG participants were stuck on learning difficulties and did not know how to proceed, but by asking questions and utilising the response provided by the chatbot, they were able to continue the learning process. For example: Participant 3 asked “How to resize jpg image in swiftUI?“ and participant 9 asked “How to put a rectangle image?“, while participant 11 copy pasted the code to the chatbot and ask what is the error. Participants 3, 9 and 11 mentioned the questions generated with explanation and code examples provided by the chatbot prompted them to start writing their code and trigger more questions to continue their learning process. Another observable interactions in the EG group is to get more explanation or confirmation. Participants 1 (EG) and 3 (EG) asked “What is swiftUI?”, Participant 1 further inquiry about modifier in swiftUI. Participant 9 (EG) asked “what is struct in simple concept?” then tried to connect the new information with previous knowledge by asking “is it the same concept of class in OOP?”. Moreover, Participant 5 (EG) mentioned “it explain what the function of did, howand what the syntax was that it gets clear examples on the syntaxes”. The CG's participants exhibited fewer observed instances requiring assistance from the assigned chatbot. This might suggest that ChatGPT was better and its ability to provide solutions is a notable feature, however, that caused the learning process to stop whereas our approach was designed to encourage further learning and discovery which led to a deeper level of learning. This was exhibited by the few observable instances from the CG that demonstrated exploration and requests for Explanation & Clarification from the chatbot highlighting a shift in the learning process's focus towards only problem-solving over knowledge acquisition. The participants in the EG were observed to be able to manage their learning difficulties and have more attempt in trying to understand which aligned with Boyd’s concept of reflective learning that mentioned reflective learning is the ability to navigate struggles(Boyd & Fales, 1983). Further analysis of participants interactions with the chatbot and instructor revealed more observable instances of participants in EG engaging in higher levels of Reflective Thinking behaviour: Reflection and Critical Reflection levels in Kember’s Reflective Thinking Scale. In contrast, the CG participants remained at the lower levels of Reflective Thinking, specifically Habitual Action and Understanding level. Kember describes Habitual Action as making no attempt to reach understanding while learning (Leung & Kember, 2003), which is the lowest level of Reflective Thinking, commonly found among novice learners. This behaviour aligns with observed instances where learners follow instructions without a full understanding of the programming concept and less likely to seek more explanation or attempt to understand. For example, Participant 10(CG) and Participant 12(CG) were noted to
be asking questions solely to complete the assigned task, rather than attempting to gain a deeper understanding. Participant 9(EG) was observed copying and pasting the code provided in the instructions and the code provided by our chatbot. Similarly, Participant 8(CG) received an immediate answer to solve the given task as a response from the ChatGPT. 13 out of 14 participants showed the behaviour of the Understanding level, where participants attempt to understand a concept or topic. Kember mentioned that at this stage, the concepts that learners are trying to understand are still abstract and lack personal relevance, making it difficult to apply in real-world contexts (Kember et al., 2000; Leung & Kember, 2003). Both the CG and the EG exhibit this behaviour as the learning progress continues. The observable instances include attempts at improvisation or exploration of code, comparisons with previous session’s material and try to improve provided code using trial and error to complete the task. As the learning process continued, the learning approach of the participants started to differ. For example, Participant 1 (EG) explored outside the scope of the given task with responses provided by our chatbot as part of a further learning exploration. Participant 3 (EG) noted that the chatbot's response was helpful in starting coding, and avoiding getting stuck. Participant 5 (EG) was able to obtain more explanation and improve their learning strategy. These examples illustrate the chatbot's impact on the participants' learning process, specifically in helping them get unstuck and obtain explanation and clarification, demonstrating behaviours typical of Kember’s reflection level. These findings also suggest that questioning skills are critical for achieving the reflection level and sustaining the learning process. Critical reflection is characterised by a shift in perspective(Kember et al., 2000; Leung & Kember, 2003)(Kember et al., 2000; Leung & Kember, 2003). The participants from the EG exhibited a change in perspective, utilising the chatbot for learning purposes and seeking guidance instead of simply asking for answers. For example: Participant 9 (EG) mentioned “I will definitely adjust the way I'm using AI. Yeah, instead of giving me the answers directly. Give me the steps and guidance”. Participant 11(EG) initially struggled and said “the answers it (our chatbot) gave was like not the answers that I was looking for”. In the last session, Participant 11 (EG) described that the example code provided by our chatbot works as reference code, to help to try solving the task “the chatbot like made me think more deeper than ChatGPT would, because ChatGPT would like give you an answer and this chatbot would give you like a reference. So I think this chatbot was better for learning”. The CG participants did not exhibit an observable change of perspective and few instances of asking for more explanation or clarification. The EG demonstrated more instances of the higher level of reflective thinking through their ability to navigate learning struggles by asking for exploration. This finding aligns with previous research suggesting that developing reflective thinking skills requires supportive activities (Guo, 2022). 6 CONCLUSIONS This study contributes to an increasing knowledge of how to utilise generative AI chatbots to support learners. This study indicates that generative AI has the potential to improve students’ reflective thinking in learning programming. The findings suggest that a specifically tuned generative AI chatbot facilitates reflective thinking.