Evaluating a continuing medical education program: New World Kirkpatrick Model Approach
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Liao, Shih-Chieh; Hsu, Shih-Yun Article Evaluating a continuing medical education program: New World Kirkpatrick Model Approach International Journal of Management, Economics and Social Sciences (IJMESS) Provided in Cooperation with: International Journal of Management, Economics and Social Sciences (IJMESS) Suggested Citation: Liao, Shih-Chieh; Hsu, Shih-Yun (2019) : Evaluating a continuing medical education program: New World Kirkpatrick Model Approach, International Journal of Management, Economics and Social Sciences (IJMESS), ISSN 2304-1366, IJMESS International Publishers, Jersey City, NJ, Vol. 8, Iss. 4, pp. 266-279, https://doi.org/10.32327/IJMESS/8.4.2019.17 This Version is available at: https://hdl.handle.net/10419/213026 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/3.0/
266 International Journal of Management, Economics and Social Sciences 2019, Vol. 8(4), pp. 266 – 279. ISSN 2304 – 1366 http://www.ijmess.com Evaluating A Continuing Medical Education Program: New World Kirkpatrick Model Approach * Shih-Chieh Liao1 Shih-Yun Hsu2 1 School of Medicine, China Medical University, Taichung City, Taiwan 2 College of Intelligence, National Taichung University of Science and Technology, Taichung City, Taiwan The New World Kirkpatrick (NWKM) four-level model is a new vision of the Kirkpatrick Model. NWKM adds new elements to recognize the complication of the educational program background and to evaluate the effectiveness of continuing education. This study used data collected from subjects, distributed to 393 participants enrolled in an acupuncture training program in Taiwan from 2010 to 2017, to explore the implication of NWKM for evaluating the effectiveness of continuing medical education and to discuss the connection and transition among the four levels of NWKM. Exploratory factor analysis was used to address that the items in the survey were grouped in different categories and mapped onto the four levels of the NWKM. Path analysis was used to describe the directed dependencies among the levels of NWKM. The results of path analysis showed that a positive relationship exists between any two levels, but direct effects can be observed only between two consecutive levels. It means that L4 outcomes can only be directly predicted by L3, but neither L1 nor L2. L4 is the ultimate outcome of evaluating the effectiveness of continuing education, but it is hard to achieve. This research concluded that L3 is the key to evaluate continuing medical education. Keywords: Continuing medical education, new world Kirkpatrick model, curriculum evaluation, acupuncture, exploratory factor analysis JEL: I19, I21 The purpose of continuing education is to promote the employee’s professional abilities to enhance the efficiency of the employing organization (Noe, 2016). To understand and increase the effectiveness of continuing education, different methods are used to examine the design, development, implementation, and outcome of the training programs (Wang and Wilcox, 2006). Researchers believe that a comprehensive evaluation of a training program should include appraisal before training, curriculum design and development, and after training (Goldstein and Ford, 2002). Different types of process data or outcome data are gathered depending on the chosen evaluation approaches (Blanchard and Thacker, 2007). Kirkpatrick’s four-level model (hereafter, KM), one of the most recognized project evaluation frameworks, emphasizes that clinical outcomes are the highest level of impact that educational interventions can achieve. Without evaluating the effectiveness of an educational program, Wong and DOI :10.32327/IJMESS . 8 . 4 .201 9 . 17 Manuscript received June 1, 2019; revised October 18, 2019; accepted Novem ber 27, 2019. © The Author(s); CC BY-NC; Licensee IJMESS *Corresponding author: [email protected]om
International Journal of Management, Economics and Social Sciences 267 Holmboe (2016) argue that few medical educational programs can successfully improve the treatment outcomes of patients. Therefore, there has been a recent call for application of KM model in evaluating the effectiveness of continuing medical education (Moreau, 2017; Sultan et al ., 2019; Yardley and Dornan, 2012). Although KM remains the most commonly used model for evaluating continuing education and training (Rafiq, 2015), there are some criticisms that the original KM has faced. First, the links between the levels are not strong and a causal relationship cannot be assumed (Alliger and Janak, 1989; Bates, 2004; Dixon, 1990). Second, as the level increases, their actual usage in continuing education evaluation decreases (Chevalier, 2004; Van Buren and Erskine, 2002). Third, the original KM does not provide an evaluator with insights into the underlying mechanisms that inhibit or facilitate the achievement of the educational program’s results (Parker et al ., 2011). In response to the criticism, Kirkpatrick and Kirkpatrick (2016) modified the original KM to develop the New World Kirkpatrick Model (hereafter, NWKM). NWKM adds new concepts to recognize the complication of the educational program background and to improve the authority for completely evaluating the effectiveness of continuing education. Therefore, the aim of this study is to apply NWKM to evaluate the effectiveness of continuing medical education as well as to understand the relations among the four level outcomes of NWKM. In other words, the major research question is to examine how lower level outcomes of NWKM might directly or indirectly predict higher level outcomes of NWKM. LITERATURE REVIEW The Original Kirkpatrick Model KM was developed by Kirkpatrick (1959) to evaluate the effectiveness of continuing education (Praslova, 2010). As a type of outcome data evaluation, KM emphasizes the understanding of the training outcomes such as satisfaction toward the instructor, knowledge or skill gained, attitude or performance changed, and improved gains of the organization (Werner and DeSimone, 2011). These outcomes are graded into 4 levels depending on the amount of time required to achieve. Level 1 (L1), the reaction level, concerns the trainee’s satisfaction toward the instructor and the curriculum. Level 2 (L2), the learning level, refers to the trainee’s learning of professional knowledge or skill. Level 3 (L3), the behavior level, is the changes in the trainee’s behavior or performance. Level 4 (L4), the result level, is the improved efficiency of the organization contributed by the trainee as a result of the continuing education program (Alliger and Janak, 1989; Kirkpatrick, 1998) (See Figure 1). Original Kirkpatrick Model: Problems and Criticism
Liao & Hsu 268 Source: Kirkpatrick (1998) Figure 1. Levels of Kirkpatrick’s Original Evaluation Model It would be best if information from all four outcome levels of the original KM could be gathered and analyzed. However, due to man power and financial concerns, behavior and result levels of outcomes are not as often obtained as those of reaction and learning levels (Geber, 1995). The 2002 State-ofthe-Industry Report (SIR) conducted by the American Society for Training and Development (ASTD) shows that 78 percent of the institutes surveyed evaluated the reaction level outcome (L1), 32 percent evaluated learning level outcome (L2), 19 percent evaluated behavior level outcome (L3), and only 7 percent evaluated result level outcome (L4) (Arthur et al ., 2003; Van Buren and Erskine, 2002). In medical education, most of KM implications of evaluating educational program reveal at L2 (learning) (Sultan et al ., 2019; Yardley and Dornan, 2012). It is not clear whether a high satisfaction rate in L1 (reaction) and L2 (learning) outcome would bring about an equally high satisfaction in L3 (behavior) and L4 (result) outcome. Noe (2016) argued that L1 (reaction) and L2 (learning) level outcome cannot be considered an index of training translation. In other words, results of L1 (reaction) and L2 (learning) evaluation cannot necessarily predict trainees’ performance and attitude changes, how trainees apply the learning to problem solving at work, or what influences the training program might bring to the institute. Emphasizing L1 (reaction) and L2 (learning) and neglecting L3 (behavior) and L4 (result) might create happy participants but would not really bring about institute development. The New World Kirkpatrick Model Based on the original KM, the New World Kirkpatrick Model (NWKM) redefines the 4 levels of outcomes
International Journal of Management, Economics and Social Sciences 269 and provides new explanations (Kirkpatrick and Kirkpatrick, 2016). In the NWKM (Figure 2), two new views are proposed. First, the original KM claims that the evaluations of continuing education can be separated into four levels, but the NWKM questions that the links between the levels are unclear. The NWKM, hence, argues that L1 (reaction) and L2 (learning) should be considered to be one larger category while L3 (behavior) and L4 (result) is the other. In other words, the NWKM contends that the correlation between L1 (reaction) and L2 (learning) and the correlation between L3 (behavior) and L4 (result) would be higher than the correlation between L2 (learning) and L3 (behavior) (Kirkpatrick and Kirkpatrick, 2016). Second, the definitions of different level outcomes need to be determined backwards from L4 (result) to L1 (reaction). As an outcome model of evaluation, the original KM observes and measures the outcome of continuing education so as to understand the training effectiveness and ultimately to improve the overall performance of an institute. L4 (result) should indicate most important and desired outcome of training. In the NWKM, L4 (result) outcome is decided first and then the rest of the levels are defined following it. In this way, results of evaluation can provide better information and washback for the design and conduction of continuing education. Source: Kirkpatrick & Kirkpatrick (2016) Figure 2. The New World Kirkpatrick Model Conceptual Framework Kirkpatrick’s original model (KM) is widely used for evaluating continuing education. New World Kirkpatrick Model (NWKM) expands the scope of the original KM by adding concepts and process measures to enable educators to interpret the results of evaluation, but with the aim of proving educational programs (Gandomkar, 2018). There are two major differences between the original KM
Liao & Hsu 270 and NWKM. 1. In the NWKM, the outcomes of L4 is decided first and then the rest of the levels are defined following it. 2. The original KM claims that the four evaluation levels are separate, but NWKM argues that L1 and L2 should be considered to be one larger category while L3 and L4 is the other. Based on the conceptual framework and the purpose of this study, this study proposes two hypotheses: H1: L1 (reaction) and L2 (learning) might be better considered as one category and L3 (behavior) and L4 (result) as the other. H2: L1 (reaction), L2 (learning), and L3 (behavior) outcomes might directly predict L4 (result). METHODOLOGY -Research Design In order to apply NWKM to evaluate the effectiveness of continuing medical education as well as to understand the relationships among the outcomes of NWKM’s different levels, the participants of an acupuncture training program offered by China Medical University were surveyed with endorsement from the Health Department of Taiwan. The acupuncture training program is specially designed for physicians and dentists who are trained in Western medicine. The program contains a total of 192 hours of training on theoretical and philosophical perspectives of Chinese medicine and acupuncture, acupuncture skills, and bedside teaching. Table 1 (see Appendix-I) shows the course contents and number of hours for each topic. Upon finishing the program, trainees are allowed to take the acupuncture specialization certification test. After successfully passing the test, they can apply acupuncture in combination with their Western medicine medical practice. -Measurement Based on the research objective, a questionnaire was developed and face and content validities were ensured by three experts, including one medical educator (who is the main program designer and also a Chinese medicine physician with more than 30 years of clinical and medical continuing education experience) and two human resource experts (one of those is the first author and both experts have more than 15 years of human resource management experiences). The questionnaire contained 25 5point Likert scale items (5 = Strongly agree, 1 = Strongly disagree). The items were designed following
International Journal of Management, Economics and Social Sciences 271 the guidelines of NWKM and related literature (Alliger and Janak, 1989; Kirkpatrick, 1998; Kirkpatrick and Kirkpatrick, 2016; Werner and DeSimone, 2011). To further improve content and face validities, we asked the trainees, who were enrolled in the acupuncture training program a year before, to highlight any issue in questionnaire items. After collecting responses from study subjects, we revised the content and wording of the survey based on their feedback. -Research Ethics The study received an exemption recommendation (CMUH REC No. CRREC-107-078) from the Research Ethics Committee, China Medical University and Hospital, Taichung, Taiwan. A signed informed consent to participate was obtained from each participant and the rights about confidentiality, anonymity, voluntary withdrawal from study, and disposal of material containing personal information after the completion of the study were explained and assured to study subjects. -Participants A total of 393 (295 male) trainees enrolled in 14 different sessions of the acupuncture training program offered by China Medical University during the years between 2010 and 2017 were selected as study subjects. All of the 393 trainees were invited to participate in this research. The questionnaire was distributed to all the participants at the end of each session of the program. All of the study objectives, methods, and procedures of data collection were explained to the participants at the time of the survey. Data Analysis Among the 393 participants, 159 (124 male) valid surveys were collected with an effective response rate of 40.46 percent. Survey data were analyzed statistically by using exploratory factor analysis (EFA), chi-square test, one-way ANOVA, Pearson correlation, and path analysis. Due to the reason that NWKM has never been applied in evaluating an acupuncture training program, EFA was used to categorize participants’ responses into a small number of main factors which could later be mapped to the four levels of NWKM. Chi-square was used to test whether the gender difference exists between male and female participants. One-way ANOVA was used to test if there is significant difference in the participants’ satisfaction rate. Pearson correlation was used to understand the strength of every two consecutive NWKM levels (Lenhard and Lenhard, 2014). Finally, path analysis was conducted to check both direct and indirect effects among the NWKM levels. According to Kirkpatrick and Kirkpatrick (2016), the role of human resource experts is critical in successfully implementing the model due to the reason that it requires expertise to assess the learning effectiveness and analyze final results of the training. Therefore, in the development of the survey
Liao & Hsu 272 and the process of data analysis, we consulted the main program designer (also a Chinese medicine physician and educator), two human resource experts, and two participants to describe the learning transfer processes. Factor Analysis Principle Component Analysis with Varimax rotation method was conducted on the 25 items. The Kaiser-Meyer-Olkin (KMO) statistic was 0.899 and Bartlett's chi-square value was 3566.853 (df = 300; p < 0.01, a meritorious interpretation). Based on the Scree plot and eigenvalue criteria, three main factors were found explaining 65.84 percent of the variance. After consulting with two human resource experts, three items were deleted. The remaining 22 items had a Cronbach’s alpha value of 0.940. Based on the same factor analysis procedures described above, the new KMO statistic was increased to 0.903 (Bartlett's χ2 =3244.538; df = 231; p < 0.01, a marvelous interpretation) and the 3 factors explained 72.05 percent of the variance (see Table 2-Appendix-II). RESULTS When mapped onto NWKM, Factor 1 was considered to be L1 (reaction) outcome level with 12 items (mean = 4.204, SD = 0.517). Factor 2 was L2 outcome level with 4 items (mean = 4.326, SD = 0.563). Factor 3 could be mapped onto L3 (behavior) and L4 (result). Since NWKM has 4 levels, we divided Factor 3 into L3 (behavior) with 3 items (mean = 4.072, SD = 0.742) and L4 (result) also with 3 items (mean = 3.675, SD = 0.864). As the outcome level goes up from L1 to L4 (result), the participants’ satisfaction rate goes down. One-way ANOVA revealed that other than L1 (reaction) and L2 (learning), there were significant differences in the participants’ satisfaction rate in between any two consecutive NWKM levels ( p < 0.05). The relationship among the four levels in NWKM All of the Pearson correlation coefficients were significant at p < 0.01. L1 (reaction) to L2 (learning) ( r12 = 0.638), L1 (reaction) to L3 (behavior) ( r13 = 0.423), L1 (reaction) to L4 (result) ( r14 = 0.416), L2 (learning) to L3 (behavior) ( r23 = 0.496), L2 (learning) to L4 (result) ( r24 = 0.431),and L3 (behavior) to L4 (result) ( r34 = 0.827). Comparisons of correlations revealed that the correlation between L1 (reaction) and L2 (learning) was higher than that between L2 (learning) and L3 (behavior), and the correlation between L3 (behavior) and L4 (result) was higher than that between L2 (learning) and L3 (behavior) ( p < 0.05). Path Analysis The results of path analysis showed that L1 (reaction) had direct effect on L2 (learning) (β = 0.638, p
International Journal of Management, Economics and Social Sciences 273 p < 0.001) and L3 (behavior) (β = 0.179, p = 0.047). L2 (learning) had direct effect on L3 (behavior) (β = 0.382, p < 0.001), and L3 (behavior) had direct effect on L4 (result) (β = 0.800, p < 0.001). No direct effect was found from either L1 (reaction) (β = 0.095, p = 0.109) or L2 (learning) (β = - 0.026, p = 0.677) to L4 (result). L1 (reaction) also had an indirect effect of 0.638*0.382=0.244 on L3 (behavior) through L2 (learning), making a total of effect of 0.244+0.179=0.423. The indirect effect from L1 (reaction) to L4 (result) through L2 (learning) and L3 (behavior) was 0.638*0.382*0.800=0.195; the indirect effect from L2 (learning) to L4 (result) through L3 (behavior) was 0.382*0.800=0.306. Only L3 (behavior) had direct effect on L4 (result). Figure 3 shows the results of path analysis. * p<0.05 Source: Study Analysis Figure 3. Path Analysis DISCUSSION Based on the results, hypothesis one is confirmed. L1 (reaction) and L2 (learning) might be better considered as one category while L3 (behavior) and L4 (result) are considered as the other. Hypothesis two, however, is partially confirmed. It was found that only L3 (behavior) can directly predict L4 (result), and L1 (reaction) and L2 (learning) can only indirectly predict L4 (result). The results are discussed below. L1 and L2 is considered as one category and L3 and L4 as the other Among the 3 factors identified by factor analysis, Factor 1 could be mapped to L1 (reaction) of KM, which described the satisfaction toward the instructor and the curriculum, and Factor 2 mapped to L2 (learning), referring to knowledge and skill growth. Factor 3 was mapped to both L3 (behavior) and L4