Psychometric properties of 12-item self-administered World Health Organization disability assessment schedule 2.0 (WHODAS 2.0) among general population and people with non-acute physical causes of disability : systematic review
Full text
This is a self-archived version of an original article. This version may differ from the original in pagination and typographic details. Author(s): Title: Year: Version: Copyright: Rights: Rights url: Please cite the original version: CC BY-NC 4.0 https://creativecommons.org/licenses/by-nc/4.0/ Psychometric properties of 12-item self-administered World Health Organization disability assessment schedule 2.0 (WHODAS 2.0) among general population and people with non-acute physical causes of disability : systematic review © 2021 Taylor & Francis Accepted version (Final draft) Saltychev, Mikhail; Katajapuu, Niina; Bärlund, Esa; Laimi, Katri Saltychev, M., Katajapuu, N., Bärlund, E., & Laimi, K. (2021). Psychometric properties of 12-item self-administered World Health Organization disability assessment schedule 2.0 (WHODAS 2.0) among general population and people with non-acute physical causes of disability : systematic review. Disability and Rehabilitation, 43(6), 789-794. https://doi.org/10.1080/09638288.2019.1643416 2021
1 PSYCHOMETRIC PROPERTIES OF 12-ITEM SELF-ADMINISTERED WORLD HEALTH ORGANIZATION DISABILITY ASSESSMENT SCHEDULE 2.0 (WHODAS 2.0) AMONGST GENERAL POPULATION AND PEOPLE WITH NON-ACUTE PHYSICAL CAUSES OF DISABILITY – SYSTEMATIC REVIEW Short title: Psychometrics of self-administered 12-item WHODAS 2.0 ABSTRACT Objective WHODAS 2.0 is a unified scale to measuring disability across diseases, countries, and cultures. The objective was to explore the available evidence on the psychometric properties of 12-item self-administered WHODAS 2.0 amongst a general population and people with non-acute physical causes of disability. Methods Five databases Medline, Embase, Web of Science, Scopus and PsycINFO were searched for papers related to the validity, reliability, responsiveness, minimal clinically important difference or minimal detectable change of 12item self-administered WHODAS 2.0. In order to avoid missing any potentially relevant studies, the search clauses were left as generic as possible and the refining search was conducted manually. As the review was focusing on chronic physical disorders and general adult population, major psychiatric diagnoses, acute traumas, other acute conditions (e.g. postpartum or pregnancy), hearing loss, progressive neurological disorders, and age <19 years were excluded. The relevancy of the studies was assessed by two independent reviewers. Results The 14 out of 191 observational studies were considered relevant. The sample sizes varied from 80 up to 31,251 participants. Great diversity was observed in the participants’ health problems. The Cronbach’s alpha was high – up to 0.96. The correlations between WHODAS 2.0 and other disability scales were high. Substantial floor without ceiling effect was reported by two studies. Exploratory factor analysis resulted in a multidimensional structure – up to five factors. The discriminative ability and test-retest reliability of the scale was good. Conclusions It seems, that the 12-item self-administered WHODAS 2.0 is internally consistent and a reliable scale demonstrating overall good correlation with other measures of disability. However, it appears that it is a multidimensional scale and its total score may represent different combinations of several contributing factors. Thus, the 12-item WHODAS 2.0 can be more reliable when creating a person’s functional profile formed by the 12 individual item scores instead of a single total sum. KEYWORDS disability evaluation; international classification of functioning, disability and health; functioning; psychometrics; reproducibility of results; consistency; floor effect; ceiling effect; whodas
2 INTRODUCTION The World Health Organization Disability Assessment Schedule 2.0 (WHODAS 2.0) is an ambitious attempt made by WHO to introduce a unified scale to measuring disability across diseases, countries, and cultures [1, 2]. WHODAS is based on the International Classification of Functioning, Disability and Health (ICF), producing standardized numeric disability levels and profiles. The comprehensive 36-item WHODAS 2.0 has been developed in order to describe six latent constructs: cognition, mobility, self-care, getting along, life activities, and participation. In theory, these six ‘sub’-constructs (or ‘domains’ in the terms of ICF) should be able to explain the broader concept of ‘general disability’. The number of indicators – WHODAS 2.0 items – varies from four to eight for each of the six latent constructs. The 12-item WHODAS 2.0 has been derived from the 36-item version to provide a briefer tool for assessing overall functioning in surveys or health-outcome studies. Two items for each of the six latent factors have been included in the 12-item WHODAS 2.0. The 12-item version has been found to be reliable, and has been reported to explain 81% of the overall variance of results of the 36-item WHODAS [1]. The total score of WHODAS 2.0 is scored either by using an item response theory or as the simple sum of scores assigned to each of the items. The WHODAS 2.0 is available in several versions: 36-, 24+12and 12-item questionnaires in self-, interviewerand proxy-administered forms. While the full 36-item version has more commonly been used, the shorter 12-item WHODAS 2.0 has raised a great interest among clinicians and researchers as an easy-to-use short indicator of disability, sometimes called a WHODAS ‘screener’. About one third of all papers identified by a recent review on WHODAS 2.0 has employed a 12-item version [3]. The psychometric properties of 36-item WHODAS 2.0 has extensively been studied [2]. Overall, it has been described as a consistent, reliable, and unidimensional tool. Instead, the knowledge on the 12-item version’s psychometrics is scarce. Previous research has often assumed that psychometric properties of the 12-item version are fully inherited from its more comprehensive 36-item form. For example, several studies have conducted confirmatory factor analysis of the 12-item WHODAS 2.0 based on a presumption of unidimensionality and hierarchical structure (one common factor ‘disability’ and six subfactors regarding different dimensions of disability) demonstrated by a 36-item [4, 5]. However, one could expect that excluding 24 out of 36 items might affect psychometrics substantially. In other words, it is uncertain how well a 12-item version is able to reproduce the psychometric properties of 36-item WHODAS 2.0. The research on psychometrics of the 12-item WHODAS 2.0 is scattered across its different forms and diverse populations of interest Firstly, the psychometrics of all three 12-item forms – self-, proxy and intervieweradministered – have sometimes been reported as the properties of a general ’12-item WHODAS’, even though the psychometrics of self-reported form might differ from proxyor interviewer-administered assessments. Secondly, the research on the subject is scattered across numerous relatively small samples with different settings and diagnostic profiles. It is true, that WHODAS 2.0 is a tool that should work, in theory, in any diseases and
3 settings, but this assumption should first be confirmed by comparing the psychometric properties of the WHODAS 2.0 between large samples involving similar conditions and analogous settings. The objective of this study was to explore the available evidence on the psychometric properties of the 12-item self-administered WHODAS 2.0 amongst a general population and people with non-acute physical causes of disability.
4 METHODS Inclusion and exclusion criteria Inclusion:Papers (including short communications and letters to editor, excluding conference proceedings, theses etc.) published in academic peer-reviewed journals. No restrictions on time of publication or language. Exclusion: Major psychiatric diagnoses, acute traumas, other acute conditions (e.g. postpartum or pregnancy), hearing loss, progressive neurological disorders, age <19 years. Databases: Medline, Embase, Web of Science, Scopus, PsycINFO. Outcome: Psychometric properties of WHODAS 2.0 understood as any property of WHODAS 2.0 related to its validity, reliability, responsiveness, minimal clinically important difference or minimal detectable change, or respective. Data sources and searches The MEDLINE (via PubMed), Embase, Web of Science, Scopus, and PsycINFO databases were searched in January 2019. The search clauses are presented in Table 1. In order to avoid missing any potentially relevant studies, the search clauses were left as generic as possible and the refining search was conducted manually. The references of identified articles and reviews were also checked for relevancy. Study selection Two independent reviewer teams (NK + EB vs. MS) screened titles and abstracts of articles and assessed the full texts of potentially relevant studies (Figure 1). Disagreements between the reviewers were resolved by consensus or by a third reviewer (KL). The methodological quality of the included trials was not rated. Data extraction The potentially relevant data were extracted from the records by one reviewer using a predefined structured form including title, first author, year of publication, country of origin, study settings, participants’ main diagnoses if specified, sample size, gender distribution, participants’ age, main psychometric measures used, main quantitative results, and the conclusions drawn by the original authors.
5 RESULTS Search results The search resulted in 191 records. Of them, 148 were excluded as duplicates and papers on hearing loss, psychiatric disorders, trauma, Huntington disease, postpartum and pregnancy, papers on 36-item version, and general commentaries. The remaining 43 records were screened based on their titles and abstracts and 13 irrelevant papers were excluded. The number of observed agreements between the reviewers was 30 (70% of the observations) and kappa was 0.23 (SE 0.16, 95% CI -0.09 to 0.55) considering the strength of agreement between the reviewers to be ‘fair’. Thirty records were assessed based on their full-texts and 16 irrelevant papers were excluded comprising 14 relevant studies potentially fit for a qualitative analysis (Figure 1). Additionally, one study that was published after the search was considered relevant into further analysis [6]. Data extraction The attempt to extract relevant data regarding a 12-item self-administered WHODAS 2.0 version from the report by Tazaki et al. [7] was unsuccessful and that study was excluded from further analysis. Tazaki et al. [7] employed five different versions of WHODAS 2.0 (12and 36-item interviewer-administered, 36-item proxy-administered, and 12and 36-item self-administered versions) and there was a discrepancy in reporting a sample size (total n=126 but 62 men and 70 women). After the selection and data extraction phases, 14 records were included into further analysis. Studied samples All of the 14 remaining papers were published after 2013 (Table 2). All of them were observational studies. Two studies focused on the elderly [8, 9] while the rest evaluated people of working age. The sizes of samples varied from 80 up to 31,251 participants. Except for one study with 98% women [10], the proportions of female participants were between 47% and 65%. A great diversity was observed in the participants’ health problems: patients waiting for an elective joint arthroplasty or neurosurgery [8, 11], general population or healthy volunteers [4, 9, 12, 13], patients with chronic musculoskeletal pain or fibromyalgia [10, 14, 15], patients with spinal cord injury [5, 16], and people reimbursed for any disabilities [17]. Psychometric properties The most common psychometric properties reported by the included studies were Cronbach’s alpha and convergent validity. The alpha estimates were usually high varying from 0.81 up to 0.96. Any pooling of the reported concurrent validity estimates was impossible as each study compared WHODAS 2.0 with different scales. However, the reported correlations between WHODAS 2.0 and other disability scales applied at the same time with WHODAS 2.0 were high in most of the studies. Floor and ceiling effects were reported by three studies. One study (the biggest sample size of 31,251) reported a substantial floor effects up to 32% (on average 20%) without a ceiling effect [12]. Another study conducted on a sample of 183 participants did not observe any floor or ceiling
6 effects [11]. Moreover, in that study, none of the participants – patients waiting for a neurosurgical procedure – reported a highest or lowest WHODAS 2.0 scores. The third study reported a significant floor effect up to 80% for all 12 items and for a total score without a ceiling effect [6]. Exploratory factor analysis or principal component analysis were employed by six studies [5, 10, 12, 14, 15, 17]. None of them reported a unidimensional structure of 12-item WHODAS 2.0. The number of factors varied from two up to five. Four studies employed confirmatory factor analysis. Only one of them reported a good model fit [4]. In one study, the hierarchical model with one common factor and six subfactors (as suggested by WHODAS 2.0 developers for 36-item version) was assessed resulting in poor fit [5]. One study reported a good fit of onefactor model but the reported root mean square error of approximation (RMSEA) was insignificant 0.079 pointing at a poor fit [11]. Another study reported a good fit of a two-factor model [15]. In one study, the discriminative ability was assessed using Karnofsky Performance Status scale as indicator of disability [11] reporting positive results. Another study assessed the discrimination ability using the item response theory [14]. That study reported discrimination of WHODAS 2.0 items being high to perfect, even though, the difficulty of items was shifted towards elevated disability rates. Such a shift implicates that a respondent should be experiencing slightly worse disability (compared with the average population rate) to achieve a 50/50 probability of giving an answer that would be interpreted by the WHODAS 2.0 as a “worse disability.” Three studies assessed test-retest reliability of the 12-item WHODAS 2.0 [9, 13, 18] reporting insignificant differences between repeated measures. In all three studies, the time interval between measures was one week.
7 DISCUSSION This systematic review of 14 observational studies evaluated the available evidence on the psychometric properties of the self-administered 12-item WHODAS 2.0 among a general adult population or people with nonacute physical causes of disability. While the spectrum of the studies was expectedly wide, some patterns could be observed. Firstly, most of the studies found WHODAS 2.0 to be internally consistent. Secondly, the scale seemed reliable in term of test-retest reproducibility even if the time interval between studied repeated measures was hardly sufficient (one week). Thirdly, the 12-item WHODAS 2.0 might have a substantial floor but not ceiling effect. Therefore, the screening ability of this WHODAS 2.0 version seems to be weak as it may not distinguish lower levels of disability severity. Fourthly, WHODAS 2.0 seems to be able to discriminate well people with other than the lowest levels of perceived disability. Fifthly, respondents might be slightly more disabled in reality than the reported level of disability implies. Finally, the biggest concern risen of this review is one regarding the factor structure of WHODAS 2.0. Instead of unidimensionality, several included studies pointed at the multidimensional structure of the scale. While unidimensionality refers to measuring a single construct (in this case, disability level), multidimensionality refers to the fact, that scale is measuring two or several different constructs. That makes the total scores of multidimensional tests hard to interpret as there is no certainty on the exact contributions of each underlying construct to the total [19]. The main weakness of this systematic review was the considerable heterogeneity of the included papers. Their study populations ranged from healthy volunteers to tetraplegics. While one of the advantages of WHODAS is comparability between different health problems, conclusions could be more reliable if there were several studies on a similar disorder in different settings and on large samples. The number of identified relevant studies was surprisingly small. The included studies assessed convergent validity of WHODAS 2.0 by comparing with a wide spectrum of different tests and scales. While those comparators were mostly valid and reliable, the small number of studies on each of them made a reliable pooling impossible. Unfortunately, no system of the assessment of systematic bias seemed to fit the purpose of the review. The uncertainty regarding the methodological quality of the included studies may substantially weaken the strength of generalization of the results. This was, however, the first attempt to evaluate systematically the properties of the 12-item self-administered WHODAS 2.0 and the review was able to deliver several generalized clinical recommendations. Only one previous review has been conducted on the topic so far [3]. Evaluating over 800 papers on the WHODAS 2.0, Federici et al. concluded that the WHODAS 2.0 shows strong correlations with several other measures of activity limitations probably due to the fact that it shares the same disability latent variable with them. This good convergent validity was in line with the findings of the present review. Concerning the factor structure of WHODAS 2.0, the conclusions of review by Federici et al. were more optimistic than the inferences of the present review that could not confirm the one-factor structure of 12-item WHODAS 2.0. The differences in the results of these two reviews may lay in the differences between their scopes. The scope of the present review was limited to a self-reported version of 12-item WHODAS 2.0 applied to a general population and people with non-acute physical
8 conditions. It is possible that the factor structures of other forms of WHODAS 2.0 are different. It is also possible that WHODAS 2.0 may behave differently when applied to populations others than studied here. It has to be noted that the majority of the papers included into the present study were published after April 2016 when the review by Federici et al. was already submitted. This review focused on a self-reported version of WHODAS 2.0. The psychometric properties of interviewerand proxy-administered forms may be different. When giving a self-reported response, a respondent may exaggerate, avoid embarrassing details, or try to confirm a guessed research question. A response may also be affected by the desire to obtain some social or financial benefit or service. On the other hand, a self-reported test may avoid the influence of interaction with an assessor. Implications for clinical practice Due to a substantial floor effect, the use of the self-administered 12-item WHODAS 2.0 as a screening tool in general population, when only mild severity of disability is expected, seems questionable. This scale may be used as an easy-to-use short questionnaire to assess the functioning profile of people with chronic physical conditions. The 12-item WHODAS 2.0 seems to be able to produce reliable repeated measures and, thus, may be used to assess the change in functioning level. Due to its multidimensional structure (measuring more than a single underlying construct), the 12-item version of WHODAS 2.0 may not be able to produce a reliable and comparable total score. Instead, the scale’s 12 items should be scored and presented separately as a profile. Recommendations for further research on 12-item WHODAS The discrimination of this scale version ability is poorly understood – only two studies are conducted on the subject so far, each employing a different statistical technique [11, 14]. The minimal clinically important difference and minimal detectable change of the scale are still unknown and should be studied separately for each of the 12 items due to a seemingly certain multidimensional structure of the 12-item version. The convergent validity should be re-tested against similar relevant standard scales. The results of the item response theory obtained from only one sample should be reproduced in different settings and populations. The test-retest reliability assessment should be repeated in different time interval between test-retest measures. A short time interval (like a one-week interval employed in the included studies) may make the carryover effects due to memory, practice, or mood more probable. Instead, longer intervals increase the probability of changes in the clinical status [20, 21]. When a reference test (gold standard) is applicable then the sensitivity and specificity of WHODAS 2.0 should be evaluated, at least, in some populations. Conclusions It seems, that the 12-item self-administered WHODAS 2.0 is internally consistent and a reliable scale demonstrating overall good correlation with other measures of disability. However, it appears that it is a multidimensional scale and its total score may represent different combinations of several contributing factors.
15 Figure 1. Search flow Medline n=29 Embase n=61 Scopus n=58 Web of Science n=26 PsycINFO n=7 n=191 Duplicates n=105 n=86 Papers on hearing loss (3), psychiatric disorders (20), trauma (3), Huntington disease (2), postpartum and pregnancy (3), papers on 36-item version and general comments (12) Assessed based on titles and abstracts n=43 Excluded n=13 Assessed based on full texts n=30 Excluded n=16 Included for the data exctraction n=14 Excluded n=1 Included for the analysis n=14 1 paper published after the search