scieee AI-readable full text Open interactive document viewer

A Generalized Ordered Logit Model to Accommodate Multiple Rating Scales

Gangl, Markus

Abstract

Rating scales are ubiquitous in the social sciences, yet may present practical difficulties when response formats change over time or vary across surveys. To allow researchers to pool rating data across alternative question formats, the article provides a generalization of the ordered logit model that accommodates multiple scale formats in the measurement of a single rating construct. The resulting multiscale ordered logit model shares the interpretation as well as the proportional odds (or parallel lines) assumption with the standard ordered logit model. A further extension to relax the proportional odds assumption in the multiscale context is proposed, and the substitution of the logit with other convenient link functions is equally straightforward. The utility of the model is illustrated from an empirical analysis of the determinants of respondents’ confidence in democratic institutions that combines data from the European Social Survey, the General Social Survey, and the European and World Values Survey series.

Full text

A Generalized Ordered Logit Model to Accommodate Multiple Rating Scales Markus Gangl 1 Abstract Rating scales are ubiquitous in the social sciences, yet may present practical difficulties when response formats change over time or vary across surveys. To allow researchers to pool rating data across alternative question formats, the article provides a generalization of the ordered logit model that accommodates multiple scale formats in the measurement of a single rating construct. The resulting multiscale ordered logit model shares the interpretation as well as the proportional odds (or parallel lines) assumption with the standard ordered logit model. A further extension to relax the proportional odds assumption in the multiscale context is proposed, and the substitution of the logit with other convenient link functions is equally straightforward. The utility of the model is illustrated from an empirical analysis of the determinants of respondents’confidence in democratic institutions that combines data from the European Social Survey, the General Social Survey, and the European and World Values Survey series. Keywords ordered logit model, rating scales, data harmonization, question format, survey data, data pooling, trend analyses, cross-nationally comparative analyses 1 School of Social Sciences (FB03), Goethe University Frankfurt am Main, Frankfurt am Main, Germany Corresponding Author: Markus Gangl, Goethe University Frankfurt am Main, School of Social Sciences (FB03), TheodorW.-Adorno-Platz 6, 60629 Frankfurt am Main, Germany. Email: [email protected]rt.de Data Availability Statement included at the end of the article Original Article Sociological Methods & Research 1–40 © The Author(s) 2023 Article reuse guidelines: sagepub.com/journals-permissions DOI: 10.1177/00491241231186655 journals.sagepub.com/home/smr Introduction Rating scales are one of the epitomes of survey research. It is a rare questionnaire indeed that would not incorporate some version of a Likert scale to tap into the intensity of respondents’agreement with some opinion or statement, ask respondents to rate their happiness or satisfaction with specific domains of their lives, or that supplies respondents with some rating scale to help them express degrees of emotional bonding with particular social groups or sentiments of trust and confidence in others and in societal institutions. The communal feature of all these various forms of rating scales in survey research is that researchers are interested in capturing respondents’location on some latent dimension. This dimension may often be conceptualized as a continuum of underlying attitudes or beliefs, yet the latent dimension is lacking any natural metric, and hence different locations on the continuum may only be approximated by providing verbal or numerical cues to respondents. These cues then imply an element of gradation—as when it may be presumed that a statement of “strongly agree”corresponds to a higher degree of affirmation than “agree,”or the choice of a happiness score of 8 to convey a higher level of contentment than the choice of a score of 5—but substantial ambiguity inevitably remains as to whether respondents are sharing a reasonably common understanding of the survey stimuli, or whether and when the number of response categories might be sufficiently large and the distance between them sufficiently evenly spaced in substantive terms to permit treating the empirical observations as satisfying a metric scaling level. In empirical research, such fine-grained methodological discussions often also seem to stem from the fact that the statistical modeling of ordinal data is something like the poor relation of the standard linear or logit regression models that social scientists are extensively familiar with. Applied researchers may be aware of the ordered logit (or probit) model that is extending fundamental principles of categorical data analysis to the case of ordinally scaled dependent variables (see Long 1997:114-47; Cameron and Trivedi 2005:519-21; Agresti 2010, Wooldridge 2010:655-59; Greene 2012:82432; Hosmer, Lemeshow, and Sturdivant 2013:289-310) or of interval regression models that find application when metric data have been recorded in response categories (e.g., income brackets) rather than as point data in the original metric (e.g., Cameron and Trivedi 2005:532-35; Wooldridge 2010:783-85), but in practice still turn back to more basic models when being confronted with ordinal outcome data. It seems fair to say that most social scientists then routinely either seek to rationalize a metric interpretation, perhaps even explicitly acknowledging the approximation, in order 2Sociological Methods & Research 0(0) to proceed with using standard linear regression on their data, or resort to identifying specific cutoff points on the ordinal outcome scale from either theoretical or empirical considerations, and then use a standard binary logit (or probit) model to analyze outcomes. And in many instances, these convenience techniques will, in fact, provide pragmatic statistical solutions that result in valid and empirically informative parameter estimates, certainly when judged against conventional inferential standards and against the typical inferential targets in quantitative social science research, where researchers are typically focused on establishing the principal existence and direction of some hypothesized effect, rather than on evaluating any sharply quantified prediction on the magnitude or range of some particular effect on some well-specified metric that would be deemed observationally compatible with a researcher’s theoretical model. Yet even when often well-founded, the social scientist’s statistical pragmatism may find its limits. With rating data, an important practical difficulty arises whenever response formats change over time or when they vary systematically across different surveys. In some such cases, it might be possible to devise data harmonization rules from, for example, noting the equivalence of certain verbal cues (“agree”and “fully agree”) that are being provided to respondents to help anchor the response scale and to then analyze the data specifically at those cutoff points or response thresholds that appear being consistently captured over time or across surveys. In other cases, it is possible to achieve data pooling and joint estimation via suitable interval regression modeling, namely when the ordinal response scale may have been merely a data recording tool to either help respondents by providing them with informative categorizations of some underlying continuous metric, or to increase item response rates by permitting respondents to choose between outcome categories (e.g., income brackets) rather than having to disclose point information on the original metric scale; multiple response formats and variability in recording categories are, in fact, straightforward to handle in the interval regression routines of standard statistical packages. Yet in one important class of situations, the applied social scientist is lacking guidance and adequate statistical tools, and that is when response formats evidently differ, when verbal cues are incompatible or unavailable, and when the scaling level is genuinely ordinal at the point of data collection because the construct of interest is lacking any natural metric. This, of course, is precisely the case of the typical Likert-type survey question that asks respondents to express their degree of agreement with some particular statement, to rate their happiness and satisfaction with different domains of life, or to state their degree of closeness and attachment to Gangl 3 some community, political party or organization, where the same question has been asked repeatedly over time or in different countries and places, but where the precise response format of the question may have changed over time or may have varied across surveys and locations. In one survey, respondents may have been asked to rate their happiness on a scale between 1 and 7, the next survey provides a scale from 0 to 10, and yet another may use four response categories that have explicit verbal labels (“excellent,”“good,”“satisfactory,”etc.) attached to them. Or respondents may have been asked to state their degree of confidence in some public institutions, yet the survey was initially using a 5-point Likert scale, then switched to an 11-point scale in some later wave, and yet another version of the questionnaire may have experimented with using a small number of verbal cues as response categories. In these and other similar constellations that frequently arise in survey research, it would be attractive to be able to pool and analyze the rating data across question formats in order to either simply increase statistical power or, perhaps more importantly, to obtain the required leverage to address broader substantive questions on, for example, historical changes or cross-country differences in outcomes and processes that cannot be addressed by using the original data sources in isolation. Yet because the rating scales in question lack any natural metric, data pooling may seem impossible or may at least seem to require that researchers be prepared to accept an inevitable degree of arbitrariness in whatever data harmonization rules they may choose to adopt. As a more principled alternative, however, it is also possible to generalize the standard ordered logit model to accommodate the presence of multiple rating scales that capture a common latent index, and to thereby resolve the apparent incommensurability of alternative response formats. In the remainder of this article, I present and discuss the resulting multiscale ordered logit model, and then illustrate its practical utility in an empirical analysis of the relationship between income inequality and citizens’trust in democratic institutions that draws on survey data from the European Social Survey (ESS), the General Social Survey (GSS), and the European and World Values Survey (EVS/WVS) series. A Generalized Ordered Logit Model to Accommodate Multiple Rating Scales Although various alternatives exist (see Agresti 2010:44-117; Fullerton 2009; Fullerton and Xu 2016; Hosmer et al. 2013:289-310), it is the so-called proportional odds model that is conventionally seen as the standard logit model for ordered outcome data (e.g., Cameron and Trivedi 2005; Clogg and 4Sociological Methods & Research 0(0) Shihadeh 1994; Long 1997; McCullagh 1980; Williams 2006, 2016). Besides retaining the straightforward interpretation and other features of the wellknown logit model for binary outcome data, the proportional odds model rests on conceptual foundations that align with the typical use of rating scales in social science surveys, and it will therefore also serve as the natural starting point for the proposed extension to a multiscale version of the model that is capable of accommodating the presence of multiple rating scale formats in the data at hand. The proportional odds model itself may be conveniently written as the threshold model Pr(Yi>j)=exp (αj+Xiβ) 1+exp (αj+Xiβ)for j=1,2,...,k−1 (1) to predict the probability that the observed response for respondent iis higher than the response category jon a rating (or otherwise ordered) scale consisting of kcategories (Williams 2006, 2016). 1 This probability to cross threshold jis modeled as a standard logit function of a covariate vector Xi, a coefficient vector β, and a set of threshold parameters or cutpoints αj. In this setup, the defining feature of the proportional odds model is that the structural component Xiβto describe the association between respondent characteristics Xiand outcomes is assumed to be exactly the same at each of the cutpoints αjdefined by adjacent response categories. When writing out the model for the conditional odds of observing respondents with characteristics Xiin a response category higher than jas Ω(Yi>j)=Pr(Yi>j|X) Pr(Yi≤j|X)= exp (αj+Xiβ) 1+exp (αj+Xiβ) 1 1+exp (αj+Xiβ) =exp (αj+Xiβ) for j =1,2,...,k−1,(2) these turn out to be exactly proportional at each cutpoint location exp (αj). By implication, the effect of a change (or difference) in any particular covariate x may then be given by the simple expression for the odds ratio ORj= Ω(Yi>j|x,xi+Δx) Ω(Yi>j|x,xi)=exp (αj+(xi+Δx)β) exp (αj+xiβ)=exp (Δx×β) for all j, (3) which is independent of the particular threshold jat which it is evaluated. In the proportional odds model, the same shift in a covariate xin other words Gangl 5 implies the exact same proportionate shift in the odds of crossing a particular response threshold, irrespective of which specific response threshold jis being considered. The exact same features may also be described in terms of positing a linear model y∗ i=α+Xiβ+ε(4) for a latent (continuous) index variable y∗ ithat is underlying the imperfect empirical observations available from respondents’choice of response category yi=jwhen αj−1≤y∗ i<αjfor j=1,2,...,k,α0=−∞ and αk=+∞.(5) The linear form of the structural model (4) implies that the same regression plane Xiβgets shifted from cutpoint to cutpoint intercepts αj, and hence the proportional odds model may also be characterized by pointing out the respective parallel regression (or parallel lines) assumption that it implies (see Agresti 2010:53-58; Fullerton and Xu 2016:9-10, 21-24; Long 1997:140-45; Williams 2006, 2016; Wooldridge 2010:658-59 for details). As a modeling device, the parallel regression assumption has the powerful implication that the structural component Xiβis independent of the precise response format of the rating scale employed to capture the underlying (continuous) index variable (see Agresti 2010:56; Long 1997:117-19). This feature is indeed the foundation for the proposed extension to a multiscale ordered logit model, yet at the same time the parallel regression assumption also tends to be seen as overly restrictive (e.g., Williams 2006, 2016), and in practice amounts to the main reason why the ordered logit model has a rather mixed reputation among social scientists. When confronted with textbook warnings along the lines of My experience suggests that the parallel regression assumption is frequently violated …When the assumption of parallel regressions is rejected, alternative models should be considered that do not impose the constraint of parallel regressions (Long 1997:145), A key problem with the parallel-lines model is that its assumptions are often violated; it is common for one or more β’s to differ across values of j; i.e., the parallel-lines model is overly restrictive (Williams 2006:60) and 6Sociological Methods & Research 0(0) The use of an ordered logit model when its assumptions are violated creates a misleading impression of how the outcome and explanatory variables are related (Williams 2016:11), empirical researchers cannot be faulted for coming to the conclusion that, at least in its standard proportional odds form, the ordered logit model is hardly worth their attention. Upon closer inspection, this widespread attitude as well as the implicit conflation of the two notions of “violated assumption”and “flawed model” it rests upon are quite misplaced, however. As can be seen from the latent variable formulation of the proportional odds model in equation (4), the effect of any covariate ximplies nothing but a location shift along the underlying attitude or rating continuum. Equation (4), in other words, is thus nothing else than a categorical data analog of the standard modeling assumption made by any researcher who decides to fit a linear ordinary least squares (OLS) regression. Like in any OLS regression, the proportional odds model can thus be understood as a regression model for the central tendency of the latent outcome distribution, except that the outcome Y∗, in this case, is not directly observed, does not have any natural metric to guide the substantive interpretation, and that model identification requires the error term to be set to the logit distribution with a mean of zero and a variance of π2/3 (or, alternatively, to the standard normal distribution if estimating an ordered probit model is being desired, see Cameron and Trivedi 2005; Long 1997; Wooldridge 2010). But when seen from this perspective, the actual meaning of the much-touted “violations”of the parallel regression assumptions becomes clearer, too: when the parallel regression assumption is violated in an empirical analysis, this is a direct indication that the association between covariates xand the mean of the conditional outcome distribution does not exhaust the statistically detectable signals in the empirical data. The “violation”of the parallel regression assumption therefore merely implies that more can be learned from the data if the researcher decides to examine different parts of the outcome distribution instead of just focusing on its central tendency. But this then is nothing like any inherent “failure”of the ordered logit model, and it clearly is something very different from the model being seen as “misrepresenting”how the outcome and explanatory variables are related. The proportional odds model (often) “misrepresents”the data in the exact same way that an OLS regression is “misrepresenting”it: both models focus on the central tendency of the outcome distribution and provide a linear regression model for it. Sometimes this is exactly what a Gangl 7 researcher wishes for, because she may have a hypothesis to test on some average group difference in outcomes. In other cases, the researcher might have more encompassing descriptive interests or her hypothesis might be more complex because it relates (also) to a group difference in the variance of the outcome distribution or specifically to a group difference in one of the tails of the outcome distribution, and then standard OLS would be inadequate and the researcher would better turn to more appropriate (conditional) quantile regression models (e.g., Koenker 2005; Koenker, Chernozhukov, He, et al. 2020), to (co)variance function regression techniques (e.g., Bloome and Schrage 2021; Western and Bloome 2009) or to other types of location-scale models (e.g., Hedeker and Nordgren 2013; Leckie, French, Charlton, et al. 2014). But in neither case would anyone ever consider faulting the OLS regression model for principally “creat[ing] a misleading impression of how the outcome and explanatory variables are related”(Williams 2016:11). Instead, one would simply note that some inferential task is beyond standard OLS regression, and then apply one of the readily available extensions of the basic model in the empirical analysis. Tellingly, the surging interest in examining various types of heterogeneities in the relationships between purported causes and effects has been accompanied by a very visible increase in the use of quantile regression and related models to ascertain not just the association between a covariate and the mean outcome, but also group differences in the shape (or variance) of the entire outcome distribution (e.g., Cheng 2014; Ebner, Kühhirt, and Lersch 2020; Lersch, Schulz, and Leckie 2020; VanHeuvelen 2018a, 2018b). Seen in this light, what is usually perceived as a disadvantage and an “overly restrictive”nature of the ordered logit model (and its probit cousin), is actually a powerful feature of key interest to substantive research. For the specific case of ordinally recorded outcome data that can be understood as an imperfect measure of some underlying continuous index, that is for the typical case of rating and other attitude data common in survey research, the proportional odds model is an elegant approach to model the central tendency of the outcome variable conditional on covariates and the assumption of a linear regression function. It is thus an analog to the standard OLS regression, except that there also is a generalized ordered logit model to relax the parallel regression assumption when required (see Fullerton and Xu 2016; Long 1997; Williams 2006, 2016), and several formal statistical tests are available to ascertain whether employing the generalized model may be indicated by some systematic signal in the empirical data (again, see Long 1997; Fullerton and Xu 2016:109ff.; Williams 2006, 2016; Brant 1990 for the well-known specification test). But the relation between the proportional 8Sociological Methods & Research 0(0) odds model and the generalized ordered logit model certainly is not one between a “failure”and an “appropriate”model, but instead between a model that exclusively focuses on establishing group differences in the location of the conditional outcome distribution and an alternative model that simultaneously addresses group differences in the central tendency and in the spread (i.e., in the shape as well as the location) of the distribution. Clearly, the suitability of choosing one over the other is not a principal matter, but one of the research priorities, substantive questions, and specific hypotheses—and unlike in the case of standard OLS regression, the ordered logit model even provides a unified framework for conducting either type of analysis. Against this background, it may have become plausible why, despite much textbook criticism, the proportional odds model is nevertheless taken as the starting point for proposing a natural extension of the ordered logit model to a multiscale setting. Indeed, it is precisely because of the assumption of parallel regressions and the associated equivalence of the model’s structural component Xiβacross any and all response thresholds recorded in the actual survey instrument or data collection effort that such an extension is readily accomplished. When the same structural component Xiβcan be assumed to govern reporting behavior (i.e., to correctly describe the association between covariates Xand observable outcomes) at each observable response category, then the exact question format of the rating scale is irrelevant and the empirical parameter estimates will not depend on the exact number or verbal cueing of response categories utilized in data collection (see Agresti 2010:56; Long 1997:117-19). But if that is the case, then the validity of the model will also not be affected if observations are being pooled across two or more surveys (or survey waves) employing somewhat different versions of a rating scale sto measure some latent index Y∗. In other words, one may obtain a workhorse for situations where data pooling would be desirable for addressing substantive questions on, for example, over-time change or cross-country differences in rating patterns by specifying the multiscale proportional odds model Pr(Yi>js|si=s)=exp (αjs+Xiβ) 1+exp (αjs+Xiβ) for js=1,2,...,ks−1 and s=1,2,...,m.(6) As with other models from the family, this multiscale model may be estimated by maximum likelihood and retains the straightforward interpretation as well as all other features of the standard ordered logit model, except for the fact Gangl 9 theoretically grounded choice of controls would be required to seriously aim at anything more. From these preludes, it is easy to summarize the substantive evidence from the multiscale regression specification M1 as indicating that macroeconomic context as well as citizens’socio-demographics matter for trust in parliament. More specifically, the effect of GDP/capita on trust is positive, while high levels of economic inequality clearly depress citizens’ trust in a core democratic institution like the national parliament. On the micro level, gender differences between male and female citizens tend to be very small on average, but higher levels of education lead to clearly higher levels of trust in the parliamentary institutions of democratic governance. The age effects indicate a U-shaped pattern of association with democratic trust, with the lowest levels of trust being found among citizens in their mid-forties, ceteris paribus. 3 And of course, as is true in any crosssectional sample, what is reported as an age effect here is likely to reflect some mixture of true life-cycle and true cohort effects, but the fundamental identification problem at the heart of any age-period-cohort (APC) model, of course, does not allow to undertake any empirically grounded attempt to distinguish between and quantify the relative importance of either temporal source of political trust—nor would any such attempt be required in the present context of a purely descriptive and associational analysis done for demonstration purpose. Instead, the characteristic achievement of the multiscale ordered logit model emerges when comparing model M1 against some more standard alternatives. Absent the multiscale specification, it would of course have been possible to fit the standard, single-scale ordered logit model on the data, or at least on those parts of the sample that originate from the same source survey and therefore share the same question format in data collection. Estimates from respective specifications are provided as models M3–M7 in Table 1, each fitting a standard ordered logit model on data from one of the original survey sources or, in the last specification M7, on pooled EVS/WVS data that share the same response format for the political trust question. When eyeballing the parameter estimates across the different models, it is evident that some quite significant heterogeneity is apparent in the determinants of citizens’trust across surveys and localities—and that the estimates obtained in the multiscale specification M1 provide something like the average over the different source surveys and over the whole sample of respondents. The effect of education, to take one example, is positive in the European data, but more so in the ESS sample than in the EVS one, but negative in the WVS sample, and quite negative in the GSS—and the multiscale parameter 16 Sociological Methods & Research 0(0) estimate of β=0.059 something like a reasonable estimate for the average effect of education on democratic trust for an overall sample that is dominated by European data. Similarly, at the macro level, there is evidence of a clear positive association between GDP per capita and democratic trust in the ESS data, a mildly positive, but non-significant association in the EVS sample, and a mildly negative, but non-significant association in the WVS sample—and then a mildly positive, but non-significant association of β= 0.156 as reported in the multiscale estimate seems like a good estimate of the average relationship across all the countries in the sample. In addition, virtually all parameters are more precisely estimated, that is, their standard errors are lower, in the multiscale specification M1 relative to its alternatives because the advantage of the larger (pooled) sample it is able to employ evidently outweighs any increase in variation that stems from pooling empirically heterogeneous data. And this exact same pattern gets repeated if one was to compare the multiscale estimates based on the two European sources (i.e., model M2) to those obtained from fitting standard ordered logit models on the ESS and EVS source data separately (i.e., to models M3 and M4)—and of course this is precisely the behavior that is to be expected from any regression model. It is a very basic regression methodology to understand that any regression coefficient reflects the (weighted) average association between Xand Yamong the sample observations, and so fitting a regression model on some pooled data will inevitably result in parameter estimates that represent the weighted average of the corresponding coefficients from the series of identical models fitted on separate (and non-overlapping) partial datasets in isolation, and these parameter estimates will typically be more precisely estimated (i.e., exhibit lower standard errors) because of the larger sample brought to the task. And it is in this exact sense that the proposed multiscale specification of the ordered logit model is shown to “work” as it should by the evidence in Table 1. It is a regression model that allows to pool rating data and that provides an estimate of the (weighted) average association between Xand Yin the full sample, despite differences in response formats across source surveys. 4 The multiscale model, in other words, substitutes a principled statistical model for any informal eyeballing that a researcher otherwise might wish to execute when trying to summarize rating scale evidence obtained from different samples and across different question formats. The benefits of adopting this type of principled approach should also be self-evident when comparing the multiscale ordered logit model to some more traditional convenience alternatives that are often being adopted to Gangl 17 avoid the ordered logit model altogether. Table 2 provides some examples, and thereby helps further illustrate their downsides relative to the main multiscale model (i.e., model M1 in Table 1). In applied research, social scientists often use standard OLS on rating data, and defend the practice by noting that substantive results more often than not tend to align with those of the ordered logit model. The same pattern is evident in the current analyses, as the linear regressions of M1 and M2 in Table 2 effectively mirror those of the corresponding ordered logit models M3 and M7 in Table 1 as far as the direction and statistical significance of the different effects are concerned. 5 But unlike with the ordered logit model, standard linear regression does not offer any constructive way forward when data pooling across different response formats may be desired. It is of course possible to adopt a rule-of-thumb harmonization protocol by, for example, distributing the four EVS/WVS response categories “evenly”across the 11-category ESS format in order to achieve data integration across these series, but whether that rule may have some empirical foundation or whether this amounts to a forced data pooling based on an entirely arbitrary methodological choice cannot adequately be decided. Similarly, it would in principle be possible to define specific thresholds of the outcome variable that are of particular interest and then fit a standard binary logistic regression on the data, this approach would provide for a more intuitive interpretation of the resulting parameter estimates that is preferred by many social scientists over the linear model or the latent variable interpretation of the ordered logit model, but it would also not provide a constructive way forward to achieve valid data harmonization. Models M3–M6 are illustrations of the point, as these report the estimates from two logit model specifications that focus on the lower and the high end of the trust distribution, respectively, and that have each been fitted separately on the ESS and EVS/WVS data. These estimates on the one hand reflect similar substantive differences between the determinants of trust in the ESS and EVS/WVS samples and provide some empirical indications that associations between covariates and trust may indeed not be constant across the entire outcome distribution on the other—a topic to which I return in the next section—, but do not permit to answer the key question about the validity of pooling the data. Is, for the lower-tail models M3 and M4, a cutoff value of 4 on the ESS scale a good equivalent to respondents stating to have “not very much”trust on the EVS/WVS item, so that pooled analysis would be defensible? Is the ESS cutoff value of 8 about the same high level of political trust as expressed by EVS/WVS respondents who state having “a great deal of”trust in parliament? 18 Sociological Methods & Research 0(0) Compared with these unanswerable questions, it may be instructive to consider how the multiscale model addresses the comparability issue by effectively sidestepping it. Seen from the starting point of a latent continuous variable that is being imperfectly observed via the (ordered) categories of some particular rating scale employed in some specific survey, the multiscale model is nothing but an extension of the standard ordered logit model that allows for the presence of multiple sets of scale location points in estimation, with one set of cutoff points corresponding to each type of question format. The empirical locations of the different cutoff points αjsare in fact being estimated as parameters of the model, and may therefore usefully be compared across response formats in order to assess what might be seen as the data harmonization rule that is implicit in the model and consistent with the empirical data. To continue the empirical example, Figure 1 provides the cutpoint locations αjsthat have been estimated for the ESS, EVS/WVS, and GSS scales in the main multiscale Figure 1. Estimated cutpoint locations for the ESS, EVS/WVS, and GSS response formats to express trust in the national parliament. Note: Inverted cutpoint estimates −αjsfrom model specification M1 in Table 1 (multiscale ordered logit model, full ESS–EVS–GSS–WVS sample). ESS =European Social Survey; GSS =General Social Survey; EVS/WVS =European and World Values Survey. Gangl 19 model (M1 in Table 1), respectively; for easier reading, Figure 1 actually displays the inverted location parameters −αjsas these correspond to a natural ordering of response categories in terms of increasing item “difficulty,”that is, to increasingly positive expressions of trust. 6 From these, it is readily apparent how the multiscale specification is implicitly answering the earlier rhetorical questions: first, on the high end, the ESS cutoff value of 8 indeed seems to index pretty much the same intensity of trust as the verbal stimulus of “a great deal”of trust in the EVS/WVSs. But, second, on the low end, the ESS cutoff value of 4 does not seem to correspond to the EVS/WVS’s verbal stimulus of having “not very much”trust. Instead, it rather is the ESS cutoff value of 3 that matches the EVS/WVS location of having “not very much”trust quite well, and Figure 1 then also suggests that the EVS/WVS category of expressing “quite a lot”of trust does not have its ready ESS equivalent, but is sitting somewhat uneasily between values 5 and 6 on the ESS scale. But that said, it is also important not to mistake the evidence of Figure 1 for a suggestion of some substantive and empirically-grounded harmonization rule that might or that even should have been adopted by the researcher. Instead, the multiscale model is better characterized as sidestepping the question of any substantive equivalence of response categories across different rating scales by combining a methodologically entirely relativist position on the “meaning”of any single response category with an additive model where the same structural component Xiβto describe the associations between covariates Xand outcomes Ygets shifted across successive cutpoint locations αjsalong the distribution of the latent outcome. The best practical illustration for this point comes once again from considering the (strategically chosen) addition of the GSS question format as the third rating scale to be integrated into the full multiscale model M1 in Table 1. Here, it does surprise the human researcher to see that the exact same verbal EVS/WVS and GSS stimuli of “a great deal”of trust are not being placed on par with each other in terms of their cutpoint locations α, but to see the EVS/WVS stimulus correspond to a cutoff value of 8 on the ESS scale, and the GSS category more to an ESS value of 9. Likewise, human researchers would probably not have equated the GSS stimulus of having “only some”trust with the ESS middle scale value of 5, and would have expected it to lie somewhere in between the EVS/WVS categories of having “not very much”and “quite a lot”of trust, but not quite to be almost the same as having “quite a lot”of trust as the empirical data seems to have it. Yet of course, these estimated cutpoint locations αdo not represent the outcome of any linguistic or substantive validation, but instead, merely 20 Sociological Methods & Research 0(0) reflect the empirical reality of the conditional outcome distribution as observed via and anchored in some particular rating scale format. The fact that the locations of the GSS stimulus “a great deal”of trust and its EVS/ WVS equivalent do not match, does not imply that the same words “mean” different things to respondents in different surveys and different countries in any substantive sense. It first and foremost means that the probability distributions differ in the sense that the share of GSS respondents who see themselves as having “a great deal”of trust is empirically smaller than the corresponding share of EVS/WVS respondents, of course averaging across all countries in the EVS/WVS sample and conditional on covariates in both cases. This relativist perspective sidesteps the question of whether the difference is substantive or methodological, that is, does not help decide whether GSS respondents differ from their EVS/WVS counterparts because Americans are truly showing less confidence in Congress than are citizens of other EVS/WVS countries in their national parliaments or because it truly is the case that the stimulus of “a great deal”of trust may indeed convey different intensities of trust in the mind of U.S. respondents relative to respondents from other countries. In the context of this current and slightly artificial analysis, it would thus not be possible to turn to the multiscale ordered logit model to answer the substantive question of whether U.S. citizens are less trustful of democratic institutions than the citizens of other countries, because the demonstration exercise has been strategically chosen to involve a perfect correlation between country and response format in the U.S. case. Hence, there is no extra degree of freedom available to estimate a country-fixed effect and the cutoff locations αfor the GSS response categories simultaneously, and all respective variation would, in fact, be attributed solely to the latter and hence become treated as a methodological nuisance parameter that is of little substantive interest to the analyst. But this, in turn, is not a bug, but indeed a feature and a decisive advantage of relying on the multiscale model: a decision on the issue of what the various response categories may “mean”and how they may compare across countries and question formats are not required at all in order to move on and permit the social scientist to use the pooled sample and to evaluate the associations between covariates Xand outcomes Ythat are of genuine interest and a matter of theoretical reflection. The elegance of the model is that it sidesteps a problem that has been plaguing (comparative) survey researchers for decades and that may ultimately prove to be intractable in some respects, but that does not actually require a solution. And of course, it is a model that would allow us to give an empirical answer to the question of whether U.S. citizens are less trustful of Gangl 21 democratic institutions than the citizens of other countries or not. In any realworld analysis, one would of course not insist on using 2018 data exclusively, one would make sure to incorporate the 2017 U.S. sample from the WVS series, and one would thereby have broken what has been a perfect correlation between country and survey instrument in the artificial setup of the present exercise. Relaxing the Parallel Regression Assumption in the Multiscale Ordered Logit Model At this point, readers may agree with the perspective that it is possible to extend the ordered logit model to a multiscale setting, while retaining a modeling framework that is well-known to social scientists and that affords flexible ways of interpreting the resulting parameter estimates either in terms of covariate effects on an underlying latent and continuous outcome or in terms of odds ratios or probability differentials of crossing specific response thresholds. At the same time, readers may likewise feel the multiscale model to still be overly restrictive insofar as it of course also shares the critical parallel regression assumption with the standard ordered logit model. And even as some criticism of that assumption may itself be rather regarded as being based on a misapprehension, there is merit in the principal insistence on methodologies that permit researchers to adequately examine issues of dispersion and (treatment) effect heterogeneity over and above central tendencies of the outcome distribution and the association between covariates and conditional mean outcomes that are the mainstay of standard regression models including the proportional odds ordered logit model. Respective interest in relaxing the parallel regression assumption may also be justified on purely empirical grounds, and, in fact, even in a somewhat artificial and restrictive setting like that of the present analysis. With dichotomous outcome measures to reflect, respectively, particularly high and low levels of political trust among citizens, a comparison of parameter estimates between ESSand EVS/WVS-based logit models M3–M6 in Table 2 suggests that several relationships may, in fact, vary systematically across the outcome distribution, like the effect of gender and GDP per capita in the ESS data, or the effect of the Gini coefficient in the EVS/WVS sample. And in principle, it is actually straightforward to address such concerns and to relax the parallel regression assumption when required. As the multiscale model is derivative of the standard ordered logit model, it also inherits the principal approach toward its generalization. Specifically, and exactly as 22 Sociological Methods & Research 0(0) with the conventional model, the natural specification of a generalized multiscale ordered logit model is Pr(Yi>js|si=s)=exp (αjs+Xiβjs) 1+exp (αjs+Xiβjs) for js=1,2,...,ks−1 and s=1,2,...,m,(8) where the constant covariate vector βthat generates the parallel regression planes in the standard model has been replaced by a cutpoint-specific covariate vector βjsthat allows associations between covariates Xand outcomes Yto freely vary at each observable response threshold (also see Fu 1998). The evident downside of this very general specification is that, as Williams (2006, 2016) correctly remarked, it may involve estimating many eventually superfluous parameters because some parameters may in fact be constant across the whole outcome distribution or at least show some more limited variation in certain parts of the distribution only, and then many parameters of the fully generalized model may, in fact, not be necessary as they are not (statistically significantly) different from each other. This consideration has brought Williams (2006, 2016) to proposing a partial proportional odds model that seeks to identify the exact minimal set of parameters that is required to describe the empirical structure of associations in a parsimonious and exhaustive way, while avoiding to estimate and reporting statistically superfluous coefficients. While this procedure has evident statistical merit insofar as it seeks to exhaust the data signal by finding a maximally parsimonious model specification to fully capture and describe it, it is also possible to approach the issue of generalizing the ordered logit model less from a data-analytic and more from a subject-matter perspective. And without intending to deny the value of Williams’(2006, 2016) alternative, this will be the approach taken here. Then, from a subject-matter perspective, it often seems less relevant to be able to fully and efficiently characterize all systematic patterns that may be apparent in the empirical data, but generalizing from the standard proportional odds formulation seems warranted whenever researchers may wish to evaluate hypotheses that extend beyond expectations about (conditional) group differences in average outcome levels. A typical case would seem to be that social scientists may harbor expectations about how some factor X would affect the shape of the outcome distribution over and above any upward or downward shift of the overall distribution that could be detected by examining conditional mean outcomes. Applied to the case at hand, one might reason that some covariates may be particularly relevant for protecting Gangl 23 citizens from disenchantment with democratic institutions, and in such cases, one would expect to observe stronger associations between these particular covariates Xand outcomes Yin the lower tail of the outcome distribution specifically, but weaker or perhaps even no statistical associations further up. As one concrete example, adequate macroeconomic performance has often been considered a necessary condition of democratic legitimacy in classical works in political sociology or among students of the history of democratic societies (e.g., Lipset 1959, 1960, 2004). Translated into statistical terms, this could be read as indicating the expectation that macroeconomic conditions mostly affect the lower tail of the trust distribution, that is, may be considered particularly decisive for determining whether someone accords at least some basic degree of confidence to the institutions of democratic governance, but may have fewer if any implications for whether someone may be expressing to trust some particular institution either “usually”or “almost all of the time.”The substantive hypothesis in this case would imply a mean shift—trust in democratic institutions would be expected to be generally higher when macroeconomic environments are good than during a recession—but even more clearly it would involve an expectation about changes in the shape of the outcome distribution. Specifically, this consideration from classical political sociology would suggest that the variance of the outcome distribution increases during macroeconomic crises (or, as the flip side of the coin, that the trust distribution is relatively more compressed under normal times), and that the increased variance comes about as the lower tail of the distribution fanning out because a certain share of the citizenry loses basic faith in the institutions of democratic governance under economic distress. To test substantive hypotheses like these, neither the fully generalized ordered logit model nor Williams’(2006, 2016) partial proportional odds model would seem to fully meet the interests of substantively-minded social scientists. The fully generalized model clearly risks to provide excessive statistical detail, whereas the partial proportional odds model is setting statistical and data-analytic priorities rather than primarily substantive ones. Against that background, another type of generalization is to specify a generalized (and of course multiscale) ordered logit model in the form of Pr(Yi>js|si=s)=exp (αjs+Xiβr) 1+exp (αjs+Xiβr) for js=1,2,...,ks−1,s=1,2,...,m, and r={js|αjs≤c1},{js|c1<αjs≤c2},...,{js|αjs〉ct} (9) 24 Sociological Methods & Research 0(0) Here, a generalization from the parallel regression model has occurred insofar as βris no longer constant across the entire outcome distribution, but is allowed to vary systematically across tuples rof cutpoints js. When tuples are defined (across scale formats) by cutpoint locations αjslying within some prespecified range cr−1<αjs≤crof the latent outcome distribution, it becomes possible to examine effect heterogeneity in βracross different zones of the outcome distribution. One obvious example would be to define asinglecutoffctowards the lower tail of the outcome distribution, and then evaluate whether covariate vectors β{js|αjs≤c}and β{js|αjs〉c}systematically differ in at least some of their elements in order to determine whether some covariates may indeed be particularly decisive at the lower end of the outcome distribution, that is, to prevent citizens’disenchantment with democratic governance in the concrete example used in the present illustration. The Empirical Example Continued: Are There Any Asymmetries in the Effects of Covariates on Citizens’Trust in Parliament? To fix ideas, it seems straightforward to continue the earlier example, and to now examine whether some covariates may indeed show asymmetries in their association with citizens’trust in the national parliament, and if so, which covariates might be important to avoid citizens turning away from the institutions of democratic governance. As before, the actual model to be estimated will be the multilevel extension Pr(Yip >js|si=s)=exp (αjs+Xiβr+up) 1+exp (αjs+Xiβr+up) for js=1,2,...,ks−1,s=1,2,...,m,p=1,2,...,q and r={js|αjs≤c1},{js|c1<αjs≤c2},...,{js|αjs〉ct} (10) of the generalized multiscale ordered logit model described in equation (9), and this model once again merely adds a context-level (country-survey wave) random effect upto reflect the hierarchical structure of the data. The resulting parameter estimates are provided in Table 3, which has the results from a simpler specification that contrasts lower-tail behavior to the associations observed at higher levels of the trust spectrum (model M1) and a second set of estimates from an expanded model (M2) that contrasts effects in the lower tail, the middle and the upper-tail of the distribution. In these generalized model specifications, I work with a cutoff of clow =logit(0.33) to define the lower tail of the distribution and chigh =logit(0.75) to define the upper Gangl 25 Seen from an SEM perspective, the multiscale ordered logit model is thus able to overcome the problem of non-commensurability in response scales precisely because the multiscale model does not require a measurement model at all, but simply reflects all response scale cutoffs exactly as they exist in the data at hand. Put differently, the multiscale model, like any standard regression model, takes the manifestly observed indicator as the outcome of interest—and then generalizes from the traditional ordered logit model by allowing for the presence of different response formats for the same rating outcome in the data. This key difference in perspectives between the multiscale model and an SEM approach can also be expressed in more substantive terms: for the running example of citizens’trust in democratic institutions, the SEM perspective is that an available set of trust statements (in government, parliament, parties, etc.) should be conceived of as observable indicators from which to distill citizens’underlying latent disposition towards democratic governance. The single-equation regression perspective is appropriate if substantive interest centers on the manifest response instead (because one really is interested in the determinants of trust in parliament specifically), and then the multiscale model provides a way forward in all those cases when the rating outcome of interest may have been collected in different response formats over time or in different primary data collection efforts. 3. All quantitative covariates enter the model in grand mean-centered form. Respondents’mean age is close to 48 years in the analysis sample. Moreover, all substantive interpretation is kept deliberately colloquial and illustrative in the present context. The methodological literature on the proper interpretation of nonlinear probability models and on the closely related issues of comparing logit and probit coefficients across groups and model specifications and of interpreting interaction terms in nonlinear probability models is literally filling volumes, and is generally concluding that reporting average marginal effects on the probability scale is the preferred metric in all these cases and models. For the purposes of the present paper, however, it should be sufficient to note that the multiscale ordered logit model of course also shares all respective features, possibilities and issues of interpretation with the whole logit family of regression models. For further background, I refer interested readers to Allison (1999), Williams (2009), Mood (2010), Karlson, Holm, and Breen (2012), Breen, Karlson, and Holm (2018), Mize, Doan, and Long (2019), and to the advanced textbook literature in the field. 4. By the same token, the multiscale specification of course allows to examine patterns of effect heterogeneity in greater detail. That some substantial degree of effect heterogeneity is present in the analysis sample is evident from Table 1 alone, and like in any standard (multilevel) regression model, one could expand on the simplistic specification adopted in this demonstration by, for example, introducing random coefficients and by then systematically considering cross-level interaction terms between the macroeconomic and the respondent-level covariates of the model. Even as this 32 Sociological Methods & Research 0(0) point is not specifically demonstrated in the empirical analysis here, it should be selfevident that increasing analysts’leverage to address and formally test for effect heterogeneity across contexts is another direct benefit of the proposed multiscale specification relative to more traditional models. 5. Given the GSS three-category response format, I do not present any separate linear regression modeling of the GSS data. Although estimation is certainly feasible, the required assumption of equidistance between categories seems to lack rather principal plausibility due to the small number of response categories in the GSS. 6. As is well known, two alternative formulations of the proportional odds model exist that are substantively entirely equivalent, but differ in the sign of the location parameters α(see Long 1997: 122-124). For didactical purposes, I prefer to build the ordered logit model from the positive probability of respondents crossing any particular response threshold, but which then suggests to invert location parameter estimates for easier interpretation. Alternatively, one could have built the model based on the probability of respondents staying below some response threshold with their recorded answer, and thereby arrive at the same set of parameters directly. 7. For a more nuanced interpretation, it should be added that the Gini coefficient is the only covariate where the null hypothesis of parallel regression lines (i.e. homogeneous effects) cannot be rejected for the data at hand. Seen from a purely statistical point of view, the effect of income inequality on trust in parliament could therefore be said to be commensurable with the standard proportional odds specification of the ordered logit model. A closer inspection of the coefficient estimates reveals, however, that the effect of income inequality is clearly negative in the lower tail of the trust distribution, but then decreasing in magnitude higher up in the distribution and even turning statistically non-significant in the upper tail of the distribution (model M2, rightmost column). From a substantive perspective, it thus seems legitimate to conclude that, as theoretically expected, some degree of effect heterogeneity is present in the data, but also that a somewhat larger sample size is likely required at the contextual level in order to obtain sufficient statistical power to establish the difference of respective coefficient estimates in formal hypothesis tests. References Agresti, Alan. 2010. Analysis of Ordinal Categorical Data. 2nd edition. Hoboken, NJ: Wiley. Allison, Paul D. 1999. “Comparing Logit and Probit Coefficients Across Groups.” Sociological Methods & Research 28(2):186–208. doi: 10.1177/004912419902 8002003 Bloome, Deirdre and Daniel Schrage. 2021. “Covariance Regression Models for Studying Treatment Effect Heterogeneity Across One or More Outcomes: Gangl 33 Understanding How Treatments Shape Inequality.”Sociological Methods & Research 50(3):1034–72. doi: 10.1177/0049124119882449 Brant, Rollin. 1990. “Assessing Proportionality in the Proportional Odds Model for Ordinal Logistic Regression.”Biometrics 46(4):1171–78. doi: 10.2307/2532457 Breen, Richard, Kristian B. Karlson, and Anders Holm. 2018. “Interpreting and Understanding Logits, Probits, and Other Nonlinear Probability Models.” Annual Review of Sociology 44:39–54. doi: 10.1146/annurev-soc-073117-041429 Cameron, A. Colin and Pravin K. Trivedi. 2005. Microeconometrics: Methods and Applications. Cambridge: Cambridge University Press. Cheng, Siwei. 2014. “A Life Course Trajectory Framework for Understanding the Intracohort Pattern of Wage Inequality.”American Journal of Sociology 120(3):633–700. doi: 10.1086/679103 Clogg, Clifford C. and Edward S. Shihadeh. 1994. Statistical Models for Ordinal Variables. Thousand Oaks, CA: Sage. Ebner, Christian, Michael Kühhirt, and Philipp Lersch. 2020. “Cohort Changes in the Level and Dispersion of Gender Ideology After German Reunification: Results From a Natural Experiment.”European Sociological Review 36(5):814–28. doi: 10.1093/esr/jcaa015 European Social Survey. 2018-2021. “European Social Survey Rounds 1-9 [Machine-Readable Data Files].”Bergen: NSD - Norwegian Centre for Research Data. European Values Study. 1981-2017. “European Values Study 1981-2017 [MachineReadable Data Files].”Cologne: GESIS Data Archive. Fu, Vincent Kang. 1998. “Sg88: Estimating Generalized Ordered Logit Models.” Stata Technical Bulletin 8:160–64. Fullerton, Andrew S. 2009. “A Conceptual Framework for Ordered Logistic Regression Models.”Sociological Methods & Research 38(2):306–47. doi: 10. 1177/0049124109346162 Fullerton, Andrew S. and Jun Xu. 2016. Ordered Regression Models: Parallel, Partial, and Non-Parallel Alternatives. Boca Raton, FL: Chapman and Hall/CRC. Greene, William H. 2012. Econometric Analysis. 7th edition. Boston, MA: Pearson. Hedeker, Donald and Rachel Nordgren. 2013. “Mixregls: A Program for Mixed-Effects Location Scale Analysis.”Journal of Statistical Software 52(12):1–38. doi: 10.18637/jss.v052.i12 Hosmer, David W., Stanley Lemeshow, and Rodney X. Sturdivant. 2013. Applied Logistic Regression. 3rd edition. Hoboken, NJ: Wiley. Inglehart, R., C. Haerpfer, A. Moreno, C. Welzel, K. Kizilova, J. Diez-Medrano, M. Lagos, Pippa Norris, E. Ponarin, B. Puranen, et al. 2014. World Values Survey: All Rounds-Country-Pooled Datafile Version. Madrid: JD Systems Institute. (https:// www.worldvaluessurvey.org/WVSDocumentationWVL.jsp). 34 Sociological Methods & Research 0(0) Karlson, Kristian B., Anders Holm, and Richard Breen. 2012. “Comparing Regression Coefficients Between Same-Sample Nested Models Using Logit and Probit: A New Method.”Sociological Methodology 42:286–313. doi: 10.1177/ 0081175012444861 Koenker, Roger. 2005. Quantile Regression. Cambridge: Cambridge University Press. Koenker, Roger, Victor Chernozhukov, Xuming He, and Limin Peng 2020. Handbook of Quantile Regression. Boca Raton, FL: CRC Press. Leckie, George, Robert French, Chris Charlton, and William Browne. 2014. “Modeling Heterogeneous Variance–Covariance Components in Two-Level Models.”Journal of Educational and Behavioral Statistics 39(5):307–32. doi: 10.3102/1076998614546494 Lersch, Philipp M., Wiebke Schulz, and George Leckie. 2020. “The Variability of Occupational Attainment: How Prestige Trajectories Diversified Within Birth Cohorts Over the Twentieth Century.”American Sociological Review 85(6):1084–116. doi: 10.1177/0003122420966324 Lipset, Seymour M. 1959. “Some Social Requisites of Democracy: Economic Development and Political Legitimacy.”American Political Science Review 53(1):69–105. Lipset, Seymour M. 1960. Political Man: The Social Bases of Politics. Garden City, NY: Doubleday. Lipset, Seymour M. 2004. The Democratic Century. Norman, OK: University of Oklahoma Press. Long, J. Scott. 1997. Regression Models for Categorical and Limited Dependent Variables. Thousand Oaks, CA: Sage. McCullagh, Peter. 1980. “Regression Models for Ordinal Data.”Journal of the Royal Statistical Society, Series B 42(2):109–42. doi: 10.1111/j.2517-6161.1980. tb01109.x Mize, Trenton D., Long Doan, and J. Scott Long. 2019. “A General Framework for Comparing Predictions and Marginal Effects Across Models.”Sociological Methodology 49:152–89. doi: 10.1177/0081175019852763 Mood, Carina. 2010. “Logistic Regression: Why We Cannot Do What We Think We Can Do, and What We Can Do About It.”European Sociological Review 26(1):67–82. doi: 10.1093/esr/jcp006 Smith, Tom W., Michael Davern, Jeremy Freese, and Stephen L. Morgan. 2019. “General Social Surveys, 1972-2018 [Machine-Readable Data File].”Chicago, IL: NORC. Solt, Frederick. 2020. “Measuring Income Inequality Across Countries and Over Time: The Standardized World Income Inequality Database.”Social Science Quarterly 101(3):1183–99. doi: 10.1111/ssqu.12795 Gangl 35 VanHeuvelen, Tom. 2018a. “Within-Group Earnings Inequality in Cross-National Perspective.”European Sociological Review 34(3):286–303. doi: 10.1093/esr/jcy011 VanHeuvelen, Tom. 2018b. “Recovering the Missing Middle: A Mesocomparative Analysis of Within-Group Inequality, 1970–2011.”American Journal of Sociology 123(4):1064–116. doi: 10.1086/695640 Western, Bruce and Deirdre Bloome. 2009. “Variance Function Regressions for Studying Inequality.”Sociological Methodology 39:293–326. doi: 10.1111/j. 1467-9531.2009.01222.x Williams, Richard. 2006. “Generalized Ordered Logit/Partial Proportional Odds Models for Ordinal Dependent Variables.”Stata Journal 6(1):58–82. doi: 10. 1177/1536867X0600600104 Williams, Richard. 2009. “Using Heterogeneous Choice Models to Compare Logit and Probit Coefficients Across Groups.”Sociological Methods & Research 37(4):531–59. doi: 10.1177/0049124109335735 Williams, Richard. 2016. “Understanding and Interpreting Generalized Ordered Logit Models.”Journal of Mathematical Sociology 40(1):7–20. doi: 10.1080/0022250X. 2015.1112384 Wooldridge, Jeffrey M. 2010. Econometric Analysis of Cross Section and Panel Data. 2nd edition. Cambridge, MA: MIT Press. World Bank. 2021. “World Development Indicators”, Washington, D.C.: World Bank. (https://data.worldbank.org/). Author Biography Markus Gangl is a Professor of Sociology at Goethe University Frankfurt. His research centers on social stratification, economic inequality, poverty, income dynamics, social mobility, labor markets and careers, and he also seeks to contribute to the development of quantitative methodology in these fields. Markus Gangl is principal investigator in the POLAR project, which examines the impact of rising economic inequality on equality of opportunity, social cohesion, and democratic orientations in Western societies. Appendix 1 Estimation of the Multiscale Ordered Logit Model With ordinal outcome data, the proportional odds model is the threshold model Pr(Yi>j)=gj(Xβ)=exp (αj+Xiβ) 1+exp (αj+Xiβ)for j=1,2,...,k−1 (11) 36 Sociological Methods & Research 0(0) to predict the probability that the observed response Yfor the respondent iis higher than the response category jon a scale consisting of kordered categories. It is widely appreciated that this ordered logit model represents the data by a common regression plane defined over the structural component Xiβ, which is being shifted across a series of cutpoints αjthat correspond to observable response categories, for example, those provided in a survey question. The realization that the proportional odds model involves an equality constraint on the coefficient vector βforms the basis for the generalized ordered logit model of Fu (1998) and Williams (2006, 2016, also see Agresti 2010:75-80), where the generalization consists of estimating k−1 entirely threshold-specific regression planes and associated coefficient vectors βjor, in case of Williams’(2006, 2016) partial proportional odds model, of finding the set of optimized threshold-specificcoefficient vectors βjthat relax the parallel regression assumption for those covariates and cutpoints where data signals indicate effect heterogeneity, but that maintain the parallel regression assumption wherever the null hypothesis of effect homogeneity cannot be refuted in corresponding Wald tests on thedataathand. The multiscale ordered logit model likewise generalizes from the standard proportional odds model, but takes the cutpoints αjrather than the coefficient vector β(j)as its point of departure. Instead of relaxing the proportionality assumption on the effects of covariates X, the multiscale model relaxes the constraint that all cutpoints αjwere to refer to observable response categories on a single rating scale, and instead allows for the presence of multiple response formats s(i.e., different question formats) to tap into the same rating dimension in the source data. In the multiscale model, the latter are represented by scale-specific cutpoint parameters αjs, so that there are m s=1(ks−1) intercept terms to the model, one for each cutpoint json each of the mdifferent rating scales sthat constitute the observed outcome data. The resulting probability model Pr(Yi>js|si=s)=gjs(Xβ)=exp (αjs+Xiβ) 1+exp (αjs+Xiβ) for js=1,2,...,ks−1 and s=1,2,...,m (12) conforms to the standard proportional odds model in form, except for being additionally conditional on the outcome Yifor respondent ibeing obtained through one particular survey instrument si=sof course. By implication, the log-likelihood function of the model is obtained by a straightforward Gangl 37 expansion of the log-likelihood function of the standard proportional odds model, namely as ln L(α,β|Y,X,s)= N i=1 m s=1 ks js=1 yisjsln[Pr(Yi=js|xi,α,β,si=s)] = N i=1 m s=1 ks js=1 yisjsln[gj−1s(Xβ)−gjs(Xβ)] (13) with yisjs=1 when outcome Yithat is observed on scale si=sis falling into the category jsfor respondent iand yisjs=0 otherwise, all, of course, assuming the simplest case of a sample of independent observations. As the above log-likelihood function can be cumbersome to program in the general setting with many scales and, possibly, widely different numbers of observed response categories, an equally consistent estimator may be provided by returning to the insight behind the generalized ordered logit model of Fu (1998) and Williams (2006, 2016). When conceiving of the multiscale model as a threshold model that applies across different scales sand cutpoints json these scales, one may obtain parameter estimates by maximizing the log-likelihood function of lnL(α,β|Y,X,s)= N i=1 m s=1 ks js=1 wisyis yisjsln[Pr(Yi>js|xi,α,β,si)] +(1 −yisjs)ln[Pr(Yi≤js|xi,α,β,si)]  = N i=1 m s=1 ks js=1 wisyis{yisjsln[gjs(Xβ)] +(1 −yisjs) ln[1−gjs(Xβ)]}, (A1.4) where yis =1 when outcome Yiis observed on scale si=sfor respondent iand yis =0 otherwise, and where the weight wis =1/(ks−1) is adjusting for differences in scale length (number of categories) across rating scales. Albeit coming at some (substantively often minor) loss of efficiency, equation A1.4 has the advantage of being consistently estimable as a standard binary logit model with clustercorrected standard errors on a suitably expanded dataset (see Clogg and Shihadeh 1994:147), and, by virtue of its practicality, it then also provides a ready starting point to further expand the basic multiscale ordered logit model to either the multilevel specification of equation 8, to the generalized version of the multiscale ordered logit model (equation [9]), or to the combination of both (i.e., equation [10] in the main text). A Stata ado to implement the multiscale ordered logit model and its extensions as described in the present paper is available from the 38 Sociological Methods & Research 0(0) Boston College Statistical Software Components (SSC) Archive and may be downloaded via the command ssc install mscologit from within Stata. Appendix 2 Question Wording in the ESS, EVS/WVS, and GSSs The present analyses are using a survey question on respondents’trust in the national parliament to illustrate the practical utility of the multiscale ordered logit model in social science research. A respective rating question has long been a standard item in survey research, often as part of a larger battery to cover respondents’trust in a range of political and public institutions, and is also part of the core questionnaires of the ESS, EVS/WVS, and GSS series. While the question wording itself is straightforward, the main difference between the survey series lies in the response format that has been chosen to capture respondents’degree of trust in the national parliament. More specifically, the exact question formats are: ESS CARD 9 Using this card, please tell me on a score of 0–10 how much you personally trust each of the institutions I read out. 0 means you do not trust an institution at all, and 10 means you have complete trust. Firstly…READ OUT… B6 …[country]’s parliament? EVS/WVS <SHOWCARD 38> Please look at this card and tell me, for each item listed, how much confidence you have in them, is it a great deal, quite a lot, not very much, or none at all? Please indicate how much confidence you have in… Q38.G Parliament No trust at all Complete trust (Refusal) (Don’t know) 0 1 2 3 4 5 6 7 8 9 10 77 88 Gangl 39 GSS 181. I am going to name some institutions in this country. As far as the people running these institutions are concerned, would you say you have a great deal of confidence, only some confidence, or hardly any confidence at all in them? L. Congress The analyses in the present article use the inverted raw scores from the EVS/WVS and GSS items in order to conform to the ESS question format as well as a standard convention that larger numbers are meant to indicate a higher degree of trust. Also, the analysis obviously needs to assume that the three items tap the same dimension of political evaluation, even while the ESS, EVS, and WVS wording solicits trust in the abstract institution of parliament, whereas the GSS item is explicitly worded to solicit trust in the people running Congress. A great deal Quite a lot Not very much None at all Dońt know No answer Not applicable Not asked in survey Countryspecific missing code 1234−1−2−3−4−5 A great deal Only some Hardly any Dońt know No answer Not applicable 123890 40 Sociological Methods & Research 0(0)