scieee AI-readable full text Open interactive document viewer

Hidden Markov models for longitudinal rating data with dynamic response styles

Colombi, Roberto,Giordano, Sabrina,Kateri, Maria

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Colombi, Roberto; Giordano, Sabrina; Kateri, Maria Article — Published Version Hidden Markov models for longitudinal rating data with dynamic response styles Statistical Methods & Applications Provided in Cooperation with: Springer Nature Suggested Citation: Colombi, Roberto; Giordano, Sabrina; Kateri, Maria (2023) : Hidden Markov models for longitudinal rating data with dynamic response styles, Statistical Methods & Applications, ISSN 1613-981X, Springer, Berlin, Heidelberg, Vol. 33, Iss. 1, pp. 1-36, https://doi.org/10.1007/s10260-023-00717-x This Version is available at: https://hdl.handle.net/10419/317724 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/ Vol.:(0123456789) Statistical Methods & Applications (2024) 33:1–36 https://doi.org/10.1007/s10260-023-00717-x 1 3 ORIGINAL PAPER Hidden Markov models forlongitudinal rating data withdynamic response styles RobertoColombi1· SabrinaGiordano2 · MariaKateri3 Accepted: 9 July 2023 / Published online: 28 September 2023 © The Author(s) 2023 Abstract This work deals with the analysis of longitudinal ordinal responses. The novelty of the proposed approach is in modeling simultaneously the temporal dynamics of a latent trait of interest, measured via the observed ordinal responses, and the answering behaviors influenced by response styles, through hidden Markov models (HMMs) with two latent components. This approach enables the modeling of (i) the substantive latent trait, controlling for response styles; (ii) the change over time of latent trait and answering behavior, allowing also dependence on individual characteristics. For the proposed HMMs, estimation procedures, methods for standard errors calculation, measures of goodness of fit and classification, and full-conditional residuals are discussed. The proposed model is fitted to ordinal longitudinal data from the Survey on Household Income and Wealth (Bank of Italy) to give insights on the evolution of households financial capability. Keywords Latent variables· Longitudinal ordinal data· Stereotype logit models * Sabrina Giordano [email protected] Roberto Colombi [email protected] Maria Kateri [email protected] 1 Department ofManagement, Information andProduction Engineering, University ofBergamo, Bergamo, Italy 2 Department ofEconomics, Statistics andFinance “Giovanni Anania”, University ofCalabria, Cosenza, Italy 3 Institute forStatistics, RWTH Aachen University, Aachen, Germany 2 R.Colombi et al. 1 3 1 Introduction Psychometric literature widely debated the different behavior patterns of respondents to rating surveys, which may introduce distortions or inaccuracies in their responses. Questions on attitudes, opinions, perceptions are usually Likert-type or rating-scale items, and the observed responses may not reflect the respondents’ true preferences but their tendency to use only a small number of the available rating scale options, governed by an underlying behavioral mechanism, known as Response Style (RS) (e.g., VanVaerenbergh and Thomas 2013, for an overview). The activation of a response style mechanism influences systematically the way interviewees use response scales, introducing bias in the responses and scale usage heterogeneity, which may impact the data quality and the validity of the results (e.g., Baumgartner and Steenkamp 2001; Roberts 2016). What is new in our approach is the interest on the longitudinal perspective where respondents are asked, at several time occasions, to give a subjective assessment about rating-scale items and their responses are indicators of a latent trait of interest (e.g., health status, environmental risk, customer satisfaction). Moreover, responses can be driven or not by RS, the RS attitude can vary dynamically and the change over time of responses and answering behaviors can depend also on individual characteristics. More precisely, in the context of longitudinal ordered categorical data analysis, the methodological contribution of the paper is a hidden Markov model (HMM) with a bivariate latent Markov chain that jointly models an unobservable trait of interest and an unobservable binary indicator of the respondent’s form of answering (response style driven or not) over time. The use of HMMs in the context of categorical longitudinal data is not new, often referred to as latent Markov models for longitudinal data (see Bartolucci et al. 2012, 2017, for a comprehensive review), but to date there does not exist any HMM-based procedure useful for modeling the evolution of an underlying response behavior over time. A further contribution of the proposed approach lies on providing a parsimonious parametrization of the probability functions of the observed responses dictated by RS. Several RSs have been identified and studied (e.g., Baumgartner and Steenkamp 2001) and here a model is introduced that enables capturing easily the most commonly encountered RS. In fact, in our approach, the observation probability functions, conditionally on the presence of RS, depend on two parameters only, but offer a great flexibility in the types of RS that can be modelled such as tendency to select categories at random (careless RS, CRS), tendency to prefer positive response categories/answer with agreement (acquiescent RS, ARS), or negative response categories (disacquiescent RS, DRS), middle/neutral categories (middle RS, MRS), or extreme categories (extreme RS, ERS). Other approaches to simultaneously tackle multiple RSs, for cross-sectional data, rely on more complex models such as multi-trait models (e.g., Wetzel and Carstensen 2015; Falk and Cai 2016) or IRT models (e.g., Böckenholt 2012; Henninger and Thorsten 2020; Zhang and Wang 2020) or latent class factor models (Kieruj and Moors 2013). 3 1 3 Hidden Markov models forlongitudinal rating data withdynamic… Furthermore, novelty is also in the use of stereotype logit models (Anderson 1984) to investigate how covariates affect the initial and transition probabilities of the latent Markov chain. To our knowledge, such parsimonious models of sound interpretation have not been previously used in the HMM framework. In summary, our approach enables: (i) the identification of groups of individuals with a different dynamics of a latent categorical trait of interest taking into account the presence of RS driven responses, (ii) the accommodation of different RSs and their change over time, (iii) the use of a very parsimonious distribution of the responses affected by RS, (iv) the introduction of covariates influencing the initial/transition probabilities of both the latent construct and the unobservable response style indicator, (v) the use of stereotype logit models for the initial/transition probabilities of the latent construct. The proposed methodology is of interest in all longitudinal surveys that model attitudes, opinions, perceptions or beliefs, that are indicators of non directly measurable and observable variables. For example, in healthcare studies, patients are asked, at several occasions, to give a subjective assessment of their health status or disability in daily living; in marketing research, customers are required to evaluate their satisfaction for services/products; in socio-economic contexts, citizens are invited to answer to what extent they agree or disagree with sensitive topics (immigration, criminality, gender gap); in environmental studies, interviewees are asked to reveal their perception of the impact of climate changes and environmental risk. In all these cases, the presence of RS cannot be ignored and substantive latent traits need to be measured taking into account effects due to RSs. To show the practical usefulness of our proposal, we investigate the evolution over time of the household financial capability (a broader term encompassing behavior, knowledge, skills and attitudes of people with regard to managing their financial resources, e.g. Zottel et al. 2013) as a latent psychological and behavioral trait that influences the household’s decision-making to face financial issues. The latent financial capability is here measured in terms of two observed indicators: the self-perceived ability to make ends meet and the self-report of perceived risk related to financial investments. These indicators have great impact on the score measuring the financial capability, as defined according to the Organization for Economic Cooperation and Development methodology (survey OECD 2020), applied by 36 countries and in Italy implemented by the Bank of Italy (D’Alessio etal. 2020). The structure of the work is as follows. In Sect. 2, the data of our motivating problem from a survey on financial conditions of Italian households are introduced, and the issues to be tackled described. In Sect.3, the modeling of different response style effects in the longitudinal perspective through HMMs is proposed and the advantages of our approach are highlighted. In Sect. 4, latent and observation components of the proposed HMM are described in detail. Alternative HMMs are examined in Sect.5, most of them being special cases of the here presented model. Section 6 is devoted to methodological contributions on: maximum likelihood estimators of the parameters, measures of goodness of fit and classification, and full-conditional residuals. In Sect.7, the proposed model is fitted on the real 4 R.Colombi et al. 1 3 data of Sect. 2, implementing the developed estimation procedure and providing answers to the questions raised in Sect.2. Concluding remarks are given in Sect.8. Technical details on the methods to calculate the standard errors are postponed in the Appendix. 2 Motivating example Our work meets the growing interest in the households financial capability. The governments are now playing an active role in meeting the financial capability challenge. Initiative taking forward to increase capability are provided throughout starting from National Strategy for Financial Capability in UK (HM-Treasury 2007), to EU Commission (Valant 2015), National Financial Capability Study in USA (Lin et al. 2019), OECD (OECD 2020), among others. Psychological and behavioral aspects affecting people’s economic and financial decisions (studied as behavioral economics) are also inserted into practices to strengthen financial consumer protection (as agreed in the action plan endorsed by G20/OECD, Lefevre and Chapman 2017). In this direction, we propose here to model the dynamics of the households’ perception of their financial conditions, accounting for the way households disclose their perceptions, through HMM. The data are from the waves of the Survey on Household Income and Wealth (SHIW). It is conducted by the Bank of Italy every two years since the 1960s to collect information about the income, wealth and saving of Italian households. Over the years, the survey has grown in scope and now it includes also aspects of households’ economic and financial behavior, furthermore since 2004 it contains information on attitude towards financial risk. The data1 used refer to 1109 Italian households involved in all the waves from 2006 to 2016. We considered the items: R1 reveals the perception of the household’s financial ability to make ends meet based on the answers of the head of the households to the question: Is your household’s income sufficient to see you through to the end of the month.... very easily, easily, fairly easily, fairly difficultly, with some difficulty, very difficulty; R2 indicates the risk perception in managing financial investments measured through the response to the question: in managing your financial investments, would you say you have a preference for investments that offer: low returns, with no risk of losing the invested capital (risk averse); a fair return, with a good degree of protection for the invested capital (risk tolerant), good-high returns, but with a fair-high risk of losing part of the capital (risk lover). We focus on these two indicators, among others, since they strongly orient policy maker choices. In particular, insights into ability helps to: developing effective programs to educate people to manage their resources, reducing welfare dependency, and identifying vulnerable groups of the population for which targeted interventions can be designed. The OECD, in the recent survey (OECD 2020), 1 All the data are available at https:// www. banca dital ia. it/ stati stiche/ temat iche/ indag inifamig lieimpre se/ bilan cifamig lie/ distr ibuzi onemicro dati/ index. html. 5 1 3 Hidden Markov models forlongitudinal rating data withdynamic… recognises that large groups of citizens are lacking the necessary financial behavior and financial resilience to deal effectively with everyday financial management. This is particularly concerning at the time of the unfolding crisis as a result of the COVID-19 pandemic, which is likely to put considerable economic and financial pressures on individuals and test their ability to preserve their financial well-being. Moreover, to understand the financial capability it is important to comprehend how households think and feel about the risks they face, Slovic (2010). The risk perception is an important determinant of protective behavior, as in general, the success of public intervention programs is largely dependent on individual risk perception. Comprehension of the perceived risk may offer useful prompts for the design of effective investor education programs and orient towards vulnerable individuals preventive initiatives against bad financial decisions (e.g., Pidgeon 1998; Nguyen etal. 2019). Some demographical and economical characteristics that can affect the degree of financial coping of a household are the covariates gender (G): female ( 27% ), male ( 73% ); job (J): self-employee (Jse, 10% ), housekeeper/retired/student (Jhrs, 47% ), employee (Je, 43% )); children (CH): with children ( 34% ), no children ( 66% ); debts (D): with debts ( 22% ), no debts ( 78% ); savings (S): with savings ( 83% ), no savings ( 17% ); education (E): up to secondary school ( 62% ), over secondary school ( 38% ), with the reference categories being in italics and the percentages referred to the initial year 2006. The frequencies of all the 18 pairs of categories of the two items, R1 and R2 , over the six years are represented in Fig.1. The perceptions evidently change over time. The most commonly chosen responses are a fairly difficult ability to make ends meet and a risk loving behavior towards financial proposals. The choice falls frequently also on the pairs: very easily-averse, easily-averse, fairly easilyaverse. The number of households who selected these four most common responses is represented in Fig. 2, over the years, for four groups of households, identified Fig. 1 Frequencies of (R1,R2) responses over time occasion 6 R.Colombi et al. 1 3 as those with the most representative profiles (most frequent configurations of the covariates among the 94 truly observed ones in the data at hand). These plots exemplify just a portion of the data on the three dimensions: responses × covariates × time. In our context, responses R1 and R2 can be considered as the manifest expressions of the latent household’s financial capability. Moreover, we believe that an unobservable answering behavior drives respondents that can reveal their perceptions in two ways: with awareness, when their answers reflect the respondents’ true opinion, or according to a response style, when in doubt or reluctant to disclose their opinion they prefer extreme or middle points of the rating scale, or according to their inclination they focus on the positive or negative side of the rating scale. Our approach gives us various opportunities: (i) to describe simultaneously the dynamic behavior of respondents in the way of answering and in disclosing their perceived financial capability measured through the degree of difficulty or ease in matching monthly expenses with disposable income and their attitude toward risk of investments, that change over time in line with Schildberg-Hörisch (2018) and De Blasio et al. (2021); (ii) to investigate if the households feel and communicate their perceptions differently according to their demographic and socio-economic characteristics; (iii) to discriminate groups of respondents, with certain profiles, in the latent classes that identify various degrees of the latent financial capability, taking into account that they can answer with awareness or may prefer a response style. Section7 will shed light on these aspects of interest. Fig. 2 Frequencies of the most common responses over the six years for the four most present profiles of households 7 1 3 Hidden Markov models forlongitudinal rating data withdynamic… 3 Response styles inlongitudinal studies RS mechanisms lead to biased measurement of the traits of interest that may influence seriously the results of a survey and thus be responsible for non-optimal decisions. Underlying RSs affect all levels of the analysis of survey data, from being responsible for violations of the adopted model assumptions up to biased estimation of parameters and measures of interest, like correlations in crosssectional survey data, as shown, among others, by Piccolo etal. (2019) for CRS, Dolnicar and Grün (2009) for ARS and ERS, Tutz and Berger (2016) and Tutz (2021) for MRS and ERS. Approaches aiming at better estimation of the original substantive trait, by controlling for RSs, are mostly based on mixtures models and employ latent variables (e.g., Grün and Dolnicar 2016; Huang 2016; Böckenholt and Meiser 2017, among many others). Several simulation studies provide evidence that ignoring response styles implies bias on the parameter estimates, see Tutz and Berger (2016); Colombi etal. (2019, 2021), among others. Though accounting for RSs in cross sectional studies has achieved considerably attention in the literature and effective models have been proposed, dealing with RSs in longitudinal data remains challenging. Questions on whether or to what extend RSs remain stable over time are still open. This paper investigates whether RS behavior is an individual time invariant feature (Bachman and O’Malley 1984; Paulhus 1991) or it is not necessarily consistent over time, depending on the measurement situation (Weijters 2006; Aichholzer 2013). In this regard, the RS is described through time-invariant and time-specific latent factors in (Weijters et al. 2010). A recent proposal in the direction of dynamic response styles is by Soland and Kuhfeld (2020), who concluded that the stability over time of within-subject RS factors is not always justified, by comparing multidimensional nominal response models (Bolt and Johnson 2009). In the context of longitudinal data, we tackle the problem of time dependence of RSs within the latent variable context, considering two unobserved classes of responses: aware (AWR) responses, i.e. not affected by any RS, and RS driven responses, assuming that an individual can switch over time from AWR to RS type of responses and vice versa. Among the approaches to modeling longitudinal categorical data (representative sources are, for example, Molenberghs and Verbeke 2005; Hedeker and Gibbons 2006; Bergsma et al. 2009), we resort to the family of hidden Markov models for their flexibility in modeling time dependence, based on sound assumptions, and their computational tractability. The novelty of the contribution is the modeling of the temporal dynamic of rating responses with a bivariate latent Markov chain that jointly models an unobservable construct of interest and an unobservable indicator of the respondent’s form of answering (AWR or RS driven). The second latent component indeed allows us to describe how the RS behavior dynamically changes over time, in contrast to other approaches where the RS is thought as a continuous time-invariant latent trait (Billiet and Davidov 2008). 8 R.Colombi et al. 1 3 4 HMMs withtwo latent variables Consider r ordinal responses observed on n units (subjects/items) at T time occasions. In particular, let Yjit , Yjit ∈Cj={1, …,cj} , denote the j-th ordinal response variable, j∈R={1, …,r} , of the i-th unit, i∈I={1, …,n} , at the tth occasion, t∈T={1, …,T} . The responses are assumed to reflect the levels of unobservable latent constructs Lit , i∈I , t∈T , with finite discrete state space SL={1, …,k} . Furthermore, they can be observed under two latent regimes: awareness (AWR) and response style (RS) that are captured by binary latent variables Uit , i∈I , t∈T , with state space SU={1, 2} , where 1 and 2 denote the RS and AWR states, respectively. For this, Uit are called response style indicators. The presence of the above mentioned two regimes is based on the idea that respondents either manifest their true preference or select categories according to a RS (CRS, ERS, DRS, MRS, ERS). The proposal is a HMM defined by two components that describe the Markov chain of the latent variables and the conditional distributions of the responses given the latent variables. The model will be referred to as a HMM with a RS component (RS-HMM). Next subsections are devoted to specifying the two model components by parameterizing the observation probabilities and the initial/transition probabilities through suitable logit models. To avoid difficulties in interpreting the results, covariates are assumed to affect only the distribution of the latent variables. In our view, in fact, the covariate effect is captured by the latent constructs Lit which are indirectly observed trough the responses Yjit . 4.1 The latent model The latent variables Lit and Uit are independent across units and, for every unit, the process {Lit,Uit }t∈T is assumed to evolve in time according to a first order bivariate Markov chain with states (u,l), u∈SU , l∈SL . For the sequel, let always i∈I and consider states u,u∈SU and l ,  l∈S L . The latent component of the model is specified through its initial and transition probabilities. The initial probabilities ( t=1 ) of the latent bivariate process {Lit,Uit }t∈T are 𝜋i1(u,l)=P(Li1=l,Ui1=u), and the transition probabilities are 𝜋it (u,l | u,  l)=P(L it =l,U it =u | L it−1 =  l,U it−1 =u),t=2, …,T . Furthermore, denote the marginal transition probabilities for the latent variables Lit and are the transition probabilities of the latent RS indicators Uit , conditioned on the transition (  l,l) of the latent construct, called for short as conditional RS transition probabilities. (1) 𝜋L it (l | u,  l)=P(L it =l | L it−1 =  l,U it−1 =u ) (2) 𝜋 U | L it (u | l,u,  l)= 𝜋it(u,l | u,  l ) 𝜋L it (l | u,  l) 15 1 3 Hidden Markov models forlongitudinal rating data withdynamic… which each latent variable does not Granger cause the other one, is a special case of the graphical multiple HMMs introduced by Colombi and Giordano (2015). The drawback of this model is that, under the two non Granger causality conditions, the transition probabilities 𝜋it (u,l | u,  l ) do not have a closed expression and must be computed numerically as a function of the probabilities 𝜋U it (u | u ) , 𝜋L it (l | l ) and a set of k−1 odds ratios defined on the bivariate transition probabilities. See Colombi and Giordano (2015) for more details on these Granger non causality conditions and on a marginal parametrization that can be used in this context. A final note concerns one of the major limitation of the HMM approach: the Geometric sojourn-time distribution in the states, implied by the Markov assumption, is restrictive when the probabilities of leaving a state are function of the length of time spent there. This limitation suggests interesting extensions of the proposed model, based on Markov chains with order greater than one (Zucchini et al. 2017) or on hidden-semi-Markov models, at the expense of an increased number of parameters or computational cost (see Pohle etal. 2022; Yu 2010, among others). However, in our proposal, the Markov assumption is not an overly restrictive condition since the transition probabilities of both, latent construct and RS indicator, depend on time varying covariates. 6 Inference Let 𝜽 denote the vector of all the parameters of the latent and observation models. For example, in the simple case of a memoryless model with k=2 , no covariates and one response with four categories, it is: 𝜽 =(𝛼 01 ,𝛼 0 ,𝛽 21 ,𝛽 12 ,  𝛽 01 ,  𝛽 02 ,𝜑 11 ,𝜑 21 ,𝜑 31, 𝜑12,𝜑22,𝜑32,𝜙01,𝜙11,𝜙02,𝜙12). Hereafter, procedures to provide maximum likelihood estimates (MLE) of these parameters and standard errors are illustrated. 6.1 Estimation viaanEM algorithm The latent binary variable d(1) it (u,l ) is equal to 1 when the i-th unit (subject) is at time t in state (u,l) and the latent binary variable d(2) it (u,l;u,  l ) is 1 if at time t, t>1 , the i-th subject is in state (u,l) while at occasion t−1 was in (u,  l) , l ,  l∈S L , u,u∈SU . Moreover, the observable binary variable djit(yj) is equal to 1 if at time t the category yj of Yjit , j∈R , is observed on the i-th individual, i∈I . If the above binary latent variables were observable, the parameters could be estimated by maximizing the following complete log-likelihood (i.e. the joint loglikelihood of the observations and the latent variables): 16 R.Colombi et al. 1 3 where fj|1 and fj|2 are provided in (11) and (12). As the latent variables are not observable and it is not easy to maximize the marginal log-likelihood, obtained by summing the joint log-likelihood over all the possible realizations of the latent indicators, it is common, in the context of HMMs, to use the EM algorithm, to compute the maximum likelihood estimates. Details on the EM algorithm in the context of HMMs are presented in many papers and books. See Bartolucci and Farcomeni (2015) for a presentation specific to the context of longitudinal data. Every iteration of the EM algorithm is composed by two steps: the Expectation (E) step and the Maximization (M) step. With respect to our model, in the E step the following expected values are computed: where Eobs() is the expected value taken conditionally on the observed values of the responses Yjit and on the covariates and given the current value  𝜽 of the parameters. The previous expected values are computed by the Baum-Welch forward-backward algorithm (Zucchini and MacDonald 2009, Ch. 4). In the M step, the following conditional expectation of the complete log-likelihood function is maximized in order to obtain an updated  𝜽 : (15) 𝓁 ∗(𝜽)= n � i=1 k � l=1 �2 � u=1 d(1) i1(u,l) � log 𝜋L i1(l) + n � i=1 2 � u=1�k � l=1 d(1) i1(u,l)�log 𝜋U i1(u) + k �  l=1�n � i=1 T � t=2 k � l=1�2 � u=1 2 � u=1 d(2) it (u,l;u, l))�log 𝜋L it (l� l)� + 2 � u=1 k � l=1�n � i=1 T � t=2 2 � u=1�k �  l=1 d(2) it (u,l;u, l))�log 𝜋U�L it (u�u,l) � + r � j=1 k � l=1⎧ ⎪ ⎨ ⎪ ⎩ cj � yj=1�n � i=1 T � t=1 d(1) it (1, l)djit(yj)�log fj�1(yj�l)⎫ ⎪ ⎬ ⎪ ⎭ + r � j=1 k � l=1⎧ ⎪ ⎨ ⎪ ⎩ cj � yj=1�n � i=1 T � t=1 d(1) it (2, l)djit(yj)�log fj � 2(yj � l)⎫ ⎪ ⎬ ⎪ ⎭ , (16) 𝛿(1) it (u,l;  𝜽)=E obs (d (1) it (u,l)),𝛿 (2) it (u,l;u,  l;  𝜽)=E obs (d (2) it (u,l;u,  l))) , 17 1 3 Hidden Markov models forlongitudinal rating data withdynamic… Note that Q(𝜽|  𝜽) is obtained from the complete log-likelihood by replacing d(1) it (u,l ) and d(2) it (u,l;u,  l ) with their expected values (16). The six addends of (17), corresponding to the models specified by (3) or (7), (4), (5) or (9), (6) and (11-12) of Sects.4.1 and 4.2, depend on disjoint subsets of the vector 𝜽 and can be maximized separately. The maximization of the sixth addend is simple as there is a closed form for the maxima. Moreover, the first addend is equivalent to the ML estimation of the logit model (3) or its stereotype variant (7), and the third and fifth terms simplify to the estimation of k(r+1) separate logit models described by (5) or (9) and (11). A similar remark applies to the second and fourth addends and the logit models defined by (4) and (6), respectively. The terms within curled brackets correspond to the log-likelihoods of the logit models that can be maximized separately. In the first two addends, the curled brackets are omitted as only one logit model is involved. The expected values within squared brackets play the role of observed frequencies. If the model is correctly specified, the estimates of the standard errors can be based either on the matrix of second derivatives of the log-likelihood function (observed information matrix, in short OIM), see Bartolucci and Farcomeni (2015), or on the outer products of the individual contributions to the score functions (outer product information matrix, OPIM, or BHHH estimate, Berndt etal. 1974). When the model is misspecified, the information matrix equivalence does not hold and the standard errors have to be calculated using the so called Sandwich matrix (White 1982), say SDW. Alternatively, standard errors can be computed using the boostrap (BOOT) technique. Technical details are given in Appendix. All the R functions, for the estimates and standard errors (with the four mentioned methods) are available from the authors. (17) Q (𝜽� 𝜽)=Eobs(𝓁∗(𝜽)) = n � i=1 k � l=1 �2 � u=1 𝛿(1) i1(u,l; 𝜽) � log 𝜋L i1(l) + n � i=1 2 � u=1�k � l=1 𝛿(1) i1(u,l; 𝜃)�log 𝜋U i1(u) + k �  l=1�n � i=1 T � t=2 k � l=1�2 � u=1 2 � u=1 𝛿(2) it (u,l;u, l); 𝜽)�log 𝜋L it (l� l)� + 2 � u=1 k � l=1�n � i=1 T � t=2 2 � u=1�k �  l=1 𝛿(2) it (u,l;u, l); 𝜽)�log 𝜋U�L it (u�u,l) � + r � j=1 k � l=1⎧ ⎪ ⎨ ⎪ ⎩ cj � yj=1�n � i=1 T � t=1 𝛿(1) it (1, l; 𝜽)djit(yj)�log fj�1(yj�l)⎫ ⎪ ⎬ ⎪ ⎭ + r � j=1 k � l=1⎧ ⎪ ⎨ ⎪ ⎩ cj � yj=1�n � i=1 T � t=1 𝛿(1) it (2, l; 𝜽)djit(yj)�log fj � 2(yj � l)⎫ ⎪ ⎬ ⎪ ⎭ . 18 R.Colombi et al. 1 3 6.2 Goodness offit andclassification The goodness of fit testing and model selection in latent Markov models for longitudinal data is not straightforward, since standard asymptotic results for test statistics may not hold. The use of Akaike’s information criterion (AIC) or Bayesian information criterion (BIC) is a broadly used and accepted procedure. In particular, for HMM, the use of BIC dominates, even though its theoretical properties are not clear (e.g., Bartolucci etal. 2009; Zucchini and MacDonald 2009). Furthermore, there have been proposed in the literature normalized indices for assessing the overall fit of a model. For example, Bartolucci etal. (2009) used the index: for assessing the fit of the model against the independence model characterized by k=1 and no RS effects, with ∑r j=1 (cj−1 ) parameters and log-likelihood function  𝓁0 . It holds R2∈[0, 1] , with higher values indicating a better fit. Indices can be introduced for measuring the quality of classification and the distinguishability of the latent classes as well; Bartolucci etal. (2009) proposed an index based on the posterior probabilities of the latent classes, which in our set-up is: with 𝛿∗ it being, for unit i at time t, the maximum with respect to (u,l) of the posterior latent class probabilities 𝛿(1) it (u,l; 𝜽 ) , introduced in (16). Measure Sk lies between 0 and 1, where 1 represents certainty in classification and a perfect separation among latent classes, while values close to 0 indicate that most of 𝛿∗ it are close to 1/2k, that is like choosing the classes at random. This index is very suitable for our context where the observed responses are manifest realizations of the latent variables, therefore a good quality in terms of separation of the 2k latent states is crucial. In line with the literature which ignores the answering behavior, we can measure the quality of the separation of the latent construct states marginally with respect to U, so that (18) reduces to: Moreover, in our context, the distinguishability among the k states of the latent construct can be interestingly measured at the AWR and RS regimes separately. The Sk index is specified for this aim as follows: R 2=1−exp { 2[  𝓁0−𝓁(  𝜽)]∕nr }, (18) S k= ∑n i=1 ∑T t=1(𝛿∗ it −1∕2k) (1−1∕2k)nT , S L k= ∑n i=1 ∑T t=1(𝛿L it −1∕k) (1−1∕k)nT ,with 𝛿L it =max l∈SL � u ∈ S U 𝛿(1) it (u,l; 𝜽) . 19 1 3 Hidden Markov models forlongitudinal rating data withdynamic… It can also be of interest to measure the ability to discriminate between AWR and RS behaviors, regardless of the latent construct. For this aim, the measure (18) is modified as: Finally, the concern can be directed to measure how well separated are the two responding regimes, in every class of the latent construct. An insight in this sense is given by: 6.3 Residual analysis After the selection of a reasonable model according to indices of goodness of classification, and indices for judging the overall fit of the model, a residual analysis which detects features of the data not captured by the model has to be carried out. We assess the adequacy of the selected model by analysing full-conditional residuals, introduced in the context of HMMs by Buckby etal. (2020), as exvisive residuals. Full-conditional residuals are an alternative to the forecast or predictive residuals (Buckby etal. 2020). The difference is that in full-conditional residuals, the expected values of observed counts at time t are taken given all the other observations while in forecast residuals they are taken given the observations before time t. Full-conditional residuals are more useful in evaluating goodness of fit while forecast residuals are more helpful to assess the predictive accuracy of the model. In the application that follows, we use Pearson full-conditional residuals, whose technical details are given below. To simplify the notation, let x it =(x (U)� i ,z (U)� it ,x (L)� i ,z (L)� it ) � be the set of covariates for individual i at time t, i∈I , t∈T . Let Dt={x1,x2,…,xdt} be the set of different configurations of covariates observed at time t∈T and D=∪ tDt . Moreover C is the set of the c = ∏j c j different configurations of the responses. For every vector yit,yit ∈C , of the r responses of unit i at time t, we define the rest of yit as Y− it ={yi1,yi2…,yit−1,yit+1,…,yiT }. For y∈C, i∈I , t∈T , the indicator S L�RS k= ∑n i=1 ∑T t=1(𝛿 L�RS it −1∕k) (1−1∕k)nT ,with 𝛿L�RS it =max l∈SL 𝛿 (1) it (1, l; 𝜽) ∑l∗∈SL 𝛿(1) it (1, l∗; 𝜽) , S L � AWR k= ∑ n i=1 ∑ T t=1(𝛿L�AWR it −1∕k) (1−1∕k)nT ,with 𝛿L � AWR it =max l∈SL 𝛿(1) it (2, l; 𝜽) ∑l ∗ ∈ S L 𝛿(1) it (2, l∗; 𝜽) . S U k= ∑n i=1 ∑T t=1(𝛿U it −1∕2) (1−1∕2)nT ,with 𝛿U it =max u∈SU � l∈ S L 𝛿(1) it (u,l; 𝜽) . S U � L=l k= ∑n i=1 ∑T t=1(𝛿 U�L itl −1∕2) (1−1∕2)nT ,with 𝛿U � L itl =max u∈SU 𝛿 (1) it (u,l; 𝜽) ∑u ∗ ∈ S u 𝛿(1) it (u∗,l; 𝜽) ,l∈SL . 20 R.Colombi et al. 1 3 dit(y)=1 if yit =y,dit(y)=0 otherwise, is defined and summing over units the counts n t(y,x)= ∑ i∶x it =xdit(y ) are obtained for every x∈Dt . We introduce a residual for every x∈Dt , t∈T , and y∈C by comparing the previous counts with their expected values defined below. Let fit(y|D,Y− it ) , y∈C , i∈I, t∈T be the joint probability density function (pdf) of the responses given the covariates and the rest of yit . The computation of this pdf is described by Buckby etal. (2020), by Zucchini and MacDonald (2009) in the related context of pseudo residuals, and can be obtained as a by product of the Baum-Welch algorithm. Starting from these pdf, we define the following conditional expected values of the counts nt(y,x) : 𝜇 t(y,x)= ∑ i∶x it =xfit(y � D,Y − it ) , for every x∈Dt , t∈T , y∈C . Accordingly, the following full-conditional Pearson residuals are introduced: for every x∈Dt , t∈T , y∈C . Plotting full-conditional Pearson residuals is an useful tool to investigate the lack of fit of the model and to highlight particular features of the data. Standardizing these residuals is possible in theory but the computation of the standard errors is not an easy analytical and computational task. This could be done by the methods used in Titman (2009) for HMMs in continuous time but, an in depth-study is needed to asses the feasibility in presence of many residuals. For every time occasion t, t∈T , and every observed covariate configuration x , t he squared full-conditional Pearson residuals sum to the corresponding Pearson’s chisquared statistic 𝜒2 t (x)= ∑y∈C 𝜌 t (y,x) 2. In this paper, the averages of full-conditional Pearson residuals over the c response configurations, i.e. 𝜒2 t(x) c ,x∈Dt,t∈T , are used to summarize the comparison of the estimated cell probabilities under the assumed model with the observed proportions. In applications of multivariate responses, practical interest may lie on univariate responses Yj or bivariate responses (Yj,Yj � ) , with j≠j′ , j,j�∈R . In such cases, residuals (19) can be marginalized to: where is a configuration of the responses of interest and the associated set of indices. The consideration of marginalized residuals is also useful in case of sparsity in response configurations. 7 SHIW data analysis We applied the proposed models to the panel data from the Survey on Household Income and Wealth described in Sect.2 to answer the questions raised there. The household’s financial capability (or condition) is the latent trait of interest measured trough the ability to make ends meet R1 and the perceived financial risk R2 , with (19) 𝜌 t(y,x)= n t (y,x)−𝜇 t (y,x) √ 𝜇 t (y,x) , 21 1 3 Hidden Markov models forlongitudinal rating data withdynamic… covariates gender (G), job (J), children (CH), debts (D), savings (S), education (E), cf Sect.2. 7.1 Model selection Models based on different hypotheses on the latent transition probabilities are compared in Table1, each one considered for an increasing number of latent states. When the initial or transition probabilities depend on the covariates are said to be heterogeneous otherwise they are homogeneous. The compared models differ in their computational complexity as the number of parameters varies. In particular, as expected, models with unrestricted effect of covariates are computationally more intensive than models with a parallel effect. In the models of Table 1, the initial probabilities 𝜋L i1 (l ) are modelled through stereotype logits (models M3, M4, M7, M8) when a stereotype model or a parallel baseline logit model is used for the transition probabilities of the latent construct, otherwise they are modelled by unrestricted logit models (models M1, M2, M5, M6). RS initial probabilities are always assumed to depend on the covariates to capture heterogeneity in the answering behavior at the beginning. The minimum BIC corresponds to the model M8 with k=4 states defined by stereotype models (7) for the latent construct initial probabilities and parallel baseline logit models (9) with scores 𝜈l l=1 , l≠  l , for the transition probabilities, and no covariate effects on the RS conditional transition probabilities, specified in (6). Model M8 with k=3 is the second best according to the BIC criterion. Both the models M8 with k=3 and k=4 have a very high value of R2 , so they fit similarly and well enough the data at hand, stressing that the dependence of the responses on time and covariates is supported by the data. Nevertheless, there is evidence of quite overfitting for Model M8 ( k=4 ), since the conditional response probabilities in two states are not easily distinguishable. Looking at the goodness of the classification of units into the latent classes, measured by the Sk index (18), it results that model M8 with k=3 has S3=0.757 greater than S4=0.712 obtained for k=4 , thus the simple model seems to better separate the latent classes. Moreover, the results of all the variants of the Sk index, i.e. measures SL k , SU k , SL|RS k , SL|AWR k , SU|L=l k with l∈SL (Sect. 6.2), illustrated in Table 2, confirm the superiority of M8 with k=3 over the analogous model with k=4 in terms of distinguishing the states of the latent financial capability within the two groups of AWR and RS respondents, and also the greater ability to distinguish the AWR and RS behaviors, marginally and conditionally on the latent classes l, except for l=1 only. Therefore, the latent construct - the households financial capability - is reasonably chosen with three states meaning that households can be grouped according to whether they feel financially confident ( l=1 ), financially fair ( l=2 ), financially distressed ( l=3 ). The choice of model M8 implies that the transition probabilities of the latent construct are well described under the parsimonious model, where the effects of covariates on the transition probabilities depend on 22 R.Colombi et al. 1 3 the previous latent state  l but do not change over the current latent state l. Other less restrictive hypotheses do not fit better, also when combined with a smaller number of latent states. The chosen model is then compared with the latent Markov models, with three states and six states, proposed by Bartolucci etal. (2012), say B, that are Table 1 The maximum value of the log-likelihood function (loglike), the number of states k, the number of parameters, BIC, Sk and R2 values are reported for models defined by different hypotheses on the transition probabilities Model Hypotheses on transition probabilities k loglike n. par. BIC Sk R2 𝜋L it (l|  l ) 𝜋U|L it (u|l,u ) M1 Unrestricted logit models Heterogeneous, m.r.s.i. 2 − 15119.64 70 30730.07 0.707 0.805 3 − 14716.95 129 30338.34 0.699 0.864 4 − 14497.80 204 30425.89 0.692 0.888 5 − 14309.60 295 30687.50 0.713 0.906 M2 Unrestricted logit models Heterogeneous 2 − 14773.11 86 30149.18 0.798 0.857 3 − 14447.88 153 29968.47 0.760 0.893 4 − 14289.60 236 30233.85 0.720 0.903 5 − 14191.18 335 30731.12 0.715 0.915 M3 Stereotype logit models Heterogeneous 2 – – – – – 3 − 14465.32 129 29835.08 0.764 0.892 4 − 14330.35 176 29894.68 0.717 0.904 5 − 14282.73 227 30157.01 0.681 0.908 M4 Parallel baseline logit models Heterogeneous 2 − 14773.11 86 30149.18 0.798 0.857 3 − 14466.80 126 29817.02 0.768 0.891 4 − 14350.83 168 29879.55 0.716 0.902 5 − 14307.23 212 30100.84 0.691 0.906 M5 Unrestricted logit models Homogeneous, m.r.s.i. 2 − 15340.33 49 31024.21 0.695 0.762 3 − 14860.61 101 30429.35 0.690 0.845 4 − 14625.49 169 30435.88 0.645 0.875 5 − 14457.99 253 30689.81 0.660 0.892 M6 Unrestricted logit models Homogeneous 2 − 14833.29 58 30073.23 0.800 0.849 3 − 14512.86 111 29803.96 0.759 0.887 4 − 14364.24 180 29990.50 0.702 0.901 5 − 14265.30 265 30388.58 0.691 0.909 M7 Stereotype logit models Homogeneous 2 – – – – – 3 − 14532.10 87 29674.19 0.760 0.885 4 − 14413.66 120 29668.66 0.695 0.897 5 − 14354.00 157 29808.76 0.661 0.902 M8 Parallel baseline logit models Homogeneous 2 − 14833.29 58 30073.23 0.800 0.849 3 − 14534.76 84 29658.46 0.764 0.885 4 − 14427.96 112 29641.18 0.713 0.895 5 − 14370.93 142 29737.46 0.697 0.900 23 1 3 Hidden Markov models forlongitudinal rating data withdynamic… alternative to our model but do not account for a distinct latent variable representing the answering behavior. The comparison with the B model with k=3 ( loglike =14883.72 , npar =85 , BIC =30363.39 ) validates the idea that an underlying binary latent variable that distinguishes AWR and RS respondents is coherent with the data at hand, thereby strengthening empirically the usefulness of our approach. Moreover, the comparison with the alternative B model with k=6 ( loglike =−14612.20 , npar =322 , BIC =31482.02 ) confirms that the restrictions hypothesized in our model on the transition probabilities and on the RS probability functions are reasonable for the analyzed data. To complete the assessment of the chosen model we carry out a residual analysis. Fig4 illustrates the 3×6×6 box plots of the full-conditional Pearson residuals (hereafter residuals), described in (19), calculated within the 6 time occasions for every combination of the categories of the two responses. Box plots, within each time occasion and responses configuration, correspond to different covariate profiles. All the residuals, across all time occasions, have very small values around zero with median =−0.210 , Q1=−0.437 , Q3=−0.025 , mean =−0.011 , sd =0.93 . In particular, 95.6% of them are between -2 and 2. Overall, 11 % are out of whiskers in the box plots of Fig4 while only 22 residuals in total are greater than 5 (2.5 ‰). The maximum residual corresponds to the profile of a male, with a job (selfemployee or employee), no children, no debts, no savings and a low educational level. A slightly larger dispersion appears for residuals (s. Fig4) corresponding to the choice fairly easily for R1 combined with all possible responses on risk perception R2 (averse, tolerant and lover). Furthermore, we calculated the averages of the full-conditional Pearson’s residuals (Sect.6.3); Fig5 illustrates the box plots of these averages, for every covariate configuration at every time occasion. We observe quite small values overall, with the exception of two points having average of Pearson’s residuals greater or equal to 4, both at time t=3 (one of them not shown on the plot for better visualization purposes). They correspond both to female respondents with quite opposite profiles. In one of them they are not self-employee, with no children-debs-savings and a low education, while in the other, they have a job (selfemployee or employee), children-debs-savings and a high educational level. Table 2 Results of indices of quality of classification illustrated in Sect.6.2 for models M8 with k=3 and k=4 states of the latent construct k Sk SL k SU k SL|RS k SL|AWR k SU|l=1 k SU|l=2 k SU|l=3 k SU|l=4 k 3 0.757 0.824 0.747 0.847 0.837 0.823 0.813 0.821 4 0.712 0.783 0.738 0.801 0.805 0.872 0.736 0.785 0.823 24 R.Colombi et al. 1 3 7.2 Model interpretation Fig 6 allows us to characterize the answers of AWR and RS respondents in the three latent states, for both response variables. The top panels of Fig 6 illustrate the response probability functions of the perceived household’s financial ability to make ends meet R1 in the three stata of the latent construct for the AWR (colored bars) and RS (grey bars) regimes. According to Sect. 4.2, the estimates  𝜙011 and  𝜙111 reported in Table3 imply that the probability function of RS respondents in the state l=1 has mode at the middle category ( c1∕2 ) fairly easily, at the middle point ( c1∕2+1 ) fairly difficulty in the state l=2 , and at the extreme ( c1 ) very difficulty in l=3 , respectively. This means that individuals, in the group of financially safer households ( l=1 ), when in doubt about their perceived capability, tend to choose with more chance the middle category (MRS) fairly easily on the optimistic side of the scale, the uncertain households with a fair capability ( l=2 ) instead take refuge Fig. 4 Box plots of residuals for occasion time and configurations of responses 31 1 3 Hidden Markov models forlongitudinal rating data withdynamic… variables, in order to relax the assumption that, at a given time point t, RS affects all response variables or none, (iii) allowing covariate effects in the observation component, (iv) modelling unobserved heterogeneity on the transition probabilities. Under assumption A2 and if the conditional RS transition probabilities are homogeneous, the hypothesis 𝜋U|L (u | l,u)=d u (u ) of time invariance of the RS indicator constrains 2k parameters on the frontier of the parametric space. A test based on the log likelihood ratio statistic can be used but in this case the asymptotic distribution of the statistic is a mixture of chi-squared distributions known as chi bar squared distribution. The test can be easily implemented as shown in Bartolucci (2006) and Colombi and Forcina (2016) who dealt with related problems. Point (ii) above can be based on the approach of graphical HMMs by Colombi and Giordano (2015). Regarding (iii), Assumption B4 can be relaxed by modelling the observation probabilities as function of individual covariates as an alternative to the presence of covariate effects on the latent component. This can be the case when the main interest is on the observed responses and the latent variable serves to account for time dependence and respondent’s unobserved heterogeneity not explained by RS. If this is the focus of modelling, alternative models exist often referred to as Markov-switching logistic models and Markov-switching dynamic logistic models (e.g. Frühwirth-Schnatter 2006). Mixed HMMs and finite mixture of HMMs (e.g. Maruotti 2011; Bartolucci etal. 2012) are further classes of models in this direction that include time invariant random effects to model unit specific heterogeneity. It is also worth mentioning the models with multiple time varying random effects by Farcomeni (2015). Fig. 7 Individual transition probabilities of the latent financial capability averaged over time 𝜋 L i (l|  l ) 32 R.Colombi et al. 1 3 Finally, an additional extension can involve (continuous or discrete) random effects (e.g. Altman 2007; Bartolucci et al. 2012) on the transition probabilities to deal with time invariant unobserved individual heterogeneity (point iv). These effects would extend our model with additional individual-specific characteristics that remain constant throughout the observed period. Incorporating discrete random effects would allow us to account for group-level heterogeneity that affects the transition probabilities. Appendix: Standard errors We discuss three approaches to the estimation of the standard errors of the maximum likelihood estimator  𝜽 of 𝜽 . Hereafter, the upper index (m), m∈{L,U} , will be omitted from the vectors of covariates x(m) i and z (m) it to simplify the notation. Let yi , i∈I , be a realization of the Tr observable variables Yjit , j∈R , t∈T, collected in the vector Yi . The joint probability function of Yi , conditioned on the vector of covariates xi,zi ( zi is obtained by stacking the zit,t>1 ), is denoted by q(yi|xi,zi;𝜽) . The log-likelihood function of the observations yi , i∈I , is: and the vector of the score functions is: The calculation of standard errors can be based on OIM, OPIM, SDW methods, as mentioned in Sect.6.1. We here sketch briefly some technical details of the three methods, an alternative approach is based on the well known parametric bootstrap technique. The OIM can be computed using the Oakes identity (Oakes 1999): 𝓁 (𝜽)= n ∑ i=1 log q(yi | xi,zi;𝜽) , s (𝜽)= n ∑ i=1 𝜕log q(yi | xi,zi;𝜽) 𝜕𝜽= n ∑ i=1 si(𝜽) . Table 7 Homogenous conditioned RS transition probabilities 𝜋U|L(u|l , u) u  l u RS AWR RS Confident 0.981 0.019 Fair 0.996 0.004 Distressed 1 0 AWR Confident 0.123 0.877 Fair 0.072 0.928 Distressed 0.056 0.944 33 1 3 Hidden Markov models forlongitudinal rating data withdynamic… as shown by Bartolucci and Farcomeni (2015). The first term inside the square brackets is easy to compute, using the outputs of the last M step, as it is block diagonal with blocks given by the Hessian matrices of the six addends of (17). The computation of the second term inside the square brackets is more demanding as it requires the derivatives with respect to  𝜽 of log 𝛿 (1) it (u,l;  𝜽 ) and log 𝛿 (2) it (u,l;u,  l;  𝜽 ) . These derivatives can be obtained as an output of the Baum-Welch forward-backward algorithm as described by Bartolucci and Farcomeni (2015). For every element  𝜃h of  𝜽 , the terms of − 𝜕 2 Q ( 𝜽 | 𝜽 ) 𝜕 𝜃 h 𝜕𝜽� |  𝜽=𝜽 are obtained by the derivatives with respect to 𝜽′ of the six addends of (17) if 𝛿(1) it (u,l; 𝜽 ) and 𝛿(2) it (u,l;u, l; 𝜽 ) are replaced by 𝛿 (1) it (u,l; 𝜽)𝜕log 𝛿 (1) it (u,l;  𝜽 ) 𝜕 𝜃 h and 𝛿 (2) it (u,l;u, l; 𝜽)𝜕log 𝛿 (2) it (u,l;u,  l;  𝜽 ) 𝜕 𝜃 h , respectively. Notice that, when the necessary expected values and derivatives are obtained from the Baum-Welch forward-backward algorithm, the computation of the standard errors require repeated calls to a function that estimates logit models. Matrix J(𝜽) is estimated by  J=J(  𝜽) where  𝜽 is the MLE of 𝜽. The standard errors of the maximum likelihood estimators are estimated by the square roots of the diagonal elements of  J −1. If the RS-HMM is correctly specified, estimates of the standard errors can be also derived from the OPIM matrix I (𝜽)= ∑i s i (𝜽)s i (𝜽) � . The matrix I(𝜽) is estimated by  I=I(  𝜽) and the estimated standard errors of the maximum likelihood estimators are the square roots of the diagonal elements of  I −1. The matrix  I is easier to compute than  J , due to the effort needed to compute 𝜕2 Q(𝜽 | 𝜽) 𝜕 𝜽𝜕𝜽� |  𝜽=𝜽 . Remind that the HMM with a RS component is misspecified if there does not exist a 𝜽 such that 𝜏(y|x,z)=q(y|x,z;𝜽) with probability 1 where, for every x , z , 𝜏(y|x,z) is the true probability function generating the data. In this case,  𝜽 is a pseudo maximum likelihood estimator which is a consistent estimator of the pseudo-true value 𝜽 0=argmin𝜽 ( Ex,zEy𝜏(y | x,z)log 𝜏(y | x,z) q(y | x,z;𝜽) ) , see Vuong (1989) and White (1982). When the RS-HMM is misspecified, estimated standard errors of the pseudo maximum likelihood estimators are given by the square roots of the diagonal elements of  J −1  I  J −1. The estimators of the standard errors, obtained in this way, are robust in the sense that they are consistent, independently from the correct specification of the model. The matrix n  J −1  I  J −1 is a consistent estimator of the SDW matrix A( 𝜽 0)−1 B ( 𝜽 0) A ( 𝜽 0)−1 where A(𝜽0)=−Ex,zEy 𝜕 2 log q(y | x,z;𝜽 ) 𝜕𝜽𝜕𝜽 � and B (𝜽0)=Ex,zEy 𝜕log q ( y | x,z;𝜽 ) 𝜕𝜽 𝜕log q ( y | x,z;𝜽 ) 𝜕𝜽 � . The SDW matrix plays a central role in testing problems on misspecified models (Vuong 1989). As all models are possibly misspecified, the estimator of the standard errors based on the sandwich matrix should be always used in practice. However, computational complexity and numerical instability problems make the use of estimates based on the matrix  I more practical in the case of the model considered here. J (𝜽)=− 𝜕 2 𝓁(𝜽) 𝜕𝜽𝜕𝜽 �=− [ 𝜕 2 Q(𝜽 | 𝜽) 𝜕𝜽𝜕𝜽 � |  𝜽=𝜽 +𝜕 2 Q(𝜽 | 𝜽) 𝜕  𝜽𝜕𝜽 � |  𝜽=𝜽], 34 R.Colombi et al. 1 3 Funding Open access funding provided by Università della Calabria within the CRUI-CARE Agreement. The author Sabrina Giordano received partial financial support by Mur PRIN 2022, grant number 2022XRHT8R - The SMILE project (Statistical Modelling and Inference to Live the Environment) and by COST Action CA19130 “Fintech and Artificial Intelligence in Finance - Towards a transparent financial industry” (FinAI), funded by COST (European Cooperation in Science and Technology). Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http:// creat iveco mmons. org/ licen ses/ by/4. 0/. References Agresti A (2010) Analysis of ordinal categorical data. John Wiley & Sons, USA Aichholzer J (2013) Intra-individual variation of extreme response style in mixed-mode panel studies. Soc Sci Res 42(3):957–970 Altman RM (2007) Mixed hidden markov models: an extension of the hidden Markov model to the longitudinal data setting. J Am Stat Assoc 102(477):201–210 Anderson JA (1984) Regression and ordered categorical variables. J Roy Stat Soc B 46(1):1–22 Bachman JG, O’Malley PM (1984) Yea-saying, nay-saying, and going to extremes: black-white differences in response styles. Public Opin Q 48(2):491–509 Bartolucci F (2006) Likelihood inference for a class of latent Markov models under linear hypotheses on the transition probabilities. J Roy Stat Soc B 68:155–178 Bartolucci F, Farcomeni A (2015) Information matrix for hidden Markov models with covariates. Stat Comput 25:515–526 Bartolucci F, Farcomeni A, Pennoni F (2012) Latent Markov Models for Longitudinal Data. CRC Press, USA Bartolucci F, Lupparelli M, Montanari GE (2009) Latent Markov model for longitudinal binary data: an application to the performance evaluation of nursing homes. Ann Appl Statistics 3:611–636 Bartolucci F, Pandolfi S, Pennoni F (2017) LMest: an R package for latent Markov models for longitudinal categorical data. J Stat Softw 81(4):1–38 Baumgartner H, Steenkamp JBEM (2001) Response styles in marketing research: a cross-national investigation. J Mark Res 38:143–156 Bergsma W, Croon M, Hagenaars JA (2009) Marginal models: for dependent, clustered, and longitudinal categorical data. Springer, Berlin Berndt ER, Hall BH, Hall RE, Hausman JA (1974) Estimation and inference in nonlinear structural models. In Annals of Economic and Social Measurement, Volume 3, number 4, pp. 653–665. NBER Billiet JB, Davidov E (2008) Testing the stability of an acquiescence style factor behind two interrelated substantive variables in a panel design. Soc Methods & Res 36(4):542–562 Böckenholt U (2012) Modeling multiple response processes in judgment and choice. Psychol Methods 17(4):665–678 Böckenholt U, Meiser T (2017) Response style analysis with threshold and multi-process IRT models: a review and tutorial. Br J Math Stat Psychol 70(1):159–181 Bolt D, Johnson T (2009) Applications of a MIRT model to self-report measures: addressing score bias and DIF due to individual differences in response style. Appl Psychol Meas 33(5):335–352 Buckby J, Wang T, Jiancang Z, Obara K (2020) Model checking for hidden Markov models. J Comput Graph Stat 29(4):859–874 Colombi R, Forcina A (2016) Testing order restrictions in contingency tables. Metrika 79:73–90 Colombi R, Giordano S (2012) Graphical models for multivariate Markov chains. J Multivar Anal 107:90–103 Colombi R, Giordano S (2015) Multiple hidden Markov models for categorical time series. J Multivar Anal 140:19–30 35 1 3 Hidden Markov models forlongitudinal rating data withdynamic… Colombi R, Giordano S, Gottard A, Iannario M (2019) Hierarchical marginal models with latent uncertainty. Scand J Stat 46(2):595–620 Colombi R, Giordano S, Tutz G (2021) A rating scale mixture model to account for the tendency to middle and extreme categories. J Educ Behav Statistics 46(6):682–716 D’Alessio G, De Bonis R, Neri A, Rampazzi C (2020) Financial literacy in Italy: results from the 2020 Bank of Italy survey. Occasional Papers, Bank of Italy 588:1–46 De Blasio G, De Paola M, Poy S, Scoppa V (2021) Massive earthquakes, risk aversion, and entrepreneurship. Small Bus Econ 57(1):295–322 De Santis S, Bandyipadhyay D (2011) Hidden Markov models for zero-inflated Poisson counts with an application to substance use. Stat Med 30:1678–1694 Dolnicar S, Grün B (2009) Response style contamination of student evaluation data. J Mark Educ 31(2):160–172 Falk CF, Cai L (2016) A flexible full-information approach to the modeling of response styles. Psychol Methods 21(3):328–347 Farcomeni A (2015) Generalized linear mixed models based on latent Markov heterogeneity structures. Scand J Stat 42(4):1127–1135 Frühwirth-Schnatter S (2006) Finite mixture and Markov switching models, vol 425. Springer, Berlin Ghahramani Z, Jordan M (1997) Factorial hidden Markov models. Mach Learn 29:245–273 Grün B, Dolnicar S (2016) Response style corrected market segmentation for ordinal data. Mark Lett 27(4):729–741 Hedeker D, Gibbons D, Robert (2006) Longitudinal Data Analysis. Wiley Henninger M, Thorsten M (2020) Different approaches to modeling response styles in divide-by-total item response theory models (part 1): a model integration. Psychol Methods 25:560–576 HM-Treasury (2007) Financial capability: the Government’s long-term approach. HM Stationery Office, ISBN 978-1-84532-247-2 Huang HY (2016) Mixture random-effect IRT models for controlling extreme response style on rating scales. Frontiers in Psychology7, Article 1706 Kieruj ND, Moors G (2013) Response style behavior: question format dependent or personal style? Quality & Quantity 47(1):193–211 Koski T (2001) Hidden Markov models for bioinformatics, Volume 2. Springer Science & Business Media Lefevre AF, Chapman M (2017) Behavioural economics and financial consumer protection. OECD Working Papers on Finance, Insurance and Private Pensions42 Lin J, Bumcrot C, Ulicny T, Mottola G, Walsh G, Ganem R, Kieffer C, Lusardi A (2019) The state of US financial capability: the 2018 National Financial Capability Study. FINRA Investor Education Foundation Maruotti A (2011) Mixed hidden Markov models for longitudinal data: an overview. Int Stat Rev 79(3):427–454 Molenberghs G, Verbeke G (2005) Models for discrete longitudinal data. Springer, Berlin Nguyen L, Gallery G, Newton C (2019) The joint influence of financial risk perception and risk tolerance on individual investment decision-making. Accounting & Finance 59:747–771 Oakes D (1999) Direct calculation of the information matrix via the em algorithm. J Roy Stat Soc B 61(2):479–482 OECD (2020) OECD/INFE international survey of adult financial literacy competencies. https:// www. oecd. org/ finan cial/ educa tion/ launc hofth eoecd infeg lobal finan ciall itera cysur veyre port. htm Paulhus DL (1991) Measurement and control of response bias. In: Robinson JP, Shaver PR, Wrightsman LS (eds) Measures of personality and social psychological attitudes, Chapter2. Academic Press, New York, pp 17–60 Piccolo D, Simone R, Iannario M (2019) Cumulative and Cub models for rating data: a comparative analysis. Int Stat Rev 87(2):207–236 Pidgeon N (1998) Risk assessment, risk values and the social science programme: why we do need risk perception research. Reliability Eng Syst Safety 59(1):5–15 Pohle J, Adam T, Beumer LT (2022) Flexible estimation of the state dwell-time distribution in hidden semi-markov models. Comput Statistics & Data Anal 172:107479 Pohle J, Langrock R, Schaar Mvd, King R, Jensen FH (2021) A primer on coupled state-switching models for multiple interacting time series. Statistical Modelling 21(3):264–285 Roberts C (2016) The SAGE Handbook of Survey Methodology, Volume36, Chapter Response styles in surveys: understanding their causes and mitigating their impact on data quality, pp. 579–598. SAGE 36 R.Colombi et al. 1 3 Schildberg-Hörisch H (2018) Are risk preferences stable? J Econom Perspect 32(2):135–54 Slovic P (2010) The feeling of risk: new perspectives on risk perception. Earthscan, New York Soland J, Kuhfeld M (2020) Do response styles affect estimates of growth on social-emotional constructs? Evidence from four years of longitudinal survey scores. Multivariate Behavioral Research, Published online: 07 Jul 2020 Titman A (2009) Computation of the asymptotic null distribution of goodness-of-fit tests for multi-state models. Lifetime Data Anal 15:519–533 Tutz G (2021) Hierarchical models for the analysis of likert scales in regression and item response analysis. Int Stat Rev 89(1):18–35 Tutz G, Berger M (2016) Response styles in rating scales: simultaneous modeling of content-related effects and the tendency to middle or extreme categories. J Educ Behav Statistics 41:239–268 Valant J (2015) Improving the financial literacy of European consumers. EPRS, European Parliamentary Research Service, Members’ Research Service PE557 Van Vaerenbergh Y, Thomas TD (2013) Response styles in survey research: a literature review of antecedents, consequences, and remedies. Int J Public Opinion Res 25(2):195–217 Volant S, Bérard C, Martin-Magniette M-L, Robin S (2014) Hidden markov models with mixtures as emission distributions. Stat Comput 24(4):493–504 Vuong QH (1989) Likelihood ratio tests for model selection and non-nested hypotheses. Econometrica 57(2):307–333 Weijters B (2006) Response Styles in Consumer Research. Ph. D. thesis, Ghent University Weijters B, Geuens M, Schillewaert N (2010) The stability of individual response styles. Psychol Methods 15(1):96–110 Wetzel E, Carstensen CH (2015) Multidimensional modeling of traits and response styles. Eur J Psychol Assess 33(5):352–364 White H (1982) Maximum likelihood estimation of misspecified models. Econometrica 50(1):1–25 Yu S-Z (2010) Hidden semi-markov models. Artificial intelligence 174(2):215–243 Zhang Y, Wang Y (2020) Validity of three IRT models for measuring and controlling extreme and midpoint response styles. Front Psychol 11:271 Zottel S, Perotti V, Bolaji-Adio A (2013) Financial capability surveys around the world: why financial capability is important and how surveys can help. Technical report, The World Bank Zucchini W, MacDonald IL (2009) Hidden markov models for time series: an introduction using R. CRC Press Zucchini W, MacDonald IL, Langrock R (2017) Hidden Markov Models for Time Series: an Introduction using R. CRC Press, USA Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.