scieee AI-readable full text Open interactive document viewer

Measuring offending: Field experiments and improving the accuracy of self-reports of delinquent behavior

Gomes, Hugo Miguel dos Santos

Abstract

The body of knowledge on the causes and correlates of offending behavior is completely reliant on the quality of crime measures. However, methodological research on the assessment of offending behavior is very scarce. This doctoral dissertation aimed to assess the state of the art of crime measurement and to improve the accuracy of self-reports of offending (SRO). Chapter I describes a review of the advantages and limitations of the three main methodological techniques, i.e. official records, observation, and SRO. Considering the advantages of observation methods presented in this chapter, especially when applied within field experimental designs, we have carried out the systematic review presented in Chapter II. In this review, we have discussed the benefits of field experiments in the study of the etiology of offending. However, field experiments are very rarely used in the study of offending behavior, where SRO are the most widely used measurement method. In Chapter III, we have carried out a systematic review of methodological experiments testing potential sources of bias in SRO, providing relevant information to improve the accuracy of SRO. Taking into consideration the inconsistent results from methodological studies using SRO and other sensitive topics regarding the benefits of self-administration, we set out to assess the sensitivity of questions about offending behavior. In Chapter IV, we have developed a multidimensional assessment of question sensitivity and asked a total of 249 students to rate the sensitivity of several behavioral variables, which included offending behaviors. Results demonstrated that questions about offending behavior are perceived as highly sensitive. Further, we have included an experimental manipulation that allowed us to show that questions about offending behavior occurring over a distant time period are perceived as less sensitive than questions about recent offending. Chapter V presents two methodological experiments with a 2 (interviewer-administered vs. self-administered) × 2 (paper-andpencil vs. computer interviews) factorial design. The first experiment was carried out in Portugal (N = 181), and the second was a replication study with students from a University in Florida (N = 154). Findings showed an increased odds of reporting offending behavior in self-administered surveys, suggesting that SRO provide more accurate estimates of offending behavior using self-administered surveys. Finally, we have included a general discussion on the main findings from this dissertation, highlighting the major contributions and implications on behavioral assessment.

Full text

Universidade do Minho Escola de Psicologia Hugo Miguel dos Santos Gomes July 2021 Measuring offending: Field experiments and improving the accuracy of self-reports of delinquent behavior Hugo Miguel dos Santos Gomes Measuring offending: Field experiments and improving the accuracy of self-reports of delinquent behavior UMinho|2021 Hugo Miguel dos Santos Gomes July 2021 Measuring offending: Field experiments and improving the accuracy of self-reports of delinquent behavior Work supervised by Professor Ângela Maia and Professor David P. Farrington Doctoral Thesis PhD in Applied Psychology Universidade do Minho Escola de Psicologia ii DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Atribuição-NãoComercial-SemDerivações CC BY-NC-ND https://creativecommons.org/licenses/by-nc-nd/4.0/ iii ACKNOWLEDGMENTS “If I have seen further it is by standing on the shoulders of Giants” Isaac Newton (1675) This doctoral dissertation would not have been possible without the great assistance and support from the kind people I have had the pleasure to deal with during this journey. I would like to thank my supervisors, Professor Ângela Maia, Professor David Farrington, and Professor Marvin Krohn. I could not have asked for better scientific and personal advisers. Through our time working together I was exposed to your brilliancy, work ethic, and caring for others. You have become role models that I aspire to live up to during my career. To all my friends and colleagues from the University of Minho, the University of Cambridge, and the University of Florida, that I have had the pleasure to meet along this doctoral journey. This doctoral dissertation was supported by the Fundação para a Ciência e a Tecnologia (FCT - SFRH/BD/122919/2016) and by the Fulbright Commission Portugal, which allowed me to become a Ph.D. visiting student at the Institute of Criminology, University of Cambridge, and a Fulbright Scholar at the Department of Sociology and Criminology & Law, University of Florida. Above all, I would like to thank my wife Joana for her tireless support and patience during my doctoral work. iv STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho. v MEASURING OFFENDING: FIELD EXPERIMENTS AND IMPROVING THE ACCURACY OF SELF-REPORTS OF DELINQUENT BEHAVIOR ABSTRACT The body of knowledge on the causes and correlates of offending behavior is completely reliant on the quality of crime measures. However, methodological research on the assessment of offending behavior is very scarce. This doctoral dissertation aimed to assess the state of the art of crime measurement and to improve the accuracy of self-reports of offending (SRO). Chapter I describes a review of the advantages and limitations of the three main methodological techniques, i.e. official records, observation, and SRO. Considering the advantages of observation methods presented in this chapter, especially when applied within field experimental designs, we have carried out the systematic review presented in Chapter II. In this review, we have discussed the benefits of field experiments in the study of the etiology of offending. However, field experiments are very rarely used in the study of offending behavior, where SRO are the most widely used measurement method. In Chapter III, we have carried out a systematic review of methodological experiments testing potential sources of bias in SRO, providing relevant information to improve the accuracy of SRO. Taking into consideration the inconsistent results from methodological studies using SRO and other sensitive topics regarding the benefits of self-administration, we set out to assess the sensitivity of questions about offending behavior. In Chapter IV, we have developed a multidimensional assessment of question sensitivity and asked a total of 249 students to rate the sensitivity of several behavioral variables, which included offending behaviors. Results demonstrated that questions about offending behavior are perceived as highly sensitive. Further, we have included an experimental manipulation that allowed us to show that questions about offending behavior occurring over a distant time period are perceived as less sensitive than questions about recent offending. Chapter V presents two methodological experiments with a 2 (interviewer-administered vs. self-administered) × 2 (paper-andpencil vs. computer interviews) factorial design. The first experiment was carried out in Portugal ( N = 181), and the second was a replication study with students from a University in Florida ( N = 154). Findings showed an increased odds of reporting offending behavior in self-administered surveys, suggesting that SRO provide more accurate estimates of offending behavior using self-administered surveys. Finally, we have included a general discussion on the main findings from this dissertation, highlighting the major contributions and implications on behavioral assessment. Keywords: Field experiments; Modes of administration; Offending; Self-report; Sensitive questions vi MEDIDAS DE CRIME: UM CONTRIBUTO PARA AS EXPERIÊNCIAS DE CAMPO E PARA A OTIMIZAÇÃO DOS AUTORRELATOS DE COMPORTAMENTO DELINQUENTE RESUMO O conhecimento acerca das causas do comportamento delinquente está totalmente dependente da qualidade das medidas de crime. No entanto, a investigação metodológica sobre as medidas do comportamento delinquente é muito limitada. A presente dissertação teve como objetivo avaliar o estado da arte da avaliação de crimes, bem como otimizar a precisão dos autorrelatos de comportamento delinquente (ACD). O Capítulo I apresenta uma revisão da literatura sobre as vantagens e desvantagens das três principais técnicas de medida de crime, i.e. registos oficiais, observação e ACD. Tendo em conta as vantagens dos métodos de observação, especialmente quando aplicados em experiências de campo, realizámos a revisão sistemática apresentada no Capítulo II. Nesta revisão, discutimos os benefícios das experiências de campo no estudo da etiologia da delinquência. No entanto, estas experiências apenas raramente são utilizadas no estudo do comportamento delinquente, onde os ACD são o método mais utilizado. No Capítulo III, realizámos uma revisão sistemática da literatura sobre as experiências metodológicas que testam potenciais fontes de enviesamento nos ACD, fornecendo informações relevantes para a otimização dos ACD. Tendo em conta a inconsistência entre os estudos metodológicos usando ACD e relatos de tópicos sensíveis em relação aos benefícios da autoadministração, no Capítulo IV, criámos uma avaliação da sensibilidade das questões e recrutámos 249 estudantes universitários para realizarem uma avaliação da sensibilidade dos ACD. Os resultados demonstraram que questões sobre crimes são tópicos altamente sensíveis. Adicionalmente, incluímos uma manipulação experimental que nos permitiu demonstrar que questões sobre crimes ocorridos há mais tempo são percebidas como menos sensíveis do que questões sobre crimes recentes. O Capítulo V apresenta duas experiências metodológicas com um design fatorial de 2 (entrevista cara-a-cara vs. autoadministração) x 2 (papel-elápis vs. computador). A primeira experiência foi realizada em Portugal ( N = 181) e a segunda consiste numa replicação com estudantes de uma universidade da Flórida ( N = 154). Os resultados destas experiências revelaram um aumento no relato de comportamentos delinquentes no formato de questionários autoadministrados, sugerindo que os ACD fornecem estimativas de crime com maior precisão em condições de autoadministração. Por fim, incluímos uma discussão geral sobre as principais conclusões desta dissertação, destacando os seus principais contributos e implicações. Palavras-chave: Autorrelatos; Crime; Experiências de campo; Modos de administração; Questões sensíveis vii TABLE OF CONTENTS INTRODUCTION.............................................................................................................1 Measures of offending behavior ............................................................................................................2 Observation methods within field experiments ........................................................................................4 Self-report methodology ......................................................................................................................6 Sensitive questions .........................................................................................................................6 Modes of administration ..................................................................................................................8 Measurement bias in self-reports of offending ................................................................................... 10 The present dissertation .................................................................................................................... 12 References ...................................................................................................................................... 15 CHAPTER I. MEASURING OFFENDING: SELF-REPORTS, OFFICIAL RECORDS, SYSTEMATIC OBSERVATION AND EXPERIMENTATION ................................................ 26 Abstract .......................................................................................................................................... 27 Introduction..................................................................................................................................... 28 Definition and units of measurement ............................................................................................... 28 Measure of crime ......................................................................................................................... 29 Comparing official records and self-reports of offending ..................................................................... 32 Scaling-up factor .......................................................................................................................... 33 Criminal career research ............................................................................................................... 34 Self-reports of offending ................................................................................................................ 37 Alternative methods for measuring crime ......................................................................................... 40 Conclusions .................................................................................................................................... 42 References ...................................................................................................................................... 45 CHAPTER II. FIELD EXPERIMENTS ON DISHONESTY AND STEALING: WHAT HAVE WE LEARNED IN THE LAST 40 YEARS?.............................................................................. 51 Abstract .......................................................................................................................................... 52 Introduction..................................................................................................................................... 53 Experimental approach ................................................................................................................. 53 Laboratory versus field experiments ................................................................................................ 54 Field experiments in the study of deviance ....................................................................................... 55 Theoretical framework: factors influencing deviance .......................................................................... 56 The present study ........................................................................................................................ 57 xiv LIST OF TABLES Table 1. Descriptive information on 60 studies in the systematic review ..................................63 Table 2. Summary of field experiments in the Fraudulent/ dishonest behavior category. .........85 Table 3. Summary of field experiments in the Stealing category ..............................................93 Table 4. Summary of field experiments in the Keeping money category ................................ 100 Table 5. Summary of field experiments in the Shoplifting category ....................................... 104 Table 6. Descriptive information on studies in the systematic review .................................... 127 Table 7. Main findings of experiments in the systematic review ............................................ 130 Table 8. Average question sensitivity of behavioral items ..................................................... 158 Table 9. Mean comparisons of question sensitivity by recall period ...................................... 159 Table 10. Average question sensitivity of behavioral items for the American pilot study ........ 162 Table 11. Demographic characteristics by experimental manipulations (experiment 1) ......... 177 Table 12. Experiment 1: Prevalence of offending and variety by modes of administration (left) and by modes of data collection (right) .................................................................................. 178 Table 13. Demographic characteristics by experimental manipulations (experiment 2) ......... 184 Table 14. Experiment 2: Prevalence of offending and variety by modes of administration (left) and by modes of data collection (right) .................................................................................. 185 1 INTRODUCTION 2 The study of the causes and correlates of offending has generated a large body of knowledge about the etiology of criminal behavior. From a developmental and life-course perspective, the acquired knowledge about the patterns of offending, risk and protective factors, as well as the effect of life events, allowed a comprehensive theoretical understanding of the development of offending (e.g., Farrington, 2005; Farrington et al., 2019; Gibson & Krohn, 2012; Moffitt, 1993). This knowledge allows the prediction of future offending and plays a major role in the development of early prevention strategies and effective interventions (e.g., Fagan et al., 2019; Farrington, 2021; Farrington & Coid, 2003; Rijo et al., 2020; Zara & Farrington, 2016). However, knowledge about the development of offending behavior is completely reliant on the quality of crime measures. Inaccurate or biased measures of offending behavior will inevitably result in misleading conclusions about the predictors and patterns of offending and, in turn, result in poor policies and interventions (Livingston, 2013; Pepper & Petrie, 2003). This makes it very important for researchers to use the best possible practices for measuring offending behavior. Nevertheless, the assessment of criminal behavior is particularly demanding and there is a ceiling to the accuracy of crime estimates (Krohn et al., 2012; Sullivan & McGloin, 2014). Offending behavior is not only a sensitive and socially undesirable matter, it involves illegal practices that are punishable by law and people naturally try to conceal it. All these aspects of offending add to the already challenging task of assessing human behavior, making crime measurement an inherently difficult task (Osgood et al., 2002). Measures of offending behavior In the present dissertation, we started by asking a fundamental research question. “What are the main measures of offending behavior?” In order to provide an answer to this question, we have carried out a review of the literature on the major crime measurement methodologies, reviewing their advantages and limitations (Gomes et al., 2018). In this review of the literature, presented in Chapter I, we concluded that there are three main methodological techniques of crime assessment. First, official records, which consist of the consultation of officially recorded information by the police, prisons, and/or the courts regarding the practice of crimes. Second, researchers may use direct and indirect observation techniques to assess offending behavior. Third, self-reports of offending (SRO), where people are asked whether they have practiced several types of offenses (Maxfield & Babbie, 2009). Because observation techniques are very difficult to implement, official records and SRO are the two most widely used measurement methods in the study of criminal behavior (Piquero et al., 2014). However, there is considerable controversy about the best measures of crime, as well as the best conditions in which to collect such data. 3 For many years, research on criminal behavior relied mostly on data obtained from official records (Thornberry & Krohn, 2000). However, many researchers criticized this methodology, mainly because official records seriously underreported the true amount of offending behavior (e.g., Murphy et al., 1946) and because the obtained criminal data varied depending on whether the officially recorded information was provided by the police, courts, and/or prisons, which could lead to completely different conclusions (Sellin, 1931). Farrington and Jolliffe (2004) made the similar observation that only part of the total crimes committed are reported to the police, from which only a part are recorded by the police, out of which only a fraction result in convictions, and so on in a successive funneling process. In this discussion regarding the accuracy of crime measurements provided by different records (i.e., police, judicial, or penal statistics), Sellin (1931, p. 346) made a very important observation that “the value of a crime rate for index purposes decreases as the distance from the crime itself in terms of procedure increases”. If we apply the ‘Sellin’s dictum’ onto the broader aspects of crime measurement, observation of offending behavior may be regarded as the most valuable assessment, where the behavior is assessed directly without any funneling or other biasing aspects described above. In our review (Gomes et al., 2018), we identified some studies using direct field observations to assess offending behavior, such as shoplifting (e.g., Buckle & Farrington, 1984, 1994). Others used indirect field observation methods to assess offending by creating opportunities for people to steal coins left in telephone booths (e.g., Bickman, 1971; Franklin, 1973) or money from apparently ‘lost letters/wallets’ (e.g., Farrington & Knight, 1979, 1980; Hornstein et al., 1968; Merritt & Fowler, 1948). However, field observation of offending behavior is a challenging task, mainly because offending is unpredictable and offenders actively try for their offenses not to be observed (Buckle & Farrington, 1984; Gomes et al., 2018). For all these reasons, studies using observation methods to test hypotheses relating to the causes of offending behavior are very scarce (Farrington et al., 2020). At the same time, the limitations of the criminal data provided by official records led researchers to apply the self-report technique to assess offending behavior. In 1943, Porterfield published the first study using the self-report methodology to measure delinquent behavior. But it was Short and Nye’s work on SRO across socio-economic status that fully displayed the potential of the self-report technique in etiological studies (Nye et al., 1958; Short & Nye, 1958), and which revolutionized researchers’ opinion on the utility and feasibility of SRO (Thornberry & Krohn, 2000). Following Short and Nye’s work, SRO became more and more used and, in the next decade, Hirschi (1969) developed a highly influential study on the etiology of delinquent behavior solely based on the self-report methodology. 4 Still, many researchers continued to cast doubt about the ability of respondents to provide useful information regarding their own criminal behavior through self-reports (e.g., Gibbons, 1979). This motivated a large body of research on the psychometric qualities of SRO that still stands until today (e.g., Ahonen et al., 2020; Auty et al., 2015; Farrington, 1973; Farrington et al., 2014; Gold, 1966; Hindelang et al., 1981; Huizinga & Elliott, 1986; Jolliffe et al., 2003; Kazemian & Farrington, 2005; Piquero et al., 2014; Yan & Cantor, 2019). These studies repeatedly showed SRO as a valid and reliable measure of delinquent behavior, making self-reports one of the most used measurement methods in the contemporary study of criminal behavior (Jolliffe et al., 2003). The gradual improvement and the widespread application of SRO completely revolutionized our knowledge about delinquent behavior (Thornberry & Krohn, 2000). From being regarded as a taboo topic by early scholars, self-reports came a long way into being considered “the most significant methodological innovation to date in our pursuit of understanding criminal behavior” (Krohn et al., 2012, p. 23). The literature reviewed in Chapter I (Gomes et al., 2018) shows that the validity of crime measures is bounded by a definite ceiling, and that perfect assessment of offending behavior is beyond the reach of contemporary measurement methods (e.g., Krohn et al., 2012; Sullivan & McGloin, 2014). Each methodology presents its own set of advantages and limitations, whereby a mixed-methods approach might result in the best assessment of the offending phenomenon. Nevertheless, researchers and practitioners must consider the specific qualities of each measurement technique and select the method(s) that best fit their research questions (for a discussion see Gomes et al., 2018). Observation methods within field experiments According to the literature included in our review of offending measures (Gomes et al., 2018), observation techniques provide the most valid information. Observation is the data source closest to the actual offending behavior, which eliminates many potential biasing factors. Through observations, researchers are able to assess the behavior of participants in the real world without them being aware that their behavior is being assessed. These characteristics are very important because they make it possible to test cause-and-effect relationships within naturalistic field experiments (Farrington, 1979). Field experiments combine the benefits of the experimental design and the external validity of testing hypotheses in the real world. Contrary to the cross-sectional and longitudinal studies, the experimental design makes it possible to test cause-and-effect relationships through the manipulation of variables under strictly controlled conditions (Christensen, 1985; Zimny, 1961). This makes experiments crucial for the development of scientific knowledge because they provide unambiguous conclusions about 5 the variables affecting human behavior. On the other hand, field experiments overcome the limitations of the artificiality of laboratory experiments where participants are aware that their behavior is being scrutinized (Farrington, 1980; Harrison & List, 2004). In the laboratory, the research setting may influence participants’ behavior in multiple ways (e.g., social desirability), which compromises its internal validity (Levitt & List, 2007). In the particular case of offending and deviant behaviors, this concern is especially relevant because people naturally try to conceal undesirable behaviors (Gomes et al., 2018). Considering these limitations, naturalistic field experiments provide the greatest internal and external validity (Farrington, 1979). In 1979, Farrington carried out a review of field experiments on deviance with special reference to dishonesty. This review included field experiments using multiple techniques to observe unaware participants acting dishonestly. For example, researchers left apparently ‘lost’ coins and observe whether or not members of the public dishonestly claimed them (e.g., Farrington & Kidd, 1977; Feldman, 1968; Korte & Kerr, 1975). Some experiments included in this review were able to actually observe offending behavior, such as theft (e.g., Diener et al., 1976; Steinberg et al., 1977). Faced with the scarcity of this robust design, Farrington (1979, p. 242) concluded by expressing his hope “that psychologists will have the ingenuity, determination, and social responsibility to meet the challenge of experiments on deviance”. Despite the benefits of observation of offending in real-life settings, especially when applied in experimental designs, most research on the causes of offending behavior is nonexperimental and field experiments are rare in social science (Franzen & Pointner, 2013; Gomes et al., 2018). However, multiple naturalistic field experiments have been conducted by behavioral economists (e.g., Harrison & List, 2004; Levitt & List, 2009). Several of these real-world experiments use field observations that are very relevant to the study of offending (Farrington et al., 2020), such as stealing and monetary dishonesty (for a review see Rosenbaum et al., 2014). Kerschbamer et al. (2016), for example, used computers with prearranged defects to study fraud in the computer repair price. Cohn et al. (2019) studied civic honesty in 40 countries by using apparently ‘lost’ wallets, providing the opportunity to members of the public to steal. Balafoutas et al. (2013) resorted to GPS data to test the dishonest behavior of taxi drivers by comparing the chosen route to the estimated correct fare. In order to provide a review of the field methods used to assess participants’ deviant and dishonest behavior in the real world, we have carried out a systematic review of field experiments seeking to study the causes of offending or monetary dishonesty that have been reported since the review of Farrington (1979). This systematic review, presented in Chapter II (Gomes et al., 2021a), illustrates the potential of field experiments to study the causes of offending and dishonest behavior in the real world, 6 which we hope will inform and motivate more researchers to apply such methods in the study of the causes of offending. However, the field experimental design is still very rarely used in the study of offending behavior, which is dominated by the self-report methodology. Self-report methodology SRO are the most widely used method of measuring criminal and deviant behavior (Gomes et al., 2018). However, despite the large effort to establish the validity of SRO, especially comparing data obtained using self-reports to official records (e.g., Clark & Tifft, 1966; Hardt & Peterson-Hardt, 1977; Kulik et al., 1968; Schore et al., 1979), much less attention has been given to the study of measurement biases and cognitive processes associated with the disclosure of offending behavior. Survey researchers, on the other hand, have developed a large body of knowledge on the processes underlying survey responses and how questions shape participants’ answers (e.g., Schwarz, 1999). Multiple cognitive processes are involved in providing information about one’s own behavior. Prior to providing an accurate estimation, survey respondents have to comprehend the question, recall relevant information, and compute a judgment through adding, averaging, and combining behavioral information (Schwarz, 1999; Tourangeau et al., 2000). Measurement error may occur in all of these processes. Asking questions about sensitive behaviors adds a further layer of potential bias because respondents may deliberately edit their answers in order to avoid disclosing socially undesirable information (Bradburn et al., 1979; Sudman & Bradburn, 1974; Tourangeau & Yan, 2007). Sensitive questions Over the past decades, researchers have used self-report questionnaires to study increasingly sensitive topics (Tourangeau & Yan, 2007). Tourangeau and colleagues (Tourangeau et al., 2000; Tourangeau & Yan, 2007) provided a three-dimensional definition of question sensitivity (i.e., intrusiveness, threat of disclosure, and social desirability). First, intrusiveness refers to questions that are themselves an invasion of privacy. Respondents may feel that these questions are inappropriate and none of the researcher’s business, whether or not the respondents have themselves engaged in such behavior. For example, respondents may feel that a question about stealing is an invasion of privacy, regardless if they have ever stolen something. Threat of disclosure, on the other hand, refers to the respondent’s concern about their truthful answers becoming known to a third party. In this case, the question’s sensitivity is dependent on the respondent’s previous behavior. A question about stealing, for example, 7 poses no threat of disclosure for someone who has never engaged in such illegal practice. However, respondents who have stolen in the past may fear potential consequences if their honest answers become known by their employer, their parents, etc. Third, social desirability reflects the extent to which a question elicits socially desirable answers. Considering that stealing is a socially undesirable behavior, a question about stealing may be regarded as sensitive because the socially desirable answer would be to deny this practice. These specific features of sensitive questions may compromise response accuracy by decreasing the likelihood of participants providing truthful answers to questions about sensitive behaviors (Tourangeau & Yan, 2007). In fact, evidence suggests that much of the misreporting found in self-reports of sensitive topics is a consequence of a motivated process of respondents editing their answers (Tourangeau & Yan, 2007). According to the motivated misreporting hypothesis, respondents who have engaged in socially undesirable behaviors will deliberately edit their responses in a socially desirable way in order to provide a positive image of themselves (Sudman & Bradburn, 1974; Tourangeau et al., 2000). Further, as the topics of the questions become more sensitive, the respondents’ motivation to edit their answers increases, progressively compromising response quality (Tourangeau & Yan, 2007). One of the most replicated effects of asking sensitive questions is the tendency of respondents to systematically underreport socially undesirable behaviors (Krumpal, 2013; Tourangeau et al., 2000). Methodological experiments have provided evidence that survey respondents underreport sensitive behaviors such as food intake (e.g., Wehling & Lusher, 2019), risky sexual behaviors (e.g., Giguère et al., 2019), substance use such as cigarettes (Liber & Warner, 2018), alcohol (e.g., Kabashi et al., 2019; Littlefield et al., 2017; Vinikoor et al., 2018), and other drugs (e.g., Druckman et al., 2015; Gerdtz et al., 2020; Kirtadze et al., 2018; Palamar et al., 2021), as well as deviant and criminal behaviors (e.g., Clark & Tifft, 1966; Wolter & Laier, 2014). Further, and in accordance with the motivated misreporting hypothesis, Hser (1997) found that underreporting is more evident for highly sensitive drugs (e.g., cocaine and opiates) than for less sensitive drugs (e.g., marijuana). In trying to circumvent the tendency to underreport socially undesirable behaviors, survey researchers have implemented data collection strategies to improve participants’ willingness to report sensitive information. The bogus pipeline, for example, consists of attaching a device to the participants that they believe can detect false reports. This technique results in an increased rate of self-reported sensitive behavior (e.g., Strang & Peterson, 2020). Similarly, randomized response techniques such as the item count technique (e.g., Wolter & Laier, 2014) or the unmatched count technique (e.g., Dalton et al., 1994), where participants’ reports of behavior are indirectly estimated without asking them to explicitly 8 reveal their sensitive behavior, consistently result in higher rates of disclosure than traditional direct selfreports (Druckman et al., 2015; Kirtadze et al., 2018). The systematic tendency of respondents to underestimate the true prevalence of sensitive behaviors, as well as the consistently higher rates of sensitive behavior obtained in conditions where the threat of disclosure is reduced (i.e., randomized response techniques) and honesty requests are heightened (i.e., bogus pipeline) cannot be explained by chance. Further, if these effects resulted from comprehension or memory faults, the response errors would be expected to be found in both directions (i.e., over and underreports). However, inaccurate responding occurs systematically in the socially desirable direction. These findings are solid evidence that respondents to sensitive questions deliberately edit their answers (Bradburn et al., 1979; Tourangeau et al., 2000). Taking into account the tendency of respondents to underreport the true amount of sensitive behaviors, survey researchers often use the ‘more is better’ assumption to determine which research method provides the most accurate reports. Even though this is just an assumption and researchers should use an external criterion for self-reported information whenever possible (e.g., biomarkers of drug use), the ‘more is better’ assumption is very useful in the study of behaviors where no gold standard can be applied, such as offending behavior. Using this assumption, survey researchers are able to experimentally compare different methods, such as different modes of administration. The modalities that result in higher reporting rates of socially undesirable behavior are assumed to be the most likely to yield accurate results (Tourangeau & Yan, 2007). Modes of administration Modes of administration are key fundamental features of the self-report methodology that can have a substantial impact on the quality of behavioral reports (Richman et al., 1999; Tourangeau & Yan, 2007). Survey information may be collected using very different types of modes of administration. Two of the most relevant variables in administration modalities are 1. whether or not respondents provide their answers to an interviewer (i.e., self-administration); and 2. whether the questions are presented on a piece of paper or on a computer. The combination of these variables provides four modes of administration that are the most typically used in behavioral assessment, i.e., paper-and-pencil personal interviews (PAPI), computer-assisted personal interviews (CAPI), paper-and-pencil self-administered questionnaires (SAQ), and computer-assisted self-administered interviews (CASI) (Thornberry & Krohn, 2000). 9 Methodological research shows that the self-administration of surveys significantly affects participants’ responses to sensitive questions (Sudman & Bradburn, 1974). Experimental studies comparing interviewer-administered and self-administered questionnaires consistently find in higher rates of admissions of socially undesirable behaviors in self-administered conditions (e.g., Aquilino, 1994; Butler et al., 2009; Jobe et al., 1997; Kreuter et al., 2008; Lee et al., 2019; Robertson et al., 2018; Schober et al., 1992; Turner et al., 1992). Tourangeau and Yan (in press) reviewed seven methodological experiments (54 effect sizes) on the effect of modes of administration on self-reports of illicit drug use and estimated that self-administration caused an increase of about 30% in drug use admissions. The findings of mode effects in reporting sensitive information are consistent with the motivated misreporting hypothesis. Face-to-face interviews require participants to verbally disclose socially undesirable information to a third person. Under self-administered conditions, respondents provide their answers directly on a piece of paper or on the computer, removing the interviewer from the data collection process and mitigating the concerns with self-image. In turn, self-administration of surveys provides an increased perception of confidentiality and anonymity which results in an increased willingness to provide socially undesirable information (Schwarz et al., 1991; Sudman & Bradburn, 1974; Tourangeau & Yan, 2007). Furthermore, the benefits of self-administration tend to be higher for more sensitive topics (Tourangeau et al., 2000; Tourangeau & McNeeley, 2003; Tourangeau & Yan, in press). Methodological experiments testing the effects of self-administration on reports of illicit drug use typically find that the mode effect is larger for reports of cocaine than for marijuana use (e.g., Aquilino, 1994; Schober et al., 1992; Turner et al., 1992). Similarly, Richman et al. (1999) carried out a meta-analysis with 61 methodological experiments (673 effect sizes) and found evidence that self-administration causes an increase in the likelihood of participants reporting sensitive behaviors (e.g., illegal drug use, risky sexual behavior, etc.), while for low sensitivity topics such as job satisfaction and personality scales reports remained similar through the different modes of administration. In line with these findings, authors such as Bradburn et al. (2004) have suggested that the disclosure of socially undesirable information regarding current behavior is more threatening than disclosing behavioral information that may have occurred in a distant past. In fact, there is evidence that the benefits of self-administration tend to be higher when asking questions about recent behavior compared with questions about behavior that may have occurred in the distant past (Tourangeau et al., 2000; Tourangeau & McNeeley, 2003; Tourangeau & Yan, in press). In their experiments, Turner et al. (1992), as well as Schober et al. (1992), included questions about illicit drug use over the lifetime, the previous year, and the previous month. In these experiments, the benefits of self-administration over 16 school students. Public Opinion Quarterly, 70 (3), 354–374. https://doi.org/10.1093/poq/nfl003 Buckle, A., & Farrington, D. P. (1984). An observational study of shoplifting. British Journal of Criminology, 24 (1), 63-73. https://doi.org/10.1093/oxfordjournals.bjc.a047425 Buckle, A., & Farrington, D. P. (1994). Measuring shoplifting by systematic observation: A replication study. Psychology, Crime and Law, 1 (2), 133-141. https://doi.org/10.1080/10683169408411946 Butler, S. F., Villapiano, A., & Malinow, A. (2009). The effect of computer-mediated administration on selfdisclosure of problems on the Addiction Severity Index. Journal of Addiction Medicine, 3 (4), 194203. https://doi.org/10.1097/ADM.0b013e3181902844 Christensen, L. B. (1985). Experimental methodology (3rd ed.). Allyn & Bacon. Clark, J. P., & Tifft, L. L. (1966). Polygraph and interview validation of self-reported deviant behavior. American Sociological Review, 31 (4), 516-523. https://doi.org/10.2307/2090775 Cohn, A., Maréchal, M. A., Tannenbaum, D., & Zünd, C. L. (2019). Civic honesty around the globe. Science, 365 (6448), 70-73. https://doi.org/10.1126/science.aau8712 Dalton, D. R., Wimbush, J. C., & Daily, C. M. (1994). Using the unmatched count technique (UCT) to estimate base rates for sensitive behavior. Personnel Psychology, 47 (4), 817-829. https://doi.org/10.1111/j.1744-6570.1994.tb01578.x Denniston, M. M., Brener, N. D., Kann, L., Eaton, D. K., McManus, T., Kyle, T. M., Roberts, A. M., Flint, K. H., & Ross, J. G. (2010). Comparison of paper-and-pencil versus Web administration of the Youth Risk Behavior Survey (YRBS): Participation, data quality, and perceived privacy and anonymity. Computers in Human Behavior, 26 (5), 1054-1060. https://doi.org/10.1016/j.chb.2010.03.006 Diener, E., Fraser, S. C., Beaman, A. L., & Kelem, R. T. (1976). Effects of deindividuation variables on stealing among Halloween trick-or-treaters. Journal of Personality and Social Psychology, 33 (2), 178-183. https://doi.org/10.1037/0022-3514.33.2.178 Dodou, D., & de Winter, J. C. (2014). Social desirability is the same in offline, online, and paper surveys: A meta-analysis. Computers in Human Behavior, 36 , 487-495. https://doi.org/10.1016/j.chb.2014.04.005 Druckman, J. N., Gilli, M., Klar, S., & Robison, J. (2015). Measuring drug and alcohol use among college student‐athletes. Social Science Quarterly, 96 (2), 369-380. https://doi.org/10.1111/ssqu.12135 17 Enzmann, D., Kivivuori, J., Marshall, I. H., Steketee, M., Hough, M., & Killias, M. (2018). A global perspective on young people as offenders and victims: First results from the ISRD3 study . Springer. https://doi.org/10.1007/978-3-319-63233-9 Fagan, A. A., Hawkins, J. D., Catalano, R. F., & Farrington, D. P. (2019). Communities that Care: Building Community engagement and capacity to prevent youth behavior problems . Oxford University Press. https://doi.org/10.1093/oso/9780190299217.001.0001 Farrington, D. P. (1973). Self-reports of deviant behavior: Predictive and stable? Journal of Criminal Law and Criminology, 64 (1), 99-110. https://doi.org/10.2307/1142661 Farrington, D. P. (1979). Experiments on deviance with special reference to dishonesty. In L. Berkowitz (Ed.), Advances in experimental social psychology (Vol. 12, pp. 207-252). Academic Press. https://doi.org/10.1016/S0065-2601(08)60263-4 Farrington, D. P. (1980). External validity: A problem for social psychology. In R. F. Kidd & M. J. Saks (Eds.), Advances in applied social psychology (Vol. 1, pp. 184-186). Lawrence Erlbaum. https://doi.org/10.4324/9781315803005 Farrington, D. P. (2005, Ed.). Integrated developmental and life-course theories of offending: Advances in criminological theory (Vol. 14). Routledge. https://doi.org/10.4324/9780203788431 Farrington, D. P. (2021). The developmental evidence base: Prevention. In D. A. Crighton & G. J. Towl (Eds.), Forensic Psychology (3rd ed., pp. 263-293). Wiley. Farrington, D. P., & Coid, J. W. (Eds.). (2003). Early prevention of adult antisocial behaviour . Cambridge University Press. https://doi.org/10.1017/CBO9780511489259 Farrington, D. P., & Jolliffe, D. (2004). England and Wales. In D. P. Farrington, P. A. Langan, & M. Tonry (Eds.), Cross-national studies in crime and justice (pp. 1–38). Bureau of Justice Statistics. http://www.ojp.usdoj.gov/bjs Farrington, D. P., Kazemian, L., & Piquero, A. R. (Eds.). (2019). The Oxford handbook of developmental and life-course criminology . Oxford University Press. https://doi.org/10.1093/oxfordhb/9780190201371.001.0001 Farrington, D. P., & Kidd, R. F. (1977). Is financial dishonesty a rational decision? British Journal of Social and Clinical Psychology, 16 (2), 139-146. https://doi.org/10.1111/j.20448260.1977.tb00209.x Farrington, D. P., & Knight, B. J. (1979). Two non-reactive field experiments on stealing from a ‘lost’ letter. British Journal of Social and Clinical Psychology, 18 (3), 277-284. https://doi.org/10.1111/j.2044-8260.1979.tb00337.x 18 Farrington, D. P., & Knight, B. J. (1980). Stealing from a “lost” letter: Effects of victim characteristics. Criminal Justice and Behavior, 7 (4), 423-436. https://doi.org/10.1177/009385488000700406 Farrington, D. P., Lösel, F., Braga, A. A., Mazerolle, L., Raine, A., Sherman, L. W., & Weisburd, D. (2020). Experimental criminology: Looking back and forward on the 20th anniversary of the Academy of Experimental Criminology. Journal of Experimental Criminology, 16 , 649–673. https://doi.org/10.1007/s11292-019-09384-z Farrington, D. P., Ttofi, M. M., Crago, R. V., & Coid, J. W. (2014). Prevalence, frequency, onset, desistance and criminal career duration in self-reports compared with official records. Criminal Behaviour and Mental Health, 24 (4), 241-253. https://doi.org/10.1002/cbm.1930 Feldman, R. E. (1968). Response to compatriot and foreigner who seek assistance. Journal of Personality and Social Psychology, 10 (3), 202-214. https://doi.org/10.1037/h0026567 Franklin, B. J. (1973). The effects of status on the honesty and verbal responses of others. The Journal of Social Psychology, 91 (2), 347-348. https://doi.org/10.1080/00224545.1973.9923060 Franzen, A., & Pointner, S. (2013). The external validity of giving in the dictator game. Experimental Economics, 16 (2), 155-169. https://doi.org/10.1007/s10683-012-9337-5 Gerdtz, M., Yap, C. Y., Daniel, C., Knott, J. C., Kelly, P., & Braitberg, G. (2020). Prevalence of illicit substance use among patients presenting to the emergency department with acute behavioural disturbance: Rapid point‐of‐care saliva screening. Emergency Medicine Australasia, 32 (3), 473480. https://doi.org/10.1111/1742-6723.13441 Gibbons, D. C. (1979). The criminological enterprise: Theories and perspectives . Prentice-Hall. Gibson, C. L., & Krohn, M. D. (Eds.). (2012). Handbook of life-course criminology: Emerging trends and directions for future research . Springer. https://doi.org/10.1007/978-1-4614-5113-6 Giguère, K., Béhanzin, L., Guédou, F. A., Leblond, F. A., Goma-Matsétsé, E., Zannou, D. M., Affolabi, D., Kêkê, R. K., Gangbo, F., Bachabi, M., & Alary, M. (2019). Biological validation of self-reported unprotected sex and comparison of underreporting over two different recall periods among female sex workers in Benin. Open Forum Infectious Diseases, 6 (2), 1-6. https://doi.org/10.1093/ofid/ofz010 Gnambs, T., & Kaspar, K. (2015). Disclosure of sensitive behaviors across self-administered survey modes: A meta-analysis. Behavior Research Methods, 47 (4), 1237-1259. https://doi.org/10.3758/s13428-014-0533-4 19 Gomes, H. S., Farrington, D. P., Defoe, I. N., & Maia, Â. (2021a). Field experiments on dishonesty and stealing: What have we learned in the last 40 years?. Journal of Experimental Criminology . Advance online publication. https://doi.org/10.1007/s11292-021-09459-w Gomes, H. S., Farrington, D. P., Krohn, M. D., Cunha, A., Jurdi, J., Sousa, B., Morgado, D., Hoft, J., Hartsell, E., Kassem, L., & Maia, Â. (2021c). The impact of modes of administration on selfreports of offending: A two methodological experiment replication [Manuscript submitted for publication]. School of Psychology, University of Minho. Gomes, H. S., Farrington, D. P., Krohn, M. D., & Maia, Â. (2021b). How sensitive are self-reports of offending?: The impact of recall periods on question sensitivity [Manuscript submitted for publication]. School of Psychology, University of Minho. Gomes, H. S., Farrington, D. P., Maia, Â., & Krohn, M. D. (2019). Measurement bias in self-reports of offending: A systematic review of experiments. Journal of Experimental Criminology, 15 (3), 313339. https://doi.org/10.1007/s11292-019-09379-w Gomes, H. S., Maia, Â., & Farrington, D. P. (2018). Measuring offending: Self-reports, official records, systematic observation and experimentation. Crime Psychology Review, 4 (1), 26-44. https://doi.org/10.1080/23744006.2018.1475455 Gold, M. (1966). Undetected delinquent behavior. Journal of Research in Crime and Delinquency, 3 (1), 27-46. https://doi.org/10.1177/002242786600300103 Hardt, R. H., & Peterson-Hardt, S. (1977). On determining the quality of the delinquency self-report method. Journal of Research in Crime and Delinquency, 14 (2), 247-259. https://doi.org/10.1177/002242787701400210 Harrison, G. W., & List, J. A. (2004). Field experiments. Journal of Economic Literature, 42 (4). 10091055. https://doi.org/10.1257/0022051043004577 Hindelang, M. J., Hirschi, T., & Weis, J. G. (1981). Measuring delinquency . Sage. Hirschi, T. (1969). Causes of delinquency . Transaction. Hornstein, H. A., Fisch, E., & Holmes, M. (1968). Influence of a model’s feeling about his behavior and his relevance as a comparison other on observers’ helping behavior. Journal of Personality and Social Psychology, 10 (3), 222–226. https://doi.org/10.1037/h0026568 Hser, Y. I. (1997). Self-reported drug use: Results of selected empirical investigations of validity. In L. Harrison & A. Hughes (Eds.), The validity of self-reported drug use: Improving the accuracy of survey estimates (NIDA Research Monograph No. 167, pp. 320-343). U.S. Department of Health and Human Services. https://www.ojp.gov/pdffiles1/Digitization/167339-167359NCJRS.pdf 20 Huizinga, D., & Elliott, D. S. (1986). Reassessing the reliability and validity of self-report delinquency measures. Journal of Quantitative Criminology, 2 (4), 293-327. https://doi.org/10.1007/BF01064258 Jobe, J. B., Pratt, W. F., Tourangeau, R., Baldwin, A. K., & Rasinski, K. A. (1997). Effects of interview mode on sensitive questions in a fertility survey. In L. Lyberg, P. Biemer, M. Collins, E. de Leeuw, C. Dippo, N. Schwartz, & D. Trewin (Eds.), Survey measurement and process quality (pp. 311329). John Wiley & Sons. https://doi.org/10.1002/9781118490013.ch13 Jolliffe, D., & Farrington, D. P. (2014). Self-reported offending: Reliability and validity. In G. Bruinsma, & D. Weisburd (Eds.), Encyclopedia of criminology and criminal justice (pp. 4716-4723). Springer. https://doi.org/10.1007/978-1-4614-5690-2_648 Jolliffe, D., Farrington, D. P., Hawkins, J. D., Catalano, R. F., Hill, K. G., & Kosterman, R. (2003). Predictive, concurrent, prospective and retrospective validity of self‐reported delinquency. Criminal Behaviour and Mental Health, 13 (3), 179-197. https://doi.org/10.1002/cbm.541 Kabashi, S., Vindenes, V., Bryun, E. A., Koshkina, E. A., Nadezhdin, A. V., Tetenova, E. J., Kolgashkin, A. J., Petukhov, A. E., Perekhodov, S. N., Davydova, E. N., Gamboa, D., Hilberg. T., Lerdal. A., Nordby, G., Zhang, C., & Bogstrand, S. T. (2019). Harmful alcohol use among acutely ill hospitalized medical patients in Oslo and Moscow: A cross-sectional study. Drug and Alcohol Dependence, 204 , 107588. https://doi.org/10.1016/j.drugalcdep.2019.107588 Kazemian, L., & Farrington, D. P. (2005). Comparing the validity of prospective, retrospective, and official onset for different offending categories. Journal of Quantitative Criminology, 21 (2), 127-147. https://doi.org/10.1007/s10940-005-2489-0 Kerschbamer, R., Neururer, D., & Sutter, M. (2016). Insurance coverage of customers induces dishonesty of sellers in markets for credence goods. Proceedings of the National Academy of Sciences, 113 (27), 7454-7458. https://doi.org/10.1073/pnas.1518015113 Kirtadze, I., Otiashvili, D., Tabatadze, M., Vardanashvili, I., Sturua, L., Zabransky, T., & Anthony, J. C. (2018). Republic of Georgia estimates for prevalence of drug use: Randomized response techniques suggest under-estimation. Drug and Alcohol Dependence, 187 , 300-304. https://doi.org/10.1016/j.drugalcdep.2018.03.019 Kleck, G., & Roberts, K. (2012). What survey modes are most effective in eliciting self-reports of criminal or delinquent behavior? In L. Gideon (Ed.), Handbook of survey methodology for the social sciences (pp. 417-439). Springer. https://doi.org/10.1007/978-1-4614-3876-2_24 21 Knapp, H., & Kirk, S. A. (2003). Using pencil and paper, Internet and touch-tone phones for selfadministered surveys: Does methodology matter? Computers in Human Behavior, 19 (1), 117134. https://doi.org/10.1016/S0747-5632(02)00008-0 Korte, C., & Kerr, N. (1975). Response to altruistic opportunities in urban and nonurban settings. The Journal of Social Psychology, 95 (2), 183-184. https://doi.org/10.1080/00224545.1975.9918701 Kreuter, F., Presser, S., & Tourangeau, R. (2008). Social desirability bias in CATI, IVR, and Web Surveys: The effects of mode and question sensitivity. Public Opinion Quarterly, 72 (5), 847-865. https://doi.org/10.1093/poq/nfn063 Krohn, M., Thornberry, T., Bell, K., Lizotte, A., & Phillips, M. (2012). Self-report surveys within longitudinal panel designs. In D. Gadd, S. Karstedt, & S. Messner (Eds.), The Sage handbook of criminological research (pp. 23-35). Sage. https://dx.doi.org/10.4135/9781446268285.n2 Krohn, M. D., Waldo, G. P., & Chiricos, T. G. (1974). Self-reported delinquency: A comparison of structured interviews and self-administered checklists. Journal of Criminal Law and Criminology, 65 (4), 545-553. https://doi.org/10.2307/1142528 Krumpal, I. (2013). Determinants of social desirability bias in sensitive surveys: A literature review. Quality & Quantity, 47 (4), 2025-2047. https://doi.org/10.1007/s11135-011-9640-9 Kulik, J. A., Stein, K. B., & Sarbin, T. R. (1968). Disclosure of delinquent behavior under conditions of anonymity and nonanonymity. Journal of Consulting and Clinical Psychology, 32 (5, Pt1), 506509. https://doi.org/10.1037/h0026260 Lee, H., Kim, S., Couper, M. P., & Woo, Y. (2019). Experimental comparison of PC web, smartphone web, and telephone surveys in the new technology era. Social Science Computer Review, 37 (2), 234-247. https://doi.org/10.1177/0894439318756867 Levitt, S. D., & List, J. A. (2007). What do laboratory experiments measuring social preferences reveal about the real world?. Journal of Economic Perspectives, 21 (2), 153-174. https://doi.org/10.1257/jep.21.2.153 Liber, A. C., & Warner, K. E. (2018). Has underreporting of cigarette consumption changed over time? Estimates derived from US National Health Surveillance Systems between 1965 and 2015. American Journal of Epidemiology, 187 (1), 113-119. https://doi.org/10.1093/aje/kwx196 Littlefield, A. K., Brown, J. L., DiClemente, R. J., Safonova, P., Sales, J. M., Rose, E. S., Belyakov, N., & Rassokhin, V. V. (2017). Phosphatidylethanol (PEth) as a biomarker of alcohol consumption in 22 HIV-infected young Russian women: Comparison to self-report assessments of alcohol use. AIDS and Behavior, 21 (7), 1938-1949. https://doi.org/10.1007/s10461-017-1769-7 Livingston, M. (2013). Assessment of adolescent alcohol use: Estimating and adjusting for measurement bias (Publication No. 3729218) [Doctoral dissertation, University of Florida]. ProQuest Dissertations Publishing. Martins, P., Mendes, S., & Fernandez-Pacheco, G. (2015, September 2-5). Cross-cultural adaptation and online administration of the Portuguese Version of ISRD3 [Paper presentation]. 15th Annual Conference of the European Society of Criminology, Porto, Portugal. Maxfield, M. G., & Babbie, E. R. (2009). Basics of research methods for criminal justice and criminology (2nd ed.). Cengage Learning. Merritt, C. B., & Fowler, R. G. (1948). The pecuniary honesty of the public at large. The Journal of Abnormal and Social Psychology, 43 (1), 90–93. https://doi.org/10.1037/h0061846 Moffitt, T. E. (1993). Adolescence-limited and life-course-persistent antisocial behavior: A developmental taxonomy. Psychological review, 100 (4), 674-701. https://doi.org/10.1037/0033295x.100.4.674 Murphy, F. J., Shirley, M. M., & Witmer, H. L. (1946). The incidence of hidden delinquency. American Journal of Orthopsychiatry, 16 (4), 686–696. https://doi.org/10.1111/j.19390025.1946.tb05431.x Nye, F. I., Short, J. F., & Olson, V. J. (1958). Socioeconomic status and delinquent behavior. American Journal of Sociology, 63 (4), 381-389. https://doi.org/10.1086/222261 Osgood, D. W., McMorris, B. J., & Potenza, M. T. (2002). Analyzing multiple-item measures of crime and deviance I: Item response theory scaling. Journal of Quantitative Criminology, 18 (3), 267-296. https://doi.org/10.1023/A:1016008004010 Palamar, J. J., Salomone, A., & Keyes, K. M. (2021). Underreporting of drug use among electronic dance music party attendees. Clinical Toxicology, 59 (3), 185-192. https://doi.org/10.1080/15563650.2020.1785488 Pepper, J. V., & Petrie, C. V. (Eds.). (2003). Measurement problems in criminal justice research: Workshop summary . National Academy Press. https://doi.org/10.17226/10581 Piquero, A. R., Schubert, C. A., & Brame, R. (2014). Comparing official and self-report records of offending across gender and race/ethnicity in a longitudinal study of serious youthful offenders. Journal of Research in Crime and Delinquency, 51 (4), 526–556. https://doi.org/10.1177/0022427813520445 23 Porterfield, A. L. (1943). Delinquency and its outcome in court and college. American Journal of Sociology, 49 (3), 199–208. https://doi.org/10.1086/219369 Potdar, R., & Koenig, M. A. (2005). Does audio-CASI improve reports of risky behavior? Evidence from a randomized field trial among young urban men in India. Studies in Family Planning, 36 (2), 107116. https://doi.org/10.1111/j.1728-4465.2005.00048.x Richman, W. L., Kiesler, S., Weisband, S., & Drasgow, F. (1999). A meta-analytic study of social desirability distortion in computer-administered questionnaires, traditional questionnaires, and interviews. Journal of Applied Psychology, 84 (5), 754-775. https://doi.org/10.1037/00219010.84.5.754 Rijo, D., Miguel, R. R., Paulo, M., & Brazão, N. (2020). The effects of the growing pro-social program on early maladaptive schemas and schema-related emotions in male young offenders: A nonrandomized trial. International Journal of Offender Therapy and Comparative Criminology, 64 (13-14), 1422-1442. https://doi.org/10.1177/0306624X20912988 Robertson, R. E., Tran, F. W., Lewark, L. N., & Epstein, R. (2018). Estimates of non-heterosexual prevalence: The roles of anonymity and privacy in survey methodology. Archives of Sexual Behavior, 47 (4), 1069-1084. https://doi.org/10.1007/s10508-017-1044-z Rosenbaum, S. M., Billinger, S., & Stieglitz, N. (2014). Let’s be honest: A review of experimental evidence of honesty and truth-telling. Journal of Economic Psychology, 45 , 181-196. https://doi.org/10.1016/j.joep.2014.10.002 Schober, S. E., Caces, M. F., Pergamit, M. R., & Branden, L. (1992). Effect of mode of administration on reporting of drug use in the National Longitudinal Survey. In C. F. Turner, J. T. Lessler, & J. C. Gfroerer (Eds.), Survey measurement of drug use: Methodological studies (pp. 267–276). National Institute on Drug Abuse. Schore, J., Maynard, R., & Piliavin, I. (1979). The accuracy of self-reported arrest data . Mathematica Policy Research. Schwarz, N. (1999). Self-reports: How the questions shape the answers. American Psychologist, 54 (2), 93-105. https://doi.org/10.1037/0003-066X.54.2.93 Schwarz, N., Strack, F., Hippler, H. J., & Bishop, G. (1991). The impact of administration mode on response effects in survey measurement. Applied Cognitive Psychology, 5 (3), 193-212. https://doi.org/10.1002/acp.2350050304 Sellin, T. (1931). The basis of a crime index. Journal of Criminal Law and Criminology, 22 (3), 335-356. https://doi.org/10.2307/1135784 24 Short, J. F., & Nye, F. I. (1958). Extent of unrecorded juvenile delinquency tentative conclusions. The Journal of Criminal Law, Criminology, and Police Science, 49 (4), 296-302. https://doi.org/10.2307/1141583 Steinberg, J., McDonald, P., & O'Neal, E. (1977). Petty theft in a naturalistic setting: The effects of bystander presence. The Journal of Social Psychology, 101 (2), 219-221. https://doi.org/10.1080/00224545.1977.9924010 Strang, E., & Peterson, Z. D. (2020). Use of a bogus pipeline to detect men’s underreporting of sexually aggressive behavior. Journal of Interpersonal Violence, 35 (1-2), 208-232. https://doi.org/10.1177/0886260516681157 Sudman, S., & Bradburn, N. M. (1974). Response effects in surveys: A review and synthesis . Aldine Publishing Company. Sullivan, C. J., & McGloin, J. M. (2014). Looking back to move forward: Some thoughts on measuring crime and delinquency over the past 50 years. Journal of Research in Crime and Delinquency, 51 (4), 445-466. https://doi.org/10.1177/0022427813520446 Thornberry, T. P., & Krohn, M. D. (2000). The self-report method for measuring delinquency and crime. In D. Duffee (Ed.), Measurement and analysis of crime and justice (pp. 33–84). U.S. National Institute of Justice. Tourangeau, R., & McNeeley, M. E. (2003). Measuring crime and crime victimization: Methodological issues. In J. V. Pepper, & C. V. Petrie (Eds.), Measurement problems in criminal justice research: Workshop summary (pp. 10-42). National Academy Press. Tourangeau, R., Rips, L. J., & Rasinski, K. (2000). The psychology of survey response . Cambridge University Press. https://doi.org/10.1017/CBO9780511819322 Tourangeau, R., & Yan, T. (2007). Sensitive questions in surveys. Psychological Bulletin, 133 (5), 859883. https://doi.org/10.1037/0033-2909.133.5.859 Tourangeau, R., & Yan, T. (in press). Reporting issues in surveys of drug use. Substance Use and Misuse . Trau, R. N., Härtel, C. E., & Härtel, G. F. (2013). Reaching and hearing the invisible: Organizational research on invisible stigmatized groups via web surveys. British Journal of Management, 24 (4), 532-541. https://doi.org/10.1111/j.1467-8551.2012.00826.x Turner, C. F., Lessler, J. T., & Devore, J. W. (1992). Effects of mode of administration and wording on reporting of drug use. In C. F. Turner, J. T. Lessler, & J. C. Gfroerer (Eds.), Survey measurement of drug use: Methodological studies (pp. 177-219). National Institute on Drug Abuse. 25 Vinikoor, M. J., Zyambo, Z., Muyoyeta, M., Chander, G., Saag, M. S., & Cropsey, K. (2018). Point-of-care urine ethyl glucuronide testing to detect alcohol use among HIV-hepatitis B virus coinfected adults in Zambia. AIDS and Behavior, 22 (7), 2334-2339. https://doi.org/10.1007/s10461-018-20308 Wehling, H., & Lusher, J. (2019). People with a body mass index⩾ 30 under-report their dietary intake: a systematic review. Journal of health psychology, 24 (14), 2042-2059. https://doi.org/10.1177/1359105317714318 Wolter, F., & Laier, B. (2014). The effectiveness of the item count technique in eliciting valid answers to sensitive questions. An evaluation in the context of self-reported delinquency. Survey Research Methods, 8 (3). 153-168. https://doi.org/10.18148/srm/2014.v8i3.5819 Yan, T., & Cantor, D. (2019). Asking survey questions about criminal justice involvement. Public Health Reports, 134 (1_suppl), 46S-56S. https://doi.org/10.1177/0033354919826566 Zara, G., & Farrington, D. P. (2016). Criminal recidivism: Explanation, prediction and prevention . Routledge. Zimny, G. H. (1961). Method in experimental psychology . Ronald Press Company. https://doi.org/10.1037/14006-000 32 in court and college students. However, the groundbreaking results of Nye et al. (1958), about the minimal differences in the prevalence of delinquent behaviour between different socioeconomic strata, revealed the true potential of the SRO technique (Krohn et al., 2010). These works drastically changed criminologists’ opinions of SRO, and Hirschi (1969) developed the Social Control Theory based on this methodology. In 1973, Farrington published the first review of the literature on the psychometric qualities of SRO surveys and concluded that this technique had predictive validity. Although it should not replace entirely the officially recorded data, Farrington (1973, p. 109) suggested that “the most accurate measure of deviant behaviour may yet prove to be some combination of official records and a self-report questionnaire”. Hindelang et al. (1981) studied this technique and produced a highly influential book called ‘ Measuring delinquency ’ that was a milestone in the use of the self-report methodology in criminological research, demonstrating that SRO were a valid measure of crime and delinquency. Since then, the self-report method became ‘one of the most important innovations in criminological research in the 20th century’ (Thornberry & Krohn, 2000, p.34) and since then criminological knowledge (e.g., criminal patterns, delinquency theories, etc.) has relied almost exclusively on data obtained by the selfreport methodology (Cops et al., 2016). Comparing official records and self-reports of offending Since the development of SRO, researchers have been interested in comparisons between official records and SRO data and have used official records as a standard to study the criterion validity of SRO (Hindelang et al., 1979, 1981). The idea was that, if the two methods measure the same construct, they should be positively correlated. As a matter of fact, Hindelang et al. (1981) found considerable concordance between official records and SRO, which led them to the conclusion that people generally admit their criminal practices. Other authors, such as West and Farrington (1977), also found an association between SRO and officially recorded offending. However, to better understand this relation, we should consider the finding by Farrington (1977) that after criminal convictions – or public labelling – there is an increase in self-reported offending, so convictions could make known offenders more likely to admit their delinquent behaviour. Nevertheless, several researchers have found that SRO significantly predict future convictions among unconvicted people (e.g., Farrington, 2003), which indicates the validity of SRO. It might be expected that SRO would provide higher estimates of offending since this technique was developed with the objective of overcoming the limitations inherent in official records (Farrington et 33 al., 2007). Despite the associations described above, researchers looked deeper into the differences between the results obtained by these two methods and their implications for criminological knowledge. In this article, we will focus on the primary differences in conclusions derived from SRO and official records in measuring criminal behaviour. Scaling-up factor As stated earlier, the primary limitation of official records of crime is that they provide an underestimate of offending. In an attempt to estimate the real number of crimes per conviction, researchers developed the ‘scaling-up factor’, which is “estimated by comparing convictions and selfreported offences of the same people at the same ages” (Theobald et al., 2014, p. 265). Considering the males in the Cambridge Study in Delinquent Development (CSDD; n = 411), at the ages of 15–18, 27–32, and 42–47, Farrington et al. (2013) estimated a scaling-up factor of 39 selfreported offences per conviction. In the Pittsburgh Girls Study (PGS), based on a sample of 2,450 girls between ages 12 and 17, a scaling-up factor of five self-reported offences was found for every police charge (Ahonen et al., 2017). In a longer follow-up of the PGS, this rose to a factor of 12 between ages 11 and 19 (Jennings et al., 2018). In this latter study, 33% of low-rate official offenders (with one to four police charges) and 27% of high-rate official offenders (with five or more police charges) self-reported no offences. This highlights possible gender differences in the scaling-up factor. In the Pittsburgh Youth Study (PYS, n = 506), with boys aged between 13 and 17, Farrington et al. (2007) found a scaling-up factor of 80. Later, in the same PYS, with boys aged between 13 and 24, Theobald et al. (2014) found a scalingup factor of nine self-reported offences for each conviction. Moreover, the evidence seems to suggest that the scaling-up factor changes throughout the life course. Indeed, Theobald et al. (2014) found that this factor increased from 8 at ages 13–15 to 14 at ages 22–24. In the CSDD, younger males (aged 15–18) had a ratio of 47 self-reported offences for each conviction, older males (aged 27–32) had a ratio of 33 and the oldest males (aged 42–47) had the lowest ratio of 17 (Farrington et al., 2013). It is possible that the relationship between the scaling-up factor and age is curvilinear. Clearly, more research on this is needed. The self-reported offences per conviction ratio also seems to vary as a function of types of crime. For example, in the CSDD, burglary and theft of vehicles had the lowest scaling-up ratios, 6 and 9, respectively, compared with theft from work and drug offences that had alarming ratios of 1,463 and 4,160 self-reported offences per conviction, respectively (Farrington et al., 2013). In the PYS, Farrington et al. (2007) reported that property offences had the lowest scaling-up factor (15), followed by violent 34 offences (154) and, finally, drug offences had the highest ratio (424). Moreover, Theobald et al. (2014) compared serious (5) and moderate (16) thefts, as well as serious (11) and moderate (13) violence offences, and found that in both cases the scaling-up factor was higher for serious offence types (Theobald et al., 2014). Another interesting result is that the scaling-up factor seems to change as a function of race. This was found by Ahonen et al. (2017) in the PGS, where African American girls (7) had a much higher ratio than Caucasian girls (2). This difference was even higher at younger ages. At age 13, African Americans had a ratio more than five times higher (27 vs. 5), a difference that gradually decreased with age (at age 17, 5 vs. 2). The results with boys in the PYS followed a different trend, where Caucasian boys had a ratio of 10, slightly higher than African American boys (8). Theobald et al. (2014, p. 274) interpreted this result by explaining that African Americans are more exposed to risk factors, and that this result “does not necessarily mean that the police or the courts are biased against African American boys”, although more research on this topic is clearly needed. Implications The discussion on the scaling-up factor clearly shows the different estimates obtained from the two measures of crime, SRO and official records. Moreover, these different estimates of criminal behaviour varied differently with age, type of offences, race, etc. An obvious consequence is the likelihood of drawing different conclusions from the different research methods. Criminal career research One other topic where different estimates of criminal behaviour from official records and SRO might have major implications is in the research on criminal careers. Authors such as Blumstein et al. (1986, 1988) demonstrated the importance of criminal career research. Understanding the sequence of offences over time of particular offenders allows us to understand the beginning of offending (i.e., the age of onset), the maintenance of criminal behaviour (i.e., persistence), the moment when they stop offending (i.e., desistance) – and, thus, criminal career duration – as well as knowledge about changes in criminal behaviour, such as specialization or diversification of criminal acts, escalation or de-escalation of the seriousness of crimes, etc. Because criminal career research requires exact information about the dates of offences, the majority of studies have based the measurement of criminal behaviour on official records, rather than on SRO, a fact that some authors have considered potentially misleading (Farrington et al., 2003). Therefore, 35 Farrington et al. (2003) suggested that SRO of offending might add value to criminal career research, with a more accurate estimate of the total number of crimes. This, and subsequent studies that based criminal career research on both methods, faced the problems of different estimates when based on official data and when based on SRO. Age of onset Considering data from the Seattle Social Development Project (SSDP; n = 808), Farrington et al. (2003) found that the first offence reported in the surveys preceded, on average, by 2.4 years the first crime in the official data (i.e., court referral). More exactly, in this study, while the average age of onset based on official records was age 15.1, the average age of onset based on SRO was at age 12.7. Moreover, this study estimated that, on average, an offender commits 26 offences before the first crime is officially recorded. Concordant results were found by Loeber et al. (2003) in the OJJDP Study Group on Very Young Offenders, where the age of onset based on self-reported serious delinquency was at the age of 11.9, whereas the average age of onset based on official records (i.e., court contact) happened 2.6 years later, at the age of 14.5 years. Kazemian and Farrington (2005) analysed data from the CSDD and found similar results. They found that the age of onset based on SRO was, on average, at 11.9, whereas the age of onset was, on average, 16.9 based on official records (5 years later). Moreover, Kazemian and Farrington (2005) noticed a relationship between the seriousness of offences and the agreement between the two estimates of the age of onset. The difference between the estimates of the age of onset based on SRO and official records became less pronounced as the seriousness of crimes increased. For example, these authors found a difference of 1.6 years for theft of vehicles (age of onset: SRO = 15.2; official records = 16.8) compared to a difference of 12 years for vandalism (age of onset: SRO = 10.7; official records = 22.7). The fact that serious offences are more likely to result in court convictions, compared with minor offences, may explain these results. Criminal career duration To date, we have discussed how SRO provide a much higher estimate of the number of crimes and indicate that criminal activity starts much earlier than according to official records of crime. An obvious consequence seems to be that criminal career duration should be longer if studied with SRO compared with official criminal records. In fact, some authors have found such a result. For example, Le 36 Blanc and Fréchette (1989; n = 470) found a duration of 5.23 years for the criminal career if based on convictions, but a career more than twice as long of 10.76 years based on SRO. Farrington et al. (2014) published a very informative paper that addressed these questions about career research. Using the data from the CSDD (between ages 8 and 48), these authors showed that, while the average age of onset in SRO was at 10, the first conviction did not happen on average until 19. Similarly, the age of desistance in SRO was at 35, whereas in official records it happened much sooner in life, at the age of 25. There was an average criminal career duration of 25 years according to SRO, compared with an average duration of 6 years based on convictions, a 19-year difference (Farrington et al., 2014). Implications Criminal career research provides a great example of the different, and at times contrasting, conclusions that could be derived from different methods for measuring criminal behaviour. The reader should keep in mind that this is only a part of the problem. There are also differences in criminal patterns. For example, Kazemian and Farrington (2005) found that whereas SRO data indicated a pattern where individuals start with a minor offence and gradually commit more serious offences, the results based on official records of crime showed the opposite serious-to-minor pattern. On the other hand, although some authors argued that criminal features, such as prevalence and frequency, vary similarly with age (e.g., Hirschi & Gottfredson, 1983), others opposed this idea and argued that the age–crime curve was driven by prevalence, while frequency was pretty constant with age (Blumstein et al., 1988). When testing this hypothesis, Farrington et al. (2003) in the SSDP discovered that both prevalence and frequency increased with age in SRO, but only prevalence increased in official records, whereas offending frequency stayed constant with age (Figure 2). Moreover, in the PYS, Farrington et al. (2007) found that, if based on SRO, the frequency of offences per offender increased with age during adolescence, but, based on official records, the frequency of offences per offender seemed to remain constant across adolescence. This led the authors to the conclusion that “in attempting to explain offending, researchers should always measure both self-reports and official records” (Farrington et al., 2007, p. 246). It thus seems clear that different conclusions may be obtained from these two methods of collecting criminal data. This is a question that should be considered seriously because it may result in different theoretical and policy implications (Kazemian & Farrington, 2005). 37 Figure 2 Prevalence and frequency of offending according to different sources Note . Source: Farrington et al. (2003). Self-reports of offending Despite the conclusion of several authors that the best estimate of crime may be achieved by a combination of both SRO and official records (e.g., Farrington, 1973; Farrington et al., 2003), the results presented in this article led some researchers to the conclusion that “official records are biased and yield distorted information about the true characteristics of offenders” (Farrington et al., 2007, p. 229). Others concluded that the “prevalence and mean frequency of self-reported offending is a better indicator of actual delinquent behaviour than is being charged by the police or the frequency of police charges” (Loeber et al., 2015, p. 163). All these reasons favouring SRO over official records certainly do not mean that SRO are without limitations. Self-reports of human behaviour can be affected by multiple factors. The format of the questionnaire, the wording of questions, the response format, modes of administration, etc. have been previously shown to impact self-reports of behaviour and attitudes (Schwarz, 1999). Furthermore, due to the sensitive and potentially incriminating nature of criminal behaviour, we have reason to believe that SRO would be even more sensitive to these biasing factors (Thornberry & Krohn, 2000). Despite the concern about these potential biases (as shown throughout the present article), current knowledge does not allow us to know which factors impact self-reports in general, to what extent these factors impact SRO in particular and how to control or minimize their effects. 0 5 10 15 11 12 13 14 15 16 17 Official Records Prevalence Frequency 0 10 20 30 40 50 60 70 11 12 13 14 15 16 17 Self-Reports Prevalence Frequency 38 According to our literature review, we can divide the major concerns about self-reports into three main categories: 1) questionnaire design; 2) modes of administration; and 3) testing effects. Questionnaire design The way the questionnaire is designed presents several features that constitute potential sources of bias. In 1973, Farrington revealed a concern about the phrasing of the questions. He noted that the vast majority of survey questions on offending were presented to the participants in the same direction, and proposed phrasing questions both positively and negatively, as a way to minimize the potential for acquiescence response bias. However, positive phrasing still stands today as the norm in measuring delinquent behaviour, e.g. ‘Have you ever in your life broken into a building to steal something?’ (ISRD3 Working Group, 2013) and ‘Have you ever stolen something worth more than 50 euros’ (Sanches et al., 2016). Enzmann (2013) focused on the response formats and tested the effects of reversing the response categories ‘yes’ and ‘no’, along with a short version of the questionnaire, and omitting followup questions. The findings from this experiment showed that the short version, where ‘yes’ appears first, and with omitted follow-up questions, generated higher estimates of offending, primarily concerning minor crimes. Response order effects have been previously described as resulting from primacy effects or social desirability; for example, participants may see the first response option as the most natural answer (Enzmann, 2013; Schwarz et al., 1991). The impact of follow-up questions is particularly important, because researchers are interested in much more information, rather than only knowing whether a person did or did not commit a certain crime (e.g., how many times; with whom; where it happened, etc.). Since follow-up questions are contingent on affirmative answers, participants might learn to answer negatively in order to avoid further questions and minimize the length of the interview (Thornberry, 1989). It may be that participants obey the law of least effort. Modes of administration Administration methods of SRO have long been a concern of researchers (e.g., Gold, 1966). Hindelang et al. (1981) developed a study where participants were randomly assigned to four conditions (i.e., non-anonymous questionnaire, anonymous questionnaire, non-anonymous interview and anonymous interview). The results showed small to no significant differences between the four methods, leading the researchers to the conclusion that SRO are largely independent of the modes of administration. This finding resulted in a substantial decrease in methodological studies of delinquency 39 surveys. It seemed that the validity of the self-report technique had been established and that further methodological research was unnecessary (Jolliffe & Farrington, 2014). Fifteen years later, Tourangeau and Smith (1996) summarized multiple studies of method effects, concluding that participants are generally more willing to report illegal activities using self-administration methods, rather than admitting them to interviewers. This experimental study compared three conditions: computer-assisted personal interviewing (CAPI), computer-assisted self-administered interviewing (CASI), and audio computer-assisted self-administered interviewing (ACASI). The findings provided evidence for the presence of method effects, showing that self-administration resulted in higher reports of sexual behaviour and drug use. More recently, other researchers were also able to demonstrate the impact of method effects on reports of offending. Denniston et al. (2010) conducted an experiment comparing paper-and-pencil versus web administration, and found that participants in the paper-and-pencil condition reported higher perceived privacy and anonymity. Wright et al. (1998) compared self-reports of smoking, alcohol and drug use in computer-assisted versus paper-and-pencil conditions. The results showed higher reports in the computer-assisted condition. Moreover, these authors found an interaction between method effects and the age of the participants, since adolescent participants were more sensitive to method effects. Lucia et al. (2007) showed that paper-and-pencil administration yielded significantly higher reports of delinquency than the internet condition in three out of 22 comparisons (i.e., selling soft drugs, vandalism and theft from the person). Thornberry and Krohn (2000) published a review of SRO, concluding that computer-assisted selfinterviews with audio elicited higher rates of delinquency. Whether or not different modes of administration impact participants’ reports of delinquent behaviour is still a debatable subject. Moreover, we seem to know very little about the underlying processes that might explain these effects, the factors that interact with method effects (e.g., age, sensitivity of questions) or even the direction of impact. Testing and panel effects Testing and panel effects are a serious threat to longitudinal studies of delinquent behaviour. Considering the prominent place of longitudinal designs in criminological research (Krohn et al., 2012), ensuring the quality of its data should be a priority in this field of science. As described by Thornberry (1989, p. 351), testing effects are “any alterations of a subject’s response to a particular item or scale caused by the prior administration of the same item or scale”, whereas panel effects refer to “a more 40 general reaction to being re-interviewed”, rather than a specific reaction to questionnaire characteristics (Thornberry, 1989, p. 361). In his study, Thornberry (1989) was able to find a decrease in the prevalence of delinquent behaviour as a function of the number of prior interviews. Other researchers also demonstrated reductions in participants’ reports of delinquency in prospective longitudinal studies that were inconsistent with the age–crime curve (e.g., Bosick, 2009; Lauritsen, 1998). However, whether these results are due to testing/panel effects is still questionable, and other features could explain such declines, for example, scale construction, item-specific age–crime curves or selective attrition (Bosick, 2009). More recently, Krohn et al. (2012) reviewed the role of surveys within longitudinal studies and appealed for the importance of further investigating the potential for testing and panel effects. Future directions for self-report methods This list of biasing effects, by no means exhaustive, may constitute a real threat to the quality of SRO and, by extension, to the validity of criminological knowledge. Nonetheless, these potential biases of the self-report technique should not be seen as reasons not to use questionnaires to measure delinquency. On the contrary, it is urgent to develop experimental studies to test to what extent these effects might contaminate the quality of results and to try to understand the ways in which they interact with each other. By doing this, we can strive to obtain ever better results, closer and closer to the reality of offending. Furthermore, developing knowledge about the best way to survey participants could also facilitate standardized self-report measures of offending. Alternative methods for measuring crime Considering the literature reviewed in this paper, even if we accept that the self-report technique provides a better indicator of offending behaviour than official records, we are led to the conclusion that survey methods are ‘also rather biased and indirect measures of offending’ (Buckle & Farrington, 1984). If anything, the previous discussion on the biases of SRO shows that there is still much work to do in order to fully understand how different methodological features interfere with the reporting of criminal behaviour. In response, some researchers have turned to observation methods to measure criminal behaviour. Obviously, most criminal acts are unpredictable, some are very difficult to observe directly (e.g., white collar crime), and offenders try to conceal their illicit activity. However, direct observation can be a very useful technique in specific domains. For example, Konecni et al. (1976) carried out systematic 41 observation of drivers’ behaviour and found that younger males were more likely to violate a red light in an intersection. Buckle and Farrington (1984) systematically observed shoppers and were able to observe nine out of 503 (1.8%) customers shoplifting. In a later study, Buckle and Farrington (1994) used the same technique and observed this illegal behaviour by 15 out of 988 (1.5%) customers. The findings from these studies included information about the personal characteristics of shoplifters (e.g., mostly males and more likely to be over 55 years old), about the offence itself (mostly small low-cost items were stolen) and about offenders’ behaviour during the offence (most checked if they were being observed by anyone before shoplifting) and after the offence (most purchased other goods to allay suspicion). Despite some concerns about this technique, namely regarding the generalization of results, these authors concluded that this method of measuring shoplifting was valid, and that it should be used more frequently. Also on the theme of shoplifting, Buckle et al. (1992) tested the technique of systematic counting. In this study, researchers repeatedly and systematically counted the number of specific items in each shop daily in order to detect item loss. By comparing items missing with items purchased in a total pool of 29 stores, the researchers found that 10.9% of items leaving the store were stolen. In the worst store, more than one-third of minor items were stolen. The authors analysed the qualities of this technique and concluded that systematic counting to measure shoplifting produced valid results. Shortly after, Farrington et al. (1993) carried out an experiment to evaluate the effectiveness of three situational interventions in preventing shoplifting, using systematic counting as the behavioural measure. This study provided evidence that one of these situational programs (i.e., electronic tagging) caused significant decreases in shoplifting over time. In the late 1970s and early 1980s, researchers were interested in field experiments in criminology, which resulted in some interesting methods of measuring deviant behaviour. Farrington and Kidd (1977) conducted a field experiment with the purpose of studying the decision process in committing financially dishonest behaviour. Basically, these authors provided random citizens the opportunity to dishonestly accept a lost coin, asking them if the supposed lost coin was actually theirs. Out of 84 participants, 31 claimed the coin dishonestly. Using this ‘dishonesty’ measure, Farrington and Kidd (1977) were able to conclude that people would act more dishonestly when the coin was less valuable (10p versus 50p) and when the experimenter was female rather than male. Interestingly, the cost had no effect on dishonesty if the experimenter was female. In another field experiment, Farrington et al. (1980) interviewed youth and asked them to participate in a coin-sorting test with the implicit purpose of providing them with an opportunity to steal. The final results showed that 10 out of a total 25 participants stole 48 Jennings, W. G., Loeber, R., Ahonen, L., Piquero, A. R., & Farrington, D. P. (2018). An examination of developmental patterns of chronic offending from self-report records and official data: Evidence from the Pittsburgh Girls Study (PGS). Journal of Criminal Justice, 55 , 71-79. https://doi.org/10.1016/j.jcrimjus.2017.12.002 Jolliffe, D., & Farrington, D. P. (2014). Self-reported offending: Reliability and validity. In G. Bruinsma, & D. Weisburd (Eds.), Encyclopedia of criminology and criminal justice (pp. 4716-4723). Springer. https://doi.org/10.1007/978-1-4614-5690-2_648 Kazemian, L., & Farrington, D. P. (2005). Comparing the validity of prospective, retrospective, and official onset for different offending categories. Journal of Quantitative Criminology, 21 (2), 127-147. https://doi.org/10.1007/s10940-005-2489-0 Konecni, V. J., Ebbesen, E. B., & Konecni, D. K. (1976). Decision processes and risk taking in traffic: Driver response to the onset of yellow light. Journal of Applied Psychology, 61 (3), 359-367. https://doi.org/10.1037/0021-9010.61.3.359 Krohn, M., Thornberry, T., Bell, K., Lizotte, A., & Phillips, M. (2012). Self-report surveys within longitudinal panel designs. In D. Gadd, S. Karstedt, & S. Messner (Eds.), The Sage handbook of criminological research (pp. 23-35). Sage. https://dx.doi.org/10.4135/9781446268285.n2 Krohn, M. D., Thornberry, T. P., Gibson, C. L., & Baldwin, J. M. (2010). The development and impact of self-report measures of crime and delinquency. Journal of Quantitative Criminology, 26 (4), 509525. https://doi.org/10.1007/s10940-010-9119-1 Lauritsen, J. L. (1998). The age-crime debate: Assessing the limits of longitudinal self-report data. Social Forces, 77 (1), 127-154. https://doi.org/10.1093/sf/77.1.127 Le Blanc, M., & Fréchette, M. (1989). Male criminal activity from childhood through youth: Multilevel and developmental perspectives . Springer-Verlag. https://doi.org/10.1007/978-1-4612-3570-5 Loeber, R., Farrington, D. P., Hipwell, A. E., Stepp, S. D., Pardini, D., & Ahonen, L. (2015). Constancy and change in the prevalence and frequency of offending when based on longitudinal self-reports or official records: Comparisons by gender, race, and crime type. Journal of Developmental and Life-Course Criminology, 1 (2), 150-168. https://doi.org/10.1007/s40865-015-0010-5 Loeber, R., Farrington, D. P., & Petechuk, D. (2003). Child delinquency: Early intervention and prevention (NCJ 186182). U.S. Office of Juvenile Justice and Delinquency Prevention. https://www.ojp.gov/pdffiles1/ojjdp/186162.pdf Lucia, S., Herrmann, L., & Killias, M. (2007). How important are interview methods and questionnaire designs in research on self-reported juvenile delinquency? An experimental comparison of Internet 49 vs paper-and-pencil questionnaires and different definitions of the reference period. Journal of Experimental Criminology, 3 (1), 39-64. https://doi.org/10.1007/s11292-007-9025-1 Maxfield, M. G., & Babbie, E. R. (2009). Basic s of research methods for criminal justice and criminology (2nd ed.). Cengage Learning. McLaughlin, E., & Muncie, J. (Eds.). (2001). The Sage dictionary of criminology . Sage. Murphy, F. J., Shirley, M. M., & Witmer, H. L. (1946). The incidence of hidden delinquency. American Journal of Orthopsychiatry, 16 (4), 686–696. https://doi.org/10.1111/j.19390025.1946.tb05431.x Nye, F. I., Short, J. F., & Olson, V. J. (1958). Socioeconomic status and delinquent behavior. American Journal of Sociology, 63 (4), 381-389. https://doi.org/10.1086/222261 Osgood, D. W., McMorris, B. J., & Potenza, M. T. (2002). Analyzing multiple-item measures of crime and deviance I: Item response theory scaling. Journal of Quantitative Criminology, 18 (3), 267-296. https://doi.org/10.1023/A:1016008004010 Piquero, A. R., Schubert, C. A., & Brame, R. (2014). Comparing official and self-report records of offending across gender and race/ethnicity in a longitudinal study of serious youthful offenders. Journal of Research in Crime and Delinquency, 51 (4), 526-556. https://doi.org/10.1177/0022427813520445 Porterfield, A. L. (1943). Delinquency and its outcome in court and college. American Journal of Sociology, 49 (3), 199-208. https://doi.org/10.1086/219369 Rosenbaum, S. M., Billinger, S., & Stieglitz, N. (2014). Let’s be honest: A review of experimental evidence of honesty and truth-telling. Journal of Economic Psychology, 45 , 181-196. https://doi.org/10.1016/j.joep.2014.10.002 Sanches, C., Gouveia-Pereira, M., Marôco, J., Gomes, H. S., & Roncon, F. (2016). Deviant behavior variety scale: Development and validation with a sample of Portuguese adolescents. Psicologia: Reflexão e Crítica, 29( 31), 1-8. https://doi.org/10.1186/s41155-016-0035-7 Schwarz, N. (1999). Self-reports: How the questions shape the answers. American Psychologist, 54 (2), 93-105. https://doi.org/10.1037/0003-066X.54.2.93 Schwarz, N., Strack, F., Hippler, H. J., & Bishop, G. (1991). The impact of administration mode on response effects in survey measurement. Applied Cognitive Psychology, 5 (3), 193-212. https://doi.org/10.1002/acp.2350050304 Sellin, T. (1931). The basis of a crime index. Journal of Criminal Law and Criminology, 22 (3), 335-356. https://doi.org/10.2307/1135784 50 Sellin, T. (1938). Culture conflict and crime: A report of the subcommittee on delinquency of the Committee on Personality and Culture . Social Science Research Council. Theobald, D., Farrington, D. P., Loeber, R., Pardini, D. A., & Piquero, A. R. (2014). Scaling up from convictions to self‐reported offending. Criminal Behaviour and Mental Health, 24 (4), 265-276. https://doi.org/10.1002/cbm.1928 Thornberry, T. P. (1989). Panel effects and the use of self-reported measures of delinquency in longitudinal studies. In M. W. Klein (Ed.), Cross-national research in self-reported crime and delinquency (pp. 347-369). Springer. https://doi.org/10.1007/978-94-009-1001-0_16 Thornberry, T. P., & Krohn, M. D. (2000). The self-report method for measuring delinquency and crime. In D. Duffee (Ed.), Measurement and analysis of crime and justice (pp. 33–84). U.S. National Institute of Justice. Tourangeau, R., & Smith, T. W. (1996). Asking sensitive questions: The impact of data collection mode, question format, and question context. Public Opinion Quarterly, 60 (2), 275-304. https://doi.org/10.1086/297751 West, D. J., & Farrington, D. P. (1977). The delinquent way of life . Heinemann. Wilson, J. Q., & Herrnstein, R. J. (1985). Crime and human nature . Free Press. Wright, D. L., Aquilino, W. S., & Supple, A. J. (1998). A comparison of computer-assisted and paper-andpencil self-administered questionnaires in a survey on smoking, alcohol, and drug use. Public Opinion Quarterly, 62 (3), 331-353. https://doi.org/10.1086/297849 51 CHAPTER II FIELD EXPERIMENTS ON DISHONESTY AND STEALING: WHAT HAVE WE LEARNED IN THE LAST 40 YEARS? Manuscript Published in: Gomes, H. S., Farrington, D. P., Defoe, I. N., & Maia, Â. (2021). Field experiments on dishonesty and stealing: What have we learned in the last 40 years?. Journal of Experimental Criminology . Advance online publication. https://doi.org/10.1007/s11292-021-09459-w 52 FIELD EXPERIMENTS ON DISHONESTY AND STEALING: WHAT HAVE WE LEARNED IN THE LAST 40 YEARS? Abstract Objectives. Field experiments combine the benefits of the experimental method and the study of human behavior in real-life settings, providing high internal and external validity. This article aims to review the field experimental evidence on the causes of offending. Methods. We carried out a systematic search for field experiments studying stealing or monetary dishonesty reported since 1979. Results. The search process resulted in 60 field experiments conducted within multiple fields of study, mainly in economics and management, which were grouped into four categories: Fraudulent/ dishonest behavior, Stealing, Keeping money, and Shoplifting. Conclusions. The reviewed studies provide a wide variety of methods and techniques that allow the realworld study of influences on offending and dishonest behavior. We hope that this summary will inspire criminologists to design and carry out realistic field experiments to test theories of offending, so that criminology can become an experimental science. Keywords: Field experiments; Naturalistic experiments; Stealing; Dishonesty; Systematic review 53 Introduction The main aim of this article is to encourage criminologists to carry out naturalistic field experiments to investigate the causes of offending. Theories of offending are usually tested in crosssectional or longitudinal studies. However, in trying to isolate the effect of a particular variable on offending, these methods can only attempt to control for other measured extraneous influences. Because of numerous unknown and unmeasured variables that might influence offending, these methods have low internal validity. In contrast, a randomized field experiment that manipulates influences on offending has higher internal validity, because the randomized design’s logic controls for all measured and unmeasured extraneous influences on offending, providing that a large number of units are randomly assigned (Weisburd, 2003). Information that is relevant to criminological theories can and should be drawn from the many experiments on prevention and intervention in criminology (see e.g., Robins, 1992), but conclusions can be drawn more directly by testing theories in naturalistic field experiments. This article presents a systematic review of field experiments on dishonesty and stealing that have been published in the 40 years since the seminal review by Farrington (1979). Remarkably, most of these experiments have been carried out by economists rather than by criminologists, and most have been designed to test ideas of rational decision-making influenced by subjectively expected benefits, costs, and probabilities. We believe that most criminologists are not familiar with this body of knowledge from the economics literature, and so we present summaries of all the experiments. We hope that these summaries will inspire criminologists to design and carry out realistic field experiments to test theories of offending, so that criminology can become a more experimental science (see e.g., Farrington, 2008). Experimental approach Experiments are the most important technique in developing scientific knowledge. The experimental approach implies the manipulation of variables under strictly controlled situations, which allows for the study of cause-and-effect relationships to provide unambiguous conclusions about the variables that influence behavior. The potential of experiments is very important in the study of criminal behavior because they can provide conclusive evidence about factors influencing offending, as well as predicting and preventing future offending behaviors. However, the vast majority of research in the field of criminology and criminal behavior is nonexperimental. Many researchers have pointed out the limitations of the experimental approach, mainly referring to the artificiality of laboratory settings which may contaminate the experiment, yielding inconclusive 54 results that are not easily generalizable to the real world. Field experiments, on the other hand, are very useful techniques that overcome these limitations by testing cause-and-effect relationships in real-world settings. In 1979, Farrington carried out a review of field experiments on deviance, urging researchers to carry out more of these realistic experiments. Recently, despite the limited number of field experiments developed in the field of criminology, several realistic field experiments on stealing and dishonesty have been carried out by behavioral economists (Farrington et al., 2020). The goal of the present article is to systematically review the field experiments on deviant behavior that have been carried out in the last 40 years, after the publication of Farrington (1979). In doing so, we explore the experimental designs, measurement techniques, and main findings of relevant studies in the interest of increasing the use of this robust technique in criminology. The scientific process or method is a systematic approach to acquiring knowledge that, through objective observation and hypothesis testing, enables an ever-growing body of knowledge (Christensen, 1985). Within the multiple types of studies and tools that constitute the scientific method, the experimental approach stands out because it allows for the testing of cause-and-effect relationships. Experiments can be described as “objective observation of phenomena which are made to occur in a strictly controlled situation in which one factor is varied and the others are kept constant” (Zimny, 1961, p. 35). The ability to control extraneous variables and precisely manipulate the independent variable (or variables) are key to arriving at unambiguous causal conclusions and provide pathways to ever more impactful treatments with fewer negative side effects, as well as cost-benefit estimations. Laboratory versus field experiments In this regard, laboratory experiments are the main experimental technique, since they maximize control. Therefore, the primary advantage of laboratory experiments is internal validity. In the laboratory, researchers are able to account for and minimize the influence of extraneous stimuli in an attempt to control the effect of environmental factors irrelevant to the study. However, the gains in internal validity conferred by the laboratory control come at the cost of artificial and sterile settings. This, in turn, may influence the results and reduce the study’s external validity, limiting the relevance for predicting behavior in the field as well as generalizability to the real world (Farrington, 1980; Harrison & List, 2004). On the other hand, field experiments are not subject to this artificiality problem, since they are carried out in real-life settings. Therefore, the main advantage of field experiments is external validity. Compared with cross-sectional and longitudinal studies, field experiments have high internal validity. However, the limited ability in some cases to control for extraneous variables in naturalistic environments 55 may cause a reduction in the internal validity of field experiments (Christensen, 1985). In some field experiments, this lack of control over the factors influencing behavior opens the possibility for alternative explanations, which may compromise the study of causal relationships (Pierce & Balasubramanian, 2015). A further potential disadvantage of field experiments is selection bias in the random selection of participants (Christensen, 1985). For example, a field experiment designed to study dishonest behavior of people buying journals may be affected by selection bias, since this sample (i.e., journal customers) may not represent the population of interest, namely, the offender population (e.g., Pruckner & Sausgruber, 2013). However, laboratory research may also be subject to selection bias because experiments, especially in psychological and social science research, are generally carried out with undergraduate students as participants, further limiting the external validity of laboratory experiments (Farrington, 1979). Finally, researchers must consider that in laboratory experiments, people are aware that their behavior is being scrutinized. This makes laboratory experiments subject to multiple sources of bias, such as social desirability, thus compromising their internal validity (Levitt & List, 2007). This is especially relevant in the study of deviance. The study of deviant behavior brings about additional concerns, because it is a highly sensitive topic that people try to conceal, possibly due to guilt, shame, or fear of repercussions (Gomes et al., 2019). Taking this into account, naturalistic field experiments on deviance, carried out in real-life contexts in which participants are unaware that their behavior is being studied, may offer the greatest internal and external validity of all methods (Farrington, 1979). Field experiments in the study of deviance Despite the apparent consensus on the relevance of experiments in the development of criminological knowledge and crime prevention practice, most research on deviance is nonexperimental, and naturalistic field experiments are still scarce in social science (Franzen & Pointner, 2013; Gomes et al., 2018). A quick search for the terms “crim* OR delinq*” in the Scopus database (i.e., article title, abstract, and keywords) results in a total of 267,523 documents up to 2018. On the other hand, the same search including the term “experiment” results in a total of 11,005 documents, which represents 4.11% of the studies. The same search for “field experiment,” however, results in only 239 documents, which represents 2.17% of all experiments and less than 0.1% (0.09%) of criminological research. Hence, field experiments on deviance are sorely needed. 56 In 1979, Farrington carried out a pioneering review of field experiments on deviance, with special reference to dishonesty. In that review, studies were included where members of the public were given the opportunity to dishonestly claim money, referring to such techniques as the lost coin where the experimenters pretend to pick up money (e.g., Farrington & Kidd, 1977; Feldman, 1968; Korte & Kerr, 1975) or leave coins in a telephone booth (e.g., Bickman, 1971; Franklin, 1973); experiments that provided opportunities for members of the public to engaged in offending behavior, for example, theft of candies (e.g., Diener et al., 1976), theft of shampoo out of a purposely forgotten expensive shampoo bottle (Steinberg et al., 1977), taking bags without paying (Lenga & Kleinke, 1974), and stealing money out of lost letters and/or wallets (e.g., Farrington & Knight, 1979, 1980; Hornstein et al., 1968). However, Farrington (1979) noted that, despite the wide variety of deviance that was studied, there were no studies on vandalism or property damage. He mentioned the famous study by Zimbardo (1969) but concluded that it did not meet the criteria for an “experiment” because of its inadequate control of independent and extraneous variables. Theoretical framework: factors influencing deviance Farrington (1979) proposed that engaging in the above-described dishonest behaviors can be considered a risky decision-making process. Therefore, a relevant specific theory would include the evaluation of the benefits and costs that follow from the choice to commit dishonest behavior (Farrington, 1979). Hence, in Farrington’s work (1979), the subjective expected utility (SEU) perspective was used as the main theoretical framework (see also Farrington & Knight, 1980). The SEU theory suggests that, in situations of risk (i.e., uncertainty), a decision about the alternative choices is based on (1) utility (i.e., subjective benefit or attractiveness), (2) subjective costs, and (3) their associated probabilities. Thus, each alternative choice has a total SEU, and, in the end, the decision-maker chooses the option with the highest SEU (Farrington & Knight, 1980). At the same time, Farrington (1979) also noted that solely focusing on costs and benefits is too simplistic to predict complex human behavior such as deviance. However, it is useful to start off with a simple and testable theory, and identify which results cannot be explained by it to subsequently determine in which ways it needs modifying or extending, rather than starting off with a complex theory that is less testable. Accordingly, in the current review, we explore whether the manipulation of benefits and costs predicted dishonest behavior in field experiments. Similar to Farrington (1979), in the current review, financial gains in some form are regarded as “benefits for the perpetrator.” Additionally, factors such as the suffering of other persons (victims) because of the actions of the perpetrator are regarded as “costs 57 for the other.” Of note, Farrington (1979) described conditions where the victims were less deserving (e.g., stealing from a young rich person) as less unpleasant and thus “low cost,” whereas conditions where victims were more deserving (e.g., stealing from an old poor person) were regarded as more unpleasant and thus “high cost.” We use the same definitions in the current review. Finally, Farrington (1979) also demonstrated that the likelihood of apprehension (i.e., costs for the perpetrator ) is also a relevant predictor of deviance. In the current review, we divide costs into costs for the other (i.e., costs for the victim) and costs for the self (i.e., costs for the perpetrator). The present study Farrington (1979) highlighted the benefits of naturalistic experimentation and expressed his wish “that psychologists will have the ingenuity, determination, and social responsibility to meet the challenge of experiments on deviance” (Farrington, 1979, p. 242). In order to provide criminology researchers with an updated review of the field experimental evidence relevant to the study of deviance, the present article aims to systematically review field experiments seeking to study the causes of offending or monetary dishonesty that have been reported since the review of Farrington (1979). We focus on field experiments on deviance that included financial dishonesty, as this overlaps most with an experimental way of studying delinquency (cf., Farrington, 1979). Unlike the review of Farrington (1979), the current review only includes studies with deviance as an outcome measure (whereas Farrington, 1979 also included studies that investigated deviance as an independent variable). In order to provide relevant information to researchers who might consider developing field experiments to test their hypotheses, the present review of field experiments on deviance will focus on the methods used to assess participants’ deviant or dishonest behavior in the field. Moreover, inspired by Farrington (1979), who concluded that many field experiments were motivated by cost-benefit theories such as SEU, we additionally coded the studies on whether they investigated independent variables that are related to benefits and costs (i.e., costs for the self and costs for the other ). In other words, we explore whether studies that manipulated these benefit and cost variables found that increases in benefits increases deviance, while decreases in costs increases deviance. 64 Furthermore, these field experiments presented multiple and creative methodologies in attempts to answer to different research questions that we were able to group into four different main topics, namely, fraudulent/dishonest behavior ( k = 21), stealing ( k = 16), keeping money ( k = 9), and shoplifting ( k = 14). Detailed information about all of these studies is presented in the results chapter. Fraudulent/dishonest behavior Within the fraud category, we have included field experiments that used a dependent variable related to illegal or dishonest behavior resulting in monetary or personal gain. This resulted in multiple types of measures of deviance, from low seriousness dishonesty such as sellers’ overcharging or methods usually applied in laboratory experiments such as the coin toss or the dice roll tasks, to more serious offensive practices such as insurance fraud (see Table 2). Five studies reported field experimental evidence related to overcharging. Balafoutas et al. (2013) carried out a naturalistic field experiment designed to study fraudulent behavior of taxi drivers. In this study, confederates posed as passengers and the taxi driver’s perception about the passenger was manipulated by the way passengers spoke and dressed, showing different degrees of familiarity with the city. By using GPS data, researchers were able to precisely record the chosen route and compare it to an estimated correct fare for the given distance, with the difference measuring the amount of overcharging. They found that taxi drivers more frequently overcharged passengers unfamiliar with the city, taking them on an average detour that more than doubled the length of the journey of familiar passengers. Conrads et al. (2015), as well as Dugar and Bhattacharya (2017), developed field experiments in order to study dishonesty in real-life pay-per-weight pricing markets. In these two studies, the purchased goods were weighted by the researchers after the transaction and the actual weight compared to the weight reported by the sellers. Conrads et al. (2015) employed this methodology to study overcharging occurring in candy stores. In this experiment, the authors found that overcharging occurred in 38% of purchases, though the apparent status of the buyer (high vs. low) and the quantity of candy bought (high vs. low) did not impact sellers’ dishonesty. In the case of Dugar and Bhattacharya (2017), overcharging was studied in fish markets. Results showed that most sellers overcharged (89%). Moreover, these results also showed how overcharging varied as a function of the potential economic benefit (i.e., the type and size of fish). The remaining two field experiments on overcharging (Jesilow & O’Brien, 1980; Schneider, 2012) had confederates visiting auto repair garages and submitting a test vehicle for repair with a prearranged set of defects. Findings from the study of Schneider (2012) showed that mechanics recommended 65 unnecessary repairs in 33% of visits. Moreover, when the researcher presented himself as one-time business, the total amount of repair cost increased significantly, compared with possible repeated business. Jesilow and O’Brien (1980) resorted to a similar methodology to study the effectiveness of deterrence interventions. In this experiment, the authors matched two areas by the degree of auto fraud and then subjected the experimental area to a deterrence intervention that included broadcasts of the existence of a state agency to which the public could report questionable repair dealers and a letter sent to the repair garages reminding them of the law and the consequences of violation. The opportunity for fraudulent behavior was created by having female confederates enter the repair facilities requesting the shops to test their car batteries (i.e., the “battery test”). Findings showed that the percentage of shops wrongly recommending a new battery in the post intervention phase was much higher in the control group compared to the intervention group. Similar methodologies were employed to study insurance fraud. In the field experiment developed by Tracy and Fox (1989), confederates visited random auto body repair shops and obtained estimates of repair costs. In this case, experimenters manipulated whether the car was or was not being covered by insurance, as well as the sex of the driver. The results showed much higher repair estimates for insured vehicles, showing that the auto shops would inflate the prices in the insured condition. Furthermore, these results also showed, not only a sex-of-the-driver effect, where the estimated repair costs were much higher to female drivers, but also a sex-coverage interaction in which the male-female differences were even greater in the non-covered condition, suggesting that male drivers were better able to “get a break.” More recently, Kerschbamer et al. (2016) also studied insurance fraud in computer repair shops. Confederates entered the repair shop and submitted manipulated test computers for repair. Results clearly showed a much higher average repair price when confederates were covered by insurance, compared to when they were not covered. Taking into consideration that insured clients who are victims of theft have the opportunity to boost their losses in order to achieve monetary gains, three studies focused on insurance fraud to study the impact of deterrent letters on insurance customers (Blais & Bacher, 2007; Shu et al., 2012; Tremblay et al., 2000). Tremblay et al. (2000) manipulated whether the claimants received a deterrent or permissive letter, as well as whether the claim regulation was carried out on the telephone or by having regulators visiting the insurer’s home. Findings showed main effects of claim regulation, where settling the claim over the telephone led to higher losses per claim. Interaction effects also showed that the permissive letter increased the average claim amounts only when the claims were settled remotely by 66 telephone, while the deterrent letter decreased the average amount claimed only in the face-to-face condition. Blais and Bacher (2007) also applied measures of insurance fraud, in this particular case to study the effects of the threat of legal sanctions on offending behavior of insurance customers. In this study, insurance companies randomly assigned reports of property theft to the control (business as usual) or experimental group. Claimants in the experimental group received a deterrent letter reminding them of the sanctions associated with claim padding. Findings showed that the deterrent letter decreased the likelihood of claim padding. In Shu et al.’s (2012) field experiment, the authors manipulated the policy review form by making insurance customers report the current odometer mileage of their insured cars and sign it either at the beginning or at the end of the form. Seeing that a lower odometer mileage indicated a lower risk of accidents and, thus, lower insurance premiums, participants were expected to underestimate their car’s mileage. Results of this field experiment showed that customers who signed at the beginning provided about 10% higher mileage estimates than those who signed at the end. Nagin et al. (2002) designed a field experiment to study the effects of monitoring of employees’ fraudulent behavior. Participants in this experiment were telephone solicitors at a call center, and their salary increased with the number of successful solicitations (i.e., contributions from potential donors). Given this incentive to claim higher solicited donations, the company monitored for falsely reported donations (i.e., “bad calls”). In this field experiment, the audit rate for bad calls that were reported back to employees was manipulated. Results showed that a perceived reduction in monitoring was quickly followed by more fraudulent behavior by employees in the number of bad calls. List and Momeni (2017, 2020) carried out two field experiments to study workers’ fraudulent behavior. In these field experiments, the authors employed online workers through MTurk (i.e., an online labor market platform) to perform a transcription task for payment. Workers had to transcribe 10 scanned images of short German texts. If the images were unreadable, workers could report and skip that image, moving on to the next image. This provided an opportunity for workers to misreport readable images as unreadable, allowing them to get the payment with less effort. A different way to behave fraudulently in this experiment was to take the upfront payment without completing the job. In the first experiment, List and Momeni (2017) paid 10% of the total payment upfront, and manipulated the total wage (i.e., $0.90; $1.20; $1.26) and Corporate Social Responsibility (CSR) by making a charity donation on behalf of the firm or on behalf of the workers. Results showed that the decrease in the wage and the increase of the expenditure on CSR, especially when framed as a pro-social act on behalf of the workers, caused an increase in the number of employees acting dishonestly. In List and Momeni’s (2020) second field 67 experiment, the authors followed the same design and manipulated the amount of upfront payment (i.e., 0%, 10%, 50%, and 90% of the total pay). Findings showed that, compared to the baseline condition, all conditions with upfront payment decreased dishonesty. On the other hand, within the upfront conditions, larger upfront payments were related to increases in the workers’ dishonest behavior. Olken (2007) developed a field experiment in order to study the impact of monitoring in the fraudulent behavior of villagers in Indonesia. In this experiment, funds were awarded to villages for the construction of roads. The information provided about government auditing (i.e., “external audits”) and direct participation in the monitoring process by villagers was manipulated. In order to assess fraudulent behavior in the construction of roads, core samples of the roads after the projects were completed were dug up, and the quantity of materials used was estimated. The difference between the amount the village claimed to have spent on the project and the engineers’ estimated price was the measure of fraudulent behavior in this study. Findings showed that, contrary to direct participation which did not affect village fraud, increasing the probability of external audits caused a substantial reduction in missing funds in the project. Bertrand et al. (2007) carried out a field experiment in order to study whether the allocation of driver’s licenses in India was influenced by a candidate’s willingness to pay. In this experiment, driver’s license candidates were randomly assigned to the control group, given free driving lessons, or given a large financial reward (i.e., the bonus group) if they obtained the driver’s license in 32 days, two days longer than the minimum legal time of 30 days. Furthermore, upon obtaining the driver’s license, participants were invited to a final session and enrolled in a surprise practical driving test in order to assess their driving skills. Results showed that participants in the bonus group were more likely to make extralegal payments and to obtain licenses without really knowing how to drive. Green (1985) studied fraudulent behavior by auditing homes which had a “basic” cable service but which stole premium cable television signals with an unauthorized descrambler. This field experiment aimed to study general deterrence hypotheses by sending people known to be stealing signals a written legal threat, providing an amnesty period to rectify the situation without being prosecuted. The cable terminals were re-audited immediately after and 6 months after the intervention. Results of this experiment showed that about two-thirds of the original violators stopped stealing cable signals and this effect was maintained during the follow-up period. The remaining five field experiments included in this category resorted to techniques frequently used in laboratory settings to study dishonest behavior, namely the dice roll task (Chytilová & Korbel, 2014; Okeke & Godlonton, 2014; Siniver & Yaniv, 2018) and the coin toss task (Bucciol & Piovesan, 68 2011; Houser et al., 2016). Experiments using these methodologies ask participants to roll a dice or toss a coin and report the outcome, knowing that different outcomes result in different rewards. These tasks are usually performed in unmonitored conditions, in order to assure participants that only they know the true outcome, creating an opportunity for them to act dishonestly for financial gain. These methods are unable to assess individual dishonest behavior, but the comparison of reported outcomes and the baseline distribution makes it possible to measure cheating at the aggregate level (see Rosenbaum et al., 2014). Chytilová and Korbel (2014) used the dice roll task with school students in order to study whether group settings influence dishonest behavior. The reward for completing a questionnaire was equal to the dice outcome, with the exception of the number “6” which would result in no payoff. Students rolled the dice either individually or in groups of three. Groups could also be determined randomly (exogenous groups) or students formed the groups themselves (endogenous groups). The main findings of this study showed that students in group settings (independently of the exogenous or endogenous formation) were more likely to act dishonestly. Okeke and Godlonton (2014) applied the dice roll technique to study whether pro-social preferences lead to dishonest behavior. These authors recruited female interviewers to carry out interviews in the community. Interviewers visited households and distributed discounted price vouchers. Interviewers were supposed to ask the interviewees to roll the dice, and the amount of the voucher depended on the score they rolled. The misallocation of price vouchers was the measure of interviewer dishonesty. Results of this experiment showed that interviewers were more likely to allocate higher value vouchers to the poorest interviewees. In the Siniver and Yaniv (2018) field experiment, participants were recruited after purchasing and scratching scratch cards at selling kiosks in order to study the effect of winning and losing in the lottery on dishonest behavior. Participants were asked to carry out the dice roll task under a cup and the monetary reward was determined by the participants’ report of the outcome. Results showed that lottery losers acted more dishonestly than lottery winners. Furthermore, the higher the lottery losses, the higher the dishonest behavior. Regarding experiments using the coin toss task, Bucciol and Piovesan (2011) studied children’s dishonest behavior by asking summer campers to toss a fair coin in private. Depending on the reported outcome, they earned a prize, thus providing an incentive for them to act dishonestly. In this experiment, researchers manipulated whether or not they mentioned the possibility of cheating to the participants, requesting the experimental group not to cheat. Results showed that participants cheated somewhat in both groups and throughout the different ages (from 5 to 15 years), although boys cheated more than 69 girls. Nevertheless, the honesty request made to the experimental group reduced dishonest behavior by 16%. Houser et al. (2016) used the coin toss technique to study the dishonest behavior of parents when the payoff was a toy for the child or cash for the parent. The presence of the parent’s child in the room during the coin toss was also manipulated in order to study whether the presence of the child would increase scrutiny and thus lessen dishonest behavior. Accordingly to the authors’ predictions, parents were more likely to act dishonestly to benefit their child than to benefit themselves. Also, dishonest behavior was expected to be higher when the child was not present. However, this effect was only found when the daughter was present, and parents’ dishonest behavior did not change in the presence of their sons. Stealing A total of 16 field experiments were included in the stealing category (Table 3). In this category, field studies used two main methodologies. The first was the “lost” letter technique (and some adapted versions of this technique), which consists of leaving stamped, addressed, and apparently lost letters in determined places, typically containing a sum of money. The failure to return a “lost” letter containing money was defined as stealing. The second group of methods in this category used multiple techniques that provided the opportunity for participants to steal things such as pens, newspapers, and money. Within the studies using the “lost” letter technique, the research conducted by Gabor and Barker (1989) used “lost” letters in order to study the prevalence of dishonesty in Canada. In their field experiment, researchers planted letters under the windshield wipers of cars of selected participants, with a note stating “found near your car.” These envelopes contained a coin and a letter either appearing to be a personal and trivial letter or an official letter stating that the value of the coin was $150. Overall, about one-quarter of sample failed to return the “lost” letter. However, the stated value of the coin failed to significantly impact the stealing of the letter. In agreement with previous experiments, participants’ sex had little effect on stealing, contrary to their age, where younger participants were less likely to return the apparently lost letter. The study conducted by Cohn et al. (2019) reported three large-scale field experiments conducted in 40 countries, using an adaptation of the “lost” letter technique to study civic honesty, by providing participants with the opportunity to return or steal a “lost” wallet. In these field experiments, confederates approached an employee at the counter (e.g., banks, hotels) and said that he/she found a wallet on the street and asked the employee to take care of it. Wallets included the “owners’ personal information” 70 which allowed the employee to voluntarily return the “lost” wallet. In the first field experiment, the authors manipulated the money in the wallet, either no money or $13.45 USD. Overall, results showed that citizens were much more likely to return the “lost” wallets with money than without. Moreover, despite dishonesty rates varying from 86% to 24% of cases, analyses showed that in none of the 40 countries the money condition increased significantly the likelihood to steal the “lost” wallet. In their second field experiment, Cohn et al. (2019) tested the same hypothesis with a larger amount of money contained in the “lost” wallet (i.e., $94.15 USD). Dishonesty rates decreased even further with the big money condition, showing that the honest return rates for the “lost” wallet were higher when the larger amount of money was added. In the third field experiment, the authors manipulated whether the wallets with money included or did not include a key, in order to study the effect of an item valuable to the owner. Results of this last study showed that adding the key increased the return rates of the “lost” wallet, suggesting people’s concern for harm to the owner. A further adaptation of the “lost” letter technique to study stealing was used in three studies (Keizer et al., 2008; Keuschnigg & Wolbring, 2015; Lanfear, 2018). This methodology consisted of leaving an envelope, visibly containing money, either hanging out of or nearby to a mailbox, and observing passerby behavior. Keizer et al. (2008) carried out six field experiments in order to study whether setting cues of violation of a contextual norm (e.g., graffiti in an anti-graffiti area) impacted deviant behavior. For the purposes of the present review, we are only going to focus on the last two field experiments referring to stealing, since the previous field experiments focused on littering and trespassing. In the fifth and sixth field experiments, “lost” letters visibly containing cash were left hanging out of a mailbox. The authors manipulated whether the setting was or not covered with graffiti (i.e., experiment 5) and whether or not there was litter on the floor around the mailbox (i.e., experiment 6). Results of both field experiments showed an increased odds of stealing the “lost” letter in the disorder conditions. Keuschnigg and Wolbring (2015) carried out three field experiments that sought to replicate Keizer et al.’s (2008) field experiments on littering and stealing, and added an adaptation of these experiments to jaywalking. Similar to the previous study, only the field experiment on stealing falls within the scope of the present review. In the stealing experiment, apparently lost letters were left in front of a mailbox with visible cash in them. The authors manipulated the amount of money in the envelope (€5, €10, or €100). Also, the area surrounding the mailbox was either kept clean or there were two heavily wrecked bicycles next to the mailbox. Results clearly replicated the previous experiment, showing an increased odds of stealing the “lost” letter in the physical disorder condition. Furthermore, this spillover effect of the norm violation on stealing behavior was influenced by the amount of cash contained in the 71 envelopes, where the effect was the strongest when the envelopes contained the €5 note (i.e., people steal more in the disorder condition), weaker for the €10 note, and disappeared completely with the €100 note, showing that “once stakes are high, the relevance of environmental cues diminishes” (Keuschnigg & Wolbring, 2015, p. 120). Lanfear (2018) carried a similar field experiment to study some key features of the broken windows theory. In this experiment, local physical disorder was manipulated by the addition or not of both litter and graffiti, and using the adapted “lost” letter technique with envelopes containing a $5 bill left near the mailbox. This experiment failed to replicate the results of Keizer et al. (2008) as well as Keuschnigg and Wolbring (2015). Local disorder failed to impact passerby behavior on stealing the “lost” letter. Nevertheless, evidence indicated that in the disorder condition, participants were less likely to act pro-socially by mailing the “lost” letter. The remaining 11 field experiments included in the stealing category used multiple methodologies that created an opportunity for participants to steal. Castillo et al. (2014), for example, sent out envelopes to Lima, Peru, from two cities in the USA via normal mail services. Researchers manipulated whether or not the envelopes contained cash, as well as the sender’s name, i.e., a foreign name (i.e., J. Tucker, M. Scott) or a local name (i.e., M. Sosa, L. Cordova). This methodology was developed to study whether the very nature of the mail influences stealing behavior of the people who handle the mail. Results showed that the envelopes containing money were much more likely to be lost. Furthermore, the mail was much more likely to be lost if the sender’s last name matched the recipient’s last name (i.e., a local name). Belot and Schröder (2015), as well as Greenberg (2002), created the opportunity for participants to steal cash. In Belot and Schröder’s (2015) field experiment, the authors recruited students for a paid job of identifying the provenance of euro coins collected in different countries. Contrary to what participants were led to believe, a fixed number of coins was given to each participant, allowing the researchers to count the cash and assess the number of stolen coins. The authors manipulated whether or not participants were monitored, as well as incentives associated with monitoring, where participants’ mistakes in the coin sorting task were penalized either mildly or harshly. Results showed that about 10% of participants stole coins, though monitoring or incentives had no impact on participants’ stealing. Greenberg (2002) used a sample of employees of a financial services company and asked them to complete a survey regarding working conditions in exchange for a payment. After completing the task, participants walked into an unsupervised room where they found a bowl of pennies, from which they should count the $2 USD that was due to them. The researchers knew the total number of pennies that were in the bowl, allowing them to figure out whether or not the participant stole coins. These employees 72 belonged to two different locations, in one of which an ethics program was in place. Furthermore, the authors also manipulated whether the payment was coming from either personal funds or the company. The results showed that participants attending the corporate ethics program had a lower likelihood of stealing coins, and that participants stole more often when the money was said to come from a company. Greenberg (1990) carried a field experiment that also focused on employee theft, in this particular case, concerning the inventory of a firm. Employees of several manufacturing plants were either or not subjected to a 15% pay reduction during a period of time. The groups receiving the wage cuts were divided into two groups. One group received an adequate explanation for the wage cut by the company president, while the other group was in the inadequate-explanation condition. Employee theft was assessed by the percentage of unaccounted inventory lost. Results revealed that employee theft increased during the pay reduction period. Furthermore, the theft rate in the inadequate-explanation condition was much higher than in the adequate-explanation condition. Cohn et al. (2014) also studied the impact of wage cuts on employee theft. In this specific case, hired workers were asked to sell promotional cards. While selling these promotional cards, workers were supposed to collect information from the customer. This created the opportunity for willing workers to steal the cash sales and fake customer information. Incorrect customer information could be checked by the research team. Sales were carried out in groups of two, and, in the first phase, all workers were given the same hourly wage. In the second phase, the wage either stayed the same, both group members received a 25% wage cut, or only one group member suffered the 25% cut. Results showed that the wage cut created an increased likelihood of employee theft, but this only happened for the employees who were directly affected by the wage cuts in the unilateral condition. Widner (1998) developed a field experiment in order to assess the effectiveness of a series of intervention techniques aimed to reduce the theft of petrified wood in a national park. The three interventions tested in this study included a uniformed volunteer, deterrent signs, and a signed pledge, and each was randomly in place for 10 days. The theft of petrified wood was assessed by direct field observation carried out by the research team. Using this methodology, researchers were able to observe a theft rate of 2.1% in the control condition, and this reduced to about 1.4% in the intervention conditions. These results revealed that the three interventions were effective in the reduction of theft, when compared to the control condition. Furthermore, these interventions showed no differential effectiveness. Schlüter and Vollan (2015) developed a field experiment where they studied the theft of flowers in a farmer’s field using an honor system. This was an unattended system that allowed the customer to enter the flower field, cut the intended flowers themselves, and pay the respective sum in a cashbox, 73 relying entirely on the honesty of customers. Researchers left a message near the cashbox which varied between legal threats, moral persuasion, and referencing a family business or a consulting firm. Theft of flowers was assessed through direct observation carried out by the researchers through a semitransparent window by counting the flowers and the respective payment into the cashbox. Findings suggested no main effect of the legal or moral messages. However, flower theft increased when the flower field was framed as a company business, compared to the family business condition. Two other field experiments used the honor payment system in order to study stealing, in their case of newspapers (Geller et al., 1983; Pruckner & Sausgruber, 2013). In order to ensure unmonitored transactions, experimenters placed just one paper in the sales booth and checked for payments at specific intervals. If the newspaper had been taken, the cashbox would be emptied recording the amount paid (Pruckner & Sausgruber, 2013). In the field experiment conducted by Geller et al. (1983), two anti-theft sign messages were implemented; one appealed to moral, internal control and the second showed a legal threat. Results supported the effectiveness of both messages in reducing newspaper theft. Similarly, in the second field experiment on theft of newspapers using the honor system (Pruckner & Sausgruber, 2013), the authors also tested the impacts of a moral, a legal, and a neutral control message. Findings revealed that about two-thirds of customers stole the newspaper, and those who paid did so by depositing much less than the indicated price (i.e., €0.60). The treatments showed no effects on newspaper theft. However, the appeal for honesty in the moral condition caused an average increase on the amount paid, compared to both control and legal treatments. The final two field experiments included in the stealing category focused on university students. In the experiment conducted by Cagala et al. (2014), students were randomly allocated to two groups with different levels of monitoring during a university exam. Students in both groups were provided with a high-quality pen that they were supposed to deliver in the post-exam phase, where the levels of monitoring were the same throughout the experimental groups. Results showed that the monitoring in the exam phase caused an intertemporal spillover effect, where participants in the low monitoring group were much more likely to steal the pen. Finally, Wortley and McFarlane (2011) carried out a field experiment in a university library and created the opportunity for students to steal a photocopying card. Researchers left a photocopying card unattended on a library table and observed passerby behavior from a distance. Researchers manipulated ownership of the card by using either a signed or an unsigned card, and manipulated guardianship, by placing the card either next to library books (giving the impression that the owner was nearby) or on its own. Both variables of symbolic territoriality (i.e., signed cards/next to 80 & Momeni, 2020; Tremblay et al., 2000). As for costs for the self , all of the seven studies that manipulated this factor showed that a lower likelihood of apprehension predicted higher levels of fraud. Stealing The three studies in the stealing category that manipulated “costs for the other” found significant effects. One of these studies showed that theft increased when payment came from a company (lower costs) compared with personal funds (higher costs) (Greenberg, 2002). A second study showed that, when the owner’s name was not signed on a photocopying card in a library (i.e., low costs), it was stolen more often (Wortley & McFarlane, 2011). The other study showed that, when “lost” wallets contained a personal item valuable to the owner (i.e., high cost for the other), the return rates of the “lost” wallet increased (Cohn et al., 2019). Next, “costs for the self” was investigated in three studies. Two of these studies found significant deterrent effects of monitoring; namely, Cagala et al. (2014) found that high monitoring during the exam phase decreased pen theft in the post-exam phase, whereas Widner (1998) found that having anti-theft interventions decreased petrified wood theft. The other study did not find that monitoring decreased the theft of coins (Belot & Schröder, 2015). Finally, four field experiments manipulated the amount of benefits to the self. Castillo et al. (2014) found that letters containing money increased mail theft. Keuschnigg and Wolbring (2015) found a significant effect in interaction with another variable (i.e., disorder environmental cues). One other study found that “lost” letter theft was not affected by the apparent value of the contained coins (Gabor & Barker, 1989). Moreover, two field experiments carried out by Cohn et al. (2019) found an opposite effect compared to our hypothesis, where the higher the amount of money in a “lost” wallet, the less stealing was committed by employees (i.e., in such cases, the employees at the counters more often mailed the wallets back to the hotel guests). Keeping money Concerning the costs for the other manipulation, we only located one study (Gabor et al., 1986) that fitted this description. Gabor et al. (1986) investigated cashiers’ dishonesty in keeping the change of customers in chain stores (low costs for the other) versus family stores (high costs for the other). However, unexpectedly, the chain stores condition did not lead to more cashiers’ dishonesty regarding keeping the change of customers. As for costs for the self , the only such study (Armantier & Boly, 2011) in the keeping 81 money category showed that low (versus high) monitoring, coupled with punishment if caught, led to increases in accepting a bribe. Finally, we located five studies in the keeping money category that manipulated benefits. Three of those studies (Armantier & Boly, 2011; Newman, 1979; Rabinowitz et al., 1993) consistently showed that higher benefits predicted more instances of participants keeping or accepting money that was not theirs (i.e., picking up a dropped coin; acceptance of a bribe; keeping due change). However, although the remaining two studies also found significant effects, the effects were in the opposite direction compared to our hypothesis. In Azar et al. (2013), customers of a restaurant received extra change after paying, and the amount of extra change was manipulated. Higher amounts of extra change actually decreased the instances in which customers kept the “extra” change. Similarly, in Yuchtman-Yaar and Rahav (1986), bus drivers gave passengers extra change and the amount of extra change was manipulated. For females, higher amounts of extra change increased dishonesty, but for males, higher amounts of extra change actually decreased keeping the extra change. It is of note is that both studies that found the opposite effect for benefits originated from Israel. Shoplifting For the shoplifting category, regarding components of the SEU theory, we only found studies that manipulated costs for the self . The intervention study of DiLonardo and Clarke (1996) investigated security measures to prevent shoplifting and showed that ink tags (versus electronic article surveillance; EAS) reduced shoplifting. Of note is that both ink tags and EAS increased the chances of apprehension (i.e., costs for the self ) compared with a condition without security measures. Thus, the intervention in DiLonardo and Clarke (1996) would have been a more stringent test of the costs for the self hypothesis, if its security conditions had been compared to a condition in which no security measures were used. On the other hand, Hayes and Downs (2011), Hayes et al. (2011, 2012, 2019), and Johns et al. (2017) compared control conditions to anti-shoplifting interventions (i.e., CCTV, keeper or safer boxes, protective product display, or anti-theft wire wraps) and showed that these interventions reduced the stores’ theft rates. McNees et al. (1980b) showed that an anti-shoplifting intervention directed to elementary school students reduced the rates of theft, though these findings were not maintained over time. Finally, Farrington et al. (1993) conducted a series of experiments and showed that electronic tagging reduced shoplifting, and this effect was maintained over time; however, a uniformed guard did not affect shoplifting. 82 Benefits versus costs for the other versus costs for the self In sum, the above-described results show that when the chance of apprehension ( costs for the self ) is low, more dishonest behavior takes place. This pattern of findings was found in all the seven studies on fraud, in two out of the three studies on the stealing, in the study on the keeping money, and in all eight studies on the shoplifting category. Concerning costs for the other , there were too few studies that manipulated this factor in order to draw strong conclusions for each category. In the shoplifting category, there were no studies that manipulated costs for the other . Across the categories, four out of eight studies that manipulated costs for the other found that when costs are low for the victims (e.g., an insurance company versus an individual), then dishonest behavior increases. Finally, when it comes to benefits, the studies across the different categories consistently showed that high benefits predicted dishonest behavior, as seven of the 11 studies found such significant effects. However, of note is that, in the fraudulent behavior category, only two studies manipulated benefits (and both studies found significant effects), and in the shoplifting category, no study manipulated benefits. Seven of the 10 studies that found significant results (i.e., three studies for the stealing category, five studies for the keeping money category, and two studies for the fraudulent category) found an effect in the hypothesized direction showing that more benefits led to more stealing and keeping money. However, in the remaining three cases, the opposite pattern of effect was reported: fewer benefits predicted more dishonest behaviors of perpetrators when the studies manipulated the amount of extra change given to customers or when the experiment manipulated how much money was in a lost wallet. Perhaps the relation between the benefits and the probability of dishonest behavior follows an inverted-U shape. Past, present, and future The current review shows that researchers in many different parts of the world have carried out field experiments to study financial dishonesty. Such cultural diversity is very welcome, in order to determine to what extent theories are generalizable. Of course, legal definitions of deviance (e.g., theft) might vary substantially across cultures. Such discrepancies should be kept in mind when interpreting the findings of the studies highlighted in this review. However, a further advantage of the field experimental methodology to study offending and dishonest behavior is the fact that the reviewed experiments focused on naturally recurring behaviors, and are generally independent of the legalistic definitions of offending. Within the present study, in order to review the field experimental evidence relevant to the study of deviance, including the field experiments on stealing and dishonesty developed by behavioral 83 economists, we have systematically reviewed the field experimental studies on stealing and monetary dishonesty. However, in doing so, we have not included the field experiments on other types of deviant behavior that might be of interest to the study of criminal behavior, such as littering, jaywalking, or vandalism (e.g., property damage). Hence, readers should bear in mind that the findings in the present review might not be generalizable to other types of deviant behavior, and we encourage researchers to investigate these topics in the future. Especially in the study of vandalism, this type of deviance should be relatively easy to investigate in field experiments, considering that (1) it often happens in public view and (2) it is less ethically sensitive compared to other types of deviancy (e.g., theft, sexual assault, physical assault) (Farrington, 1979). For example, vandalism experiments could be conducted in areas where vandalism already takes place in public view (Zimbardo, 1969). Therefore, researchers would need to worry less about ethical considerations associated with providing individuals with the opportunity to act in a deviant manner, which is typically the case in field experiments on deviance. Costs and benefits were the focus of this review, in part because these are immediate situational factors that are suitable for manipulation in short-term experiments (Farrington & Knight, 1980). However, it should be noted that dishonesty is a complex behavior, which cannot solely be explained by such immediate situational factors. Future studies should also attempt to vary other non-situational variables (e.g., impulsivity), as well as social environmental factors (e.g., the presence of peers) (Defoe et al., 2019; Defoe, in press). Studies that manipulate the social context remain rare in field experiments in the criminology literature. However, the few available studies suggest that the immediate social context also plays a role (Farrington, 1979). Conclusion Our review shows that it is worthwhile for criminologists to study influences on offending using field experiments within a SEU framework. This review clearly demonstrates that variations in the benefits and costs (particularly the likelihood of apprehension) associated with a dishonest act are important predictors of offending. Specifically, higher levels of financial benefits and lower probabilities of apprehension predict higher levels of dishonesty. Interestingly, some studies found that fewer benefits led to more stealing. More research is needed on why this effect is sometimes in the opposite direction, and why higher benefits sometimes lead to less dishonesty. Perhaps in such cases, there could be an interaction with costs and benefits that are driving the effects. Therefore, future studies are also encouraged to investigate potential interactions between costs and benefits. 84 The present review shows how immediate situational influences on dishonesty (e.g., costs and benefits) can be manipulated in field experiments to better understand the causes of stealing and dishonesty. Although many economists have undertaken this challenge, such experiments in criminology remain rare (for an overview see Clarke, 1995; Clarke & Cornish, 1985). However, field experiments on financial dishonesty overlap considerably with everyday delinquency, and hence, such experiments could be a powerful tool for criminologists (Farrington, 1979; Farrington et al., 2020). In fact, targeting immediate situational factors that predict crime could be just as successful as prevention programs that solely target individual characteristics (e.g., impulsivity). We conclude that criminologists should seek to carry out naturalistic field experiments on offending to investigate theories and explanations of offending. 85 Table 2 Summary of field experiments in the Fraudulent/ dishonest behavior category Study Participants Design Measure Main findings SEU Balafoutas et al. (2013) Greece Taxi drivers (348 taxi rides) Task: taxi ride. Manipulation - taxi driver’s perception of customers: Information about the city: Local vs. Non-local natives; Information about the tariff system: Native vs. Foreigner; Income: Low vs. High income. Overcharging Non-local natives increased overcharging. Foreigner customers increased overcharging. Customer’s perceived income did not impact overcharging. Costs for the other Bertrand et al. (2007) India 822 driver’s license candidates Task: Obtaining a driver’s license. Manipulation: Prize: Bonus (large financial reward if obtained in 32 days) vs. Free driving lessons vs. Control. Extra-legal payments Bonus group members are more likely to make extra-legal payments to obtain licenses. *Benefits Blais and Bacher (2007) Canada Insurance customers (765 claims) Task: Insurance companies randomly assigned claims of property theft to study groups. Manipulation: Deterrence: Conventional vs. Deterrent letter. Insurance fraud (i.e. claim padding) The deterrent letter decreased fraudulent behavior. *Costs for the self 86 Bucciol and Piovesan (2011) Italy 160 children attending a summer camp Task: Summer campers were asked to carry out the coin toss task as a typical camp activity. Manipulation: Honesty request: Control vs. Explicit request to refrain from cheating. Coin toss task. Honesty request reduced cheating. NA Chytilova and Korbel (2014) Czech Republic 226 school students Task: Students were recruited for a task. Reward would be determined by the dice roll task. Manipulation: Setting: Individual vs. Groups of three; Group formation: Exogenous (randomly formed groups) vs. Endogenous (groups formed by themselves). Dice roll task. Group settings increased dishonesty. Group formation did not impact students’ dishonesty. NA Conrads et al. (2015) Germany Candy sellers in 50 markets (200 observations) Task: Confederates entered the market and bought a bag of candy. Manipulation: Status of the buyer: Wealthy vs. Poor; Quantity of candy bought: High (150g) vs. Low (50g). Overcharging Overcharging in 38% of purchases. Both status of buyer and quantity of candy did not affect overcharging. Costs for the other Dugar and Bhattacharya (2017) Fish sellers in 10 markets (160 observations) Task: Overcharging. Overcharging in 89% of purchases. NA 87 India Confederates entered markets and purchase a pre-determined quantity of fish. Manipulation: Size of fish: Small (less expensive) vs. Large (more expensive); Type of fish: Rohu (less expensive) vs. Catla (more expensive). Within less expensive type of fish, large fish increased overcharging. Within more expensive type of fish, small fish increased the probability of overcharging. Green (1985) USA 67 subjects found to be stealing cable television signals Pretest / Posttest design. (1) Researchers identified houses that illegally tempered with terminals; (2) a deterrent letter threatening criminal prosecution was sent; (3) a re-audit was developed after the letter was sent; and (4) follow-up audit six months after. Stealing cable television signal. The deterrent letter decrease cable crime. Deterrent effect lasted at least six months. *Costs for the self Houser et al. (2016) USA 249 parent-child pairs Task: Parents of 3-6 year-old children were recruited for a task. Reward would be determined by the coin toss task. Manipulation Scrutiny: Parent is alone vs. Child in the room during the coin toss; Moral cost: Low (prize pack for the child) vs. High ($10 for the parent). Coin toss task. Low moral cost increased cheating. No scrutiny increased cheating. NA 88 Jesilow and O'Brien (1980) USA 145 auto shops Pretest / Posttest design Task: Female confederates entered repair facilities and requested to test the car batteries because they their cars would not start. Manipulation: Intervention: Control vs. Deterrent intervention (deterrent announcements and letters). Fraud in the vehicle repair price. Deterrent intervention decreased fraudulent behavior. *Costs for the self Kerschbamer et al. (2016) Austria 61 computer repair shops Task: Confederate entered computer repair shops and asked for a repair. Computers were manipulated with a destroyed RAM module. Manipulation: Insurance: Control vs. Insurance Fraud in the computer repair price. Average repair price in Control and Insurance groups was 70.17€ and 128.68€, respectively. Insurance increased fraudulent behavior. *Costs for the other List and Momeni (2017) USA and India 3,022 hired workers through MTurk Task: Participants were contracted online to transcribe 10 images. Before starting, workers reported whether the image was readable. If not readable, the transcription was not necessary. Dishonesty Decrease in wage increased dishonest behavior. Implementation of CSR, especially on behalf of the workers, increased dishonest behavior. NA 89 Manipulation: Wage: Low ($0.90) vs. Medium ($1.20) vs. High ($1.26); Corporate Social Responsibility (CSR): Charity donation on behalf of the firm vs. On behalf of the workers. List and Momeni (2020) USA 2,000 hired workers through MTurk Task: Participants were contracted online to transcribe 10 images. Before starting, workers reported whether the image was readable. If not readable, the transcription was not necessary. Manipulation: Upfront payment: 0% vs. 10% vs. 50% vs. 90% of the total pay. Dishonesty Upfront payment decreased dishonesty, when compared to 0% upfront. Considering upfront conditions, the higher the upfront payment the higher the probability to behave dishonestly. *Benefits Nagin et al. (2002) USA Employees of a large call center company working Task: Employees called potential donors and request contributions. Payment was a base salary and a bonus for the number of successful solicitations. Manipulation Reported monitoring to employees: audit rates varied. Fraud (i.e. “Bad calls”). A reduction in monitoring increased fraud. *Costs for the self 96 USA services company payment in private from a bowl of pennies. Manipulations: Victim of theft: Organization vs. Individual (money was being paid from personal funds); Corporate ethics program: Control vs. Office in which an ethics program in place. Working in an office in which there was no ethics program in place increased employee theft. Keizer et al. (2008) Netherlands 203 members of the public Task: Participants passed by a mailbox and noticed an envelope visibly containing a 5€ note hanging out of a mailbox. Manipulation: Norm violation: Control (clean) vs. DisorderGraff (mailbox covered with graffiti) vs. DisorderLitter (litter on the ground). Adapted “lost” letter technique. Graffiti disorder increased stealing compared to control. Litter disorder increased stealing compared to control. NA Keuschnigg and Wolbring (2015) Germany 270 members of the public Task: Participants passed by a mailbox and noticed an envelope visibly containing money in front of a mailbox. Manipulation: Adapted “lost” letter technique. Disorder condition increased stealing. Disorder effect was stronger for 5€ condition and marginally significant for 10€ condition. *Benefits (but only significant in an interaction with the disorder manipulation) 97 Amount money: 5€ vs. 10€ vs. 100€. Norm violation: Control (clean) vs. Disorder (two heavily wrecked bicycles next to the mailboxes). Within 100€ condition, disorder did not affect stealing. Lanfear (2018) USA 2786 members of the public Task: Participants passed by a mailbox and noticed an envelope visibly containing a $5 note near the mailbox. Manipulation: Norm violation: Control (clean) vs. Disorder (graffiti and litter) Adapted “lost” letter technique. Norm violation did not affect stealing. Disorder condition decreased pro-social behavior (i.e., mailing the dropped envelope). NA Pruckner and Sausgruber (2013) Austria Newspaper customers (120 observations) Task: Newspaper transactions in booths on the streets via an “honor system” where costumers are supposed to make a payment without monitoring. Manipulation: Reminder: Control (“The paper costs €0.60.”) vs. Legal (“… Stealing a paper is illegal”) vs. Moral (“… Thank you for being honest”). Theft of newspapers. Legal reminder did not affected newspaper theft. Moral reminder did not affected newspaper theft. Moral reminder increased the average amount paid. NA Schlüter and Vollan (2015) Flower customers (336 observations) Task: Theft of flowers. Both reminder messages (i.e., legal NA 98 Germany In an unattended flower field, customers picked and paide for the flowers via an “honor system”. Manipulation: Reminder: Control (No message) vs. Legal (threatening message) vs. Moral (thankful message); Who is asking: Control vs. Family vs. Business (consulting firm). and moral) did not affect theft of flowers. Family business condition decreased theft of flowers. Widner (1998) USA National Park visitors (40 days of observation) Task: The behavior of visitors was directly observed. Manipulation: Anti-theft interventions: Control vs. Uniformed Volunteer vs. Sign (depicting the progressive loss of petrified wood) vs. Pledge (visitors signed an anti-theft pledge before entering the park). Theft of petrified wood. All interventions decreased theft of petrified wood, when compared to control. No differences between intervention effectiveness were found. *Costs for the self Wortley and McFarlane (2011) Australia University students (2,098 minutes of observation) Task: In a University library, students passed by an unattended photocopy card. Manipulation: Territoriality ownership: Signed (“M. Smith”) vs. Unsigned card; Theft of photocopy cards. Unsigned cards increased card theft. Card on its own increased card theft. *Costs for the other 99 Territoriality guardianship: Card next to two library books vs. Card on its own. Note . SEU = Availability of Costs or Benefits Manipulation; NA = not applicable; * = the manipulation of costs or benefits was significant. 100 Table 4 Summary of field experiments in the Keeping money category Study Participants Design Measure Main findings SEU Alem et al. (2018) Tanzania 225 farmers Participants received an amount of money on their phone. Then, received an SMS asking to return the money. Manipulation: Message frame: Control (neutral message) vs. Kindness (gift of 25%) vs. Guilt. Keeping money wrongly received. Kindness framed message reduced unethical behavior compared to control. Guilt inducing message reduced unethical behavior compared to control. NA Armantier and Boly (2011) Burkina Faso 247 adults with university degrees or enrolled at a university Task: Participants were recruited to grade a set of 20 exam papers. The 11th paper came with a bribe and a request to find few mistakes. Manipulations: Amount of bribe – No bribe vs. Low bribe vs. High bribe; Wage – Low vs. High wage; Monitoring – Low vs. High monitoring. Acceptance of bribe. High bribe amount increased acceptance of bribe. High wage decreased acceptance of bribe. Monitoring and punishment decreased acceptance of bribe. *Benefits *Costs for the self 101 Azar et al. (2013) Israel 192 customers at a restaurant Task: After paying, customers received extra change. Manipulation: Extra change: Low ($3) vs. High ($12) amount of change. Keeping extra change. High amount of extra change decreased unethical behavior. Repeated customers returned the extra change more often than one-time customers. *Benefits (but the effect is in the opposite direction) Gabor et al. (1986) Canada Cashiers at 125 convenience stores Task: A confederate bought a newspaper ($0.30) with a single dollar bill and left without awaiting the change. Manipulation: Sex of confederate: female vs. male; Type of store: chain type vs. family store. Keeping due change. Male confederates increased cashiers’ dishonesty. Type of store did not affect dishonesty. Costs for the other Gire and Williams (2007) USA Colleges’ members (80 lost bills) Task: Money (one dollar note) was left at the campuses. Manipulation Type of college: Military vs. Nonmilitary; Type of setting: Public (sidewalk) vs. Private (bathroom). “Lost” dollar bill. Within military colleges, private setting increased dishonesty. Within nonmilitary colleges, type of setting did not affected dishonesty. NA Newman (1979) 80 university students and Task: Picking a “dropped” coin. City site increased dishonesty. *Benefits 102 UK adult members of the public A female confederate dropped a coin while approaching an unsuspecting participant. Manipulation: Amount of money: Low (2p) vs. High (10p); Site: University campus vs. City (shopping area). Higher value of coin increased dishonesty. Rabinowitz et al. (1993) Austria 96 female souvenir shop cashiers. Task: Confederates purchased two postcards costing 4 shillings ($.36 USD) and left without awaiting the request or return of the money. Manipulation: Sex of confederate: female vs. male; Payment: Overpayment (+1) vs. Underpayment (-1 shilling). Keeping due change. Payment did not affect dishonesty. Female confederates increased dishonesty. *Benefits Yap et al. (2013) USA 88 members of the public Task: Participants were recruited for a study in exchange for $4. While making the payment, the experimenter “accidentally” handed $8. Manipulation: Keeping extra money. Holding an expansive pose increased dishonesty. NA 103 Postural expansiveness: Hold an expansive pose vs. Hold an contractive pose for 1 min. Yuchtman-Yaar and Rahav (1986) Israel 328 bus passengers Task: Bus drivers gave passengers extra change. Manipulation: Incentive: Low (extra change was 7% of the fare) vs. High (extra change was 25% of the fare). Keeping extra change. Within females, higher incentives increased dishonesty. Within males, higher incentives decreased dishonesty. *Benefits (but in opposite direction for males) Note . SEU = Availability of Costs or Benefits Manipulation; NA = not applicable; * = the manipulation of costs or benefits was significant. 104 Table 5 Summary of field experiments in the Shoplifting category Study Participants Design Measure Main findings SEU Carter et al. (1980) Sweden Customers of a grocery store Multiple baseline design. Intervention: Public identification: signs and red dots alerting customers for frequently shoplifted items. Shoplifting. Public identification reduced shoplifting. NA Carter and Holmberg (1993) Sweden Customers of a grocery store Pretest / Posttest design. Intervention: Public identification: signs and red dots alerting customers for frequently shoplifted items. Shoplifting. Public identification reduced shoplifting. NA Carter et al. (1988) Sweden Employees of a grocery store. Multiple baseline design. Intervention: Product identification: Oral presentation, list of target items, and data on losses graphed biweekly the lunchroom. Employee theft. Information on product identification reduced employee shoplifting. NA DiLonardo and Clarke (1996) Customers of 4 stores Pretest / Posttest design. Intervention: Replacement of EAS (i.e. Electronic Article Surveillance) with ink tags. Shoplifting Ink tags reduced shoplifting when compared to EAS condition. *Costs for the self 105 Farrington et al. (1993) UK Customers of 9 stores Pretest / Posttest design. Intervention: Anti-shoplifting intervention: Control vs. Electronic tagging vs. Store redesign vs. Uniformed guard. Shoplifting Electronic tagging reduced shoplifting, maintained over time. Store redesign reduced shoplifting, but was not maintained over time. Uniformed guard did not affect shoplifting. *Costs for the self Hayes and Blackwood (2006) USA Customers of 21 stores Pretest / Posttest design. Intervention: Source-tagged products: Control vs. 50% (half the products received a hidden EAS) vs. 100% (all products received a hidden EAS). Shoplifting. Electronic Article Surveillance did not affect shoplifting when compared to control. NA Hayes and Downs (2011) USA Customers of 47 stores Randomized Controlled Trial Intervention: Anti-shoplifting intervention: Control vs. In-aisle closed-circuit television (CCTV) public view monitor vs. Inaisle CCTV dome vs. Keeper/safer box. Shoplifting All three interventions reduced shoplifting *Costs for the self Hayes et al. (2011) Customers of 10 stores Randomized Controlled Trial Shoplifting Keeper/safer boxes reduced shoplifting *Costs for the self 112 Greenberg, J. (1990). Employee theft as a reaction to underpayment inequity: The hidden cost of pay cuts. Journal of Applied Psychology, 75 (5), 561–568. https://doi.org/10.1037/00219010.75.5.561 Greenberg, J. (2002). Who stole the money, and when? Individual and situational determinants of employee theft. Organizational Behavior and Human Decision Processes, 89 (1), 985–1003. https://doi.org/10.1016/S0749-5978(02)00039-0 Harrison, G. W., & List, J. A. (2004). Field experiments. Journal of Economic Literature, 42 (4). 10091055. https://doi.org/10.1257/0022051043004577 Hayes, R., & Blackwood, R. (2006). Evaluating the effects of EAS on product sales and loss: Results of a large-scale field experiment. Security Journal, 19 (4), 262–276. https://doi.org/10.1057/palgrave.sj.8350025 Hayes, R., & Downs, D. (2011). Controlling retail theft with CCTV domes, public view monitors and protective containers: A randomized controlled trial. Security Journal, 24 (3), 237–250. https://doi.org/10.1057/sj.2011.12 Hayes, R., Downs, D. M., & Blackwood, R. (2012). Anti-theft procedures and fixtures: A randomized controlled trial of two situational crime prevention measures. Journal of Experimental Criminology, 8 (1), 1–15. https://doi.org/10.1007/s11292-011-9137-5 Hayes, R., Johns, T., Scicchitano, M., Downs, D., & Pietrawska, B. (2011). Evaluating the effects of protective Keeper boxes on ‘hot product’ loss and sales: A randomized controlled trial. Security Journal, 24 (4), 357–369. https://doi.org/10.1057/sj.2011.2 Hayes, R., Strome, S., Johns, T., Scicchitano, M., & Downs, D. (2019). Testing the effectiveness of antitheft wraps across product types in retail environments: A randomized controlled trial. Journal of Experimental Criminology, 15 (4), 703–718. https://doi.org/10.1007/s11292-019-09365-2 Hornstein, H. A., Fisch, E., & Holmes, M. (1968). Influence of a model’s feeling about his behavior and his relevance as a comparison other on observers’ helping behavior. Journal of Personality and Social Psychology, 10 (3), 222–226. https://doi.org/10.1037/h0026568 Houser, D., List, J. A., Piovesan, M., Samek, A., & Winter, J. (2016). Dishonesty: From parents to children. European Economic Review, 82 , 242–254. https://doi.org/10.1016/j.euroecorev.2015.11.003 Jesilow, P., & O’Brien, M. J. (1980). Deterring automobile repair fraud - A field experiment (NIJ Reference Service No. 89242). National Institute of Justice. https://www.ncjrs.gov/pdffiles1/Digitization/89242NCJRS.pdf 113 Johns, T., Hayes, R., Scicchitano, M., & Grottini, K. (2017). Testing the effectiveness of two retail theft control approaches: An experimental research design. Journal of Experimental Criminology, 13 (2), 267–273. https://doi.org/10.1007/s11292-017-9284-4 Keizer, K., Lindenberg, S., & Steg, L. (2008). The spreading of disorder. Science, 322 (5908), 1681– 1685. https://doi.org/10.1126/science.1161405 Kerschbamer, R., Neururer, D., & Sutter, M. (2016). Insurance coverage of customers induces dishonesty of sellers in markets for credence goods. Proceedings of the National Academy of Sciences, 113 (27), 7454-7458. https://doi.org/10.1073/pnas.1518015113 Keuschnigg, M., & Wolbring, T. (2015). Disorder, social capital, and norm violation: Three field experiments on the broken windows thesis. Rationality and Society, 27 (1), 96–126. https://doi.org/10.1177/1043463114561749 Korbel, V. (2013). Children and cheating: A field experiment with individuals and teams [Master’s thesis, Charles University in Prague]. Thesis Repository of Charles University in Prague. https://dspace.cuni.cz/bitstream/handle/20.500.11956/58563/DPTX_2011_2_11230_0_3 55601_0_125761.pdf?sequence=1&isAllowed=y Korte, C., & Kerr, N. (1975). Response to altruistic opportunities in urban and nonurban settings. The Journal of Social Psychology, 95 (2), 183-184. https://doi.org/10.1080/00224545.1975.9918701 Lanfear, C. C. (2018). Disorder in the neighborhood: A large-scale field experiment on disorder, norm violation, and pro-social behavior [Master’s thesis, University of Washington]. Research Works Archive of the University of Washington. http://hdl.handle.net/1773/40974 Lenga, M. R., & Kleinke, C. L. (1974). Modeling, anonymity, and performance of an undesirable act. Psychological Reports, 34 (2), 501–502. https://doi.org/10.2466/pr0.1974.34.2.501 Levitt, S. D., & List, J. A. (2007). What do laboratory experiments measuring social preferences reveal about the real world?. Journal of Economic Perspectives, 21 (2), 153-174. https://doi.org/10.1257/jep.21.2.153 List, J. A. (2007). Field experiments: A bridge between lab and naturally occurring data. Advances in Economic Analysis & Policy, 5 (2), 1–47. https://doi.org/10.2202/1538-0637.1747 List, J. A., & Momeni, F. (2017). When corporate social responsibility backfires: Theory and evidence from a natural field experiment (Working paper 24169). National Bureau of Economic Research. https://www.nber.org/papers/w24169.pdf 114 List, J. A., & Momeni, F. (2020). Leveraging upfront payments to curb employee misbehavior: Evidence from a natural field experiment. European Economic Review, 130 , 103601. https://doi.org/10.1016/j.euroecorev.2020.103601 McNees, P., Gilliam, S. W., Schnelle, J. F., & Risley, T. R. (1980a). Controlling employee theft through time and product identification. Journal of Organizational Behavior Management, 2 (2), 113–119. https://doi.org/10.1300/J075v02n02_04x McNees, M. P., Kennon, M., Schnelle, J. F., Kirchner, R. E., & Thomas, M. M. (1980b). An experimental analysis of a program to reduce retail theft. American Journal of Community Psychology, 8 (3), 379–385. https://doi.org/10.1007/BF00894349 Nagin, D. S., Rebitzer, J. B., Sanders, S., & Taylor, L. J. (2002). Monitoring, motivation, and management: The determinants of opportunistic behavior in a field experiment. American Economic Review, 92 (4), 850–873. https://doi.org/10.1257/00028280260344498 Newman, C. V. (1979). Relation between altruism and dishonest profiteering from another’s misfortune. The Journal of Social Psychology, 109 (1), 43–48. https://doi.org/10.1080/00224545.1979.9933637 Okeke, E. N., & Godlonton, S. (2014). Doing wrong to do right? Social preferences and dishonest behavior. Journal of Economic Behavior & Organization, 106 , 124–139. https://doi.org/10.1016/j.jebo.2014.06.011 Olken, B. A. (2007). Monitoring corruption: Evidence from a field experiment in Indonesia. Journal of Political Economy, 115 (2), 200–249. https://doi.org/10.1086/517935 Pierce, L., & Balasubramanian, P. (2015). Behavioral field evidence on psychological and social factors in dishonesty and misconduct. Current Opinion in Psychology, 6 , 70–76. https://doi.org/10.1016/j.copsyc.2015.04.002 Pruckner, G. J., & Sausgruber, R. (2013). Honesty on the streets: A field study on newspaper purchasing. Journal of the European Economic Association, 11 (3), 661–679. https://doi.org/10.1111/jeea.12016 Rabinowitz, F. E., Colmar, C., Elgie, D., Hale, D., Niss, S., Sharp, B., & Sinclitico, J. (1993). Dishonesty, indifference, or carelessness in souvenir shop transactions. The Journal of Social Psychology, 133 (1), 73–79. https://doi.org/10.1080/00224545.1993.9712120 Ramos, J., & Torgler, B. (2012). Are academics messy? Testing the broken windows theory with a field experiment in the work environment. Review of Law & Economics, 8 (3), 563-577. https://doi.org/10.1515/1555-5879.1617 115 Robins, L. N. (1992). The role of prevention experiments in discovering causes of children’s antisocial behavior. In J. McCord & R. E. Tremblay (Eds.), Preventing antisocial behavior: Interventions from birth through adolescence (pp. 3–18). Guilford Press. Rosenbaum, S. M., Billinger, S., & Stieglitz, N. (2014). Let’s be honest: A review of experimental evidence of honesty and truth-telling. Journal of Economic Psychology, 45 , 181-196. https://doi.org/10.1016/j.joep.2014.10.002 Schlüter, A., & Vollan, B. (2015). Flowers and an honour box: Evidence on framing effects. Journal of Behavioral and Experimental Economics, 57 , 186–199. https://doi.org/10.1016/j.socec.2014.10.002 Schneider, H. S. (2012). Agency problems and reputation in expert services: Evidence from auto repair. The Journal of Industrial Economics, 60 (3), 406–433. https://doi.org/10.1111/j.14676451.2012.00485.x Shu, L. L., Mazar, N., Gino, F., Ariely, D., & Bazerman, M. H. (2012). Signing at the beginning makes ethics salient and decreases dishonest self-reports in comparison to signing at the end. Proceedings of the National Academy of Sciences, 109 (38), 15197–15200. https://doi.org/10.1073/pnas.1209746109 Siniver, E., & Yaniv, G. (2018). Losing a real-life lottery and dishonest behavior. Journal of Behavioral and Experimental Economics, 75 , 26–30. https://doi.org/10.1016/j.socec.2018.05.005 Steinberg, J., McDonald, P., & O'Neal, E. (1977). Petty theft in a naturalistic setting: The effects of bystander presence. The Journal of Social Psychology, 101 (2), 219-221. https://doi.org/10.1080/00224545.1977.9924010 Thurber, S., & Snow, M. (1980). Signs may prompt antisocial behavior. The Journal of Social Psychology, 112 (2), 309–310. https://doi.org/10.1080/00224545.1980.9924336 Tracy, P. E., & Fox, J. A. (1989). A field experiment on insurance fraud in auto body repair. Criminology, 27 (3), 589–603. https://doi.org/10.1111/j.1745-9125.1989.tb01047.x Tremblay, P., Bacher, J. L., Tremblay, M., & Cusson, M. (2000). Inflated claims of theft and tolerance threshold of insurers: Experimental analysis of situational deterrence. Canadian Journal of Criminology-Revue Canadienne de Criminologie, 42 (1), 21–38. Weisburd, D. (2003). Ethical practice and evaluation of interventions in crime and justice: The moral imperative for randomized trials. Evaluation Review, 27 (3), 336–354. https://doi.org/10.1177/0193841X03027003007 116 Weisburd, D. (2005). Hot spots policing experiments and criminal justice research: Lessons from the field. The ANNALS of the American Academy of Political and Social Science, 599 (1), 220–245. https://doi.org/10.1177/0002716205274597 Widner, C. J. (1998). Re ducing and understanding petrified wood theft at Petrified Forest National Park [Doctoral dissertation, Virginia Tech]. VTechWorks of Virginia Tech. http://hdl.handle.net/10919/40250 Wortley, R., & McFarlane, M. (2011). The role of territoriality in crime prevention: A field experiment. Security Journal, 24 (2), 149–156. https://doi.org/10.1057/sj.2009.22 Yap, A. J., Wazlawek, A. S., Lucas, B. J., Cuddy, A. J., & Carney, D. R. (2013). The ergonomics of dishonesty: The effect of incidental posture on stealing, cheating, and traffic violations. Psychological Science, 24 (11), 2281–2289. https://doi.org/10.1177/0956797613492425 Yuchtman-Yaar, E., & Rahav, G. (1986). Resisting small temptations in everyday transactions. The Journal of Social Psychology, 126 (1), 23–30. https://doi.org/10.1080/00224545.1986.9713565 Zimbardo, P. G. (1969). The human choice: Individuation, reason, and order versus deindividuation, impulse, and chaos. In W. J. Arnold & D. Levine (Eds.), Nebraska Symposium on Motivation (Vol. 17, pp. 237–307). University of Nebraska Press. Zimny, G. H. (1961). Method in experimental psychology . Ronald Press Company. https://doi.org/10.1037/14006-000 117 CHAPTER III MEASUREMENT BIAS IN SELF-REPORTS OF OFFENDING: A SYSTEMATIC REVIEW OF EXPERIMENTS Manuscript Published in: Gomes, H. S., Farrington, D. P., Maia, Â., & Krohn, M. D. (2019). Measurement bias in self-reports of offending: A systematic review of experiments. Journal of Experimental Criminology, 15 (3), 313-339. https://doi.org/10.1007/s11292-019-09379-w 118 MEASUREMENT BIAS IN SELF-REPORTS OF OFFENDING: A SYSTEMATIC REVIEW OF EXPERIMENTS Abstract Objectives. Self-reported offending is one of the primary measurement methods in criminology. In this article, we aimed to systematically review the experimental evidence regarding measurement bias in selfreports of offending. Methods. We carried out a systematic search for studies that (a) included a measure of offending, (b) compared self-reported data on offending between different methods, and (c) used an experimental design. Effect sizes were used to summarize the results. Results. The 21 pooled experiments provided evidence regarding 18 different types of measurement manipulations which were grouped into three categories, i.e., Modes of administration, Procedures of data collection, and Questionnaire design. An analysis of the effect sizes for each experimental manipulation revealed, on the one hand, that self-reports are reliable across several ways of collecting data and, on the other hand, self-reports are influenced by a wide array of biasing factors. Within these measurement biases, we found that participants’ reports of offending are influenced by modes of administration, characteristics of the interviewer, anonymity, setting, bogus pipeline, response format, and size of the questionnaire. Conclusions. This review provides evidence that allows us to better understand and improve crime measurements. However, many of the experiments presented in this review are not replicated and additional research is needed to test further aspects of how asking questions may impact participants’ answers. Keywords: Bias; Delinquency; Experiment; Measurement; Methodology; Modes of administration; Offending; Question design; Self-reports; Systematic review 119 Introduction The measurement of crime is at the heart of criminology. Every research question which includes a measurement of offending behavior is reliant on the quality of the measurement technique. Similarly, the validity of research findings is limited by the validity of the measurement itself. Traditionally, the most widely used methods of measuring crime are official records and self-reports of offending (SRO) (for reviews, see Gomes et al., 2018a; Thornberry & Krohn, 2000). Both measurements have their strengths and weaknesses, though there is evidence that self-report measures provide better estimates of the prevalence and mean frequency of delinquent behavior (e.g., Loeber et al., 2015). SRO were first introduced in an attempt to overcome the limitations of official records of crime (Nye & Short, 1957; Porterfield, 1943). Since then, self-reports have become the most widely used technique in criminal behavior research, becoming “one of the most important innovations in criminological research in the 20th century” (Thornberry & Krohn, 2000, p. 34). However, the great number of studies on the validity of SRO seen in the 1960s and 1970s (e.g., Clark & Tifft, 1966; Farrington, 1973; Hardt & Peterson-Hardt, 1977; Kulik et al., 1968; Schore et al., 1979) decreased after the publication of the influential book Measuring Delinquency (Hindelang et al., 1981), which seemed to have established the validity of SRO once and for all (Jolliffe & Farrington, 2014). Despite the scarcity of recent studies on the validity of SRO, psychological research on self-reports of sensitive behavior has increased remarkably (for reviews, see Schwarz, 1999; Tourangeau et al., 2000). Sensitive questions are commonly defined by an invasion of privacy, which may pose a threat of disclosure, and by the need for socially undesirable answers (Tourangeau & Yan, 2007). As a result, when faced with sensitive questions, participants tend to systematically underreport behaviors that are considered socially undesirable (e.g., Krumpal, 2013). While attempting to improve the measurement accuracy of sensitive questions, researchers have been developing experiments using different measurement techniques and comparing their behavioral estimates. Since participants are expected to underreport sensitive information, researchers usually apply the “more is better” hypothesis, assuming that the procedure that provides the highest prevalence is the most accurate method (Tourangeau & Yan, 2007). Despite generally accepted through the sensitive question literature, this assumption could be threatened by the possibility that some individuals may overreport some forms of deviant behavior. However, literature does seem to support that overreporting is a less prevalent problem than underreporting. Studies comparing official records and SRO (mainly arrests) show medium to high agreement between the two methods (e.g., Krohn et al., 2013; Piquero et al., 2014), although indicating 120 a higher frequency with SRO (e.g., Auty et al., 2015; Maxfield et al., 2000). Official records’ databases may be incomplete, and this may overestimate the true amount of overreporting (Daylor et al., 2019). On a slightly different note, Clark and Tifft (1966) interviewed students with and without a polygraph in order to study the validity of SRO. Findings from this study showed that participants were three times more likely to underreport deviant behavior than to overreport. Therefore, for the purposes of examining bias in self-report techniques, we focus on underreporting of offending behavior. Several aspects of data collection have been shown to minimize response bias and to improve the quality of participants’ responses to sensitive questions. For instance, evidence suggests that privacy is an important aspect of disclosure. Ong and Weiss (2000), for example, found that students’ reports of cheating in school were much higher in an anonymous condition (74%) compared to a confidential condition (25%). Similar results were obtained regarding substance use by postpartum women (Beatty et al., 2014) or undergraduate students’ reports of sexual behavior (Durant et al., 2002). Like anonymity, many other variables seem to affect participants’ willingness to report sensitive information, for example, setting effects, e.g., school vs. home (Biglan et al., 2004); bystander effects, e.g., the presence of a parent (Moskowitz, 2004); and response format, e.g., closed vs. open-ended questions (Tourangeau & Smith, 1996). One key variable that has been shown to affect participants’ responses is mode of administration. Research on mode effects is extensive and sometimes yields conflicting results. For example, while some studies found a higher prevalence of drug, cigarette, and alcohol use in self-administered modes (e.g., surveys), compared to other-administered modes (e.g., interviews) (Gribble et al., 1998, 2000), others found no significant differences in reports of alcohol use (e.g., Sobell & Sobell, 1981) or cigarette smoking (e.g., Moskowitz, 2004). Other studies even found higher reports of alcohol use in interviews compared to self-administered modes (Cutler et al., 1988; Rehm & Spuhler, 1993). Despite the apparently conflicting results, literature reviews suggest that modes of administration affect self-reports (Richman et al., 1999) and that the benefit of self-administration increases as a function of item sensitivity (Turner & Miller, 1997) and the recency of the behavior (Tourangeau & McNeeley, 2003). Unfortunately, research on sensitive questions commonly includes questions about income, voting, sexual behaviors, and drug use (Tourangeau & Yan, 2007), and only very rarely are self-reports of offensive behavior included. Kleck and Roberts (2012), for example, reviewed experiments on mode effects of self-reports of delinquent behavior and, from a total of 27 studies, only 6 included measures of offending behavior; “most findings in this area pertain to illegal drug use, and it is possible they do not apply to other kinds of criminal behavior” (Kleck & Roberts, 2012, p. 438). Considering the 121 abovementioned definition of item sensitivity, surveys of offending behavior should be considered as highly sensitive; people naturally try to conceal their offenses, which often involve feelings of guilt and shame, and participants might fear potential incriminating consequences of their reports. Therefore, while sensitive questions research should more often include items about offending behavior, knowledge derived from item sensitivity research should be considered with caution by crime researchers and results should be replicated and further explored within criminological experiments. In this article, we systematically review findings regarding potential sources of bias in collecting data on SRO. In this review, we rely only on experimental studies that compared estimates of offending from different methods of data collection, in order to gather evidence on measurement techniques, where differences are caused by the data collection method itself and the potential for confounding variables is minimized. From this systematic review of experiments, we intend to summarize the available information about the best ways of collecting SRO. Methods Search strategy In order to maximize the number of experiments included in this systematic review, the literature search was developed in four steps. In a first step, we carried out a systematic search for experiments conducted until June 2018 by entering selected keywords into 30 data bases, i.e., Scopus, EBSCOhost (Anthropology Plus, Bibliography of Asian Studies, British Education Index, Business Source Ultimate, Child Development and Adolescent Studies, Criminal Justice Abstracts, eBook Collection (EBSCOhost), Education Abstracts (H.W. Wilson), Educational Administration Abstracts, ERIC, Global Health, GreenFILE, Library, Information Science and Technology Abstracts, PsycARTICLES, PsycINFO, Russian Academy of Sciences Bibliographies, Teacher Reference Center); Elsevier (ScienceDirect); Wiley InterScience; Web of Science (Web of Science Core Collection, Current Contents Connect, Derwent Innovations Index, Korean Journal Database, Medline, Russian Science Citation Index, and Scielo Citation Index); ProQuest; Ethos. The literature search was carried out using the following keywords: (“self-report” or “selfreported” or “self-reporting” or “self-interview” or “self-interviewing” or “self-administered” or “selfadministration”) and (antisocial* or delinquen* or crim* or offend* or devian* or violen* or aggressi* or arrest* or convict*) and (bias* or missing* or nonrespons* or “under-report” or “over-report” or underreport* or overreport*) and (experiment*). 128 Hindelang et al. (1981) USA 13,842 adolescents Mode of administration and Anonymity Lifetime prevalence Year prevalence Delinquent behavior (69 items grouped into Contacts with the criminal justice system; Serious crimes; General delinquency; Drug offenses; and School and family). Horney & Marshall (1992) USA 700 convicted male offenders Questionnaire design Year prevalence Magnitude of self-reported offending frequency, i.e. Lambda (Burglary, robbery, theft, auto-theft, forgery, fraud, assault, and drug deals). King et al. (2012) USA 245 adolescents patients In-person followup Year prevalence Risk factors for suicidal behavior (including aggressive/delinquent behavior) Kivivuori et al. (2013) Finland 924 adolescent students Supervision Lifetime prevalence Year prevalence Delinquent behavior (graffiti drawing, vandalism at school, vandalism elsewhere, shoplifting, stealing at school, motor vehicle theft, other theft, breaking and entering, fighting, beating up someone, robbery, drunken driving, illegal downloading) Knapp and Kirk (2003) USA 352 undergraduate students Mode of administration Lifetime prevalence Sensitive questions (Have you ever written on a restroom wall?, Have you ever used someone else’s credit card (number) without their permission?, and Have you ever been in jail?) Krohn et al. (1974) USA 321 undergraduate students Mode of administration and Interviewer Year prevalence Delinquent behavior (Drunken driving, Fighting, Petty theft, Grand larceny, Property damage, and Illegal entry) Lucia et al. (2007) Switzerland 1,203 adolescent students Mode of administration and Reference period Lifetime prevalence Year prevalence Delinquent behavior (Driving without license, Shoplifting [more than €35], Shoplifting [less than €35], Breaking into a car, Harassing somebody in the street, Theft at school, Theft at home, Fare dodging, Vehicle theft, Theft of an object from a vehicle, Assault, Threats with gun/knife, Racket [extortion], Robbery, Arson, Selling soft drugs, Selling hard drugs, Graffiti, Vandalism, Theft from the person) Potdar and Koenig (2005) India 900 male undergraduate students and 600 Mode of administration Lifetime prevalence Risk behavior (Carrying a weapon/gun and Engaged in abusive, violent behavior after drinking) 129 male residents in slums. Strang and Peterson (2020) USA 93 young, community men Bogus Pipeline Lifetime prevalence Sexual aggression (Verbal coercion tactics, Drugs and alcohol tactics, and Force tactics of sexual assault) Trapl et al. (2013) USA 275 adolescent students Mode of administration Lifetime prevalence Sensitive behaviors (Shoplifting) Turner et al. (1998) USA 1,672 adolescent males Mode of administration Year prevalence 30-day prevalence Risk behavior (Threatened to hurt someone; Carried a gun; In physical fight; Pulled knife or gun on someone; and Carried a knife or razor) van de LooijJansen et al. (2006) Netherlands 704 adolescent students Anonymity Lifetime prevalence Health indicators (Aggressive behavior; Vandalism and stealing; Violent delinquent behavior; and Carrying a weapon) van de Looij‐ Jansen and de Wilde (2008) Netherlands 532 adolescent students Mode of administration Year prevalence Health indicators (Aggressive behavior; Vandalism and stealing; and Carrying a weapon) Walser and Killias (2012) Switzerland 1,197 adolescent students Supervision Lifetime prevalence Year prevalence Delinquent behavior (Assault; Group fight; Robbery; Sexual assault; Burglary; Shoplifting; Bicycle theft; Other theft; Vandalism; Carrying a weapon; Drug dealing; and Any delinquency) 130 Table 7 summarizes the results of these experiments, organized by measurement manipulations. For each manipulation, we provided information regarding the OR effect sizes (i.e., OR, 95% confidence intervals, Z statistics, and p value). Because most comparisons are made with few cases, additionally to ORs, we also reported the number of statistically significant differences found in individual item comparisons (when available). Since this review includes results from several different manipulations, Table 7 provides information on the experimental manipulation under analysis (experimental condition A vs. experimental condition B). Considering the calculation of OR effect sizes to be the odds of reporting offending behavior in condition A divided by the odds of reporting offending in condition B, an OR > 1 indicates higher reports in condition A, while an OR < 1 indicates higher reports in condition B, and an OR = 1 indicates a null effect. For example, in the first line of Table 7, we present the comparison of personal interview (i.e., condition A) vs. self-administered questionnaire (i.e., condition B) (Krohn et al., 1974); an OR = 0.70 indicates that the odds of reporting deviant behavior in the interview (i.e., condition A) were decreased by 30% relative to the questionnaire (i.e., condition B). Table 7 Main findings of experiments in the systematic review Study Comparison ( p < .05) OR 95% CI z p Modes of administration Personal Interview (PI) vs. Self-Administered Questionnaire (SAQ) (k = 3) Krohn et al. (1974) 0 of 6 0.70 [0.34, 1.45] -0.96 .336 Hindelang et al. (1981) - 0.97 [0.92, 1.04] -0.90 .398 Potdar and Koenig (2005) 0 of 2 0.83 [0.36, 1.91] -0.45 .656 Random model 0.97 [0.92, 1.03] -0.95 .341 Personal Interview (PI) vs. Audio Computer-Assisted Self-Interview (ACASI) (k = 1) Potdar and Koenig (2005) > PI (1 of 4) 1.23 [0.84, 1.80] 1.05 .293 Self-Administered Questionnaire (SAQ) vs. Computer-Assisted Self-Interview (CASI) (k = 10) Beebe et al. (1998) >SAQ (2 of 5) 1.42 [0.91, 2.20] 1.55 .122 Knapp and Kirk (2003) 0 of 3 1.11 [0.59, 2.09] 0.33 .742 Beebe et al. (2006) 0 of 2 1.06 [0.65, 1.71] 0.22 .823 Brener et al. (2006) >CASI (2 of 5) 0.84 [0.70, 0.99] -2.06 .040 Hamby et al. (2006) > CASI (1 of 4) > SAQ (1 of 4) 0.93 [0.49, 1.77] -0.21 .835 Lucia et al. (2007) > CASI (2 of 40) > SAQ (5 of 40) 1.12 [0.80, 1.56] 0.66 .507 131 van de Looij‐Jansen and de Wilde (2008) > CASI (1 of 3) 0.81 [0.59, 1.10] -1.36 .174 Eaton et al. (2010) > CASI (5 of 7) 0.90 [0.77, 1.04] -1.42 .157 Trapl et al. (2013) 0 of 1 1.10 [0.58, 2.10] 0.30 .764 Baier (2017) > SAQ (1 of 10) 0.94 [0.71, 1.24] -0.47 .635 Random Model 0.92 [0.84, 1.01] -1.85 .064 Self-Administered Questionnaire (SAQ) vs. Audio Computer-Assisted Self-Interview (ACASI) (k = 3) Turner et al. (1998) > ACASI (4 of 5) 0.69 [0.51, 0.93] -2.42 .015 Potdar and Koenig (2005) - 1.05 [0.47, 2.36] 0.12 .902 Trapl et al. (2013) 0 of 1 1.14 [0.60, 2.19] 0.41 .685 Random Model 0.82 [0.59, 1.14] -1.20 .232 Computer-Assisted Self-Interview (CASI) vs. Audio Computer-Assisted Self-Interview (ACASI) (k = 1) Trapl et al. (2013) 0 of 1 1.04 [0.54, 1.99] 0.11 .914 Self-Administered Questionnaire (SAQ) vs. Automated Touch-Tone Telephone (TACASI) (k = 1) Knapp and Kirk (2003) 0 of 3 1.18 [0.72, 1.93] 0.66 .510 Computer-Assisted Self-Interview (CASI) vs. Telephone Audio Computer-Assisted Self-Interview (TACASI) (k = 1) Knapp and Kirk (2003) 0 of 3 1.06 [0.54, 2.07] 0.17 .863 Procedures of Data Collection Supervision by teachers vs. Supervision by researchers (k = 2) Walser and Killias (2012) > teacher (2 of 22) 1.04 [0.83, 1.31] 0.35 .726 Kivivuori et al. (2013) > research (2 of 26) 0.87 [0.58, 1.31] -0.66 .508 Random Model 1.00 [0.82, 1.22] -0.02 .981 Non-anonymous vs. Anonymous (k = 2) Hindelang et al. (1981) - 0.98 [0.92, 1.04] -0.71 .481 van de Looij-Jansen et al. (2006) > Anonym. (3 of 4) 0.67 [0.51, 0.88] -2.89 .004 Random Model 0.83 [0.58, 1.20] 1.00 .319 No-Disclosure vs. Disclosure (k = 1) Beebe et al. (2006) > No Discl. (1 of 2) 1.69 [0.99, 2.88] 1.93 .053 Home setting vs. School setting (k = 1) Brener et al. (2006) > school (5 of 5) 0.75 [0.63, 0.89] -3.27 .001 ‘Conservative’ interviewer vs. ‘Hip’ interviewer (k = 1) Krohn et al. (1974) > ‘Hip’ (2 of 6) 0.54 [0.27, 1.08] -1.75 .080 No in-person follow-up vs. In-person follow-up (k = 1) King et al. (2012) 0 of 1 0.62 [0.36, 1.06] -1.76 .079 Bogus pipeline (BPL) vs. Control group (k = 1) 132 Modes of administration In the first category, we included all the experimental manipulations regarding the methods through which participants provide their answers to the offending questions. In this review, experiments considered the following: (a) personal interviews (PI), where questions are delivered in face-to-face interviews and answers are provided orally to an interviewer; (b) self-administered questionnaires (SAQ), where participants are given a paper-and-pencil questionnaire which they complete on their own; (c) computer-assisted self-interviews (CASI), where participants are given a questionnaire on a computer screen which they complete on their own directly onto a computer; (d) audio computer-assisted selfinterview (ACASI), where questionnaires are presented on a computer screen and participants can listen to audio records of the questions and provide their answers directly onto the computer; and (e) telephone audio computer-assisted self-interview (TACASI), where participants are contacted via telephone, listen to audio records of the questions, and provide their answers on the telephone which are recorded via automated software. PI vs. SAQ Three studies compared results of SRO collected under PI and SAQ (Hindelang et al., 1981; Krohn et al., 1974; Potdar & Koenig, 2005). The pooled effect sizes presented virtually null ORs, slightly in favor of SAQ but with no statistical significance. The overall analysis under a random model suggested Strang and Peterson (2020) > BPL (2 of 8) 2.18 [0.82, 5.81] 1.55 .121 Questionnaire design Response Format: 2-options vs. 7-options (k = 1) Hamby et al. (2006) > 7-options (2 of 4) 1.19 [0.63, 2.25] 0.52 .602 Long vs. Short questionnaire (k = 1) Enzmann (2013) > Short (5 of 24) 0.89 [0.74, 1.06] -1.31 .192 Standard vs. Month-by-month reporting (k = 1) Horney and Marshall (1992) 0 of 8 0.98 [0.69, 1.39] -0.13 .900 Reference Period: “12 months” vs. “Since October 2003” (k = 1) Lucia et al. (2007) 0 of 20 1.04 [0.62, 1.74] 0.15 .878 Note . The “Comparison ( p < .05)” column shows the number of statistically significant differences found in individual item comparisons (when available). > = higher estimates, e.g. “> PI (1 of 2)” = 1 of 2 item comparisons presented significantly higher estimates of self-reported offending in the Personal Interview. 133 no significant differences between data collected with these two methods (OR = 0.97, 95% CI [0.92, 1.03], z = −0.95, p = .341). PI vs. ACASI Only one study compared PI and ACASI (Potdar & Koenig, 2005). Results mainly favored PI (OR = 1.23, 95% CI [0.84, 1.80], z = 1.05, p = .293), though it did not reach statistical significance ( p > .05). SAQ vs. CASI The analysis of SAQ vs. CASI was the most replicated comparison in the present review, with 10 studies (Baier, 2017; Beebe et al., 1998, 2006; Brener et al., 2006; Eaton et al., 2010; Hamby et al., 2006; Knapp & Kirk, 2003; Lucia et al., 2007; Trapl et al., 2013; van de Looij-Jansen & de Wilde, 2008). An analysis of the individual effect sizes showed that 5 comparisons favored CASI, though only one reached statistical significance with an OR of 0.84 (Brener et al., 2006), while of the 5 comparisons favoring SAQ none reached statistical significance. On average, the mean effect slightly favored CASI over SAQ (OR = 0.92, 95% CI [0.84, 1.01], z = −1.85, p = .064), though with only marginal significance ( p < .10). SAQ vs. ACASI Three studies provided comparisons of offending behavior collected with SAQ or ACASI (Potdar & Koenig, 2005; Trapl et al., 2013; Turner et al., 1998). One out of the three ORs presented statistically significant results in favor of the ACASI mode (OR = 0.69, p = .015). Considering random effects, the average effect size showed an OR = 0.82 favoring ACASI but with no statistical significance (OR = .82, 95% CI [0.59, .136], z = −1.20, p = .232). CASI vs. ACASI Trapl et al. (2013) conducted the sole experiment comparing SRO obtained through CASI and ACASI. Despite participants reporting slightly higher estimates of lifetime shoplifting under the CASI mode of data collection, results of this experiment showed a nonsignificant OR effect size (OR = 1.04, 95% CI [0.54, 1.99], z = 0.11, p = .914). SAQ vs. TACASI 134 Knapp and Kirk (2003) carried out the unique experiment that compared SAQ and TACASI. Results showed slightly higher estimates of offending in the SAQ mode of administration, though with no statistical significance (OR = 1.18, 95% CI [0.72, 1.93], z = 0.66, p = .510). CASI vs. TACASI Similar to the previous results, the experimental comparison between CASI and TACASI (Knapp & Kirk, 2003) showed a nonsignificant effect size (OR = 1.06, 95% CI [0.54, 2.07], z = 0.173, p = .863). Procedures of data collection The second category of manipulations takes into account different procedures applied in the data collection that might influence the participants’ SRO. This category accounts for seven out of the total 18 manipulations, which included manipulations in Supervision of data collection ( k = 2), Anonymity ( k = 2), Characteristics of the Interviewer ( k = 1), Setting of data collection ( k = 1), Disclosure of information ( k = 1), In-person follow-up ( k = 1), and Bogus pipeline ( k = 1). Supervision Two studies compared supervision by the participants’ teacher with supervision by the researchers during the completion of the questionnaire with CASI methodology (Kivivuori et al., 2013; Walser & Killias, 2012). In general, results showed slightly higher estimates in the condition where participants were supervised by researchers, though not reaching statistical significance. On average, random effects showed no statistically significant differences between the two methods (OR = 1.00, 95% CI [0.82, 1.22], z = − 0.02, p = .981). Anonymity From the pooled experiments, two studies focused on the issue of anonymity in SRO. Hindelang et al. (1981) used both anonymous/non-anonymous questionnaires and anonymous/non-anonymous interviews (where contact between interviewer and interviewee was prevented by a screen). Results showed no statistically significant differences, with an OR of 0.98 ( p = .481). In the experiment of van de Looij-Jansen et al. (2006), participants received questionnaires with their names on them (i.e., confidential group) vs. questionnaires with no identifying information (i.e., anonymous condition). In this case, results showed higher SRO in the anonymous condition (OR = 0.67, p = .004). The average effect 135 size favored anonymous procedures, showing a reduced odds by 17% of reporting offending behavior in the non-anonymous condition, though with no statistically significant effects (OR = 0.83, 95% CI [0.58, 1.20], z = − 1.00, p = .319). Disclosure Beebe et al. (2006) conducted an experiment studying the effect of disclosure of self-reported information. This experiment compared results of two groups. In one group, participants were told that their responses would only be seen by the researchers and in a second group, participants were told that a summary report would be given to their health care provider. Findings showed an increased odds by 69% of reporting offending behavior in the no-disclosure condition, though statistical significance reached only a marginal level (OR = 1.69, 95% CI [0.99, 2.88], z = 1.93, p = .053). Setting Brener et al. (2006) developed an experiment to test differences between data collection at home vs. data collection at school. Results considerably favored data collection at schools, with a reduced odds of reporting offending behavior by 25% in a home setting (OR = 0.75, 95% CI [0.63, 0.89], z = −3.27, p = .001). Characteristics of the interviewer Krohn et al. (1974) carried out an experiment to test the hypothesis that the characteristics of the interviewer might influence the reports of offending. The two experimental conditions included interviewers with a conservative appearance, dressed formally and closely trimmed hair (i.e., “conservative” interviewers) vs. a group of interviewers casually dressed and with long hair (i.e., “hip interviewers”). Findings showed that the odds of reporting delinquent behavior decreased by 46% with the “conservative” interviewer, though the statistical test revealed to be only marginally significant, i.e., p < .10 (OR = 0.54, 95% CI [0.27, 1.08], z = −1.75, p = .080). In-person follow-up King et al. (2012) conducted the unique experiment comparing self-reports of aggressive/delinquent behavior of adolescent patients seeking medical emergency services who were randomly allocated to two groups. The control group had no in-person follow-up, but in the experimental group, participants were told about a subsequent session of in-person follow-up where they would receive 136 feedback on their answers. Results from this experiment showed a decreased odds of self-reports by approximately 38% in the control group (i.e., no in-person follow-up), and once again, z statistics showed only marginally significance at a level of p < .10 (OR = 0.62, 95% CI [0.36, 1.06], z = −1.76, p = .079). Bogus pipeline Finally, Strang and Peterson (2020) carried out an experiment to test the effects of a bogus pipeline in reporting sexual aggressive behavior. In the control group, participants were attached to a physiological measurement device and were told that it was to “determine the level of anxiety prior to starting the questionnaire.” In the bogus pipeline group, participants were attached to the same physiological measurement device and were told it was “similar to a polygraph or lie detector test” and “that the machine was being attached to encourage honest responding.” Overall, despite non-significant results from z statistics, findings showed an increased odds ratio of 2.18 of reporting sexual aggression (including verbal coercion, use of drugs and alcohol tactics, and force) in the bogus pipeline condition (OR = 2.18, 95% CI [0.82, 5.81], z = 1.55, p = .121). Moreover, individual item comparisons revealed that men in the bogus pipeline condition showed 6.5 times greater odds of reporting illegal sexual assault (OR = 6.49, 95% CI [1.78, 23.69], z = 2.83, p < .01). Questionnaire design In the third category, we grouped the experimental manipulations of the design of the questionnaire itself. This category accounts for four out of the total 18 manipulations, which included manipulations in response format ( k = 1), response format and follow-up questions ( k = 1), Month-bymonth reporting ( k = 1), and reference periods ( k = 1). Response format One study focused on the response format (Hamby et al., 2006). In this experiment, self-reports of partner violence perpetration were given in two different formats: (a) a dichotomous response format (i.e., yes and no) and (b) a 7-category response format (i.e., once, twice, 3 to 5 times, 6 to 10 times, 11 to 20 times, more than 20 times, and never). The average effect size showed nonsignificant effects of the response manipulation (OR = 1.19, 95% CI [0.63, 2.25], z = 0.52, p = .602). However, results varied considerably according to the types of crimes. Self-reports of psychological aggression (OR = 0.80, 95% CI [0.31, 2.04]) and physical assault (OR = 0.77, 95% CI [0.41, 1.46]) were slightly higher in the 137 dichotomous condition, but not statistically significant (p > .05). For self-reports of sexual coercion (OR = 3.58, 95% CI [1.34, 9.58]) and injury (OR = 3.35, 95% CI [1.03, 10.89]), results were significantly higher in the 7-option response condition ( p < .05). Response format and follow-up questions Enzmann (2013) developed a cross-sectional experiment testing a shorter version of the ISRD-2 questionnaire. The two experimental conditions were as follows: (a) a standard ISRD-2 questionnaire (i.e., long version), with five follow-up questions for each offending item, and a no-yes response pattern; (b) a short version of the ISRD-2 questionnaire, with only one follow-up question, and a yes-no response pattern. The effect size showed a slight decrease in chances of reporting delinquent activity in the long version by 11%, though without statistical significance (OR = 0.89, 95% CI [0.74, 1.06], z = −1.31, p = .192). However, individual item comparison showed statistically significant higher reports in the short version in 5 out of 24 comparisons. Standard vs. Month-by-month reporting Horney and Marshall (1992) carried out an experiment comparing standard interviewing methods in the RAND Second Inmate Survey (Chaiken & Chaiken, 1982) and a Month-by-month reporting interview to measure Lambda (i.e., individual offending frequency). Results showed little difference between the two methods (OR = 0.98, 95% CI [0.69, 1.39], z = −0.13, p = .900). Reference period Finally, Lucia et al. (2007) conducted the only experiment found in the present systematic review that attempted to study the potential effects of different instructions regarding the recall period. In this experiment, authors manipulated the instructions about the reference period: (a) “During the last 12 months,” and (b) “Since the school vacation of October 2003” (which corresponded to a 12-month period). The results showed similar estimates of delinquent behavior in both conditions (OR = 1.04, 95% CI [0.62, 1.74], z = 0.15, p = .878). Discussion Despite the wide use of the self-report methods in criminology, many researchers have shared their concerns about the quality of this methodology and how several contextual features may impact