Preregistration of experiment "Response Priming with metacontrast-masked traffic sign stimuli" (student project)
Abstract
Preregistration for the analysis of the behavioral data in a student project (BA theses) on response priming by metacontrast-masked traffic sign stimuli, with Thomas Schmidt as the supervisor.
Full text
1 University of Kaiserslautern-Landau (RPTU) Center for Cognitive Science, Visual Attention and Awareness Laboratory Head: Prof. Dr. Thomas Schmidt, Dipl.-Psych., FpsyS Preregistration on zenodo.org Date: Monday, December 15, 2025 Authors: Thomas Schmidt, Nils Thomé, Hannah Kappes, Elena Metz, Anne Nußbaum, Elias Pätz Preregistration: Masked perception of Ampelmännchen stimuli General approach We report how we determined our sample size, all data exclusions (if any), all manipulations, and all measures in the study (“21 word solution”). We do not report conditions, variables, analyses, or participants selectively without transparently revealing the reasoning behind the selection. We avoid the use of analysis techniques with excessive and unwarranted user-degrees of freedom. In our style of experimentation, we subscribe to the psychophysical tradition of “small-N / large-Ni designs” (Smith & Little, 2018) that relies on the careful analysis of individual observers instead of extensive averaging across many observers. We therefore prefer to draw statistical power from a large number of trial repetitions rather than the number of participants (Arend & Schäfer, 2019; Baker et al., 2020). We follow the principle that research articles must be methodologically transparent without requiring additional knowledge of supplementary materials or preregistration documents. We share data with natural persons for scientific purposes in accordance with German data protection laws (Landesdatenschutzgesetz Rheinland-Pfalz).
2 Novel or preexisting data? No data have been collected for this study yet. Purpose of the experiment, major hypotheses and expectations We measure response priming effects in response times and error rates. On each trial, participants are presented with a brief prime display (an Ampelmännchen 1 shown either in red or green) followed by a target display (another Ampelmännchen). Trials are congruent if prime and target afford the same response (e.g., both red, or both showing the same action such as standing vs. walking), and incongruent otherwise. Participants are instructed to respond as quickly and accurately as possible to the target by pressing one of two response keys. In one task (Target ID), participants are asked to discriminate whether the target Ampelmännchen is standing or walking, again responding as quickly and accurately as possible. In a second task (Prime ID), they are asked to identify the prime Ampelmännchen as accurately as possible without time pressure. The prime–target SOA is constant by 10 ms. All stimuli are composed of hexagonal elements that are masked by a hexagonal grid of lines („honeycombs“). In this way, we aim to achieve effective masking of the Ampelmännchen primes. The gray scale of the hexagonal elements is varied in three steps in order to modulate the degree of masking. The color of the hexagonal masking grid remains constant. The prime target SOA is also constant at 100 ms. We measure full priming functions in which response times and error rates are plotted as a function of prime–target congruency (incongruent vs. congruent conditions) and hexagon color. We also measure full masking functions of prime discrimination accuracy as a function of hexagon color. Additionally, we examine response times for effects 1 A distinctive German pedestrian-signal figure displayed in traffic lights, shown either standing or walking.
3 between compatible color mappings (red: standing, green: walking) and the reverse, incompatible mapping. Our major question is whether response priming is modulated by using masked Ampelmännchen displays. We aim to demonstrate that response priming effects remain robust, even when masking strength differs substantially across conditions. The central hypothesis is that previously presented Ampelmännchen systematically bias the processing of subsequently presented ones, revealing how early masked visual information influences later perceptual decisions. Design, independent variables, and dependent variables; variables that are not part of the analysis plan We apply a completely crossed two-dimensional repeated-measures design with factors of gray scales (3 levels) and prime–target congruency (2 levels). Each participant completes every condition across two experimental sessions. All analyses follow this same factorial structure. The dependent variables are mean response times and error rates. In addition to analyzing mean error rates, we apply hazard analysis and conditional accuracy functions, in which physical time is subdivided into time bins and both the hazard rate of responding and the accuracy of responses are plotted as a function of time. This applies to both the color discrimination task and the action discrimination task (standing vs. walking). No other dependent variables are included. Apart from the major independent variables, there are additional variables required for counterbalancing and the avoidance of experimental artifacts. These include the specific identity of the prime Ampelmännchen and the target Ampelmännchen (e.g., particular figure rendering or position within the grid), as well as the specific combination of color and action on each trial. These balancing variables are not part of the analysis plan and are averaged over.
4 The same holds for all participant variables, such as gender, age, handedness, or other individual characteristics, which are not analyzed separately and are averaged across all analyses. Criteria for data trimming, discarding of data or participants, or selective analyses We rarely have to exclude participants from analysis after data are collected. We only do so if a participant clearly does not follow instructions. Indicators of that are very high error rates (close to chance level) on the target identification task or monotonic use of the same response key. Apart from that, sometimes participants’ data cannot be used because of equipment malfunction, mistaken instructions, or other unforeseen events. In such cases, we try to recover as much of the valid data as we can. However, when a participant decides to abort an experiment prematurely we discard all data from that person. Finally, we exclude trials with problems in experimental timing, such as skipped monitor frames as measured by the timing feedback from the Psychophysics Toolbox. In experiments involving color discrimination, we screen participants for color deficiencies by using Ishihara plates. A frequently used strategy for data analysis in unconscious perception research is to selectively analyze trials or participants with certain performance outcomes (e.g., trials with low visibility ratings or participants not exceeding certain prime discrimination criteria). We are strongly opposed to such practices, which lead to strong statistical distortions and artifacts like regression to the mean, and we never use them. Planned analyses, variable transformations The major analysis strictly follows the three-factorial structure of the independent variables. We apply three-factorial repeated-measures analysis of variances (ANOVA; factors Congruency, SOA) on response times as well as error rates. Practice blocks are discarded. Error rates p are transformed into logits according to the formula logit = log
5 [p/(1-p)], after replacing perfect error rates of 0 and 1 with values of 1/2r and 1 - 1/2r, where r is the number of stimulus repetitions per participant. In other words, half of a response is added to or subtracted from any perfect score to prevent division by zero. In fully crossed repeated-measures designs, we adjust standard error bars for intersubject variability (Loftus & Masson, 1994). Because the statistical models for repeated measures only use the interaction between the participant factor and the effect of interest as an error term, the main effect of the participant factor does not enter the error variance and can be discarded. Error bars should reflect the resulting increase in power to aid the graphical interpretation of the data. We use a very simple adjustment method (Bakeman & McArthur, 1996) where each participant’s data pattern is vertically shifted by an additive constant until all individual means are equal to the grand mean (ipsative data). This way, the intersubject variance is zero while all interactions remain intact. Then standard errors are calculated across the adjusted participants as usual. Statistical decision criteria and multiple tests We generally report statistical tests as conventionally significant at a false-positive risk of α = .05 because many readers use that criterion. However, internally we are more conservative and regard p-values between .01 and .05 as only mild evidence against the null hypothesis and would not base strong conclusions on such results. GreenhouseGeisser correction is used for all tests irrespective of the outcome of a formal sphericity test. Corrections for type-I error accumulation are handled in the following way. We follow the custom of not correcting the set of F tests within a single ANOVA model for multiple tests (usually three main effects and four interactions). We also follow the custom of not correcting F tests across the ANOVAS of different dependent variables. However, when
6 we conduct k multiple tests within the frame of an analysis (e.g., analyzing priming effects separately for each SOA), we use a Bonferroni correction of α’ = 1 - (1 - α)k. Foreseeable follow-up analyses In many experiments, discoveries in the data lead to follow-up analyses. Furthermore, reviewers often request additional analyses. Examples would be learning effects (timecourses of effects across blocks or sessions), sequential effects (e.g., response times following error trials), or delta plots. We always clearly specify which analyses are planned (follow from the design) and which ones are post-hoc (based on discoveries). Not every aspect of data analysis can be preplanned. For instance, it is always an executive decision to choose a good bin size for a histogram, a color code for a heat-map, or the minimum number of trials that allow for plotting a data point in a conditionalaccuracy or hazard diagram. Unexpected problems may lead to unbalanced designs or unusual distributions of data, forcing analyzers to adjust statistical methods. Advances in methodological knowledge and creative development of methods may even lead to the replacement of one technique by a superior one. We believe that this is generally a sign of scientific progress, not of researcher degrees of freedom running wild. We are transparent about such issues and developments in our own analysis strategies and clearly identify them in our publications. Sample size determination In multi-factor repeated-measures designs, statistical power can be calculated if all effect sizes can be predicted along with their respective error variances. In practice, however, too many terms are unknown for a meaningful power analysis. Because the number of trials per participant and condition is about as important for power as the number of participants (Arend & Schäfer, 2019; Baker et al., 2020; Smith & Little, 2018), we control measurement precision at the level of individual participants in single tasks and stimulus
7 conditions (Biafora & Schmidt, 2019; Lakens, 2024). For each task, we calculate precision as s/√r (Eisenhart, 1962), where s is a single participant's standard deviation in a given cell of the basic design (e.g., Consistency x SOA) and r is the number of repeated measures in each cell and subject. We assume standard deviations of SD = 60 ms for response times (based on our own benchmark data) and the maximal possible standard deviation of .5 for error and accuracy rates. For instance, when r = 100, we can expect a precision of 60/√100 = 6 ms in individual response times and 0.5/√100 = 0.05 (five percentage points) in error/accuracy rates. This means that an uncorrected 95% confidence interval calculated within a single observer would be able to resolve mean differences of two precision units, ≈ 2s/√r. In the Target ID experiment, r = 128, and we assume SD(RT) = 60 ms, max. SD(Error) = 0.5. Therefore, we expect a precision of 5,3 ms in response times and 4,4 percentage points in accuracy scores of each individual observer and condition. Precision thus exceeds our previous recommendations for response priming studies (r = 60, F. Schmidt et al., 2011). In the Prime ID experiment, r = 256, and we assume SD(RT) = 60 ms, max. SD(Error) = 0.5. Therefore, we expect a precision of 3,8 ms in response times and 3,1 percentage points in accuracy scores of each individual observer and condition. Precision thus exceeds our previous recommendations for response priming studies (r = 60, F. Schmidt et al., 2011). Based on these considerations, we measure 10 observers in two sessions. One Session is in the Target ID experimental condition with r = 128 trials per condition and observer. The other session is in the Prime ID experimental condition with r = 256 trials per condition and observer. In our publications, we show the data patterns of all individual observers.
8 Literature: Arend, M. G., & Schäfer, T. (2019). Statistical power in two-level models: A tutorial based on Monte Carlo simulation. Psychological Methods, 24(1), 1–19. https://doi.org/10.1037/met0000195 Bakeman, R., & McArthur, D. (1996). Picturing repeated measures: Comments on Loftus, Morrison, and others. Behavior Research Methods, Instruments, & Computers, 28(4), 584-579. Baker, D. H., Vilidaite, G., Lygo, F. A., Smith, A. K., Flack, T. R., Gouws, A. D., & Andrews, T. J. (2020). Power contours: Optimising sample size and precision in experimental psychology and human neuroscience. http://dx.doi.org/10.1037/met0000337 Biafora, M., & Schmidt, T. (2019). Induced dissociations: Opposite time courses of priming and masking induced by custom-made mask-contrast functions. Attention, Perception, & Psychophysics, 82, 1333–1354. https://doi.org/10.3758/s13414-019-01822-4 Cousineau, D. (2005). Confidence intervals in within-subject designs: A simpler solution to Loftus and Masson’s method. Tutorials in Quantitative Methods for Psychology, 1, 42–45. https://doi.org/10.20982/tqmp.01.1.p042 Eisenhart, C. (1962). Realistic evaluation of the precision and accuracy of instrument calibration systems. In H. H. Ku (Ed.), Precision Measurement and Calibration (1969), (pp. 21–48). Washington, D.C.: National Bureau of Standards. Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1). https://doi.org/10.1525/collabra.33267 Loftus, G. R., & Masson, M. E. J. (1994). Using confidence intervals in within-subject designs. Psychonomic Bulletin & Review, 1, 476-490. Panis, S., Schmidt, F., Wolkersdorfer, M. P., & Schmidt, T. (2020). Analyzing response times and other types of time-to-event data using event history analysis: A tool for mental chronometry and cognitive psychophysiology. I-Perception, 11, 1-24. doi: 10.1177/2041669520978673 Schmidt, F., Haberkamp, A., & Schmidt, T. (2011). Dos and don’ts in response priming research. Advances in Cognitive Psychology, 7, 120–131. https://doi.org/10.2478/v10053-008-0092-2 Schmidt, T. (2000). Visual perception without awareness: Priming responses by color. In T. Metzinger (Ed.), Neural correlates of consciousness (pp. 157-179). Cambridge: MIT Press. https://doi.org/10.7551/mitpress/4928.003.0014 Schmidt, T. (2002). The finger in flight: Real-time motor control by visually masked color stimuli. Psychological Science, 13, 112-118. https://doi.org/10.1111/14679280.00421
9 Schmidt, T., Niehaus, S., & Nagel, A. (2006). Primes and targets in rapid chases: Tracing sequential waves of motor activation. Behavioral Neuroscience, 120(5), 1005–1016. https://doi.org/10.1037/0735-7044.120.5.1005 Schmidt, T., & Seydell, A. (2008). Visual attention amplifies response priming of pointing movements to color targets. Perception & Psychophysics, 70, 443-455. doi: 10.3758/PP.70.3.443 Smith, P. L., & Little, D. R. (2018). Small is beautiful: In defense of the small-N design. Psychonomic Bulletin & Review, 25(6), 2083–2101. https://doi.org/10.3758/s13423-018-1451-8 Vorberg, D., Mattler, U., Heinecke, A., Schmidt, T., & Schwarzbach, J. (2003). Different time courses for visual perception and action priming. Proceedings of the National Academy of Sciences USA, 100(10), 6275–6280. https://doi.org/10.1073/pnas.0931489100