Full text
ACTION-ORIENTED VISION Ashima Keshava
Action-Oriented Vision Investigating the Role of Body Task and Environment in Shaping Oculomotor Control Dissertation zur Erlangung des Grades Doktor der Naturwissenschaften (Dr. rer. nat.) im Fachbereich Humanwissenschaften der Universität Osnabrück vorgelegt von Ashima Keshava 19 May 2025
©May 2025 - Ashima Keshava Cover: Modified from The Body Has a Mind by Daniel Friedman. https://www.flickr.com/photos/daniel_friedman/18776984929/ Licensed under the Creative Commons license.
Thesis Supervisors 1st:Prof. Dr. med. Peter König Universität Osnabrück, Germany 2nd:Prof. Dr. rer. nat. Gordon Pipa Universität Osnabrück, Germany iv
List of Publications 2024 •Keshava, A., Wächter, M. A., Boße, F., Schüler, T., & König, P. (n.d.). Low Dimensional Representations of Visuomotor Coordination in Natural Behavior In bioRxiv. (Under review at Journal of Neurophysiology) •Keshava, A., Nezami, F. N., Neumann, H., Izdebski, K., Schüler, T., & König, P. (2024). Just-in-time: Gaze Guidance in Natural Behavior. PLOS Computational Biology †Nolte, D., Vidal De Palol, M., Keshava, A., Carvajal, J. M., Gert, A.L., Von Butler, E., Komurluoglu, P., König, P. (2023). Combining EEG and eye-tracking in virtual reality: Obtaining fixationonset event-related potentials and event-related spectral perturbations.Attention, Perception & Psychophysics. 2022 †Derakhshan, S., Nezami, F. N., Wächter, M. A., Czeszumski, A., Keshava, A., Lukanov, H., De Palol, M. V., Pipa, G., & König, P. (2022). Talking Cars, Doubtful Users—A Population Study in Virtual Reality. IEEE Transactions on Human-Machine Systems, 52(4), 602–612. 2021 •Keshava, A., Gottschewsky, N., Balle, S., Nezami, F. N., Schüler, T., & König, P. (2021). Action Affordance Affects Proximal and Distal Goal-Oriented Planning. European Journal of Neuroscience, 57(9), 1546–1560. †König, S. U., Keshava, A., Clay, V., Rittershofer, K., Kuske, N., & König, P. (2021). Embodied Spatial Knowledge Acquisition in Immersive Virtual Reality: Comparison to Map Exploration. Frontiers in Virtual Reality, 2. Articles marked with •are my main research projects and form the body of the thesis. Articles marked with †are in the appendices; these articles complement and support the central themes of the dissertation, through shared methodology and conceptual alignment. v
†Czeszumski, A.*, Gert, A. L.*, Keshava, A.*, Ghadirzadeh, A., Kalthoff, T., Ehinger, B. V., Tiessen, M., Björkman, M., Kragic, D., & König, P. (2021). Coordinating With a Robot Partner Affects Neural Processing Related to Action Monitoring. Frontiers in Neurorobotics, 15, 686010. (*shared first-author) 2020 •Keshava, A., Aumeistere, A., Izdebski, K., & König, P. (2020). Decoding Task From Oculomotor Behavior In Virtual Reality. ACM Symposium on Eye Tracking Research and Applications, 1–5. 2018 Afsari, A.,Keshava, A., Ossandon, J.P., & König, P. (2018). Interindividual Differences Among Native Right-to-Left Readers and Native Left-to-Right Readers During Free Viewing Tasks. Visual Cognition 26(6), 430-441 2017 Wahn, B., Keshava, A., Sinnett, S., Kingstone, A., & König, P. (2017). Audiovisual Integration is Affected by Performing a Task Jointly Annual Conference of the Cognitive Science Society, 2017 Harkness, D. L., & Keshava, A. (2017). Moving from the what to the how and where – Bayesian models and predictive processing. In Philosophy and predictive processing(pp. 254–263). MIND Group. Invited Talks 2022 Action Affordance Affects Proximal and Distal Goal-oriented Planning 64th Conference of Experimental Psychologists (TeaP 2022) 2021 Stress testing VR Eye-Tracking System Performance Neuroergonomics Conference 2021 2022 Eye-Tracking in VR with Active Users (Workshop) Neuroergonomics Conference 2021 vi
2020 Preparatory Eye Movements Reveal Prior Knowledge and Intended Interaction in VR Neuromatch Conference 2020 Supervised B.Sc. & M.Sc. Theses 2022 Predicting Spatial Bias of Attentive Gaze During Tool Interactions in Naturalistic VR. M.Sc. Thesis of Linus Tiemann 2022 To drive or not to drive? Exploring anticipatory gaze behavior between active and passive driving (by an autonomous vehicle). M.Sc. Thesis of Johannes Maximillian Pingel 2022 Decoding Neural Sources of Proximal & Distal Action Planning: A Generalized Eigen Decomposition Analysis of EEG data during Naturalistic Task in VR. B.Sc. Thesis of Özge Özenoglu 2020 Can Anticipatory Eye Movements Reveal Tool Familiarity & Intention During Tool Interaction in VR? Short Answer: Yes. B.Sc. Thesis of Nina Gottschewsky 2020 Creating an Immersive Virtual Reality Experience Featuring Hand Interaction to Investigate the Influence of Tool Knowledge on Anticipatory Eye Fixations in a Highly Ecologically valid Environment. B.Sc. Thesis of Stephan M. Balle 2020 Exploring EEG-Dynamics of Error Processing in Human-Robot Interaction B.Sc. Thesis of Tilman Kalthoff vii
Acknowledgements My journey into academia would not be possible without the numerous people who helped and mentored me along the way. First and foremost, I would like to express my sincere gratitude towards Prof. Dr. Peter König. Under his guidance, I developed a deep appreciation for scientific inquiry and curiosity. His unwavering support, sage advice, and encouragement have been invaluable. I am profoundly grateful to him for granting me the freedom and the financial security to pursue my research interests. I am deeply grateful to Dr. Sabine König for her kind guidance through the early years of formulating research questions, writing papers, and giving me opportunities to collaborate with her. I wish to extend my heartfelt thanks to the members of the Neurobiopsychology Group, with whom I had the privilege of collaborating on various exciting projects. I want to thank Benedikt Ehinger, Anna Lisa Gert, Artur Czeszumski, Farbod Nosrat Nezami, Maximillian Wächter, Shadi Derakhshan, and Debora Nolte for all the memorable science discussions. Learning and making discoveries alongside them has been an incredibly rewarding experience, and these memories are ones I will cherish forever. The inspiration I have drawn from both those I worked with directly and those I observed from afar has shaped me as a scientist, from learning to ask the right research questions to refining details like presentation fonts or crafting impactful visualizations. Every interaction has contributed to my growth. Special thanks go to Ms. Marion Schmitz and Ms. Julia Reuter for their vital support with administrative matters and making those challenges much more manageable. I want to express my deepest gratitude to the citizens of Germany for supporting the grand enterprise of Science & Technology. I hope more countries would emulate the institutional support given to young capable scientists. I want to thank my husband, who spent innumerable hours helping me think and prioritize the most important aspects of my scientific career. My daughter, of course, is a massive inspiration to me as I observe up close the development of her beautiful mind. Last but not least, I want to thank my parents for their constant love and support. ix
List of Tables 3.1 Model estimates of effect of task complexity on trial duration . . . . . . 71 3.2 Model estimates of effect of task complexity on behavior . . . . . . . . 72 3.3 Model estimates of effect of task complexity on gaze transition behavior 80 4.1 Factor loadings on principal components during grasp onset . . . . . . 106 4.2 Factor loadings on principal components during grasp offset . . . . . . 109 xvii
List of Abbreviations GOFAI Good Old-Fashioned Artificial Intelligence . . . . . . . . 7 LGN Lateral Geniculate Nucleus . . . . . . . . . . . . . . . . . . 20 V1 Primary Visual Cortex . . . . . . . . . . . . . . . . . . . . . 20 V2 VisualArea2 .......................... 20 V3 VisualArea3 .......................... 20 V4 VisualArea4 .......................... 20 MT Middle Temporal Visual Area . . . . . . . . . . . . . . . . . 20 IT Inferior Temporal Cortex . . . . . . . . . . . . . . . . . . . . 20 FEF FrontalEyeFields ........................ 21 SEF Supplementary Eye Fields . . . . . . . . . . . . . . . . . . . 21 dlPFC Dorsolateral Prefrontal Cortex . . . . . . . . . . . . . . . . 21 ACC Anterior Cingulate Cortex . . . . . . . . . . . . . . . . . . . 21 PPC Posterior Parietal Cortex . . . . . . . . . . . . . . . . . . . . 21 IPS IntraparietalSulcus....................... 21 fMRI Functional Magnetic Resonance Imaging . . . . . . . . . 27 VR VirtualReality .......................... 5 EEG Electroencephalography . . . . . . . . . . . . . . . . . . . 27 PCA Principal Components Analysis . . . . . . . . . . . . . . . 31 EVC Expected Value of Control . . . . . . . . . . . . . . . . . . 150 dACC dorsal Anterior Cingulate Cortex . . . . . . . . . . . . . . . 150 xix
Action and Cognition 1 "Our brain is an organ of action that is directed toward practical tasks." Santiago Ramón y Cajal Advice for a Young Investigator (1897) From the moment we are born, we are immersed in a rich, dynamic world that bombards our senses. Newborns have little understanding or ability to interact with their environment. Within just a few weeks, infants begin to move their eyes and focus on brightly colored objects (Gerber et al., 2010). By two to four months of age, they develop the capacity to track early visual and motor milestones illustrate how perception and action co-develop, forming the basis for cognition moving objects through a combination of eye and head movements. By three months, they achieve sufficient eye-arm coordination to reach out and interact with nearby objects. As they grow, babies move beyond passively observing their surroundings to actively engaging with them. This dynamic interplay between the body and the environment is fundamental to cognitive development, progressing alongside an infant’s increasing ability to move and interact with the world (Piaget, 1952). Sensorimotor interactions play a central role in nearly every aspect of daily life. Tasks such as driving, cooking, or gardening require constant shifts in gaze and coordination with other actions. These activities often involve a seamless sequence of interactions with objects, where vision and action are so intricately linked that transitions between tasks feel effortless. This fluid integration of vision and action illustrates how deeply our ability to navigate and inhabit our environment depends on their interplay. Underlying this seamless interaction are complex brain processes that integrate sensory inputs and motor outputs. Consider, for example, the seemingly simple act of reaching for a cup on a shelf. First, the visual CHAPTER 1. ACTION AND COGNITION |1
system identifies the cup’s three-dimensional location, which is initially represented as a two-dimensional image on the retina. This visual information is transformed into motor commands that guide the hand to the cup. As the hand approaches, additional sensory data about the cup’s shape and size configure the fingers for grasping. Once grasped, sensory feedback about the cup’s weight and dynamics ensures precise adjustments to stabilize the hand-cup system. These sensorimotor transformations, refined through experience, form the foundation for the complex repertoire of behaviors we rely on in daily life. A significant proportion of research has been dedicated to uncovering the computational principles that underlie complex behaviors. Since the mid-20th century, disciplines such as cybernetics, ethology, and neuroscience have focused on addressing one fundamental question: How does the brain give rise to complex behaviors? This line of research seeks not only to understand how intelligent behavior emerges from a biological substrate but also to replicate it in artificial systems and machines. At the core of the problem lie traditional approaches in neuroscience, which attempt to infer what the brain does (producing behavior) from how it does it (activations of neural circuits). This distinction can be likened tothe hardware-software metaphor to understanding cognition has serious limitations the relationship between software and hardware, where researchers aim to deduce the processes (software) implemented by the brain by analyzing the activity of its physical components (hardware). A provocative study by Jonas and Kording, 2017 illustrated this challenge by using a classical microprocessor as a model organism to test whether popular neuroscience data analysis methods could reveal how the system processes information. Despite the operations of the processor being fully understood beforehand (its functionality mapped out as a flowchart), the study demonstrated that prevalent analytical approaches failed to yield meaningful insights about the system, regardless of the amount of data collected. The study fundamentally challenges the assumptions underpinning much of neuroscience, raising questions about whether traditional methods are sufficient for understanding complex systems like the brain. Recent studies have shown that traditional methodologies in neuroscience need to incorporate more behavioral analysis. Krakauer et al., 2017 have critiqued the existing reductionist approach in neuroscienceneuroscience needs more behavior and advocated for a more balanced methodology that emphasizes be2|
havioral analysis. They suggest a careful theoretical and experimental decomposition of behavior to discover component processes and the underlying algorithms that explain cognition. Furthermore, they suggest that studies on the neural implementation of behavior should generally follow, rather than precede, detailed behavioral analysis. They propose a more pluralistic approach to neuroscience, where behavioral analysis provides a better understanding of the underlying neural processes. This perspective challenges the prevailing focus on neural mechanisms in neuroscience and emphasizes the importance of rigorous behavioral analysis in understanding brain-behavior relationships. The complexity of studying behavior begins with the challenge of defining the phenomenon itself. A recent survey (Levitis et al., 2009) investigated how scientists conceptualize behavior, revealing a lack of consensus among 174 animal behavior-focused scientific society members. Drawing from the responses, the authors proposed a new definition: "Behavior is the internally coordinated responses (actions or inactions) of whole living organisms (individuals or groups) to internal and/or external stimuli." However, this definition highlights further challenges. Animals and humans are constantly in motion, and behaviors often lack a clearly defined beginning or end. Moreover, most behaviors are intricately linked to sensory events that are behaviorally significant and can only be approximated if the organism is situated in a veridical world (Krakauer et al., 2017). For example, engaging a subject in a real or simulated environment can uncover crucial behavioral principles for occupying its niche. Therefore, understanding behavior (and its underlying processes) at a level that provides meaningful neural insights necessitates focusing on naturalistic behaviors carried out by individuals in their ecological context. Edward Tolman and Nikolaas Tinbergen, prominent figures in the field adaptive behavior is best understood by taking into account the organism’s environment and goals of ethology, emphasized the need to study animal behavior within their environments. Tolman argued that animal behavior is not just a series of reflexive responses to stimuli but is, in fact, goal-directed and purposeful (Tolman, 1932). He argued that animals can acquire knowledge about their environment through exploration, even if such knowledge is not immediately evident in their actions but can be applied when needed. According to Tolman, beCHAPTER 1. ACTION AND COGNITION |3
havior should be understood holistically, considering the animal’s goals and intentions. On the other hand, Tinbergen’s framework (Tinbergen, 1951) emphasizes a comprehensive examination of behavior, highlighting the significance of observing animals in their natural habitats and conducting carefully controlled field experiments. Both Tolman and Tinbergen stress the importance of internal states and the natural environment in shaping the generation and adaptation of behavior. Despite their focus on non-human animals, their work in animal behavior has significantly influenced human neuroscience. Their methodology for analyzing behavior is especially crucial in understanding human behavior and its associated neurological processes. In understanding human behaviors, it is essential to incorporate the concept of intentional action. According to the Stanford Encyclopedia of Philosophy 1, intentional actions possess distinct characteristics that differentiate them from unintentional or reflexive behaviors. Actions are fundamentally purpose-driven, guided by the agent’s goals, and involve a degree of volitional control. They require planning and decision-making among alternatives, as well as predicting or anticipating intended outcomes. Moreover, intentional actions are closely tied to a sense of agency, i.e., the agent’s conscious awareness of performing the action and its goals. In this sense, many behaviors and movements fall outside the category of intentional actions, as they lack purpose, control, or conscious awareness. When it comes to understanding the human brain’s role in producing intelligent behavior, much of 20th-century research has been constrained by narrow experimental paradigms that fail to capture the complexity of natural behavior. Laboratory-based experiments are often highly controlled, limiting participants’ ability to move freely or interact meaningfully with their surroundings . Visual experiences in such studies are frequently restricted toreductionist paradigms do not fully explain natural behavior static images or simplified stimuli. Even in action-oriented paradigms, participants are typically externally cued to move their eyes, heads, or hands, leaving little opportunity for natural, self-directed actions. These restrictive approaches fail to account for the rich complexity of body movements and environmental dynamics, essential for developing naturalistic and holistic accounts of cognition. Natural behavior and the neural processes that give rise to it are highly 1https://plato.stanford.edu/entries/action/ 4|
complex. By simplifying our experimental paradigms, we cannot fully explain the phenomenon that brings about this behavior. Moreover, such constrained paradigms raise concerns about the ecological validity of the studies, i.e., the extent to which experimental findings can be generalized to real-world situations. Ecological validity measures how well the condistudies with low ecological validity are not representative of the natural environment tions and results of a study represent the natural behaviors and cognitive processes that occur in daily life (Shamay-Tsoory & Mendelsohn, 2019). It is critical because it determines whether findings from controlled laboratory experiments can be applied to understand natural human behavior and cognition. The challenge here lies in balancing control and realism, where researchers must find a balance between maintaining experimental control and creating realistic conditions. Three aspects are essential in achieving balancing experimental control and realism is a critical component of studying behavior ecological validity (Shamay-Tsoory & Mendelsohn, 2019). First is the test environment’s verisimilitude, i.e., how similar the experimental setting is to real-life situations. Second, whether the stimuli used resemble those encountered in everyday life. And third, how well the elicited responses represent natural behaviors. Together, these aspects help us design experiments that can meaningfully explain real-world cognitive processes and behaviors. In recent years, technological advancements, such as wearable sensors and Virtual Reality (VR) have made it increasingly feasible to study naturalVR and portable neuroimaging, play a crucial role in studying natural behavior istic human behavior. These tools enable researchers to observe participants freely interacting with their surroundings, opening new avenues for understanding the complexity of human behavior and its neural correlates. This shift toward studying cognition in real-world contexts holds the potential to challenge longstanding assumptions and deepen our understanding of behavior and the brain. In the following sections, I elaborate on motivations for studying selfgenerated behavior within naturalistic, action-oriented paradigms. I detail the theoretical and philosophical frameworks underlying this approach, highlighting their significance and broader implications. Additionally, I discuss the role of task demands and the dimensionality of behavior creating challenges for studying cognition in its natural form. In subsequent chapters, I present my published research, which aims to understand cognitive processing within this action-oriented framework. CHAPTER 1. ACTION AND COGNITION |5
process. Cognitive processes ain’t (all) in the head! (Clark & Chalmers, 1998, p.8) The core claim made by Clark and Chalmers is that of an active externalism based on the acting in the environment to facilitate cognitive processes. These ideas, taken together, have formed the basis for the 4E theory of cognition, i.e., cognition is embodied, embedded, enacted, and extended. The theory starkly contrasts early cognitivist and representationalist views of the mind. The 4E framework suggests a more holistic understanding of cognition that integrates bodily experiences, sensorimotor skills, and interactions within different environmental contexts. In the sections below, I discuss the central claims made by this theory. Claims & Caveats of 4E Cognition The 4E cognition framework places the body, environment, and bodily action in a deeply coupled complex system, making them integral to cognitive processing. The general approach to recognizing human cognition as deeply rooted in sensorimotor processing comes with several caveats. According to (Wilson, 2002), the most prominent claims under the banner of embodied cognition are as follows: ◦Cognition is situated. ◦Cognition is time-pressured ◦We off-load cognitive work onto the environment ◦The environment is part of the cognitive system ◦Cognition is for action ◦Off-line cognition is body-based Research work often presents these claims as a single point of view. However, as the framework has gained popularity over the years, there is a need to disentangle these claims and examine their various criticisms. Wilson, 2002 systematically addresses these caveats, and I will summarize them below. Cognition is situated. According to this claim, cognitive activity occurs in the context of real-world environments and inherently involves perception 12 |1.1 EMERGENCE OF COGNITION FROM ACTION
and action. Tasks like driving or gardening illustrate how cognition relies on interacting with the environment. However, many cognitive activities, such as planning or reflecting, occur offline without direct environmental interactions. Wilson, 2002 warns that overstating this claim impedes understanding of cognitive phenomena without task-relevant inputs or outputs. Cognition is time-pressured. Cognition often operates under real-time constraints, requiring quick decisions and responses, such as in sports or fast-paced problem-solving. This creates a ’representational bottleneck’ where situations that demand fast and adaptive responses do not have the time to build a detailed model of the environment to derive a plan of action. Hence, a situated cognizer must use cheap and efficient mechanisms to generate actions. However, there are many activities, such as reading or deliberate problem-solving, where humans act without temporal urgency and often mitigate these pressures by slowing down or preparing in advance. Therefore, viewing this claim as a core principle of cognition must be done carefully. We off-load cognitive work onto the environment. In the face of a representational bottleneck, humans use the environment strategically, such as writing notes, using tools, or rearranging objects to reduce mental effort. There is already strong evidence from Tetris-like games that humans prefer to rotate shapes physically on the screen instead of mentally computing a solution (Kirsh, 1994). Such a strategy is often dubbed a "minimal memory strategy." Wilson, 2002 contends such a strategy might only apply to spatial tasks. Off-loading cognitive work, in the case of using pen and paper to solve a math problem or drawing Venn diagrams, is also situated and spatial and requires physical manipulation to map out the spatial relationships between concepts. However, unlike the Tetris example, such cognitive activity pertains to something that is not present in the immediate environment. Hence, off-loading cognitive work to the environment might be a generalpurpose strategy to preserve cognitive and bodily resources and can have far-reaching consequences for understanding cognition in general. The environment is part of the cognitive system. According to this claim, cognitive processes are not confined to the brain but involve ongoing interactions with the external environment. The suggestion is that cognition is distributed outside the brain-skull boundary, and the environment is an active component of the cognitive system. Wilson asserts that the environCHAPTER 1. ACTION AND COGNITION |13
ment might be supportive of cognition but not an obligatory component. She argues the effectiveness of environmental support varies with context and task demands, raising doubts about the universality of the claim across different cognitive domains. Future research must address these challenges to provide a more comprehensive understanding of how cognition extends beyond the brain into the surrounding environment. Cognition is for action. One of the main claims of embodied cognition is that cognition evolved to guide adaptive behavior and is fundamentally tied to action, with mechanisms like perception and memory serving situationally relevant tasks. Immense literature suggests that vision and visual processing are specified for guiding actions such as reaching or grasping (Goodale & Milner, 1992; Goodale, 2008, 2011). Similarly, working memory and semantic memory are honed toward past actions, anticipating the outcome of current actions and storing information about the world that facilitates interactions (Colby, 1998; Van der Stigchel, 2020). According to Wilson, while cognition supports action, not all cognitive activities are directly tied to immediate behavior. Abstract thought, planning, and symbolic reasoning suggest that cognition often operates beyond immediate actionoriented purposes. This view can sometimes underemphasize the complexity of cognitive processes that do not involve direct motor actions. Therefore, while cognition may certainly be action-oriented in many contexts, this claim does not fully account for all cognitive processes, especially those that are abstract, future-oriented, or theoretical. Off-line cognition is body-based. According to this claim, cognition in the absence of direct interaction with the world (offline cognition) still relies on sensorimotor systems that evolved for physical interaction with the world. This suggests that cognitive processes that do not involve real-time motor output still involve bodily mechanisms. Work in natural language, conceptual knowledge, and metaphors has shown that the sensorimotor structure can scaffold abstract concepts and are often based on bodily experiences (Gallese & Lakoff, 2005; Johnson & Grafton, 2003; Lakoff & Johnson, 2008). This is reflected in how we use everyday language, e.g., ’grasping’ a concept, ’reaching’ a conclusion, ’move’ through arguments, etc. Wilson criticizes the overgeneralization of sensorimotor involvement in processes such as mathematical and philosophical reasoning. Some studies (Barsalou, 1999a; Barsalou, 1999b) have demonstrated that perceptual symbols 14 |1.1 EMERGENCE OF COGNITION FROM ACTION
may play a larger role in abstract reasoning, and sensorimotor representations might not be as fundamental to cognition as some might suggest. The 4E approach to cognition highlights the deep entanglement of cognitive processes with the body and its interactions with the surrounding en4E cognition does not address abstract reasoning but centers real-world action at the heart of cognitive processing vironment. Yet, its relevance is not universal. It varies depending on the nature of the task at hand, such as whether it is situated or abstract and the extent to which the environment imposes constraints. Applying embodied cognition effectively, thus, requires a subtle understanding of when and where its principles must apply. Despite these contextual limitations, 4E theories have gained considerable momentum among both philosophers and cognitive scientists, offering a more integrative perspective on cognition, particularly in ecologically valid settings. This growing interest signals a shift toward a more pragmatic (Engel et al., 2013), action-oriented framework for investigating how cognition unfolds in the real world. 1.2 Cognition as Prediction Traditional views in cognitive science often treated cognition as a reactive process in response to external stimuli where sensory inputs are processed and have some motor outputs. It draws a clear distinction between sensory and motor domains, creating a perceived gap between perception and action. Traditional research portrays the external world as consisting of preexisting objects and features. Sensory processing is typically viewed as beginning with the transmission of these features by low-level neurons, progressing through hierarchical stages where increasingly complex patterns are extracted to inform subsequent decisions and actions. However, brains are in the business of doing much more than passively responding to stimuli. Newer theories of cognition, instead, emphasize the central role of predictFrom passive response to active prediction ing the sensory consequences of one’s own actions. By doing so, these theories dissolve the rigid boundary between sensory and motor processing, presenting perception and action as deeply interdependent processes. A review by Cisek and Kalaska, 2010 presents strong evidence against the modular and sequential information-processing models of cognition. Instead, current research supports a parallel and distributed model of decision-making and motor planning in the brain. Data from sensorimotor CHAPTER 1. ACTION AND COGNITION |15
regions suggest that decision processes are not confined to abstract cognitive regions but are intertwined with motor regions that also guide the execution of actions (Hoshi & Tanji, 2007). Similarly, neurophysiologicalCognition and action planning unfold in parallel, not in sequence. evidence supports parallel processing in decision making, where deciding what action to take (action selection) and planning to perform the action (action specification) are simultaneous processes rather than sequential (Goodale & Milner, 1992; Milner & Goodale, 1995). Moreover, data shows that decision-making, traditionally assumed to be located in higher cognitive centers such as the prefrontal cortex, is actually distributed across the cortex and involves motor circuits that guide actions. Neural circuits involved in motor execution are also implicated in decision variables such as payoff, risk, and reward (Glimcher & Fehr, 2013). Hence, cognitive processes such as decision-making and action planning are deeply integrated with motor processes and not simply abstract functions that precede motor output. Further, action plans are encoded simultaneously with sensory inputs. A recent study by Boettcher et al., 2021 investigated how prospective action plans are integrated into visual working memory. They explored when mo-Sensorimotor encoding is concurrent and prospective. tor preparation is activated relative to the stimulus onset. Using EEG, the study found that prospective action plans are encoded into working memory early, and sensory information is encoded simultaneously. This early action encoding (the brain’s preparation for future use of visual information) happened before the action was needed. The early encoding was present even when an intervening task discouraged immediate action preparation. The study provides compelling evidence that prospective actions are deeply intertwined with decision processes and do not have intermediate processing of sensory inputs. In dynamically changing environmental conditions, it makes sense that actions are encoded simultaneously with incoming sensory information.Predicting the future state of the environment based on pastInternal models support flexible action under uncertainty. experience of how it is likely to change over time becomes even more necessary. Predicting and anticipating the outside world allows us to anticipate the consequences of our own actions. Hence, predictions based on our internal models of the world enable flexible behavior in dynamically changing environments where sensory-motor delays can make real-time processing challenging. In the sections below, I will discuss the distinct role of prediction with respect to the body and the environment. 16 |1.2 COGNITION AS PREDICTION
Predicting Sensory Changes in the World Predictive mechanisms are crucial for motor control. Babies in their first year learn to predict a moving object’s future position (Kubicek et al., 2017; von Hofsten, 2004). This is even more salient in adults when predicting the trajectory of a bouncing ball in games like ping-pong or cricket (Land & McLeod, 2000; Mann et al., 2019). Evidence shows that neurons in the parietal cortex are responsible for extrapolating and predicting a moving object’s trajectory (Assad & Maunsell, 1995). The visual system not only learns the statistical regularities of the environment but also predicts motion. This allows the brain to anticipate fuVision is anticipatory, shaped by prediction and experience. ture sensory events more accurately based on past experiences. Vullings and Madelain, 2019 showed that saccade latencies are tuned to predictiondriven reinforcement, with faster saccades occurring when visual targets are presented at specific times. Notaro et al., 2019 demonstrated how anticipatory fixations and small saccades toward likely targets indicate the brain’s ability to learn environmental statistics and predict the next visual event. Similarly, Rao and Ballard, 1999 introduced the idea of predictive coding for object recognition in the visual system. Here, higher-level object representations inform early visual areas and allow for the prediction of future sensory events. The brain compares the incoming sensory signal to the expected outcome, and prediction errors are used to update the perceptual models. Recent studies by Kok et al., 2017 and de Lange et al., 2018 provide evidence of low-level visual activity before stimulus presentation, consistent with the predictive coding hypothesis. These studies suggest that the brain actively prepares for expected sensory input before it occurs. Predicting Consequences of One’s Own Actions. Prediction is also critical in anticipating the sensory consequences of selfgenerated actions, particularly eye and arm movements. There is compelling evidence of the predictive remapping in the lateral intraparietal cortex and other regions involved in eye movements (Duhamel et al., 1997; Melcher, 2007; Melcher & Colby, 2008). These studies demonstrate how the brain predictively adjusts visual receptive fields in anticipation of an upcoming saccade (Sommer & Wurtz, 2004, 2008). This remapping helps the CHAPTER 1. ACTION AND COGNITION |17
brain maintain visual stability by adjusting the visual representations before a movement and thus reconciling the preand post-saccadic images of a stimulus and maintaining continuity of visual perception. Predictive mechanisms also extend to somatosensory predictions during self-generated movements. For instance, during reaching and grasping,Predictive remapping stabilizes perception and guides motor control. the brain uses visual information to predict somatosensory consequences, such as tactile feedback (Flanagan et al., 2006). Object manipulation tasks typically involve a sequence of action phases—grasping, moving, and releasing, each accompanied by discrete sensory events across modalities such as vision, touch, and audition. The brain constructs action plans as a series of subgoals and predicts the sensory events that signify their successful completion. The motor system monitors task progression and dynamically adjusts subsequent motor commands by comparing these predictions with actual sensory feedback. This predictive mechanism ensures the coordination, adaptability, and precision necessary for effective object manipulation. The strength of predictive mechanisms is also modulated by cognitive resources, and aging increases reliance on sensorimotor predictions. For example, older adults who exhibit sensory depreciation show stronger somatosensory suppression during reaching tasks, and this suppression is negatively correlated with executive cognitive function (Klever et al., 2019). Hence, predictive mechanisms are enhanced to compensate for age-related sensory changes. Predictive Processing Clark, 2013 elegantly brings together the various concepts from 4E cognition and ties them with the theory of active inference (Friston, 2005) and the free energy principle (Friston, 2009). In Clark’s formulation, the brain actively predicts what will happen next and adjusts its predictions based on sensory feedback. Hence, action is deeply tied to prediction, and the brain uses self-generated movements and interactions with the environment to test its predictions. In this sense, action serves as a way to verify and adjust the brain’s models of the world. Crucially, predictive processing extends beyond information about the 18 |1.2 COGNITION AS PREDICTION
past or present; it also generates predictions about the future of the body and environment. This future-oriented approach underpins various aspects Cognition emerges as future-oriented inference. of perception, action, and cognitive control. In this framework, cognition emerges as the brain’s capacity to anticipate sensory information, refine action plans, and adapt to dynamic environments. Cognitive processes such as attention, memory, decision-making, and problem-solving are thus deeply rooted in the brain’s predictive mechanisms. In Friston et al., 2012, the authors assert that perception corresponds to hypothesis testing, and eye movements in service of visual search are optimal experiments to gather sensory data and test hypotheses or beliefs of how data are caused. Under the free energy minimization model, eye movements are a way to elucidate the hidden states of the world and thus minimize the entropy of these hidden states and their sensory consequences. In this way, actions maximize the confidence in the internally generated predictions. The brain’s capacity to generate predictions poses a problem of computational complexity. Sensory inputs are rich, multivariate, and highly complex, and predicting every feature of the incoming signals is impossible. König et al., 2013 argue that predictions are fundamentally constrained by Predictions are constrained by the body’s action repertoire. the action repertoire of the organism. This simplifies the problem and results in the reduction of the computational complexity. In this frame, the vast data of the sensorium is parsed first and foremost through the filters of the organism’s body and its interactions with the world. As actions are directly related to the organism’s survival, sensory data is processed by its relevance to set, predictable behaviors. Hence, organisms with similar sensoria but different action repertoires might have distinct views of the world. This general computational principle serves as a universal framework to understand the brain’s predictive mechanisms in light of sensorimotor coupling and interactions with the world. The predictive processing framework fundamentally reshapes our understanding of cognition, positioning it as an active, anticipatory process rather than a passive response to sensory input. Evidence from motor planning, decision-making, sensory processing, and neural plasticity demonstrates that the brain does not merely react to stimuli, it constantly generates predictions about the future state of the body and environment. This predictive ability enables efficient motor control, decision-making, and adaptive CHAPTER 1. ACTION AND COGNITION |19
learning in dynamic, uncertain conditions. Ultimately, cognition as prediction emerges from the dynamic interplay between sensorimotor processes and environmental interactions. This per-Predictive processing grounds 4E cognition. spective challenges traditional models that treat cognition as separate from action and instead emphasizes the deeply embodied and enactive nature of intelligent behavior. As research advances, integrating predictive processing with the principles of 4E cognition—embodied, embedded, enactive, and extended—promises a more comprehensive framework for understanding the brain’s fundamental role in shaping perception, action, and learning. 1.3 Gaze as a Window to Cognition Approximately 30% of the brain is allocated to visual processing (GrillSpector & Malach, 2004). The primate visual cortex consists of multiple distinct areas that are intricately interconnected within a hierarchical framework featuring several overlapping processing streams. PioneeringVisual processing is hierarchical and task-driven research by Van Essen et al., 1992 details the various cortical regions that hierarchically interpret incoming sensory information. While the lower levels of this hierarchy, such as the retina and Lateral Geniculate Nucleus (LGN), focus on analyzing the spatial and temporal frequencies of signals, initial cortical areas like Primary Visual Cortex (V1), Visual Area 2 (V2), Visual Area 3 (V3), Visual Area 4 (V4), and Middle Temporal Visual Area (MT) establish specialized processing pathways for form (e.g., V4) and motion (e.g., MT). Despite the specialization of these pathways, considerable interaction occurs among them, allowing for the integration of diverse features extracted from inputs. Moreover, the hierarchy of visual processing is not purely unidirectional. It includes feedback pathways from higher regions to lower ones, refining and enhancing processing capabilities. For instance, feedback from higher-order regions such as the Inferior Temporal Cortex (IT) can influence lower-level processing by modulating attention or supplying context for object recognition. Van Essen et al., 1992 also describe the linkage between the visual hierarchy and the motor control areas. The interaction between visual perception and motor output is particularly evident in tasks like visually guided actions (e.g., reaching for an object), where motor commands are influenced by processed visual 20 |1.3 GAZE AS A WINDOW TO COGNITION
information. The interconnectedness and feedback between the perceptual and motor systems ensures adaptability, enabling the brain to navigate complex tasks in real-time while adjusting processing based on context and task requirements. Thus, the visual system is a highly interconnected and hierarchical network that efficiently processes complex visual stimuli through both ascending and feedback pathways. These hierarchical networks are essential for handling the dynamic and complex nature of visual perception, supporting diverse tasks from object recognition to motion tracking and spatial awareness. The interplay between specialized streams, feedback mechanisms, and integration with motor control is key to understanding how the brain processes visual information in a way that is both adaptive and flexible in real-world scenarios. A crucial aspect of this interconnected system is the role of gaze movements, which serve as both an input to and an output of visual processing. Pouget, 2015 describes the current evidence of various brain regions that control eye movements. Within the frontal cortex, the Frontal Eye Fields (FEF), Supplementary Eye Fields (SEF), and Dorsolateral Prefrontal Cortex (dlPFC) are involved in the saccade and pursuit eye movements. Similarly, the Anterior Cingulate Cortex (ACC) also guides eye movements and attentional mechanisms. In the parietal cortex, Posterior Parietal Cortex (PPC) and Intraparietal Sulcus (IPS) have been shown to be involved in the control of saccades and attention. Areas like the dlPFC and ACC are often found to be involved in decision processes and assessment of risk and rewards (Bush et al., 2002; Kahnt et al., 2011; Kennerley et al., 2006; Shenhav et al., 2013; Shenhav et al., 2016). On the other hand, areas PPC and IPS are involved in visuomotor control and somatosensory integration (Breveglieri et al., 2015; Culham et al., 2006; Hamilton & Grafton, 2006; Marconi et al., 2001). Hence, the neural circuitry that is involved in various higher-level cognitive processes is also implicated in guiding eye movements. By studying gaze behavior, we gain insight into cognitive processes such as attention, decision-making, and action planning. König et al., 2016 arGaze links perception, attention, and decision-making gues that examining various spatiotemporal aspects of eye movements is a productive pursuit in deciphering the visual sampling process that supports numerous cognitive phenomena. The interaction between visual perCHAPTER 1. ACTION AND COGNITION |21
a functional understanding of natural cognition. Advances in mobile brain-body imaging, portable EEG, VR, and motion tracking offer new pathways to study cognition in more natural settings. This opens up opportunities to devise paradigms that allow participants to move freely, sense, plan, and act in situated contexts, like walking through a city (real or virtual), using tools, solving puzzles, etc., while simultaneously recording eye and body movement behavior in tandem with neural activity. Also, controlling different contextual cues allows one to arrive at cognitive functioning as-is in varied real-world scenarios. Moving away from artificial constraints will help enhance our understanding of how cognition emerges through active engagement with the environment. Some essential components are critical when designing naturalistic experiments. First, the experiments must allow for self-generated actions. Second, participants should be situated in an environment that supports continuous and reciprocal engagement. Third, experimenters need to be mindful of the dimensionality of the behavior under study, that is, considering the complexity of natural behavior and how effectively to record and analyze it. I elaborate on these points in detail below. Self-Generated Actions In the pioneering work of Held and Hein, 1963, pairs of kittens were reared in darkness from birth. One kitten (the active kitten) was allowed to move freely in a carousel-like apparatus, while the other (the passive kitten) was placed in a basket and carried by the active kitten’s movements. Both kittens received identical visual stimuli, but only the active kitten generated selfproduced movement. The authors hypothesized that active, self-produced movement is necessary for developing visually guided behaviors, as it allows the integration of sensory feedback and motor control. The active kittens developed normal visual-motor coordination, successfully performing tasks such as avoiding obstacles and judging depth. Despite identical visual exposure, the passive kittens failed to exhibit normal visually guided behavior and lacked the ability to avoid obstacles or recognize depth effectively. The study demonstrates that self-produced movement is critical for developing visually guided behavior and linking visual input with motor 28 |1.4 LINKING WORLD-BRAIN-BEHAVIOR
output through active exploration of the environment. Moving towards action-oriented paradigms will lead to better explanations of neural dynamics. For example, recent studies have shown that self-generated saccades better explains the variability of early neural responses in the visual cortex than fixation onset (Amme et al., 2024; Nolte, Schmidt, et al., 2024). Traditionally, fixation onset has been used to study Self-generated actions reveal neural dynamics more accurately neural responses to various object categories and has not been challenged. In many experiments, spontaneous and task-irrelevant behaviors (e.g., saccade, body movements) are usually ignored, which can provide insights into cognitive flexibility (Nau et al., 2024). Thus, adopting action-oriented paradigms to investigate the brain will perhaps upend many long-held notions in neuroscience and reshape our theories. As discussed above, Clark, 2013 emphasized the crucial role of self-generated actions in cognitive processing. Self-generated actions are downstream of neural processing that integrates sensory inputs, higher-order predictions, action plans, etc., (Buzsáki, 2019). Accordingly, they provide a more robust measure of neural activity that reflects natural cognition, boosting the generalizability of findings. Situatedness A situated agent is an organism or system that interacts with its environment continuously and reciprocally (Clark, 2013). Unlike the traditional cognitivist Ecological relevance ensures authentic behavior perspective, 4E cognition proposes that humans and animals do not build new representations of the environment from scratch. Instead, they rely on dynamic sensory inputs and constantly engage with their environment. This explains why humans and animals adapt fluidly to new environments. Clark’s view fundamentally challenges the internalist views of cognition and highlights the importance of studying how external resources shape thought and action. To study natural behavior, one must place organisms or agents in simulated or real environments that effectively capture the feedback loop in which the brain predicts sensory inputs, compares them to real-world signals, and adjusts actions accordingly. Studies in VR have shown promise in this area. For instance, when examining human evaluations of decisions made by selfCHAPTER 1. ACTION AND COGNITION |29
driving cars, analyzing simple questionnaire responses can lead to vastly different reactions from participants who are actually situated in a simulated vehicle making such decisions (Derakhshan et al., 2021; Faulhaber et al., 2019; Huang et al., 2024). Likewise, investigating human spatial navigation is more meaningful when individuals are immersed in a city environment that allows for active exploration rather than passively studying a city map (König et al., 2019; König et al., 2021). Therefore, studying situated agents provides superior opportunities to understand natural cognition as it occurs in the real world and creates better models of intelligent behavior. Situatedness mandates that the environment of the situated agent under study can support behavior that is ecologically varied and natural. Furthermore, the environment should offer sensory signals that are ecologically relevant to the agent and contain the same action affordances that are found in the natural world. This makes sure that the elicited behavior is not "a special case" subject to the parameters of the study design but rather a true reflection of what is observed in the real world. The Dimensionality of Behavior Natural behavior is rich and highly complex. Behavior is the ultimate expression of brain function in that all functions of cognition, perception, motor control, and learning serve the fundamental function of generating adaptive behavior. As natural behavior is complex, multidimensional, and dynamic,Understanding natural behavior requires rich data and context-aware models investigating it has several challenges. Moreover, natural behavior is hierarchical; for example, simple motor actions (e.g., stepping) form the building blocks of other behaviors such as walking, running, escaping, etc. Behavior can unfold over multiple timescales from milliseconds (e.g., a saccading eye) to minutes (e.g., singing and preening in birds). Environmental contexts, social interactions, and internal states influence real-world behavior. This makes behavior highly variable and context-dependent. Anderson and Perona, 2014 propose an automated analysis of behavior to replace human subjective scoring. The authors suggest that computational methods utilizing clustering and unsupervised techniques can identify behaviors that extend beyond pre-defined action categories. This fosters a data-driven discovery of behavioral motifs independent of pre30 |1.4 LINKING WORLD-BRAIN-BEHAVIOR
defined labels. Researchers can capture behavior across various spatial Data-driven tools reveal hidden structure in behavior and temporal scales by utilizing multiple sensors, real-time tracking, and machine learning models, thereby providing a comprehensive view of behavior. Mobbs et al., 2021 also advocates for a big-data approach to studying behavior. By implementing feature learning methods, researchers can extract relevant environmental features for tasks and behaviors, offering deeper insights into the latent variables that influence behavioral dynamics. Moreover, the evolution of individual behaviors can be linked to the corresponding neural activity. Techniques such as representation similarity analysis (Kriegeskorte et al., 2008) can elucidate the development of representational dynamics related to these behavioral state transitions. Consequently, this computational shift towards examining naturalistic behavior holds substantial promise for bridging the gaps between ethology, neuroscience, and psychology. Bialek, 2022 argues that behavior is low-dimensional and can usually be described using a small set of variables. This is not to say that behavior does not have multiple degrees of freedom. For example, human arm movement has multiple degrees of freedom, from shoulder rotation, elbow flexion, and wrist movement, each of which can be varied independently. Unlike Behavioral complexity emerges from lowdimensional control raw degrees of freedom, behavioral dimensionality describes the number of meaningful, independent components needed to explain behavior. Indeed, human arm movements can be represented by a low-dimensional superposition of independent components (Sanger, 2000). According to Bialek, 2022, due to the biomechanical constraints of the body and learned movement patterns, sensorimotor interactions are governed by low-dimensional manifolds. That is, behavioral complexity emerges from a small number of core principles. The brain and body reduce the complexity of movement through motor synergies and predictive control, making it possible to model complex behaviors using low-dimensional frameworks. Therefore, extracting the low-dimensional representations of behavior can simplify models of cognition and adaptive control. In this regard, techniques like Principal Components Analysis (PCA) or other eigen decomposition methods can help quantify behavioral dimensionality. PCA can determine how many dimensions are required to explain most of the variance in behavior. This can help distinguish between constrained behaviors vs. those requiring high-dimensional control. If moveCHAPTER 1. ACTION AND COGNITION |31
ments are highly correlated across time and space, they can be described with fewer independent components. This approach can also identify neural systems that are involved in constraining behavior. Exploring natural behavior poses distinct challenges because of its complexity, variability, and hierarchical structure. Behaviors occur over different timescales and are shaped by environmental and internal factors. Nevertheless, recent progress in computational techniques, machine learning, and sensor-based tracking has facilitated a more data-driven and quantitative method for behavior analysis. 1.5 About this thesis This dissertation examines the role of task, body and environment in shaping cognitive processing and gaze control. The research presented in the following chapters aligns with the 4E framework, which conceptualizes cognition as embodied, embedded, enactive, and extended. Through a series of studies, I employ VR and 3D motion modeling to investigate how gaze and body movements support in-situ action planning and execution. Below is an overview of the subsequent chapters. Study I: Does gaze encode real-world task parameters? This study investigates whether eye movement patterns alone can predict the task a person is performing in a fully immersive virtual reality (VR) environment. Since eye movements naturally adjust based on the task, we explored whether these patterns could be classified using machine learning techniques. With a VR headset featuring built-in eye-tracking, participants were asked to manipulate virtual cubes to recreate a model alignment. The study examined where participants looked, how long they fixated on specific cube regions, and how these gaze patterns changed according to the task. We developed a non-linear classifier model using Support Vector Machine (SVM) to predict tasks using points of fixations on interactive regionsof-interest. The model could classify tasks well above chance, reinforcing the notion that eye movements carry task-specific information. A crucial observation was that VR-based eye-tracking remains effective even when 32 |1.5 ABOUT THIS THESIS
participants have the freedom to move. In contrast to traditional studies in which participants remain seated and focus on a static screen, this experiment permitted full head and body movements, showing that oculomotor patterns remain informative in dynamic and natural environments. The ability to predict user intent could enhance the intuitiveness and immersiveness of digital experiences. This has significant implications for human-computer interaction and AI-driven gaze-based systems. This study offers compelling evidence that task-related gaze patterns can be decoded using straightforward machine learning methods, even in fully ambulatory VR environments. The study showcases the practicality of eye-tracking in naturalistic settings and the possibilities for gaze-based interaction, predictive intention recognition models, and real-world applications in virtual reality and human cognition research. Study II: How is gaze controlled during action sequences? In this study, we examined how humans utilize gaze to plan and execute actions in real-world situations. While prior research has investigated eye movements in everyday, routine tasks (such as making tea or sandwiches), this study focuses on how individuals use their gaze when undertaking a novel task that requires planning. We analyzed whether people plan several steps ahead or prefer to make decisions in the moment when action is necessary. To test this, participants performed a VR-based task, sorting objects according to specific features like color and shape. Some tasks were straightforward, requiring sorting based on a single feature, while others were more intricate, necessitating sorting based on multiple features. Throughout the task, eye, hand, and body movements were recorded to analyze how gaze is used during the planning and execution of actions. The findings show that people primarily use their gaze to search for relevant objects just before acting rather than planning long sequences of actions. Gaze was directed toward objects immediately before reaching for them and shifted only once the action is completed. Gaze behavior was predominantly engaged in "just-in-time" planning, i.e., participants planned in the moment just before the action was required. This means that people CHAPTER 1. ACTION AND COGNITION |33
do not pre-plan their actions but instead prefer to use their environment dynamically to guide their behavior step by step. Additionally, as tasks grew more complex, participants took longer to search for the next object to interact with, but this did not necessarily result in more efficient solutions. Rather than optimizing their movements, they relied on straightforward spatial strategies, such as selecting objects from the left and moving them toward the right. This indicates that humans naturally prefer to minimize mental effort rather than plan for the most efficient solution. The results provide a deeper understanding of how humans interact with their environment, how they plan actions in real time, and how gaze can reveal underlying cognitive strategies. The study also highlights the potential of VR and eye-tracking technologies for studying real-world behavior in controlled yet naturalistic settings. Study III: How do eyes-head-hands coordinate to execute actions? This study explores how the eyes, head, and hands work together to guide actions in a natural setting. While previous research on eye-hand coordination has often been conducted in controlled laboratory environments, these setups limit movement and may not fully capture how humans interact with objects in everyday life. To address this gap, we used VR to study how participants coordinate their movements when freely interacting with objects on a life-size shelf. Throughout the experiment, 3D movement data from their eyes, head, and hands were tracked to uncover the underlying coordination strategies when picking up and placing objects. We analyzed the low-dimensional structure of the 3D data using PCA. Despite the high number of degrees of freedom in eye, head, and hand movements, the data could be reduced to a few key components that explained most of the variance. This suggests that the brain optimizes movement control by reducing redundant information, using only essential coordination patterns to guide actions. The results revealed that eye, head, and hand movements are generally independent but synchronize just before an action "just-in-time". While gaze behavior was flexible and predictive, head and hand movements remained tightly coupled, meaning that people nat34 |1.5 ABOUT THIS THESIS
urally align their head movements with their reaching actions, suggesting that the brain may use a shared control system for guiding these movements. This suggests that hand and head coordination follows a more rigid, low-dimensional structure, while gaze is more dynamic and anticipatory. By using VR and 3D motion tracking, the study bridges the gap between lab-based research and real-world behavior. These findings have important implications for understanding how people naturally coordinate their actions, which can lead to better-designed visuomotor coordination experiments. This study demonstrates the power of advanced movement and behavioral analysis in uncovering the fundamental principles of real-world behavior. Study IV: Does gaze anticipate fine motor movements? This study investigates how eye movements help plan actions when interacting with tools, particularly how people fixate on different tool parts depending on their goal. The main research question is whether the way people look at tools before using them depends on the realism of their interaction. We tested this by letting subjects interact with a VR controller or moving their hands naturally. We were interested in how gaze behavior was influenced by specific and generic hand movements, familiarity with a tool, and how the tool is positioned. Participants performed a virtual reality (VR) task in which they interacted with familiar and unfamiliar tools using two different setups. In one experiment, they used VR controllers, where grasping was simulated by pressing a button. In another experiment, they used a Leap Motion sensor, which allowed them to make more natural hand and finger movements. We recorded where participants looked at the tool before interacting with it, specifically focusing on whether they fixated more on the tool’s handle or its functional (effector) part. The results showed that participants focused more on the tool’s functional part (effector) when asked to perform tool-specific movements, particularly with an unfamiliar tool. This implies that eye movements assist in gathering mechanical information about a tool’s operation when prior knowledge is absent. However, when the tool’s handle was positioned opposite the participant’s dominant hand, they spent more time observing the handle, CHAPTER 1. ACTION AND COGNITION |35
likely to plan their grasp. Interestingly, eye movements related to planning the grasp were significantly influenced in the experiment where participants utilized natural hand movements, unlike when using VR controllers. These findings suggest that how we look at tools before using them relies on both the immediate goal (grasping) and the long-term goal (using the tool for a task). When action affordances are more realistic, such as when participants can move their hands naturally, their gaze behavior adapts to accommodate both proximal (grasping) and distal (tool-use) planning. This study emphasizes the significance of examining human cognition in more naturalistic settings, where vision and action are closely intertwined. Additionally, the study has implications for experimental design in VR and human-computer interaction, as it reveals that how individuals observe objects depends on their anticipated interaction with them. Appendix In the appendix of this dissertation, I have added full texts of published works to which I have made substantial contributions, all of which are related to the themes explored in the thesis. 36 |1.5 ABOUT THIS THESIS
Does gaze encode real-world task parameters? 2 "The law of effect is roughly this. An act that is closely followed by satisfaction, pleasure, or comfort will be learned; one that is followed by discomfort, forgotten or eliminated." D.O. Hebb The Organization of Behavior (1949) This chapter has been peer-reviewed and published as a conference contribution: Keshava, A., Aumeistere, A., Izdebski, K., & König, P. (2020). Decoding Task From Oculomotor Behavior In Virtual Reality. ACM Symposium on Eye Tracking Research and Applications. CHAPTER 2. DOES GAZE ENCODE REAL-WORLD TASK PARAMETERS? |37
(a) (b) (c) Figure 2.4: Relative frequency of PORs over the reference and transfer cube. (a) Relative frequency of PORs over the three regions-of-interest; (b) Relative frequency of PORs for the three cube sizes; (c) Relative frequency of PORs for the six planes of the cubes In order to scale the training set features, we standardized each feature observations by subtracting the mean of the feature and dividing the difference by its standard deviation. During training, we applied two optimization protocols. Firstly, in order to arrive at a subset of features optimized for the best f1-score (calculated as the harmonic mean of precision and recall for each alignment type), we used recursive feature elimination (Zeng et al., 2009) on five-folds of the cross-validated training data. Secondly, using the reduced set of features, we did a grid search for the best regularization parameter C and the best kernel coefficient gamma from 10−4to 104over another five-fold cross-validated training data using radial basis kernel SVM. To minimize any bias resulting from class imbalance during training, for both the recursive feature elimination and the grid search for the best SVM parameters, we created the cross-validation folds so that we always had the same number of trials in each class label. Before testing the performance of the classi44 |2.4 RESULTS
fier, we standardized the test data as above. We fitted the whole training set using the reduced features and the optimized parameters, C and gamma, and then tested the SVM fit on the test set. In this way, we trained and tested the SVM classifiers 24 times. Taken together, the kernel-based SVM results in a mean f1-score of 0.34±0.1 overall class labels (binomial test, p<0.001). Results show that we can predict the cued alignment type for each subject well over chance level, indicating that aggregation of PORs on the regions-of-interest captures somewhat distinct information about the task the subjects performed. As the number of PORs collected on the S cubes was strikingly less than PORs on M and L cubes (Figure 4b), we performed the SVM classification as above by removing all trials which included S cubes. This resulted in a mean f1-score of 0.51±0.17 overall class labels (binomial test, p<0.001). The f1-score per alignment type is illustrated in Figure 5. This indicates that the region-of-interest based PORs on the M and L cubes were more informative about the task performed, leading to an increase in the f1-score by 0.17 from the former model with all cubes. We hearken back to the inverse Yarbus process, demonstrating that gaze patterns can predict the task performed by the subject even in a fully ambulatory experimental setup in VR. 2.5 Discussion and Conclusion We showed that there is sufficient information in the eye-tracking data to infer the task given to the subjects using oculomotor information alone. Specifically, that proportion of PORs on the regions-of-interest of the manipulated objects represented the task parameters reliably. The kernel-based SVM classification method was robust against inter-individual differences of the subjects, further suggesting that there are distinct oculomotor signatures that can facilitate task-inference. Furthermore, even though the task presented to participants was based on minor spatial differences, the SVM classification yielded above chance predictions. Also, we use the proportion of PORs for each trial as our training feature, which subsumes both the number of fixations and dwelling time on the regions-of-interest. While this method helped capture relevant information for our study, better feature extraction methods can still be used. E.g., Kanan et al., 2014 used Fisher kernelCHAPTER 2. DOES GAZE ENCODE REAL-WORLD TASK PARAMETERS? |45
Figure 2.5: Results of the kernel SVM for the four alignment types with the precision, recall, and f1-score results for the four different alignment types with S cube trials excluded; the dashed line indicates a 0.25 chance-level of prediction, error bars indicate 95% confidence interval. based feature vectors for SVM classifications, which out-performed summary statistics of eye-movement data. It should be noted that the calibration accuracy of the eye tracker can affect the classification performance of the SVM. In our study we maintained a calibration error of <1° of visual angle, however, for the different sized cubes, the angle subtended on the eye would differ based on the viewing distance. Consequently, the discrimination of the location of the PORs on the regions-of-interest on the S cube is greatly affected by larger viewing distances as compared to M or L cubes. We demonstrate this by the showing an increase in the classification performance when we remove all trials with S cubes. With this study, we also demonstrate the robustness of eye movement data collected in a fully ambulatory virtual environment. While classical lab-based studies provide a high degree of control over key physiological variables, studying human behavior in more naturalistic environments ensures high ecological validity and generalizability of the findings. Here, we emphasize the importance of studying human behavior from an embodied perspective. Even though subjects move and perform 46 |2.5 DISCUSSION AND CONCLUSION
the tasks in their idiosyncratic ways leading to high inter-individual variance, we find that task-related oculomotor information can still be captured reliably. The present study is task-specific but can find many applications, such as intention recognition in virtual reality. Our paradigm investigates tasks based on spatial alignment of objects and can be generalized to similar setups. Specifically, virtual reality prototype editors such as Boxplan 1use simple block-like structures to create complex environments. Our experimental paradigm can be directly applicable to user interaction protocols based on eye movements and help in providing the intended interaction cues. Such predictive methods can add to gaze-augmented user experience design, making it more intuitive and human-centric. Author Contributions AK, KI, PK: conceived and designed the study. AA: programmed the experiment. AA: data collection. AK: data analysis. AK: initial draft of the manuscript. AK, PK: revision and finalizing the manuscript. All authors contributed to the article and approved the submitted version. Acknowledgement We are grateful for the financial support by the German Federal Ministry of Education and Research for the project ErgoVR (Entwicklung eines Ergonomie-AnalyseTools in der virtuellen Realität zur Planung von Arbeitsplätzen in der industriellen Fertigung)-16SV8052. We would also like to thank Maximilian Wächter and Sabine König for their editorial help. 1https://www.embodied.engineering/boxplan CHAPTER 2. DOES GAZE ENCODE REAL-WORLD TASK PARAMETERS? |47
How is gaze controlled during action sequences? 3 "The eyes of men converse as much as their tongues, with the advantage that the ocular dialect needs no dictionary, but is understood all the world over." Ralph Waldo Emerson The Prose Works of Ralph Waldo Emerson (1872) This chapter has been published as a peer-reviewed article: Keshava, A., Nezami, F. N., Neumann, H., Izdebski, K., Schüler, T., & König, P. (2024). Just-in-time: Gaze Guidance in Natural Behavior. PLOS Computational Biology CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |49
3.1 Abstract Natural eye movements have primarily been studied for over-learned activities such as tea-making, sandwich-making, and hand-washing, which have a fixed sequence of associated actions. These studies demonstrate a sequential activation of lowlevel cognitive schemas facilitating task completion. However, whether these action schemas are activated in the same pattern when a task is novel and a sequence of actions must be planned in the moment is unclear. Here, we recorded gaze and body movements in a naturalistic task to study action-oriented gaze behavior. In a virtual environment, subjects moved objects on a life-size shelf to achieve a given order. To compel cognitive planning, we added complexity to the sorting tasks. Fixations aligned with the action onset showed gaze as tightly coupled with the action sequence, and task complexity moderately affected the proportion of fixations on the task-relevant regions. Our analysis revealed that gaze fixations were allocated to action-relevant targets just in time. Planning behavior predominantly corresponded to a greater visual search for task-relevant objects before the action onset. The results support the idea that natural behavior relies on the frugal use of working memory, and humans refrain from encoding objects in the environment to plan long-term actions. Instead, they prefer just-in-time planning by searching for action-relevant items at the moment, directing their body and hand to it, monitoring the action until it is terminated, and moving on to the following action. 3.2 Author summary Eye movements in the natural environment have primarily been studied for over-learned and habitual everyday activities (tea-making, sandwich-making, hand-washing) with a fixed sequence of associated actions. These studies show eye movements correspond to a specific order of actions learned over time. In this study, we were interested in how humans plan and execute actions for tasks that do not have an inherent action sequence. To that end, we asked subjects to sort objects based on object features on a life-size shelf in a virtual environment as we recorded their eye and body movements. We investigated the general characteristics of gaze behavior while acting under natural conditions. Our paper provides a comprehensive approach to preprocess naturalistic gaze data in virtual reality. Furthermore, we provide a data-driven method of analyzing the different 50 |3.1 ABSTRACT
action-oriented functions of gaze. The results show that bereft of a predefined action sequence, humans prefer to plan only their immediate actions, where eye movements are used to search for the target object to immediately act on, then to guide the hand towards it and monitor the action until it is terminated. Such a simplistic approach ensures that humans choose sub-optimal behavior over planning under sustained cognitive load. 3.3 Introduction In a pragmatic turn in cognitive science, there has been a significant push to study cognition and cognitive processing during context-dependent interactions within the constraints of the local properties of an environment (Clark, 1998; Parada & Rossi, 2020). This pragmatic shift places willed/voluntary actions and proactive control at the heart of cognitive processing, where behavior is natural and not cued externally. Moreover, Engel et al., 2013 have proposed that cognition encompasses the body, and bodily action can be used to infer cognition. Furthermore, eye movements provide a unique window to understand cognitive control in natural, real-life situations (König et al., 2016). This requires a mobile setup that allows a human subject to be recorded as they actively interact in a controlled but unconstrained environment König et al., 2018. In recent years, virtual reality (VR) and mobile sensing have offered great opportunities to create controlled, naturalistic environments. Here, subjects’ eye and body movements can be reliably measured along with their interactions with the environment Keshava et al., 2020; Keshava et al., 2023; Nolte, Vidal De Palol, et al., 2024. Studying eye movements in mobile subjects gives us a richer, veridical view of cognitive processing for willed actions. Seminal studies have investigated eye movement behavior in natural environments with fully mobile participants. In the pioneering studies of Land et al., 1999 and Hayhoe et al., 2003, subjects performed everyday activities such as making tea and sandwiches, respectively. These studies required subjects to enact a sequence of actions that involved manipulating objects one at a time to achieve a goal. Both studies showed that nearly all gaze fixations were primarily allocated in a taskoriented manner, and a negligible number of fixations were found in task-irrelevant locations. These experiments in naturalistic settings have revealed several distinct functions of information sampling during routine everyday tasks. Investigation of habitual tasks has revealed a systematic timing between visual CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |51
fixations and object manipulation. Specifically, fixations are made to target objects about 600ms before manipulation. Notably, Ballard et al., 1995 proposed a "just-intime" strategy that universally applies to this relative timing of fixations and actions. Fixations that provide information for a particular action immediately precede that action and are crucial for the fast and economical execution of the task. Moreover, Land and Hayhoe, 2001 have broadly categorized fixations while performing habitual tasks into four functional groups. "Locating" fixations retrieve visual information. "Directing" fixations acquire the position information of an object, accompany a manipulation action, and facilitate reaching movements. "Guiding" fixations alternate between manipulated objects, e.g., fixations on the knife, bread, and butter while buttering the bread. "Checking" fixations monitor whether the task constraints have been met. These fixation categories have been corroborated by Pelz and Canosa, 2001 and Mennie et al., 2007. Hence, there is a just-in-time strategy while performing common tasks where gaze is specifically allocated toward searching for a target, guiding hands, and monitoring the action as it unfolds. Based on the above observations, Land and Hayhoe, 2001 have proposed a framework that outlines the flow of visual and motor control during task execution Fig 3.1A. The process summarizes the various operations that must occur during an object-related action’, i.e., individual actions performed on separate objects to achieve the desired goal. Each schema "specifies the object to be dealt with, the action to be performed on it, and the monitoring required to establish that the action has been satisfactorily completed." (Land, 2006). In the action preparation period, the gaze control samples the information about the location and identity of the object and directs the hands to it. Subsequently, in the action execution period, the motor control system of the arms and hands implements the desired actions. Here, vision provides information about where to direct the body, which state to monitor, and determine when the action must be terminated. Thus, a ’script’ of instructions is sequentially implemented where eye movements earmark the task-relevant locations in the environment that demand specific attentional resources for that action. An implicit theme in the studies investigating behavior in natural tasks (e.g., teamaking, sandwich-making, hand-washing) is that these tasks have an organized and well-known structure. They involve specific task-relevant objects, object-related actions such as picking up the teapot, pouring water, etc., and a predefined ’action script’ for executing the tasks. Therefore, they study eye movements under strict control of a task sequence. Moreover, these tasks are over-learned as they are part of routine actions for a healthy human adult. For these habitual tasks, cognitive 52 |3.3 INTRODUCTION
Figure 3.1: Study Design. A: Schematic of motor and gaze control during performance of natural tasks. A cognitive schema selects between the object of interest, the corresponding action, and the monitoring of the action execution. If the object’s location is not in memory, a visual search is undertaken to locate the object, the body and hands are directed towards it in preparation for the action execution, and finally, the action is monitored until the behavioral goal has been achieved. B: Experimental Task. In a virtual environment, participants sorted 16 objects based on two features, color or shape, while we measured their eye and body movements. The objects were randomly presented on a 5x5 shelf at the beginning of each trial. Participants were cued to sort objects by shape and color. Trials, where objects were sorted based on just one object feature (color or shape), were categorized as EASY trials. Conversely, the trials where sorting was based on both features (color and shape) were categorized as HARD trials. All participants performed 24 trials (16 easy and eight hard) without time constraints. C: Labelling of the continuous data stream into action preparation and execution epochs. To study the function of eye movements, we divided each trial into action preparation and execution epochs. The action execution epochs start from the grasp onset till the grasp ends for each object displacement. In contrast, the action preparation epochs start from the end of the previous grasp and the onset of the current object displacement. schemas allocate gaze for specific information retrieval (locating, directing, guiding, monitoring) and are likely not executed under deliberate conscious control. Hence, it is unclear how cognitive schemas control gaze where an internal task script is unknown. Furthermore, a core feature of Land and Hayhoe’s gaze and motor control framework is its sparse use of visual working memory. In this framework, low-level cognitive schemas are activated successively as needed. For example, it states CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |53
Next, we removed grasping periods where the beginning and final locations of the objects on the shelf were the same. We calculated the interquartile range (IQR) of the participants’ number of object displacements for the two trial types (EASY and HARD). To remove the outlying object displacements in the trials, we removed the trials with 1.5 times the IQR of object displacements. We also removed those trials with fewer than three object displacements. After pre-processing, we were left with data from 48 subjects with 999 trials in total and 20.8±1.8trials per subject. Furthermore, we had eye-tracking data corresponding to 12,403 object displacements with 12.3±5.11 object displacements per trial. Using the grasp onset and end events, we divided the continuous data stream into action preparation and execution epochs. Specifically, the action preparation phase starts from the end of a previous grasp onset till the start of the immediate grasp onset. Similarly, the execution epoch started from the start of the current grasp onset to the end of that grasp. The division of the data into these epochs is depicted in Fig 3.1C. We further organized the data to examine the spatial and temporal characteristics of eye movements during the action preparation and execution epochs. We labeled the fixation based on its location on the different objects and shelf locations into eight regions of interest (ROIs). These included fixations of targets in the previous, current, and next grasping actions. Specifically, the previous target object refers to the object handled in the previous action, and the previous target shelf is the shelf where the previous target object was placed. Similarly, the current target object refers to the object picked up and placed on the target shelf in the current action, and the next target object and next target shelf are in the immediate following action in the sequence. All other fixated regions that did not conform to the above six ROIs were categorized as ’other objects’ and ’other shelves’. In this format, we could parse the sequence of fixations on the eight ROIs that are relevant for preparing and executing the object-related actions. Data Analysis 60 |3.4 METHODS
Influence of Task Complexity on Behavior We used four measures to examine the grasping behavior during the experiment. First is trial duration, i.e., the total time to finish the tasks. Here, we tested the hypothesis that task complexity affected the total duration of finishing the tasks using a two-sample independent t-test. Second is the action preparation duration, which is the time between dropping the previous object and picking up the current target object. Third is execution duration, i.e., the time to move the target object to the desired location on the shelf. We tested the hypothesis that the action epochs durations differed based on the action epoch type and the task complexity. Finally, we were interested in the number of object displacements made by the participants to complete the tasks. We tested the hypothesis that participants were optimal in their action choices in the EASY and HARD conditions. Taken together, these measures capture the influence of task complexity on overt behavior. We modeled the relationship between the action epoch duration dependent on the action epoch type (PREPARATION, EXECUTION) and the trial type (EASY, HARD). As the durations differed from the start of the trial to the end, we added the grasp index as a covariate in the model. All within-subject effects were modeled with random intercepts and slopes grouped by subject. The dependent variable was log-transformed so that the model residuals were normally distributed. The categorical variables trial type and epoch type were effect-coded (Schad et al., 2020) so that the model coefficients could be interpreted as main effects. The model coefficients and the 95% confidence intervals were back-transformed to the response scale. The model fit was performed using restricted maximum likelihood (REML) estimation (Corbeil & Searle, 1976) using the lme4 package (v1.1-26) in R 3.6.1. We used the L-BFGS-B optimizer to find the best fit using 20000 iterations. Using the Satterthwaite method (Luke, 2017), we approximated degrees of freedom of the fixed effects. The full model in Wilkinson notation (Wilkinson & Rogers, 1973) is denoted as: Log(duration)∼1 + trial_type ∗epoch_type ∗grasp_index +(1 + trial_type +epoch_type +grasp_index|Subject) CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |61
Comparison of Human vs. Model Performance To ascertain whether the action choices of the subjects were optimal, we compared their performance to a greedy search model. The model was designed to maximize the score of each possible move. The model operations are summarized in Box 1. For the EASY tasks, a move’s score depended on the number of identical object features, e.g., color in each row or column, where the score of each shelf configuration was the sum of the scores of each row. In the HARD tasks, the score of a move was calculated by summing the number of unique object features per row or column. In this way, the greedy model found the optimal moves by maximizing the score of the possible next move. Algorithm 1 Greedy Model 1: Create empty stack Sfor each shelf state 2: PUSH initial shelf configuration to S 3: while Sis not empty do 4: POP latest state siin S 5: current score = score of state si 6: for each row in sido 7: for each object in row do 8: for each row in sithat is not current row do 9: for each empty space in row do 10: snext = object in empty space 11: new score = score of snext 12: if new score > current score and move not in Sthen 13: PUSH snext to S 14: update current score to new score 15: end if 16: end for 17: end for 18: end for 19: end for 20: end while 21: Return length of Sas number of moves to final state We modeled the relationship between the number of moves made by subjects in each trial dependent on the trial type (EASY, HARD) and the solver type (MODEL, HUMAN). As before, all within-subject effects were modeled with random intercepts and slopes grouped by subject. The independent variables were also effect-coded. 62 |3.4 METHODS
The full model in Wilkinson notation (Wilkinson & Rogers, 1973) is denoted as: ObjectDisplacements ∼1 + trial_type ∗solver_type +(1 + trial_type +solver_type|Subject) Action Locked Gaze Control We analyzed the average fixation behavior aligned with the action onset. For each grasp onset in a trial, we chose the period from 3s before grasp onset to 2s after. We divided this 5s period into bins of 0.25 seconds and calculated the number of fixations on the eight ROIs described above. For each time bin, we calculated the proportion of fixations on the ROIs per trial type (EASY, HARD). We used the cluster permutation method to find the periods of significant differences between EASY and HARD trials for a given ROI. Here, we use the t-statistic as a test statistic for each time-bin, where t is defined as: t=√N∗x σ where x is the mean difference between the trial types, σis the standard deviation of the mean, and N is the number of subjects. We used a threshold for t at 2.14, corresponding to the t-value at which the p-value is 0.05 in the t-distribution. We first found the time bins where the t-value was greater than the threshold. Then, we computed the sum of the t-values for these clustered time bins, which gave a single value representing the cluster’s mass. Next, to assess the significance of the cluster, we permuted all the time bins across trials and subjects and computed the t-values and cluster mass for 1000 different permutations. This gave us the null distribution over which we compared the cluster mass derived from the real data. To account for the multiple independent comparisons for the eight ROIs, we considered the significant clusters to have a Bonferroni corrected p-value less than 0.006 (0.05/8). In the results, we report the range of the significant time bins for the eight ROIs for the two trial types and the corresponding p-values. Gaze Transition Behavior To examine the gaze transitions within the action preparation and execution epochs we created transition matrices that show the origin and destination locations of the CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |63
fixations on the eight ROIs. A series of fixations on the ROIs in a given action epoch were transformed into a 2D matrix (M) with eight rows as the origin of the fixation and eight columns as the destination. The diagonal of the matrix was set to zeros as we were not interested in re-fixations on the same ROI. Using matrix M, we calculated the net and total transitions from and to each ROI. For every transition matrix Mper action preparation and execution epoch, the net (Mnet), and total (Mtotal) transition are defined as follows: Mnet =M−MT Mtotal =M+MT If subjects make equal transitions between the ROI pairs in both directions, we can expect zero transitions in Mnet. Conversely, with strong gaze guidance towards a particular ROI, we would expect more net transitions. Hence, using the net and total transitions per action epoch, we calculated the relative net transitions (R) as: R=PMnet PMtotal We then took the mean of the relative net transitions per trial to measure asymmetric gaze guidance behavior in that trial. We modeled the linear relationship of the relative net transitions dependent on the trial type (EASY, HARD), epoch type (preparation, execution), and their interactions. All within-subject effects were modeled with random intercept and slopes grouped by subject. The categorical variables trial type and epoch type were effectcoded. The full model in Wilkinson notation (Wilkinson & Rogers, 1973) is defined as: RelativeT ransitions ∼1 + trial_type ∗epoch_type +(1 + trial_type ∗epoch_type|Subject) Latency of First Fixations on Task-relevant ROIs Finally, we also calculated the latency of the first fixation on the eight ROIs from the start of the preparation and execution epochs. We computed the median time 64 |3.4 METHODS
to first fixate on an ROI per trial. As the action preparation and execution epochs varied in duration, we normalized the time points by dividing them by the duration of the epoch. Thus, the normalized elapsed time from the start of an epoch is comparable to all epochs across trials and subjects. We modeled the linear relationship of the normalized median time to the first fixation dependent on the trial type (EASY, HARD) and the eight ROIs and their interactions. We computed two models for the action preparation and execution epochs. All within-subject effects were modeled with random intercepts and slopes grouped by subject. The categorical variables trial_type and ROI were effect coded as before, and the model coefficients could be interpreted as main effects. For both models, we chose the latency of the first fixation on the current target object as the reference factor so that the latency of the fixation of all other ROIs could be compared to it. The full model in Wilkinson notation (Wilkinson & Rogers, 1973) is defined as: Fixationtime ∼1 + trial_type ∗ROI +(1 + trial_type ∗ROI|Subject) 3.5 Results Fifty-five human volunteers performed object-sorting tasks in a virtual environment. They sorted objects on a 5x5 shelf that measured approximately 2m in width and height. The tasks were based on sorting 16 objects of four colors (red, blue, yellow, green) and shapes (cube, sphere, pyramid, cylinder). Fig 3.1B illustrates the experimental setup. We experimentally modulated the task complexity into EASY and HARD trials. In the EASY trials, we gave participants instructions such as "Sort objects so that each row has the same color or is empty," and in the HARD tasks, an example trial instructed them to "sort objects so that each row has each unique color and unique shape once." Participants were not given any other instructions on how to complete the tasks. Hence, their action choices were not cued to follow a set strategy. Participants wore an HTC Vive Pro Eye virtual reality head-mounted display (HMD) as they performed the task. Using the HMD, we measured their eye and head movements simultaneously, while the HTC hand controllers measured CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |65
their hand movements and grasping behavior. The eye and body tracking sensors were calibrated regularly throughout the experiment session. We utilized two data streams for this study: the cyclopean eye (average of the left and right eye) direction vector and hand movement position. At the outset, we rejected data from participants where the sensors malfunctioned mid-experiment. We derived the angular velocity of the cyclopean eye from the direction data. We then used a median absolute deviation-based saccade detection method to cluster gaze samples into either fixation or saccade. Following this, we removed fixations lasting less than 100 ms. Further, we rejected the fixations with a duration greater than 3.5 median absolute deviation. This step rejected 2.4% of the outlying gaze data. Next, we derived the onset and end of a grasping action from the VR hand controller’s trigger button, i.e., the time when a participant pressed the button with their right index finger to pick up an object and the time point when they released the button after moving it to a desired location. We did not consider grasps where the pickup and drop-off locations were the same. Finally, trials with fewer than three grasps were rejected. Also, we rejected the trials with more than 1.5 times the inter-quartile range (IQR) of the number of grasps by all participants. This step helped remove trials where subjects had outlying action choices. Hence, after pre-processing, we were left with 12403 grasping actions over 999 trials from 48 participants. As a final step, we segmented the data stream into the action preparation and execution epochs. As illustrated in Fig 3.1D, these epochs marked the end of the previous grasp to the start of the immediate grasp and the time from the beginning of the current grasp to its end, respectively. Hence, every grasping action had a preparation and an execution phase with concurrent gaze data. Further, using the gaze data stream, we defined regions of interest (ROI) on the target object that was grasped and the target shelf where it was moved. These ROIs were further classified into the previous action, the current action, and the following action. The regions that did not correspond to targets in the action sequence were categorized as ’other objects’ and ’other shelves.’ This step prepared the data to demonstrate a "flow" of fixations from one action-relevant region to another during both the preparation and the execution phase of the grasping action. 66 |3.5 RESULTS
Influence of Task Complexity on Natural Behavior In this study, the primary object-related action was to iteratively prepare and execute the pickup and placement of objects until they were sorted according to the task instructions. We used four measures to account for the behavioral differences between the task complexities. First is trial duration, i.e., the total time to finish the tasks. Second is the action preparation duration, the time between dropping the previous object and picking up the current target object. Third is execution duration, i.e., the time to move the target object to the desired location on the shelf. Fourth, the number of object displacements or "moves" made by the participants to complete the EASY and HARD tasks. These measures characterize the behavioral outcomes influenced by the task complexity. Fig 3.2A shows the trial durations in EASY and HARD tasks. A two-sample independent t-test showed that the duration of the two task types was significantly different (t=−10.92, p < 0.001) where EASY tasks were shorter (Mean = 51.15s, SD = ±10.21) compared to HARD tasks (Mean = 123.46 ±44.70s). Hence, the increase in task complexity markedly increased the duration of the tasks by a factor of two. At a more granular level, we were interested in how task complexity affected the average action preparation and execution duration. For EASY tasks, the mean preparation epoch duration was 2.10s(SD = ±0.78), and the mean execution epoch duration was 1.76s(SD= ±0.27). The mean preparation epoch duration for HARD trials was 4.05s(SD= ±1.99), and the mean execution epoch duration was 2.15s(SD =±0.58). Fig 3.2B depicts the mean and variance of the preparation and execution duration for the EASY and the HARD tasks. We tested the hypothesis that the complexity of HARD trials increased the action preparation and execution duration. We used a linear mixed effects model with epoch duration for a given grasp as the dependent variable and trial type (EASY, HARD) and epoch type (PREPARATION, EXECUTION) as independent variables. We also added the grasp index, i.e., an integer denoting the order of the object displacement in the trial (for example, 1st,2nd,3rd, etc.) as a covariate. The independent variables are effect-coded so that the regression coefficients can be directly interpreted as main effects. Table 3.1 summarizes the model results. Epoch duration was significantly different for the two trial types, where epochs of HARD trials were on average 1.31slonger than EASY trials. Irrespective of the trial type, the action preparation epochs were significantly longer than the execution epochs by CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |67
1.08s. An insignificant positive correlation of 0.002 with the grasp index indicates that overall, the epoch duration of initial moves did not differ from the latter ones. The preparation duration was significantly longer in the HARD trials, as shown by the interaction between the trial and epoch types. Furthermore, the difference in the slope of epoch duration and grasp index of HARD and EASY trials was insignificant. Similarly, the difference in the slope of the duration with respect to the grasp index between preparation and execution epochs was also not significant. However, the three-way interaction between trial type, epoch type, and the grasp index was 0.01 and significant, showing that the preparation duration of the HARD trials was significantly shorter at the beginning of the trial and increased linearly as the trial progressed. To summarize, the preparation and execution phases of the action were distinctly different, and task complexity specifically affected the duration of preparing actions and less so the time taken to execute them. Moreover, the significant interaction terms of the model suggest that in HARD trials specifically, participants made fast moves at the beginning of the trials and spent more time preparing their actions in the later part. 68 |3.5 RESULTS
CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |69
previous target object (dark blue trace), ’other objects’ (dark grey trace), and ’other shelves’ regions (light grey trace). The comparable proportion of fixations on previous targets and the ’other’ categories indicates a visual search for task-relevant objects/shelves at the start of the preparation period. Subsequently, as the fixations on the previous targets and ’other objects’ decrease, there is a steep increase in the proportion of fixations on the current target object (dark red trace), reaching a maximum of 68% of fixations at the point of grasp onset (time= 0s). Throughout the preparation period, the proportion of fixations on the next target object (dark green trace) and shelf (light green trace) remains comparatively small. The large proportion of fixations in the action preparation phase devoted to the current target object shortly before the hand makes contact with the object indicates just-in-time fixations used to direct and guide the hand toward the target object. In the action execution phase, up to 2s after the grasp onset, the proportion of fixations on the different ROIs reveals further action-relevant oculomotor behaviors. Firstly, with time, the proportion decreases from the current target object and increases over the current target shelf (light red trace), where approximately 40% of the fixations are allocated around 700ms after grasp onset and well before the end of the action execution at 1.91s (95%CI[1.81,2.01]). Secondly, before the end of the action execution, there is an increase in the proportion of fixations on the next target object. The presence of approximately 10% of fixations on the next target object before the end of the current action execution shows that the target for the next action is located shortly before ending the current action in a number of trials. Moreover, about 20% of the fixations are devoted to the ’other’ ROIs after 1s. This is again indicative of a search process for task-relevant objects/shelves. In a nutshell, during the action execution phase, the majority of fixations suggest two primary functions of gaze: first, to aid the completion of the current action by fixating on the shelf location in anticipation of the object drop-off, and second, to commence an object search for the next action when the target is not already decided. Importantly, the proportion of fixations on the eight ROIs in the EASY and HARD tasks does not deviate dramatically across time. Even though there are large behavioral differences in the EASY and HARD tasks, both in terms of time taken to prepare for the actions and the number of actions required, the fixation profiles time-locked to the grasp onset are comparable. Most importantly, the fixations are made just in time, where gaze shifts occur systematically in sync with the previous, current, and next action sequences. Hence, we can conclude that the average oculomotor behavior described above is driven just-in-time by overt actions and less by 76 |3.5 RESULTS
the additional cognitive load of the tasks. There are, however, slight differences in the proportion of fixations due to the task complexity. These differences point to aspects of the oculomotor behavior subject to the cognitive load imposed by task complexity. We used a cluster permutation analysis on the average fixation time course over each ROI to quantify these differences. The analysis revealed periods where the proportion of fixations significantly differed in the two task types on each of the eight ROIs. Fig 3.3 depicts the start and end of the period when the differences in a given ROI were statistically significant. An interesting characteristic of the task-specific periods of differences is that they occur early in the action preparation epoch. During HARD tasks, there are fewer fixations on the previous target object (dashed dark blue trace) and previous target shelf (dashed light blue trace) and more fixations on the other target objects (dashed dark grey trace) and shelf (dashed light grey trace). In HARD tasks, there are slightly more fixations on the next target object (dark green trace), shelf (light green trace), and on the current target object (dashed dark red trace) and target shelf (dashed light red trace). These early fixations indicate that task complexity increased the incidence of ’search’ and ’locate’ fixations. As the average action preparation epoch was longer in the the HARD trials, the analysis exemplifies the prolonged visual search undertaken due to task complexity. Similarly, concurrent with the action execution, there are more fixations on the previous target object and shelf, the next target shelf, and the other objects in the HARD tasks. More fixations on the previous target object and shelf are characteristic of ’checking’ fixations used to monitor the current task execution with respect to the task condition. Further, more fixations on the ’other’ objects indicate a greater search for the next target object in the HARD tasks. With an increase in fixations on the other object ROIs in HARD trials, later in the execution phase, there are correspondingly fewer fixations on the current and next target object. As a proportional decrease is not seen in the fixations on the current target shelf, we can infer that subjects began search for the next target object before the immediate action was terminated. Thus, specifically in the execution epochs of HARD tasks, the average fixations on the different ROIs show the functions of task monitoring and searching for the next action targets. To summarize, the results in this section support two broad findings. First, despite significant behavioral differences between the two task types, the eye movements show slight deviations. This supports the idea that eye movements are tightly CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |77
coupled to the actions and less so with the cognitive effort added by task complexity. Secondly, the observed significant differences between EASY and HARD trials are compatible with a slightly increased focus away from the previous action towards the current and the next step in the action sequence. This includes an earlier reduction of fixations on previous targets to facilitate search, the requirement for additional checking fixations after action onset, and a later attendance towards the next target object. Thus, cognitive processing maintains a structured sequence of activations corresponding to searching and locating a target object, guiding hands to it, and monitoring the task. Gaze Transition Behavior To further elucidate the concurrent cognitive processes, we focused on sequences of attended regions of interest in individual trials. Specifically, we computed transition matrices to capture fixations to and from each of the eight ROIs. Fig 3.3B1 shows the exemplar transition matrix for a preparation epoch in an EASY trial. Fig 3.3B2 shows the derived net (Mnet)and total (Mtotal)transition matrices from the example transition matrix (M). Total transitions represent the total number of fixations to and from one ROI. On the other hand, net transitions denote the asymmetry in gaze allocation from one ROI to another. For example, an equal number of fixations to and from pairs of ROIs would result in zero net transitions between those ROIs. Similarly, greater transitions from one ROI to another than vice versa would result in greater net transitions. The net transitions emphasize a circularity in the fixation behavior where subjects would exhibit more return fixations to already looked-at objects. It would show subjects’ scant use of working memory by not encoding already fixated-on objects into memory. In contrast, purposeful use of working memory would require fewer refixations, more forward fixations, and greater net transitions. We used the metric mean relative net transitions as the ratio of net transition by total transition for both action epochs of a trial. Hence, we computed the mean relative net transition for the preparation and execution epochs per trial. With maximum relative net transitions equal to 1, we can expect an equal number of the net and total transitions between ROIs, indicating forward saccades between ROIs and showing strictly action-oriented gaze guidance to ROIs, e.g., from previous to current in the action preparation phase. Such a gaze behavior would show no inter78 |3.5 RESULTS
mediate refixations on the ROIs. Conversely, the minimum relative net transition of 0would indicate an equal number of saccades from and to ROI pairs. Searching for target objects by fixating back and forth over the ROIs or switching gaze between ROI pairs to monitor the task progression would lead to 0net transitions. Hence, relative net transitions would help distinguish the function of saccades between ROI pairs as locating or searching target objects or monitoring the task. We tested the hypothesis of whether task complexity influenced the net transitions of fixations during action preparation and execution. From the previous analysis, we inferred search behavior from the proportion of fixations during the action preparation and execution phase. We hypothesized that in HARD tasks, subjects would show lower net transitions during the action preparation epochs due to increased search behavior and even lower net transitions during the execution epoch due to increased monitoring and searching for the next action targets. The data shows subtle differences in the relative net transitions for the trial types and the action epochs Fig 3.3C. In the EASY trials, the mean relative net transition for the preparation epochs was 0.69SD =±0.10 and for the execution epochs 0.68 ±0.08. In the HARD trials, the mean relative net transition for the preparation epochs was 0.59 ±0.11, whereas it was 0.64 ±0.10 in the execution epochs. We used a linear mixed effects model to fit the mean relative net transitions in a trial per action epoch and trial type. Table 3.3 details the model coefficients. A significant effect of trial type showed that HARD trials had lower relative net transitions than EASY trials. No effect of epoch type showed that overall preparation epochs did not differ from execution epochs regarding relative net transitions. However, there was a significant interaction between trial type and epoch type, showing that preparation epochs in HARD trials had significantly lower relative net transitions. Thus, increased task complexity was associated with decreased relative net transitions. Notably, with increased task complexity, the reduction in net transitions was greater in the action preparation epochs and less in the execution epochs. The analysis above lends further evidence to differential gaze guidance due to task complexity, specifically in the action preparation epochs. The higher relative net transitions in the EASY trials suggest saccades were made towards the immediate action-relevant objects with little requirement for search. The lower net transitions in the HARD tasks suggest significantly higher gaze switching to and from ROIs in search of task-relevant objects, especially during action preparation. The influence of task complexity was not likewise pronounced in the action execution epochs, indicating eye movements were engaged in a similar function of monitoring CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |79
Table 3.3: Model estimates of differences in relative net transition w.r.t task type, epoch type Model: NetTransitions ∼1 + trial_type ∗epoch_type +(1 + trial_type ∗epoch_type|Subject) Estimate 95% CI t p-value trial type -0.06 [-0.09, -0.06] -8.85 0.001*** (HARD - EASY) epoch type -0.02 [-0.04, 0.01] -1.76 0.08 (PREPARATION - EXECUTION) trial type : epoch type -0.06 [-0.08, -0.03] -4.58 < .001*** the task progression, resulting in back-and-forth saccades between the previous, current, and next target ROIs in the action sequence. Hence, with increased task complexity, there is an increase in visual search for action-relevant items during action preparation and a negligible influence during action execution. Latency of Fixations on Task-relevant ROIs To further study the temporal aspects of gaze guidance behavior during action preparation and execution, we were interested in the latency of the first fixations to the eight ROIs. As a measure of latency, we used the median time to first fixate on an ROI in each action epoch per trial. We hypothesized that if fixations are made to the task-relevant regions per a set cognitive scheme, we would see a strict temporal window where these fixations occur in the preparation and execution epochs. A cognitive schema that involves first searching for the targets, directing the hand towards the target, locating the next target, etc., would lead to first fixations on the ’other objects’, then on the current target object, and then onto the next target object. On the other hand, if fixations do not occur based on a set cognitive schema behind it, the occurrence of fixations would not follow a temporal structure. Furthermore, to make the latency comparable across the different durations of preparation and execution epochs, we normalized it by the duration of the given epoch. Hence, the normalized latency of the first fixation to the eight ROI provides a temporal measure of when a schema is initiated during the preparation and execution of an action. Here, we describe the latency of the first fixations on the eight ROIs normalized 80 |3.5 RESULTS
by the epoch duration. Fig 3.4A shows the distributions of the first fixations on the eight ROIs in the preparation phase. The distributions show a flow of fixations from the previous target object to other objects and shelves and then to the current target objects. Importantly, the first fixations on the current target object were at 0.42(SD =±0.13) in EASY trials and at 0.46(SD =±0.11) in the HARD trials. As shown in Fig 3.4A, the distribution of the fixation latencies shows a structured sequence of attending to task-relevant ROIs. Notably, the first fixation on the current target object occurs only after 40% of the preparation epoch duration has elapsed. Figure 3.4: Latency of gaze targets during action preparation and execution.A, B: Distributions of median time to first fixation on the 8 ROIs for the action preparation and execution and trial types. The latency (abscissa) is normalized by the length of the preparation and execution epoch, respectively. The ordinate is sorted in ascending order of mean latency from epoch start. We modeled the latency of first fixations dependent on ROI type and trial type using linear mixed models. We used a linear regression model to compared the latency of the first fixation on the current target object with the rest of the ROIs and their interactions with respect to the task complexity. The model shows the first fixations go from the previous target object and shelf to the other shelves and objects, onto the next targets, and finally to the current target object. Furthermore, task complexity significantly affected the latency of first fixations on the previous target object and shelf, as well as the other and next target object, compared to the current target object. In HARD tasks, the first fixations on previous targets and the ’other’ ROIs occurred significantly earlier than in the EASY tasks. Taken together, irrespective of task complexity, the action preparation epochs show a systematic progression of fixations from one ROI to another. This structured temporal sequence of fixCHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |81
ations with a defined temporal window shows the fixations on the action-relevant ROIs are not incidental and are part of a set cognitive schema. Moreover, given the task complexity, the temporal profiles of these locating/searching fixations can change and occur earlier. Most importantly, the first fixations on the current target object occur last in sequence after more than 40% of the epoch time has elapsed, lending strong evidence to the just-in-time manner in which gaze is allocated to the immediate action requirements. We similarly analyzed the latency of first fixations on the ROIs during the action execution epochs. Fig 3.4B depicts the distribution of the latency of first fixations on the different ROIs during the action execution epoch. Here, we see a similar temporal structure of first fixations that go from ’other’ objects and shelves to previous targets and have similar latencies on the next target object and the current target shelf. Here again, the fixations on the current target shelf occur after half of the epoch duration has elapsed. As before, we used a linear mixed model to model the latency of the first fixations dependent on the ROI type and trial type. We used a linear regression model to compared the latency of the first fixation on the current target object with the rest of the ROIs and their interactions with respect to the task complexity. To summarize, irrespective of the task complexity, there is a systematic sequence of first fixations on the eight ROIs during action execution. The early fixations on the previous target object and shelf indicate these fixations are related to monitoring/checking the current action execution with respect to the previously handled object. Similarly, later fixations on the next target object and shelf indicate fixations used for planning the next action even before the concurrent action has been completed. Importantly, the increased task complexity did not affect the latency of first fixations on the taskrelevant ROIs during action execution. Here again, the first fixation on the current target shelf, i.e., where the target object is placed, is fixated towards the end after 50% of the time has elapsed, strongly indicating a just-in-time nature of the actionrelevant fixations. Taken together, the latency of first fixations on the ROIs revealed the activation of specific cognitive schemas and their relative time of occurrence in the action preparation and execution epochs. Task complexity affected the timing of first fixations during action preparation and not in the execution phase. Furthermore, first fixations on the current targets occur after almost half the epoch time has elapsed, showing a just-in-time manner of gaze allocation. The overall results suggest that action sequences are predominantly decided 82 |3.5 RESULTS
in a just-in-time manner, and task complexity influences distinct aspects of actionoriented fixations. The average fixation profile over the ROIs showed that fixations are driven from previous to current and next actions. Task complexity affected the proportion of fixations early during the preparation period to facilitate the visual search for targets and during the execution period for increased action monitoring. The sequential gaze transitions between ROI pairs further demonstrated an increased visual search associated with task complexity in the action preparation period. In contrast, task complexity did not influence gaze guidance during the execution phase. Task complexity also affected the latency of the first fixations on the ROIs in the preparation epochs but not in the execution epochs. Importantly, the latency analysis emphasized the just-in-time nature of gaze guidance, where fixations are made to the immediate action targets after half the action epoch (both preparation and execution) has passed. Finally, increased action preparation durations were associated with increased visual search. However, the prolonged visual search was not associated with optimal action choices. Instead, participants chose arbitrary spatial heuristics. Thus, gaze behavior was primarily explained by lowlevel cognitive schemas of searching, locating, directing, and task monitoring in the context of immediate actions. 3.6 Discussion In this study, we examined the mechanisms of gaze control in tasks that are not over-learned and routine. Building on the work of Land et al., 1999, where eye movements were studied while preparing tea, the tasks in our study are novel because they do not have an inherent action sequence associated with them. Our experimental setup provided a way to capture oculomotor behavior for tasks that did not have a strict action sequence. By studying the dispersion and timing of fixations on the previous, current, and next action-relevant regions, we show a structured sequence of fixations that may be classified based on previous work. Moreover, we add to the current body of research by providing a data-driven approach to studying locating, directing fixations, guiding fixations, and checking fixations. Our analysis and findings generalize the occurrence of these fixations explicitly to the action sequence and completely disregard object identity. Land et al., 1999 proposed various functions of eye movements, from locating, directing, guiding, and checking during a tea-making task. Our study also shows CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |83
the general signature of these fixations time-locked to the action onset and how they are impacted by cognitive load. Importantly, our study shows the occurrence of visual search early in the action preparation phase with increased search due to increased task complexity. The directing and guiding fixations occurred just before the action onset and were unaffected by task complexity. The checking/monitoring fixations occurred concurrently with the action execution as fixations on the previous action targets. Considering the latency and the proportional change in fixations on the task-relevant regions, our study further corroborates the cognitive schemas that sequentially support the preparation and execution of manual action. Land and Furneaux, 1997 additionally elaborated on the schemas that direct both eye movements and motor actions. They proposed a chaining mechanism by which one action leads to another, where the eyes supply constant information to the motor system via a buffer. In our analysis, the temporal structure of the fixations on the ROIs lends further evidence of such a cognitive buffer that facilitates moving from one action to another where eye movements predominantly search and locate the regions necessary for immediate action. Interestingly, the increased cognitive load did not affect this simple mechanism of cognitive arbitration. Instead, we see a pronounced reliance on explicit visual search instead of encoding objects in memory. This is exemplified by the gaze transition analysis, where complex task conditions result in lower net transitions, indicating more refixations and fewer forward fixations. Thus, we show that cognitive schemas are activated to inform the current action with scant use of the visual working memory. A notable feature demonstrated in the study is the just-in-time nature of actionoriented fixations. Here, the fixations on the target object occurred shortly before the action onset. Determining the causal relationship between these fixations and manual action is difficult in this study. A question remains: are gaze and arm movements planned together, or does the hand move only after the gaze has acquired the target’s location? We did not explicitly study hand movement trajectory or velocity relative to the first fixation on the target object. However, the timing of the fixations on the different regions of interest indicates that once the target is acquired halfway through the preparation period, the remaining time is allocated to guiding the hand and body toward the target object. This is evidenced by a dip in the proportion of fixations on non-target objects and a sharp increase in fixations on the target object (Fig 3.3). This observation points to a model of greedy observer/actor model that uses visual information and plans only the immediate actions. Triesch et al., 2003 have emphasized the task-specific nature of eye movements where infor84 |3.6 DISCUSSION
mation from fixations is used for computations only just in time to solve the current goal. Our study shows that a just-in-time strategy of gaze guidance consists of a search, locate, and act routine, and cognitive planning is primarily associated with the immediate action in the sequence. Even though participants in our study spent a lot of time planning for the subsequent action under complex conditions, they were strikingly sub-optimal compared to the model solutions. There is considerable psychological evidence that supports the idea that anticipated cognitive demand plays a significant role in behavioral decision-making. Further, all things being equal, humans tend to take actions that reduce the cost of physical and cognitive effort (Kool et al., 2010). Studies in human performance in reasoning tasks and combinatorial optimization problems (MacGregor & Chu, 2011) have revealed that humans solve these tasks in a self-paced manner rather than being dependent on the size of the problem. Pizlo and Li, 2005 found that subjects do not implicitly search the problem space to plan the moves without executing them. Hence, longer task durations do not necessarily lead to shorter solution paths. Instead, they showed that humans break the problem down into component sub-tasks, giving rise to a low-complexity relationship between the problem size and time to solution. Further, they showed that subjects use simpler heuristics to decide the next action. This is in line with Gigerenzer and Goldstein, 1996, who posit that humans use simplifying strategies to gather and search information that results in sub-optimal behavior. In more complex tasks, we observed higher planning durations and greater object displacements, but the action execution duration was similar to less complex tasks. Thus, participants in our study do not necessarily demonstrate a lack of planning. Instead, the results show they divided the overall task into sub-tasks and planned which object needed to be handled and where it should be moved. We observed a clear and general spatial heuristic employed by the participants. They used a strategy of arbitrarily organizing the objects from left to right and top to bottom and predominantly leaving out the last row and column. Such a simplistic strategy might be an artifact of reading direction habit (Afsari et al., 2016; Faghihi & Vaid, 2023) or asymmetric attentional biases (Ossandón et al., 2014). Also, a case could be made that moving objects to the last row was especially associated with intrinsic motor costs. Burlingham et al., 2024 recently studied a phenomenon of motor "laziness" in naturalistic gaze selection, where eye and head motor effort contribute toward gaze guidance behavior. In this regard, rolling the eye and head downward would result in a motor effort to scan the shelf’s state. This motor cost CHAPTER 3. HOW IS GAZE CONTROLLED DURING ACTION SEQUENCES? |85
proaches offer a vital addition to traditional, highly controlled laboratory studies (Ingram & Wolpert, 2011), allowing researchers to investigate behavior in contexts that more closely mirror everyday experience. Several influential studies have pioneered this direction by enabling participants to engage in self-initiated strategizing and unconstrained interactions with real-world objects Ballard et al., 1992; Ballard et al., 1995; Ballard and Hayhoe, 2009; Hayhoe and Ballard, 2014; Land et al., 1999. While these studies provide foundational insights, their analyses tend to be primarily descriptive, examining eye, head, and hand movements in isolation without a formal model of their joint coordination. In addition, most rely on manual annotation of gaze and action from video, making the methods difficult to scale or generalize across participants. Hence, there are also methodological concerns about how we study visuomotor coordination. In this study, we ask: How do the eyes, head, and hands coordinate during natural, self-initiated movements, and how is this coordination influenced by the spatial structure of the task space? Advances in virtual reality (VR) now allow researchers to explore such questions in controlled yet immersive settings, using wearable eye and body trackers to record kinematic signals with high precision (Keshava et al., 2020; Keshava et al., 2023; König et al., 2021; Nolte, Vidal De Palol, et al., 2024). Here, we used VR to capture rich, three-dimensional (3D) data streams from the eyes, head, and hands as participants performed naturalistic sorting tasks. Unlike earlier studies, our approach uses automated preprocessing and computational modeling to uncover how the eye-head-hand effectors coordinate. Our approach allows us to scale from descriptive accounts to a more formal, computational ethology perspective (Mobbs et al., 2021) of visuomotor coordination. Natural behavior is highly dynamic and complex, especially when multiple effectors operate in 3D space. Understanding coordination mechanisms thus requires methods that can reduce this complexity while preserving its latent structure. To this end, we used low-dimensional representations to capture the shared variance of eye-head-hand movements. High-dimensional kinematic data, reflecting translations and rotations of various effectors, were projected into a latent subspace using Principal Component Analysis (PCA), which identifies patterns of coordinated variation across signals. This low-dimensional representation provides a compact summary of visuomotor behavior and helps characterize the set of states the body occupies during natural action (Bialek, 2022; Sanger, 2000; Santello et al., 1998). This approach enables a principled analysis of complex, high-dimensional behavior, offering insights into the underlying structure and dynamics of sensorimotor 92 |4.2 INTRODUCTION
coordination. We explored human volunteers’ movement trajectories while sorting objects on a 2m wide and 2m high life-size shelf in VR. They performed pick-and-place actions iteratively until a cued goal was achieved. Crucially, the participants were free to move and generate their own movements, unconstrained by time, allowing them to act as naturally as possible in the virtual environment. The tasks were generic (involving pick-and-place actions) but novel (requiring planning to sort the objects according to varying rules), making the findings relevant to visuomotor coordination in a proactive and natural context. We focussed on how eye-head-hand coordination unfolded in this context by studying the movements in a low-dimensional subspace and cross-validating how these representations predict the timing and location of upcoming action events. By addressing these questions, our study extends previous work on natural behavior through a formal and generalizable framework that captures the dynamics of visuomotor coordination. 4.3 Methods Participants 27 participants (18 females, mean age = 23.9 ± 4.6 years) were recruited from the University of Osnabrück and the University of Applied Sciences Osnabrück. Participants had normal or corrected-to-normal vision and no history of neurological or psychological impairments. They either received a monetary reward of C 7.50 or one participation credit per hour. Before each experimental session, subjects gave their informed consent in writing. They also filled out a questionnaire regarding their medical history to ascertain they did not suffer from any disorder/impairments that could affect them in the virtual environment. Once the consent was obtained, we briefed them on the experimental setup and task. The Ethics Committee of the University of Osnabrück approved the study (Ethik-37/2019). Apparatus & Procedure For the experiment, we used an HTC Vive Pro Eye head-mounted display (HMD)(110° field of view, 90Hz, resolution 1080 x 1200 px per eye) with a builtCHAPTER 4. HOW DO EYES AND HANDS COORDINATE TO PLAN ACTIONS? |93
Figure 4.1: Experimental Setup. A. In a virtual environment, 27 participants sorted 16 different objects based on two features (color or shape) while we measured their eye and body movements. The objects were randomly presented on a 5x5 shelf at the beginning of each trial and were cued to sort objects by shape and/or color. All participants performed 24 trials in total with no time limit. B. We measured the position and the direction vectors of the head in a global reference frame and the eye direction and hand position in the head reference frame. C. To understand the coupling of the different sources of data from the head, eye, and hand in different reference frames, we performed dimensionality reduction analysis on the measured signals relative to action critical events of grasp onset (when the hand touches the object to pickup and the VR controller trigger is pressed) and grasp offset (when controller trigger is released and the hand drops the object on the desired shelf after object displacement). For each object displacement, we segmented the data to 1s before and after the action events. For each subject, the data matrix was composed of G×M×T, where Gdenotes object displacements, Mdenotes the 12 features (four data streams in x, y, z coordinates), and Tdenotes each time point relative to the action events. We applied Principal Component Analysis at each time point to the G×Mmatrix and reduced its dimensionality to G×2. The resulting 2D representation across time would then elucidate the contribution of the original data sources to the principal components and their explained variance. in Tobii eye-tracker1with 120 Hz sampling rate. With their right hand, participants used an HTC Vive controller to manipulate the objects during the experiment. The HTC Vive Lighthouse tracking system provided positional and rotational tracking and was calibrated for a 4m x 4m space. We used the 5-point calibration function provided by the manufacturer to calibrate the gaze parameters. To ensure the calibration error was less than 1◦, we performed a 5-point validation after each cali1https://enterprise.vive.com/us/product/vive-pro-eye-office/ 94 |4.3 METHODS
bration. Due to the study design, which allowed many natural body movements, the eye tracker was calibrated repeatedly during the experiment after every 3 trials. Furthermore, subjects were fitted with HTC Vive trackers on both ankles, both elbows and one on the midriff. The body trackers were also calibrated subsequently to give a reliable pose estimation using inverse kinematics of the subject in the virtual environment. We designed the experiment using the Unity3D2version 2019.4.4f1 and SteamVR game engine and controlled the eye-tracking data recording using HTC VIVE Eye-Tracking SDK SRanipal3(v1.1.0.1). The experimental setup consisted of 16 different objects placed on a shelf of a 5x5 grid. The objects were differentiated based on two features: color and shape. We used four high-contrast colors (red, blue, green, and yellow) and four 3D shapes (cube, sphere, pyramid, and cylinder). The objects had an average height of 20cm and a width of 20cm. The shelf was designed with a height and width of 2m with 5 rows and columns of equal height, width, and depth. Participants were presented with a display board on the right side of the shelf where the trial instructions were displayed. Subjects were also presented with a red buzzer that they could use to end the trial once they finished the task. The physical dimensions of the setup are illustrated in Figure 4.1A. The horizontal eccentricity of the shelves from the center of the shelf extended to 89.2 cm in the left and the right directions. Similarly, the vertical eccentricity of the shelf extended to 89.2 cm in the up-down direction from the center-most point of the shelf. This ensured that the task setup was symmetrical in both the horizontal and the vertical directions. The objects on the top-most shelves were placed at a height of 190cm and on the bottom-most shelves at a height of 13cm from the ground level. Experimental Task Subjects performed two practice trials where they familiarized themselves with handling the VR controller and the experimental setup. In these practice trials, they were free to explore the virtual environment and displace the objects. After the practice trials, subjects were asked to sort objects based on the one and/or two features of the object. Each subject performed 24 trials in total, with each trial instruction (as listed below) randomly presented twice throughout the experiment. 2Unity, www.unity.com 3SRanipal, developer.vive.com/resources/vive-sense/sdk/vive-eye-tracking-sdksranipal/ CHAPTER 4. HOW DO EYES AND HANDS COORDINATE TO PLAN ACTIONS? |95
The experimental setup is illustrated in Figure 4.1A. The trial instructions were as follows: ◦Sort objects so that each row has the same shape or is empty ◦Sort objects so that each row has all unique shapes or is empty ◦Sort objects so that each row has the same color or is empty ◦Sort objects so that each row has all unique colors or is empty ◦Sort objects so that each column has the same shape or is empty ◦Sort objects so that each column has all unique shapes or is empty ◦Sort objects so that each column has the same color or is empty ◦Sort objects so that each column has all unique colors or is empty ◦Sort objects so that each row has all the unique colors and all the unique shapes once ◦Sort objects so that each column has all the unique colors and all the unique shapes once ◦Sort objects so that each row and column has each of the four colors once. ◦Sort objects so that each row and column has each of the four shapes once. Data pre-processing We measured the position and direction 3D vectors of the head and hand in global reference frame. The eye direction and position vectors were measured in the head reference frame. Figure 4.1B illustrates the vector representations of the eye, head, and hand position and orientation vectors. At the outset, we downsampled the data to 40 Hz. The sections below explain the steps we took to process the raw data and arrive at the magnitude of translation and rotation movements made by the eye, head, and hand. Gaze data Using the eye-in-head 3D gaze direction vector for the cyclopean eye, we calculated the gaze angles for the horizontal θhand vertical θvdirections. vector samples were sorted by their timestamps. The 3D gaze direction vector of each sample is 96 |4.3 METHODS
represented in (x, y, z)coordinates as a unit vector that defines the direction of the gaze in VR world space coordinates. In VR, the xcoordinate corresponds to the left-right direction, yin the up-down direction, zin the forward-backward direction. From the eye-in-head gaze direction vectors, we computed the horizontal (θh) and the vertical (θv) gaze angles in degrees using the following formula: θh=180 π·arctan x z(4.1) θv=180 π·arctan y z(4.2) HMD data From the HMD we obtained the global head position vector and the head direction vector. Using the 3D head direction vector, angular orientation in the horizontal θh and vertical θvdirections using the following formula: θh=180 π·arctan x z(4.3) θv=180 π·arctan y z(4.4) At the beginning of each trial, subjects were asked to stand still facing the shelf for 3s at a set location. Thus, we could estimate the initial position of the head from the average position of the HMD in the 3s period. Subsequently, we calculated the deviations of the HMD from the initial position arrive at the magnitude and direction of translation movements. For each time point after the initial 3s period, we subtracted the initial position from the current position in 3D coordinates. Hence, we could estimate the degree of head translation movements in the left-right and up-down directions. Hand controller data Subjects used the trigger button of the HTC Vive controller to virtually grasp the objects on the shelf and displace them to other locations. In the data, the trigger was recorded as a boolean, which was set to TRUE when subjects pressed the trigger button on the hand controller to initiate an object displacement and was reset CHAPTER 4. HOW DO EYES AND HANDS COORDINATE TO PLAN ACTIONS? |97
to FALSE when the trigger button was released, and the object was placed on the shelf. Using the position of the controller in the world space, we determined the locations from the shelf where a grasp was initiated and ended. We also removed trials where the controller data showed implausible locations in the 3D space. These faulty data were attributed to the loss of tracking during the experiment. Next, we removed grasping periods where the origin and final locations of the objects were the same on the shelf. Next, we calculated the angular position of the hand with respect to the head positions in each trial. As described above, the x-coordinate corresponds to the leftright direction, y in the up-down direction, z in the forward-backward direction. Using the 3D Cartesian coordinates (x, y, z)of the controller position (Hand(x,y,z)), HMD position (Head(x,y,z)) and HMD direction unit vector ( ~ Head(x,y,z)) in world space, we calculated the horizontal and the vertical angular position of the hand with respect to the head. The horizontal (φh) and vertical (φh) angular position of the hand was calculated as follows: φh=180 π·arccos Hand(x,z)−Head(x,z) kHand(x,z)−Head(x,z)k·~ Head(x,z)!(4.5) φv=180 π·arccos Hand(y,z)−Head(y,z) kHand(y,z)−Head(y,z)k·~ Head(y,z)!(4.6) where kk denotes the norm of the vector, and (·)denotes the dot product. Data Analysis After pre-processing, we were left with data from 27 subjects with 554 trials in total and Mean = 20.51, SD =±2.20 trials per subject. Furthermore, we had eyetracking data corresponding to 5664 grasping actions with 12.85, SD =±1.91 object displacements per trial and subject. Principal Components Analysis Given the naturalistic setting of our experimental setup, with complex movements of the head, eye, and hand, we wanted to understand the contributions and coupling of each of these effectors while performing object pickup and dropoff actions. We first epoched the data using the grasp onset and offset triggers. We selected a time window of 1s before and after the trigger. Thus, each epoch consisted of 3D data 98 |4.3 METHODS
from head position, head rotation, eye direction, and hand position at each time point spaced 0.025s apart. Each subject’s feature matrix then consisted of a 3D matrix of 80 time-points, each with 12 features, for each grasp. We explored the time-wise low-dimensional representation of this matrix using Principal Component Analysis. For each subject, the input matrix for PCA was composed of G×M×T, where Gdenotes the grasping epoch, Mdenotes the 12 features (four data streams in x, y, z coordinates) and Tdenotes each time point relative to the action events. For each time point in T, we standardized G×Mmatrix to zero mean and unit standard deviation. This was done to make sure each feature contributed equally to the PCA. We then applied PCA to the G×Mmatrix at each time point. We used the ‘sklearn.preprocessing.PCA‘ python package to compute the eigenvalues and principal components. We then used the ‘transform()‘ function to reduce the dimensionality of the matrix from G×Mto G×2. Thus, for each subject, we obtained T×2eigenvalues and 2×M×Tcoefficients of the eigenvectors. The eigenvalues were used to show the explained variance of the 2 principal components (PCs) across time. The coefficients provided the loadings of each Mfeature onto the PCs. Hence, for each subject, we could ascertain the evolution of the contribution of the features in the PCA subspace across time. Cosine Similarity To understand the contributions of the individual features to the PCA subspace, we calculated the directional similarity of the variables. For each eye-head, eyehand, and head-hand pair, we computed the cosine similarity of their horizontal and vertical components as follows: cos θ=~a ·~ b k~ak·k~ bk(4.7) Using the values of cos θwe could determine if the coefficients of the eigenvector were aligned in the same direction (cos θ= 1), opposite directions (cos θ=−1) or orthogonal (cos θ= 0) across the different time points. CHAPTER 4. HOW DO EYES AND HANDS COORDINATE TO PLAN ACTIONS? |99
Generalization & Predictions To explore the generalizability of the PCA subspace across subjects. We used Nfold cross-validation, where Ncorresponds to the number of subjects. For each fold, we divided the dataset into train and test sets, where data from N−1subjects was used for training, and the left-out dataset from 1subject was used for testing. During training, we applied PCA to A×Mmatrix at each time point of the grasping epochs, where Adenotes all grasping epochs across N−1subjects. After applying PCA, we obtained the A×2reduced matrix in PCA subspace. Using the A×2 matrix, we trained a kernel-based support vector machine (SVM) to predict the location of the action at each time point. We used the sklearn.svm.SVC python package for training and prediction. We used the default parameters for the model and did not perform any hyper-parameter optimization. During testing, we standardized the G×Mmatrix of the left-out subject, applied the PCA weights from the training set to reduce its dimensionality to G×2, and recorded the mean prediction accuracy of the trained model on the left-out data. This analysis was repeated for the grasp offset events. In this manner, we could test the generalizability and predictive power of the PCA subspace across time for grasp onset and offset events. 4.4 Results Twenty-seven healthy human volunteers performed an object-sorting task in VR on a 2m high and 2m wide life-size shelf. Participants performed 24 trials each. In each trial, participants sorted 16 randomly presented objects on the shelf based on a cued task and iteratively performed pick and place actions until the cued sorting was achieved. Figure 4.1A illustrates the experimental setup in VR. The experimental task consisted of sorting objects by their features. Each object was differentiated by color and shape. As the objects were randomly presented to the participants, they made reaching movements to different locations in space, such as reaching for objects closer to their feet or to the top of their heads. As participants performed these tasks, we simultaneously recorded their eye, head, and hand movement trajectories in 3D. In doing so, we recorded ecologically valid movements within a large spatial context. For this study, we utilized four streams of continuous data. Namely, the head position of the participants, the unit direction in which the head was oriented, the 100 |4.4 RESULTS
cyclopean eye (average of left and right eye) unit direction within the head, and the hand position in world coordinates. These 3D movement vectors were represented in (x, y, z)coordinates (Figure 4.1B). In all, we analysed data from 27 participants, 5664 grasp actions with 209.77 ±3.41 grasps per subject. At the outset, we down-sampled the data to 40 Hz so that the continuous samples were equally spaced at 25ms. The down-sampling also smoothed the movement trajectories. We segmented the data streams relative to the onset of a grasp. We chose the time window from 1s before the grasp onset to 1s after. We similarly epoched the data relative to the grasp offset with the same window lengths. This allowed us to capture the eye, head, and hand movements while reaching the object, grasping it, and guiding it to a desired location. For both grasp onset and grasp offset events, we had 80 time points starting from -1s to 1s relative to an action, where time point 0 marked the onset or offset of said action. We then used PCA to reduce the overall dimensionality of the segmented data at each time point. The original data comprised of 12 dimensions (four data streams, each with 3d coordinates). We performed a time-wise PCA on each of the 80 time points to illustrate how the low-dimensional representation of the ongoing visuomotor coordination morphed relative to an action event and how the individual eye, head, and hand orientations contributed to these low-dimensional representations across time. Figure 4.1C depicts the data segmentation steps and PCA analysis approach in this study. Complexity of Natural Behavior Participants performed complex movements, which led to simultaneous translation and rotation of the head. To understand the range of movements made by the participants, we transformed the position and rotation vectors into 3D Euler angles (see section 4.3). Figure 4.2A shows the joint distribution of the initial position of the participants at the beginning of each trial in the horizontal and vertical planes. All participants started the trial from a fixed position in VR. The mean initial position of the participants in the horizontal plane was 0.0m, SD =±0.03 and in the vertical plane was 1.63m±0.07. The variance in the vertical direction corresponds to the variance in the height of the participants. From the initial positions, we calculated the translation-based deviations of the participants during the trials. Figure 4.2B shows the bi-variate distribution of the horizontal and vertical translation head movements CHAPTER 4. HOW DO EYES AND HANDS COORDINATE TO PLAN ACTIONS? |101
Figure 4.4: PCA analysis of 3D eye direction, head direction, head position, and hand position vectors. A. Temporal profile of eigenvalues of the 12 principal components (PCs) 1s before and after the grasp onset at time 0s. The solid lines denote the mean explained variance ratio over the 27 participants by each PC, and the shaded region depicts the 95% confidence interval. B. Bivariate distribution of the contributions of the horizontal and vertical eye direction vectors on first two PCs across the time points and participants, and their respective marginal distributions. C. The joint contribution of the horizontal and vertical head direction vectors on the two PCs. E. The contribution of the horizontal and vertical hand position vectors on the two PCs across the different time points. As seen from these distributions, the horizontal components of the eye, head, and hand contributed primarily to PC2, and the vertical components contributed to PC1. 4.4.1 Cosine Similarity Analysis In the above section, we performed time-wise PCA on action-critical events of grasp onset and grasp offset using head position, head direction, eye direction, and hand position vectors. The results showed a larger explained variance of the first two PCs, where this ratio was maximum just before the action events. PCA provides a subspace where the data’s variance is maximized across the principal components. Using a vector similarity analysis, we can further explore relationships between the original variables within this subspace and identify which variables contribute similarly to the explained variance. Hence, exploration of the subspace can lead to a deeper understanding of the data structure, such as revealing groups of variables that might be part of the same underlying process or phenomenon. 108 |4.4 RESULTS
Table 4.2: Mean factor loadings of the original variables on the first two principal components during grasp offset variable PC1 PC2 Mean (SD) Mean (SD) Head Direction X 0.07 (0.06) 0.57 (0.11) Head Direction Y 0.42 (0.03) 0.08 (0.07) Head Direction Z 0.12 (0.08) 0.14 (0.11) Head Position X 0.06 (0.05) 0.19 (0.12) Head Position Y 0.30 (0.08) 0.08 (0.07) Head Position Z 0.17 (0.10) 0.14 (0.10) Eye Direction X 0.06 (0.05) 0.36 (0.13) Eye Direction Y 0.43 (0.02) 0.07 (0.06) Eye Direction Z 0.34 (0.09) 0.13 (0.12) Hand Position X 0.07 (0.06) 0.50 (0.11) Hand Position Y 0.40 (0.03) 0.06 (0.04) Hand Position Z 0.39 (0.04) 0.08 (0.07) To quantify the source of the increase in the explained variance ratio close to the action onset and offset events, we used cosine similarity analysis. Cosine similarity provides a measure of the correlation between the factor loadings in the 2D PCA subspace. Since PCA aims to identify patterns of similarity and differences across variables by transforming them into PCs based on their covariance, using cosine similarity to analyze further the orientation and correlation of variables in this transformed space complements the goals of PCA. It helps in understanding the structure and relationships between variables beyond mere dimensionality reduction. When the PC loadings point in the same direction, they indicate a high positive correlation with each other. Similarly, when they point in opposite directions, they indicate a high negative correlation. When they are orthogonal to each other, they indicate no correlation at all. Hence, using cosine similarity, we could ascertain the evolution of the correlation between the horizontal and vector components of the movement vectors before and after the action-critical events. In grasp onset epochs, we calculated the cosine similarity between the x(horizontal) and y(vertical) components of the eye, head, and hand direction vector for each time point across subjects. Figure 4.5A shows the mean cosine similarity of the horizontal and vertical components of the head and eye direction vectors and the standard deviation. At time point 0, the horizontal and vertical components had a mean similarity of 0.99 ±0.002 and 0.99 ±0.001, respectively. Between eye and hand factor loadings (Figure 4.5B), we observed a mean similarity of 0.99 ±0.002 CHAPTER 4. HOW DO EYES AND HANDS COORDINATE TO PLAN ACTIONS? |109
Figure 4.5: Cosine similarity between head, eye, and hand direction vectors for grasp onset and grasp offset events. A. Cosine similarity between the horizontal and vertical components of head and eye direction 3D vectors during grasp onset events. The solid lines denote the mean cosine similarity across 27 participants, and the shaded region denotes the standard error or mean. A cosine similarity value of 1 indicates a perfect correlation between the vectors in the PCA subspace, whereas a value of -1 denotes a perfect negative correlation. A value of 0 indicates the vectors are not related to each other. B. Cosine similarity between the horizontal and vertical components of eye direction and hand position 3D vectors during grasp onset events. C. Cosine similarity between the horizontal and vertical components of head and eye direction vectors during grasp offset events. D. Cosine similarity between the horizontal and vertical components of eye direction and hand position vectors during grasp offset events. in the horizontal direction and 0.99 ±0.001 in the vertical direction at grasp onset. Between head and hand factor loadings (Figure 4.5C), we observed a remarkable consistency throughout the action epoch where the similarity of the horizontal components was 0.99 ±0.0003 and for vertical components 0.99 ±0.001 at time 0s. Taken together, the vertical components of the eye, head, and hand factor loadings were well correlated throughout the grasp onset epoch. However, the horizontal components were correlated around 0.5s before the grasp onset, and this correlation reduced drastically shortly after the grasp event was triggered. Throughout the time course, the hand and head vectors varied in the same direction, both in the horizontal and the vertical planes, and exhibited a strong correlation. We repeated the above analysis for the grasp offset events. First, we calculated the cosine similarity between the horizontal and vertical components of the eye, head, and hand factor loadings in the PCA subspace. Figure 4.5D illustrates the average similarity of eye-head loadings over subjects across the different time points. 110 |4.4 RESULTS
At time point 0, the average cosine similarity was 0.98 ±0.004 in the horizontal direction and 0.95 ±0.01 in the vertical direction. Between eye-hand factor loadings (Figure 4.5E), the horizontal and vertical components were perfectly aligned at time 0s and showed an average similarity of 0.97 ±0.008 and 0.98 ±0.003, respectively. Between head-hand factor loadings (Figure 4.5E), we observed a striking similarity as before, where the horizontal components had a mean similarity measure of 1.00±0.003 and 0.99±0.003 for the vertical components. Taken together, there was a strong coupling between the vertical components of the eye, head, and hand. In contrast, the horizontal components were aligned in the same direction briefly before the grasp offset. Further, the head direction and the hand position vectors covaried in the same direction and showed very high correlations. The exploration of the PCA subspace with the loadings of the original variables showed interesting aspects of visuomotor coordination. Namely, the vertical components of the eye, head, and hand vectors were almost perfectly aligned in the low-dimensional space. The horizontal component of the eye direction vector, on the other hand, was only briefly oriented in the same direction as the head direction and hand position vectors at about 0.5s before the grasp onset and offset. This window of complete alignment of the vectors also coincides with the increase in the explained variance ratio of the first two PCs before the action onset. Crucially, the head direction and the hand position vectors tracked in complete unison throughout the action epochs. Thus, the similarity analysis of the effectors in the PCA subspace showed distinct coordination mechanisms for the horizontal and vertical components of the visuomotor system. 4.4.2 Generalization and Predictive Accuracy of the Lowdimensional Space To further expound on the generalizability and predictive power of these lowdimensional structures, we predicted the location of the action at each time point based on the PCA-transformed data. For each time point t, we pooled the data from N−1subjects and standardized it to zero mean and unit standard deviation. We then reduced the dimensionality of the data at each time point into a 2D space. For each time point, we trained a kernel-based support vector machine (SVM) to classify the location of the upcoming action. We used leave-one-subject cross-validation to arrive at the validated test accuracy. For each time point, we CHAPTER 4. HOW DO EYES AND HANDS COORDINATE TO PLAN ACTIONS? |111
Figure 4.6: Accuracy of 2D PCA subspace in predicting the action location at each time point of the action epoch. A. The training (blue trace) and test/validation (red trace) accuracy for -1s before and 1s after the grasp onset events. Solid trace represents mean accuracy, and the shaded region denotes 95% CI of mean. B. The training and test accuracy for 1s before and 1s after grasp offset events. C. The linear fit over peak accuracy for the rows and columns of the predicted action location. Each dot represents the peak accuracy of a subject for grasp onset (red) and grasp offset(blue) events per row and column of the predicted action location on the shelf. standardized the test data and transformed it using the PCA weights from the training data. We then computed the prediction accuracy of the trained SVM on the test data at that time-point. We performed the above steps until each of the 27 subjects’ data was used as test data for all time points ranging from 1s before and after the grasp onset. We repeated this analysis for the grasp offset events as well. Our analysis provided an aggregated prediction accuracy of the PCA-transformed training and test datasets. Thus, using cross-validation, we could generalize the information encapsulated in the PCA subspace and ensure the prediction accuracy was not affected by the peculiarities of single subjects. The prediction of the object pickup action location across the grasp onset epoch provided a greater understanding of the evolution of the PCA subspace as seen in Figure 4.6A. At the 1s before grasp onset, the prediction accuracy on the test data is low at Mean = 0.20,95%CI = [0.17,0.23]. At time 0, the accuracy increases to 0.62 [0.52, 0.72]. At 1s after the grasp onset, the prediction accuracy is reduced to 0.12 [0.10, 0.13]. This shows that the PCA subspace encodes more relevant 112 |4.4 RESULTS
information about the upcoming action close to the action onset. The maximum predictive accuracy was not at the moment of grasp at time 0, but slightly earlier at −0.20s, 95%CI = [−0.32,−0.09]. This indicates that maximum information about the eye, head, and hand coordination is available just in time for the action. Similarly, the prediction of the object dropoff action locations across the grasp offset epoch is shown in Figure 4.6B. At 1s before the grasp offset, the predictive accuracy on the test data is 0.15,95%CI = [0.13,0.17]. At time point 0, the accuracy increases to 0.38,95%CI = [0.30,0.46] and decreases to 0.08,95%CI = [0.06,0.09] at 1s after the grasp offset event. Again, the maximum predictive accuracy was at time point −0.32s, 95%CI = [−0.42s, −0.23s]before the grasp offset. This further indicates the directional components of the eye, head, and hand vectors are aligned for action just in time. With the above analysis, we generalized the predictive accuracy of the 2D PCA subspace. We determined the predictive information encapsulated in the subspace and the timing of the best prediction. For both grasp onset and grasp offset, the PCA space’s predictive accuracy increased with the approaching event and decreased after. Moreover, maximum accuracy is achieved just-in-time of the action event. This can be accounted for by the large explained variance ratio of the 2D subspace and the corresponding high correlations between the eye, head, and hand orientation vectors at that moment. Sources of Variance in Prediction During grasp onset and especially grasp offset events, the validation accuracy is lower than the training accuracy and shows high variance. We looked into the source of this variance by separating the prediction accuracy by the shelf location of the action. We used a linear model to test the hypothesis that the row and column location of the upcoming action affected the predictive accuracy of the PCA-transformed 3D movement vectors. As shown in Figure 4.6C, the maximum accuracy increased linearly from the top to the bottom row in grasp onset events (β= 0.02,95%CI = [0.01,0.03], t(633.12) = 4.55, p =<0.001). The maximum accuracy was not significantly affected during grasping objects from left to right columns (β= 0.00001,95%CI = [−0.01,0.01], t(632.91) = −0.02, p = 0.987). In the case of grasp offset events, the peak accuracy decreased significantly from top to bottom rows (β=−0.11,95%CI = [−0.13,−0.10], t(555.10) = −11.96, p =<0.001). In contrast, the peak accuracy decreased significantly but with a smaller slope across the CHAPTER 4. HOW DO EYES AND HANDS COORDINATE TO PLAN ACTIONS? |113
shelf columns (β=−0.03,95%CI = [−0.05,−0.02], t(555.15) = −3.95, p =<0.001). Hence, the predictive accuracy of the PCA-transformed features was highly dependent on where the action would happen. The results show that the accuracy for grasp onset events did not differ dramatically across the different rows and columns of the shelf. However, there was a pronounced decrease in accuracy for grasp offset events, especially for lower shelf regions. The maximum accuracy shows that the best information to predict the upcoming action is available just-in-time. However, the quality of this signal is lower for grasp offset events and significantly decreases for drop-off actions on the lower and rightward areas of the shelf, indicating a low correlation between the effectors for these locations. 4.5 Discussion Our study explored the low-dimensional representations of natural visuomotor coordination. Subjects exhibited complex translation and rotation movements with their eyes, head, and right hand by making reaching movements to pick up and place objects on a life-size shelf in VR. We applied a time-wise PCA on the position and orientation vectors of the different effectors to capture the explained variance at each time point relative to grasp onset and offset events. Our analysis showed the complex system composed of the eye, head, and hand could be well described in a 2D PCA subspace. The PCA subspace showed an increase in the explained variance ratio at grasping events (onset and offset), where more than 60% of the variance is accounted for by the first two eigenvectors. Our analysis demonstrates a dynamic and distinct coupling of the horizontal and vertical components of effectors just in time for the upcoming action. Furthermore, this coupling showed high predictive accuracy of the target location of the forthcoming action. However, the accuracy was substantially influenced by the horizontal and vertical target location on the shelf . Hence, our study demonstrates a dynamic coupling and decoupling of eye, head, and hand movement vectors that have distinct features in the horizontal and vertical axes and are dependent on the reach target location. 114 |4.5 DISCUSSION
Methodological Considerations Eye-hand or eye-head coordination is usually studied under constrained settings. In most cases, the horizontal and vertical positions or directions of the eye, head, and hand are extracted and directly correlated. In natural behavior, the complexity of the system does not afford such simplistic measures, as variables can interact with each other in non-obvious ways. Our study explored the latent relationships between the eye, head, and hand orientations in a low-dimensional space, helping to understand the underlying structure and relationships in the data relative to the action-critical events. We aimed to capture both the explained variance and the evolution of the variable loadings across time. Our analysis of the cosine similarity of the original variables in the PCA subspace revealed strong associations between the horizontal and vertical components of the eye, head, and hand vectors. PCA transforms variables into principal components based on their variances and covariances. When using cosine similarity to assess relationships between original variables based on their loadings, it’s crucial to note that PCA mainly focuses on explaining variance, not necessarily revealing direct correlations between variables. It’s important to clarify that cosine similarity measures the angle between two vectors and is a measure of orientation similarity rather than a direct measure of statistical correlation in the traditional sense (Pearson’s correlation). Cosine similarity is less sensitive to the magnitude of vectors and focuses on their direction, which could be a limitation when the scale or variability of the original variables is relevant to their interpretation. In order to avoid drawing improper inferences, we plotted the distribution of the variable loadings on the first two PCs. We can confirm that the horizontal and vertical component loadings of the eye, head, and hand had similar magnitudes on the PCs. Hence, by comparing the directionality of loadings, we could identify which variables share similar directional influences on the PCs, indicating underlying correlations that are not immediately obvious from the PCA results alone. Finally, given the present study’s naturalistic setting, various noise sources could affect the findings. The noise source could be the eye or body trackers, which could exhibit errors due to slippage (Niehorster et al., 2020) or calibration errors (Ehinger et al., 2019). We calibrated the trackers after every three trials to mitigate such errors. Moreover, as participants performed the task while wearing the VR head-mounted display, the head movements could have been cumbersome when picking up objects from the lower shelf locations. We did not direct participants to CHAPTER 4. HOW DO EYES AND HANDS COORDINATE TO PLAN ACTIONS? |115
move in any one particular manner and asked them to make movements that were comfortable for them. Nonetheless, the insights offered by our study open the door to further experimental replications to validate our findings. Independence of Movement Vectors in the Horizontal and Vertical Axes The cosine similarity analysis showed that eye, head, and hand coordination has distinct correlations between horizontal and vertical axes. While the vertical components of the eye, head, and hand were highly correlated during the investigated time windows, the eye-head and eye-hand horizontal components were correlated close to the action onset but were otherwise uncorrelated. In non-human primates, there is evidence of distinct areas in the premotor neural circuits for independent generation and control of saccadic movements in the horizontal and vertical axes (Moschovakis et al., 1996). Similar to the saccadic system, horizontal and vertical head movements are controlled by distinct circuits in the cerebellum (Shaikh et al., 2004) and the brainstem (Crawford et al., 2003). Similarly, premotor areas, primary motor cortex, and parietal cortex are differentially tuned for direction, with some studies indicating specialized subpopulations for horizontal versus vertical hand movements (Georgopoulos et al., 1982; Hadjidimitrakis et al., 2022; Kalaska et al., 1983; Lacquaniti, 1995). Hence, there is ample evidence from primate studies showing distinct circuits that operate in a coordinated manner to produce smooth, multi-directional movements. For example, horizontal and vertical control centers are simultaneously active when making oblique hand or head movements, allowing for complex movement patterns (Crawford et al., 2011). Thus, our study offers preliminary behavioral evidence of independent but coordinated eye, head, and hand movement control in the horizontal and vertical axes in humans. Synergistic Coupling of the Effectors The vertical components of the eye, head, and hand varied in the same direction and contributed significantly to the overall variance explained. Conversely, the horizontal component of the eye direction vector aligned with the head and hand horizontal components shortly before the action onset. Previous studies have shown that head and eye movements are generated by simultaneously receiving the same 116 |4.5 DISCUSSION
motor commands (Land, 1992). Here, head movements are necessary to center gaze in the orbits. Our data implies that head movements facilitate the vertical directionality of the eye, and the eye completes the last leg of the operation by making horizontal adjustments. Hence, the vertical components of the eye and head direction vectors were highly correlated, whereas the horizontal components showed a significant degree of independence. The horizontal and vertical components of the head and hand were aligned throughout the grasping epoch. Studies have corroborated this strong coupling between the head and arm with respect to the eye in unrestrained macaque monkeys (Arora et al., 2019) and humans (Smeets et al., 1996). These studies argue that head movements facilitate foveation on the target to guide the final stages of object manipulation, leading to large correlations between the head and hand movements. In a similar vein, Pelz et al., 2001 showed a strong linkage between the head and hand movement trajectories, while the eye has a synergistic relationship instead of an obligatory one with them. Possibly this strong head-arm coupling results from learned motor behaviors during feeding where the head and hand orientations are coordinated to bring food to the mouth (Hadjidimitrakis, 2020). Head and hand coordination is not commonly studied, our results suggest the strong coupling between the two must be a consequence of a common neural code that drive this behavior. Differences between Grasp Onset and Grasp Offset To validate the generalization of the PCA subspace, we predicted the grasp onset and offset locations by transforming the test data with the PCA weights of the training data. Since we cross-validated the model with unseen data, the test accuracy was expected to be lower than the training accuracy. During the grasp onset time window, the prediction accuracy across time on the test data is similar to the training data. However, the prediction accuracy for grasp offset events on the test data is considerably lower. This indicates idiosyncratic coordination to guide objects and drop them to desired locations, leading to lower explained variance and low correlation. Consequently, the 2D PCA subspace is likely less generalizable for grasp offset events. The predictive accuracy also had a larger variance for grasp offset events. Upon inspection, we found that the source of this variance is the location of the upcomCHAPTER 4. HOW DO EYES AND HANDS COORDINATE TO PLAN ACTIONS? |117
the semantic information was not readily available. This effect was enhanced when subjects were asked to perform tool-specific movements instead of a generic action of lifting the tool by the handle. The authors, hence, concluded that eye movements are used to actively infer the appropriate usage of the tools from their mechanical properties. In the study, the tools were presented as images on a screen, and participants pantomimed lifting or using the tool. While the study revealed valuable insights into anticipatory gaze control, a question remains if these results are part of natural cognition and can be reproduced in more realistic environments. Herbort and Butz, 2011 further investigated the interaction of habitual and goaldirected processes that affect grasp selection while interacting with everyday objects. They presented objects in different orientations and showed that grasp selection depended on the overarching goal of the movement sequence dependent on the object’s orientation. Belardinelli, Stepper, et al., 2016 further showed that fixations have an anticipatory preference for the region where the index finger is placed. Consequently, the location of fixations is predictive of both proximal goals of manual planning and task-related distal goals. When studying anticipatory behaviors corresponding to an action, there must be a distinction of symbolized or pantomimed vs. real actions. Króliczak et al., 2007 showed brain areas typically involved in real actions are not driven by pantomimed actions and that pantomimed grasps do not activate the object-related regions within the ventral stream. Similarly, Hermsdörfer et al., 2012 showed a weak correlation between the hand trajectories for pantomimed and actual tool interaction. These studies indicate that the realism of sensory and tactile feedback while acting (e.g., a grasp) can be an essential factor when studying anticipatory behavioral control. As research steadily moves towards a more ecological view of cognitive processing with bodily actions and interactions with the environment, there is a greater need to understand behavioral differences induced by varying degrees of realism of action affordances. Chalmers and Ferko, 2008 posit that defining levels of realism is necessary to achieve a one-to-one mapping of an experience in the virtual environment with the same experience in the real environment. Moreover, for research purposes such one-to-one mapping is necessary to avoid misrepresenting the real environment. In virtual reality (VR), realistic actions can be produced by different kinds of interaction methods. Using interfaces such as VR controllers, ego-centric visual feedback of a hand can be simulated. These interfaces usually consist of hand-held devices that are tracked in space and through which different actions are 124 |5.2 INTRODUCTION
controlled by pressing buttons. One advantage of controller-based VR interaction is the possibility of haptic feedback. However, a disadvantage is that the actual hand posture while holding the controller does not necessarily correspond to the user’s simulated hand as they perform the action. Conversely, camera-based interaction interfaces such as LeapMotion, capture the real-time movements of the user’s hand and finger gestures, like wrap grasp or pinch grasp, to control different actions in the environment. These interfaces give the user a realistic simulation of their finer hand and finger movements, while they cannot give direct haptic feedback of the gripped object. Consequently, the chosen method of interaction in VR can afford different levels of realism and could elicit different behavioral responses. In the present study, we investigated anticipatory gaze control pertaining to tool interactions in two different experiments. We were interested in the extent to which the action affordance and the environment modified active inference processes exhibited in gaze behavior. In experiment-I, subjects performed the experiment in a low realism environment and interacted with the tool models using a VR controller, which mimicked grasp in the virtual environment by pulling the index finger. In experiment-II, subjects were immersed in a high realism setting where they interacted with the tools using LeapMotion, which required natural hand and finger movements. Thus, the action affordance appeared closer to the real-world. Furthermore, in both experiments, participants interacted with 3D models of tools by lifting or using them in VR. For the stimuli set, we used familiar or unfamiliar tools as described in Belardinelli, Barabas, et al., 2016. Additionally, to differentiate between proximal and distal goal planning, we manipulated the spatial orientation of the tool handle so that they were either congruent or incongruent to the subjects’ handedness. With this experimental design, we investigate the influence of task, tool familiarity, the spatial orientation of the tool, and, notably, the impact of the action affordance on anticipatory gaze behavior. 5.3 Methods Experimental Task Subjects were seated in a virtual environment where they had to interact with the presented tool by either lifting or pretending its use. The time course of the trials CHAPTER 5. DOES GAZE ANTICIPATE FINE HAND MOVEMENTS? |125
Figure 5.1: Experimental Task. In two virtual environments participants interacted with tools in two ways (LIFT, USE). The tools were categorized based on familiarity (FAMILIAR, UNFAMILIAR) and presented to the participants in two orientations (HANDLE LEFT, HANDLE RIGHT). The two virtual environments differed based on the mode of interaction and perceived realism, wherein in one experiment, subjects’ hand movements were rendered virtually using the HTC-VIVE controllers. In the other experiment, the hands were rendered using LeapMotion, allowing finer hand and finger movements. Panel A shows the timeline of a trial. Panel B shows a subject in real-life performing the task in the two experiments. Panel C shows the differences in realism in the two experiments; TOP panels correspond to experiment with the controllers, the USE and LIFT conditions for an UNFAMILIAR and FAMILIAR tool, respectively with the tool handles presented in two different orientations. BOTTOM panels illustrate the three different conditions in a more realistic environment with LeapMotion as the interaction method. Panel D Familiar tools, from top-left: screwdriver, spatula, wrench, fork, paintbrush, trowel. Panel E Unfamiliar tools, from top-left: spokewrench, palette knife, daisy grubber, lemon zester, flower cutter, fish scaler. is illustrated in Figure 5.1A. At the start of a trial, subjects saw the cued task for 2 sec after which the cue disappeared, and a tool appeared on the virtual table. Subjects were given 3 sec to view the tool, after which there was a beep (go cue) which indicated that they could start manipulating the tool based on the cued task. Subjects were seated in a virtual environment where they had to interact with the presented tool by either lifting or pretending its use. After interacting with the tool, subjects pressed a button on the table to start the next trial. Participants For experiment-I with the HTC Vive controller’s interaction method, we recruited 18 participants ( 14 females, mean age=23.68, SD=4.05 years). For experiment-II 126 |5.3 METHODS
with the interaction method of LeapMotion, we recruited 30 participants (14 female, mean age=22.7, SD=2.42 years). All participants were recruited from the University of Osnabrück and the University of Applied Sciences Osnabrück. Participants had a normal or corrected-to-normal vision and no history of neurological or psychological impairments. All of the participants were right-handed. They either received a monetary reward of C10 or one participation credit per hour. Before each experimental session, subjects gave their informed consent in writing. They also filled out a questionnaire regarding their medical history to ascertain they did not suffer from any disorder/impairments which could affect them in the virtual environment. Once we obtained their informed consent, we briefed them on the experimental setup and task. Experimental Design and Procedure The two experiments differed based on the realism of the action affordance and the environment. Figure 5.1B illustrates the physical setup of the participants for the two experiments. In experiment-I, subjects interacted with the tool models using the HTC Vive VR controllers. While in experiment II, subjects’ hand movements were captured by LeapMotion. Figure 5.1C illustrates two exemplar trials from the experiments. We used a 2x2x2 experimental design for both experiments, with factors task, tool familiarity, and handle orientation. Factor task had two levels: LIFT and USE. In the LIFT conditions, we instructed subjects to lift the tool to their eye level and place it back on the table. In the USE task, they had to pantomime using the tool to the best of their knowledge. Factor familiarity had two levels, FAMILIAR and UNFAMILIAR, which corresponded to tools either being everyday familiar tools or tools that are not seen in everyday contexts and are unfamiliar. The factor handle orientation corresponded to the tool handle, which was presented to the participants either on the LEFT or the RIGHT. Both experiments had 144 trials per participant, with an equal number of trials corresponding to the three factors. Subjects performed the trials over six blocks of 24 trials each. We measured the eye movements and hand movements simultaneously while subjects performed the experiment. We calibrated the eyetrackers at the beginning of each block and ensured that the calibration error was less than 1 degree of the visual angle. At the beginning of the experiment, subjects performed three practice trials with a hammer to familiarize themselves with the CHAPTER 5. DOES GAZE ANTICIPATE FINE HAND MOVEMENTS? |127
experimental setup and the interaction method. Each experiment session lasted for approximately an hour. After that, subjects filled out a questionnaire to indicate their familiarity with the 12 tools used in the experiment. They responded to the questionnaire based on a scale of 5-point Likert-like scale where 1 corresponded to “I have never used it or heard about it,” and 5 referred to “I see it every day or every week.” Experimental Stimuli The experimental setup consisted of a virtual table that mimicked the table in the real world. The table’s height, width, and length were 86cm, 80cm, and 80cm, respectively. In experiment-I, subjects were present in a bare room with grey walls and constant illumination. They sat before a light grey table, with a dark grey button on their right side to indicate the end of the trial. Similarly, in experiment-II, subjects were present in a more immersive, realistic room. They sat in front of a wooden workbench with the exact dimensions of the real-world table and a buzzer on the right to indicate the end of a trial. We displayed the task (USE or LIFT) over the desk 2m away from the participants for both experiments. For both experiments, we used the tool models as presented in Belardinelli, Barabas, et al., 2016. Six of the tools were categorized as familiar (Figure 5.1D) and the other six as unfamiliar (Figure 5.1E). We further created bounding box colliders that encapsulated the tools to capture the gaze position on the tool models. The mean length of the bounding box was 34.04cm (SD=5.73), mean breadth=7.60cm (SD=3.68) and mean height= 4.17cm (SD=2.13). To determine the tool effector and tool handle regions of interest, we halved the length bounding box colliders from the center of the tool and took one half as the effector and the other half as the handle. This way we refrained from making arbitrary-sized regions-of-interest for the different tool models. Apparatus For both experiments, we used an HTC Vive head-mounted display (HMD)(110◦ field of view, 90Hz, resolution 1080 x 1200 px per eye) with a built-in Tobii eyetracker 1. The HTC Vive Lighthouse tracking system provided positional and rota1https://enterprise.vive.com/us/product/vive-pro-eye-office/ 128 |5.3 METHODS
tional tracking and was calibrated for a 4m x 4m space. For calibration of the gaze parameters, we used 5-point calibration function provided by the manufacturer. To make sure the calibration error was less than 1◦, we performed a 5-point validation after each calibration. Due to the nature of the experiments, which allowed a lot of natural head movements, the eye tracker was calibrated repeatedly during the experiment after each block of 48 trials. We designed the experiment using the Unity3D game engine 2(v2019.2.14f1) and controlled the eye-tracking data recording using HTC VIVE Eye Tracking SDK SRanipal3(v1.1.0.1) For experiment-I, we used HTC Vive controller4(version 2.5) to interact with the tool. The controller in the virtual environment was rendered as a gloved hand. When participants pulled the trigger button of the controller with their right index finger, their right virtual hand made a power grasp action. To interact with the tools, subjects pulled the trigger button of the controller over the virtual tools and the rendered hand grasped the tool handle. Similarly, in experiment-II, we used LeapMotion5(version 4.4.0) to render the hand in the virtual environment. Here, subjects could see the finer hand and finger movements of their real-world movements rendered in the virtual environment. When participants made a grasping action with their hand over the virtual tool handle, the rendered hand grasped the tool handle in the virtual environment. 5.3.1 Data pre-processing Data Rejection For both experiments, we rejected trials based on two criteria. Firstly, we rejected trials where the hand position was not recorded. Secondly, we rejected trials where the gaze position and direction vectors were not recorded or recorded as invalid. For experiment-I, we rejected 29.8% (SD=±9.6) of trials over 18 participants. For experiment-II, we rejected a mean of 35.5% (SD=±19.24) of trials over 30 participants. There was a greater rejection rate in experiment-II as the LeapMotion camera lost hand-tracking more often. 2Unity, www.unity.com 3SRanipal, developer.vive.com/resources/vive-sense/sdk/vive-eye-tracking-sdksranipal/ 4SteamVR, https://valvesoftware.github.io/steamvr_unity_plugin/articles/Quickstart.html 5LeapMotion Unity modules, https://developer.leapmotion.com/unity CHAPTER 5. DOES GAZE ANTICIPATE FINE HAND MOVEMENTS? |129
Gaze Data As a first step, using eye-in-head 3D gaze direction vector for the cyclopean eye we calculated the gaze angles in degrees for the horizontal θhand vertical θvdirections. All of the gaze data was sorted by the timestamps of the collected gaze samples. The 3D gaze normals are represented as (x, y, z)a unit vector that defines the direction of the gaze in VR world coordinates. In our setup, the x coordinate corresponds to the left-right direction, y in the up-down direction, z in the forward-backward direction. The formulas used for computing the gaze angles are as follows: θh=180 πarctan x z θv=180 πarctan y z Next, we calculated the angular velocity of the eye in both the horizontal and vertical coordinates by taking a first difference of the angular velocity and dividing by the difference between the timestamp of the samples using the formula below: ωh= ∆θh/∆t ωv= ∆θv/∆t Finally, we calculated the magnitude of the angular velocity (ω) at every timestamp from the horizontal and vertical components using: ω=qω2 h+ω2 v To classify the fixation and saccade-based samples, we used an adaptive threshold method for saccade detection described by Voloh et al., 2020. We selected an initial saccade velocity threshold θ0of 200 ◦/sec. All eye movement samples with an angular velocity of less than θ0were used to compute a new threshold θ1.θ1was three times the median absolute deviation of the selected samples. If the difference between θ1and θ0was less than 1 ◦/sec θ1was selected as the saccade threshold else, θ1was used as the new saccade threshold and the above process was repeated. This was done until the difference between θnand θn+1 was less than or equal to 1 ◦/sec. This way we arrived at the cluster of samples that belonged to fixations and the rest were classified as saccades. After this, we calculated the duration of the fixations and removed those fixations that had a duration less than 50 ms or were larger than 3.5 times the median 130 |5.3 METHODS
absolute deviation of the fixation duration. For further data analysis, we only considered those fixations that were positioned on the 3D tool models. We further categorized the fixations based on their position on the tool, i.e., whether they were located on the effector or handle of the tool. Data Analysis Odds of Fixations in favor of tool effector After cleaning the dataset, we were left with 2174 trials from 18 subjects in experiment-I and 3633 trials from 30 subjects in experiment-II. For both experiments, we analysed the fixations in the 3 second period from the tool presentation till the go cue. For the two experiments, we modeled the linear relationship of the log of odds of fixations on the effector of the tools and the task cue (LIFT, USE), the familiarity of the tool (FAMILIAR, UNFAMILIAR), and orientation of the handle (LEFT, RIGHT) and the experiment interaction method (CONTROLLER, LEAP MOTION). All within-subject factors were also modeled with random intercepts and slopes for each subject. We used effect coding (Schad et al., 2020) to construct the design matrix for the linear model, where we coded the categorical variables LIFT, FAMILIAR, RIGHT, CONTROLLER to -0.5 and USE, UNFAMILIAR, LEFT, LEAPMOTION to 0.5. This way, we could directly interpret the regression coefficients as main effects. The model fit was performed using restricted maximum likelihood (REML) estimation (Corbeil & Searle, 1976) using the lme4 package (v1.1-26) in R 3.6.1. We used the L-BFGS-B optimizer to find the best fit using 10000 iterations. Using the Satterthwaite method (Luke, 2017), we approximated degrees of freedom of the fixed effects. For both experiments, the Wilkinson notation (Wilkinson & Rogers, 1973) of the model was: logp(fixations on effector) p(fixations on handle)∼ 1 + task ∗familiarity∗ handle_orientation ∗interation_method +(1 + task ∗familiarity ∗handle_orientation|Subject) CHAPTER 5. DOES GAZE ANTICIPATE FINE HAND MOVEMENTS? |131
As we used effects coding, we can directly compare the regression coefficients of the two models. The fixed-effect regression coefficients of the two models would describe the differences in log-odds of fixations in favor of tool effector for the categorical variables task, familiarity, and handle orientation and the effect of the interaction method used in the experiment groups. Spatial bias of fixations on the tools In this analysis, we wanted to assess the effects of task, tool familiarity, and handle orientation on the eccentricity of fixations on the tools. To do this, we studied the fixations from the time when the tool was visible on the table (3s from the start of trial) till the go cue indicated when subjects could start manipulating the tool. We divided this 3s period into 20 equal bins of 150ms each. For each trial and time bin, we calculated the median distance of the fixations from the tool center. Next, we normalized the distance with the length of the tool so that we could compare the fixation eccentricity across different tools. To find the time-points where there were significant differences for the 3 conditions and their interactions, we used the cluster permutation method. Here, we use the t-statistic as a test statistic for each time-bin, where t is defined as: t=√N∗x σ and, x is the mean difference between conditions, and σis the standard deviation of the mean and N is the number of subjects. We used a threshold for t at 2.14 which corresponds to the t-value at which the p-value is 0.05 in the t-distribution. We first found the time-bins where the t-value was greater than the threshold. Then, we computed the sum of the t-values for these clustered time-bins which gave a single value that represented the mass of the cluster. Next, to assess the significance of the cluster, we permuted all the time-bins across trials and subjects and computed the t-values and cluster mass for 1000 different permutations. This gave us the null distribution over which we compared the cluster mass shown by the real data. We considered the significant clusters to have a p-value less than 0.05. In the results, we report the range of the significant time-bins for the 3 different conditions and their interactions and the corresponding p-values. 132 |5.3 METHODS
Subjective Rating of Tool Familiarity To validate that the tool categorization in our experiment design aligned with the subjective assessments of participants, we analysed the questionnaire completed by participants at the end of each experiment. We calculated the mean subjective rating of familiar and unfamiliar tools for both experiment groups. We performed a mixed-ANOVA with familiarity as a within-subject factor, the experiment group as the between-subject factor and the subjective rating of the tool as the dependent variable. Learning Effects In order to quantify the learning effects on fixation patterns due to the repeated presentation of familiar and unfamiliar tools, we computed the relative change in fixations on tool effector. For each participant, we computed the mean proportion of fixations on tool effector in the first five and last five familiar trials and unfamiliar trials. We, subsequently, calculated the percent change (C) from early trials to late trials for each tool familiarity using the following formula: C= 100 ∗Xf−Xi Xi where, Xfdenotes the mean proportion of fixations on tool effector in the last five trials, and Xidenotes the mean proportion of fixations on tool effector in first five trials for a given subject. To statistically, assess the differences between the experimental groups and the tool familiarity, we performed a mixed-ANOVA with C as the dependent variable, too familiarity as a within-subject factor and experiment interaction method as between subject factor. Difference between experiment groups To assess if participants allocated attention to the overall environment and the tools in the two experiments, we calculated the percentage of fixations allocated to the environment vs. the tool during the 3s viewing period in each trial of the two experiments. To statistically assess the difference in the mean percentage of fixations allocated to the tools vs environment, we performed a mixed-ANOVA with fixation location (tools vs environment) as a within-subject factor, the two experiments as a between-subject factor, and the percentage of fixations as the dependent variable. CHAPTER 5. DOES GAZE ANTICIPATE FINE HAND MOVEMENTS? |133
nous attraction of gaze in dyadic interactions. Attention, perception & psychophysics,86(8), 2761–2777. Hoffmann, J. (2003). Anticipatory behavioral control. In MV Butz, O Sigaud, & P Gérard (Eds.), Anticipatory behavior in adaptive learning systems: Foundations, theories, and systems (pp. 44–65). Springer Berlin Heidelberg. Holleman, GA, Hooge, ITC, Kemner, C, & Hessels, RS. (2021). The reality of “real-life” neuroscience: A commentary on shamay-tsoory and mendelsohn (2019). Perspectives on psychological science: a journal of the Association for Psychological Science,16(2), 461–465. Hoppe, D, & Rothkopf, CA. (2019). Multi-step planning of eye movements in visual search. Scientific reports,9(1), 144. Hoshi, E, & Tanji, J. (2007). Distinctions between dorsal and ventral premotor areas: Anatomical connectivity and functional properties. Current opinion in neurobiology,17(2), 234–242. Hu, Y, & Goodale, MA. (2000). Grasping after a delay shifts size-scaling from absolute to relative metrics. Journal of cognitive neuroscience, 12(5), 856–868. Huang, A, Derakhshan, S, Madrid-Carvajal, J, Nezami, FN, Wächter, MA, Pipa, G, & König, P. (2024). Enhancing safety in autonomous vehicles: The impact of auditory and visual warning signals on driver behavior and situational awareness. Vehicles,6(3), 1613–1636. Ingram, JN, & Wolpert, DM. (2011). Naturalistic approaches to sensorimotor control. Progress in brain research,191, 3–29. Itaguchi, Y. (2021). Size perception bias and Reach-to-Grasp kinematics: An exploratory study on the virtual hand with a consumer immersive Virtual-Reality device. Frontiers in Virtual Reality,2. Itti, L, & Koch, C. (2000). A saliency-based search mechanism for overt and covert shifts of visual attention. Vision research,40(10-12), 1489– 1506. 236 |BIBLIOGRAPHY
Itti, L, & Koch, C. (2001). Computational modelling of visual attention. Nature reviews. Neuroscience,2(3), 194–203. Jeannerod, M. (1981). Intersegmental coordination during reaching at natural visual objects. Attention and Performance, 153–169. Jeannerod, M. (1986). The formation of finger grip during prehension. a cortically mediated visuomotor pattern. Behavioural brain research, 19(2), 99–116. Jeannerod, M. (1988). The neural and behavioural organization of goaldirected movements. Oxford psychology series, No. 15.,283. Jeannerod, M. (2006). Motor cognition: What actions tell the self. OUP Oxford. Johansson, RS, Westling, G, Bäckström, A, & Flanagan, JR. (2001). Eyehand coordination in object manipulation. The Journal of neuroscience: the official journal of the Society for Neuroscience,21(17), 6917–6932. Johansson, RS, & Flanagan, JR. (2009). Coding and use of tactile signals from the fingertips in object manipulation tasks. Nature reviews. Neuroscience,10(5), 345–359. Johnson, SH, & Grafton, ST. (2003). From ‘acting on’ to ‘acting with’: The functional anatomy of object-oriented action schemata. Progress in brain research (pp. 127–139). Elsevier. Johnson-Frey, SH. (2004). The neural bases of complex tool use in humans. Trends in cognitive sciences,8(2), 71–78. Jonas, E, & Kording, KP. (2017). Could a neuroscientist understand a microprocessor? PLoS computational biology,13(1), e1005268. Kahnt, T, Heinzle, J, Park, SQ, & Haynes, JD. (2011). Decoding different roles for vmPFC and dlPFC in multi-attribute decision making. NeuroImage,56(2), 709–715. Kalaska, JF, Caminiti, R, & Georgopoulos, AP. (1983). Cortical mechanisms related to the direction of two-dimensional arm movements: Relations BIBLIOGRAPHY |237
in parietal area 5 and comparison with motor cortex. Experimental brain research,51(2), 247–260. Kanan, C, Ray, NA, Bseiso, DNF, Hsiao, JH, & Cottrell, GW. (2014). Predicting an observer’s task using multi-fixation pattern analysis. Proceedings of the Symposium on Eye Tracking Research and Applications, 287–290. Kennerley, SW, Walton, ME, Behrens, TEJ, Buckley, MJ, & Rushworth, MFS. (2006). Optimal decision making and the anterior cingulate cortex. Nature neuroscience,9(7), 940–947. Keshava, A, Aumeistere, A, Izdebski, K, & Konig, P. (2020). Decoding task from oculomotor behavior in virtual reality. ACM Symposium on Eye Tracking Research and Applications, (Article 30), 1–5. Keshava, A, Gottschewsky, N, Balle, S, Nezami, FN, Schüler, T, & König, P. (2023). Action affordance affects proximal and distal goal-oriented planning. The European journal of neuroscience,57(9), 1546–1560. Keshava, A, Nezami, FN, Neumann, H, Izdebski, K, Schüler, T, & König, P. (2024). Just-in-time: Gaze guidance in natural behavior. PLoS computational biology,20(10), e1012529. Kirsh, D. (1994). On distinguishing epistemic from pragmatic action. Cognitive science,18(4), 513–549. Klever, L, Voudouris, D, Fiehler, K, & Billino, J. (2019). Age effects on sensorimotor predictions: What drives increased tactile suppression during reaching? Journal of vision,19(9), 9. Kok, P, Mostert, P, & de Lange, FP. (2017). Prior expectations induce prestimulus sensory templates. Proceedings of the National Academy of Sciences of the United States of America,114(39), 10473–10478. König, P, Wilming, N, Kietzmann, TC, Ossandón, JP, et al. (2016). Eye movements as a window to cognitive processes. Journal of eye movement research,9(5). 238 |BIBLIOGRAPHY
König, P, Melnik, A, Goeke, C, Gert, AL, König, SU, & Kietzmann, TC. (2018). Embodied cognition. 2018 6th International Conference on Brain-Computer Interface (BCI), 1–4. König, P, Wilming, N, Kaspar, K, Nagel, SK, & Onat, S. (2013). Predictions in the light of your own action repertoire as a general computational principle. The behavioral and brain sciences,36(3), 219–220. König, SU, Clay, V, Nolte, D, Duesberg, L, Kuske, N, & König, P. (2019). Learning of spatial properties of a large-scale virtual city with an interactive map. Frontiers in human neuroscience,13, 240. König, SU, Keshava, A, Clay, V, Rittershofer, K, Kuske, N, & König, P. (2021). Embodied spatial knowledge acquisition in immersive virtual reality: Comparison to map exploration. Frontiers in Virtual Reality, 2. Kool, W, McGuire, JT, Rosen, ZB, & Botvinick, MM. (2010). Decision making and the avoidance of cognitive demand. Journal of experimental psychology. General,139(4), 665–682. Krakauer, JW, Ghazanfar, AA, Gomez-Marin, A, MacIver, MA, & Poeppel, D. (2017). Neuroscience needs behavior: Correcting a reductionist bias. Neuron,93(3), 480–490. Kriegeskorte, N, Mur, M, & Bandettini, P. (2008). Representational similarity analysis - connecting the branches of systems neuroscience. Frontiers in systems neuroscience,2, 4. Król, ME, & Król, M. (2020). The right look for the job: Decoding cognitive processes involved in the task from spatial eye-movement patterns. Psychological research,84(1), 245–258. Króliczak, G, Cavina-Pratesi, C, Goodman, DA, & Culham, JC. (2007). What does the brain do when you fake it? an FMRI study of pantomimed and real grasping. Journal of neurophysiology,97(3), 2410–2422. Kubicek, C, Jovanovic, B, & Schwarzer, G. (2017). The relation between crawling and 9-month-old infants’ visual prediction abilities in spatial BIBLIOGRAPHY |239
object processing. Journal of experimental child psychology,158, 64–76. Lacquaniti, F. (1995). Representing spatial information for limb movement: Role of area in the. Cerebral Cortex Scp/Oct,5, 391–409. Ladouce, S, Donaldson, DI, Dudchenko, PA, & Ietswaart, M. (2016). Understanding minds in Real-World environments: Toward a mobile cognition approach. Frontiers in human neuroscience,10, 694. Lakoff, G, & Johnson, M. (2008). Metaphors we live by. Land, M, Mennie, N, & Rusted, J. (1999). The roles of vision and eye movements in the control of activities of daily living. Perception,28(11), 1311–1328. Land, MF. (1992). Predictable eye-head coordination during driving. Nature, 359(6393), 318–320. Land, MF, & Furneaux, S. (1997). The knowledge base of the oculomotor system. Philosophical transactions of the Royal Society of London. Series B, Biological sciences,352(1358), 1231–1239. Land, MF, & Hayhoe, M. (2001). In what ways do eye movements contribute to everyday activities? Vision research,41(25-26), 3559–3565. Land, MF, & McLeod, P. (2000). From eye movements to actions: How batsmen hit the ball. Nature neuroscience,3(12), 1340–1345. Land, MF. (2006). Eye movements and the control of actions in everyday life. Progress in retinal and eye research,25(3), 296–324. Levitis, DA, Lidicker, WZ, & Freund, G. (2009). Behavioural biologists don’t agree on what constitutes behaviour. Animal behaviour,78(1), 103– 110. Lohmann, J, Belardinelli, A, & Butz, MV. (2019). Hands ahead in mind and motion: Active inference in peripersonal hand space. Vision (Basel, Switzerland),3(2). 240 |BIBLIOGRAPHY
Luke, SG. (2017). Evaluating significance in linear mixed-effects models in R. Behavior research methods,49(4), 1494–1502. MacGregor, JN, & Chu, Y. (2011). Human performance on the traveling salesman and related problems: A review. The Journal of Problem Solving,3(2), 2. Makeig, S, Gramann, K, Jung, TP, Sejnowski, TJ, & Poizner, H. (2009). Linking brain, mind and behavior. International journal of psychophysiology: official journal of the International Organization of Psychophysiology,73(2), 95–100. Malcolm, GL, & Henderson, JM. (2009). The effects of target template specificity on visual search in real-world scenes: Evidence from eye movements. Journal of vision,9(11), 8.1–13. Malcolm, GL, & Henderson, JM. (2010). Combining top-down processes to guide eye movements during real-world scene search. Journal of vision,10(2), 4.1–11. Mann, DL, Nakamoto, H, Logt, N, Sikkink, L, & Brenner, E. (2019). Predictive eye movements when hitting a bouncing ball. Journal of vision, 19(14), 28. Maravita, A, Spence, C, Kennett, S, & Driver, J. (2002). Tool-use changes multimodal spatial interactions between vision and touch in normal humans. Cognition,83(2), B25–34. Marconi, B, Genovesio, A, Battaglia-Mayer, A, Ferraina, S, Squatrito, S, Molinari, M, Lacquaniti, F, & Caminiti, R. (2001). Eye-hand coordination during reaching. I. anatomical relationships between parietal and frontal cortex. Cerebral cortex,11(6), 513–527. Matthis, JS, Yates, JL, & Hayhoe, MM. (2018). Gaze and the control of foot placement when walking in natural terrain. Current biology: CB, 28(8), 1224–1233.e5. Melcher, D. (2007). Predictive remapping of visual features precedes saccadic eye movements. Nature neuroscience,10(7), 903–907. BIBLIOGRAPHY |241
Melcher, D, & Colby, CL. (2008). Trans-saccadic perception. Trends in cognitive sciences,12(12), 466–473. Melnik, A, Schüler, F, Rothkopf, CA, & König, P. (2018). The world as an external memory: The price of saccades in a sensorimotor task. Frontiers in behavioral neuroscience,12, 253. Mennie, N, Hayhoe, M, & Sullivan, B. (2007). Look-ahead fixations: Anticipatory eye movements in natural tasks. Experimental brain research. Experimentelle Hirnforschung. Experimentation cerebrale, 179(3), 427–442. Mills, M, Hollingworth, A, Van der Stigchel, S, Hoffman, L, & Dodd, MD. (2011). Examining the influence of task set on eye movements and fixations. Journal of vision,11(8), 17. Milner, AD, & Goodale, MA. (1995). The visual brain in action. Oxford psychology series, No. 27.,248. Mnih, V, Kavukcuoglu, K, Silver, D, Rusu, AA, Veness, J, Bellemare, MG, Graves, A, Riedmiller, M, Fidjeland, AK, Ostrovski, G, Petersen, S, Beattie, C, Sadik, A, Antonoglou, I, King, H, Kumaran, D, Wierstra, D, Legg, S, & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature,518(7540), 529–533. Mobbs, D, Wise, T, Suthana, N, Guzmán, N, Kriegeskorte, N, & Leibo, JZ. (2021). Promises and challenges of human computational ethology. Neuron,109(14), 2224–2238. Moschovakis, AK, Scudder, CA, & Highstein, SM. (1996). The microscopic anatomy and physiology of the mammalian saccadic system. Progress in neurobiology,50(2-3), 133–254. Mottelson, A, & Hornbæk, K. (2017). Virtual reality studies outside the laboratory. Proceedings of the 23rd ACM Symposium on Virtual Reality Software and Technology, (Article 9), 1–10. Mullen, T, Kothe, C, Chi, YM, Ojeda, A, Kerth, T, Makeig, S, Cauwenberghs, G, & Jung, TP. (2013). Real-time modeling and 3D visualization of source dynamics and connectivity using wearable EEG. Annual 242 |BIBLIOGRAPHY
International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference,2013, 2184–2187. Najemnik, J, & Geisler, WS. (2005). Optimal eye movement strategies in visual search. Nature,434(7031), 387–391. Najemnik, J, & Geisler, WS. (2008). Eye movement statistics in humans are consistent with an optimal search strategy. Journal of vision,8(3), 4–4. Nau, M, Schmid, AC, Kaplan, SM, Baker, CI, & Kravitz, DJ. (2024). Centering cognitive neuroscience on task demands and generalization. Nature neuroscience,27(9), 1656–1667. Newell, A. (1963). A guide to the general problem-solver program GPS-2-2. Newell, A, Shaw, J, & Simon, H. (1959). Report on a general problemsolving program. IFIP Congress, 256–264. Niehorster, DC, Santini, T, Hessels, RS, Hooge, ITC, Kasneci, E, & Nyström, M. (2020). The impact of slippage on the data quality of head-worn eye trackers. Behavior research methods,52(3), 1140–1160. Nolte, D, Schmidt, V, Grasso-Cladera, A, & König, P. (2024). Investigating saccade-onset locked EEG signatures of face perception during free-viewing in a naturalistic virtual environment. bioRxiv, 2024.12.12.628113. Nolte, D, Vidal De Palol, M, Keshava, A, Madrid-Carvajal, J, Gert, AL, von Butler, EM, Kömürlüo˘ glu, P, & König, P. (2024). Combining EEG and eye-tracking in virtual reality: Obtaining fixation-onset event-related potentials and event-related spectral perturbations. Attention, perception & psychophysics. Notaro, G, van Zoest, W, Altman, M, Melcher, D, & Hasson, U. (2019). Predictions as a window into learning: Anticipatory fixation offsets carry more information about environmental statistics than reactive stimulus-responses. Journal of vision,19(2), 8. BIBLIOGRAPHY |243
O’Regan, JK, & Noë, A. (2001). A sensorimotor account of vision and visual consciousness. The Behavioral and brain sciences,24(5), 939–73, discussion 973–1031. Ossandón, JP, Onat, S, & König, P. (2014). Spatial biases in viewing behavior. Journal of vision,14(2). Parada, FJ. (2018). Understanding natural cognition in everyday settings: 3 pressing challenges. Frontiers in human neuroscience,12, 386. Parada, FJ, & Rossi, A. (2020). Perfect timing: Mobile brain/body imaging scaffolds the 4E-cognition research program. The European journal of neuroscience,54(12), 8081–8091. Parr, T, Holmes, E, Friston, KJ, & Pezzulo, G. (2023). Cognitive effort and active inference. Neuropsychologia,184(108562), 108562. Parsons, TD. (2015). Virtual reality for enhanced ecological validity and experimental control in the clinical, affective and social neurosciences. Frontiers in human neuroscience,9, 660. Pelz, J, Hayhoe, M, & Loeber, R. (2001). The coordination of eye, head, and hand movements in a natural task. Experimental brain research. Experimentelle Hirnforschung. Experimentation cerebrale,139(3), 266– 277. Pelz, JB, & Canosa, R. (2001). Oculomotor behavior and perceptual strategies in complex tasks. Vision research,41(25-26), 3587–3596. Pennartz, CMA. (2018). Consciousness, representation, action: The importance of being goal-directed. Trends in cognitive sciences,22(2), 137–153. Pezzulo, G, Hoffmann, J, & Falcone, R. (2007). Anticipation and anticipatory behavior. Cognitive processing,8(2), 67–70. Pezzulo, G, Zorzi, M, & Corbetta, M. (2021). The secret life of predictive brains: What’s spontaneous activity for? Trends in cognitive sciences, 25(9), 730–743. 244 |BIBLIOGRAPHY
Piaget, J. (1952). The origins of intelligence in children. International University. Pizlo, Z, & Li, Z. (2005). Solving combinatorial problems: The 15-puzzle. Memory & cognition,33(6), 1069–1084. Pouget, P. (2015). The cortex is in overall control of ’voluntary’ eye movement. Eye,29(2), 241–245. Priorelli, M, Stoianov, IP, & Pezzulo, G. (2024). Embodied decisions as active inference. Neuroscience, (biorxiv;2024.05.28.596181v1). Rao, RP, & Ballard, DH. (1999). Predictive coding in the visual cortex: A functional interpretation of some extra-classical receptive-field effects. Nature neuroscience,2(1), 79–87. Renninger, LW, Verghese, P, & Coughlan, J. (2007). Where to look next? eye movements reduce local uncertainty. Journal of vision,7(3), 6. Rosenbaum, DA, van Heugten, CM, & Caldwell, GE. (1996). From cognition to biomechanics and back: The end-state comfort effect and the middle-is-faster effect. Acta psychologica,94(1), 59–85. Rossion, B, Gauthier, I, Tarr, M, Despland, P, Bruyer, R, Linotte, S, & Crommelinck, M. (2000). The N170 occipito-temporal component is delayed and enhanced to inverted faces but not to inverted objects: An electrophysiological account of face-specific processes in the human brain. Neuroreport,11(1), 69–72. Rossion, B, & Jacques, C. (2008). Does physical interstimulus variance account for early electrophysiological face sensitive responses in the human brain? ten lessons on the N170. NeuroImage,39(4), 1959– 1979. Rothkopf, CA, Ballard, DH, & Hayhoe, MM. (2007). Task and context determine where you look. Journal of vision,7(14), 16.1–20. Sanger, TD. (2000). Human arm movements described by a lowdimensional superposition of principal components. The Journal of BIBLIOGRAPHY |245
Shadi Derakhshan analyzed the head-tracking data and contributed to the writing process. Ashima Keshava and Artur Czeszumski assisted with data analysis as well as the review of the manuscript, Hristofor Lukanov and Marc Vidal de Palol developed the questionnaire used in the experiment. Peter König and Gordon Pipa supervised the project and reviewed the manuscript. ◦Appendix C: Czeszumski, A.*, Gert, A. L.*, Keshava, A.*, Ghadirzadeh, A., Kalthoff, T., Ehinger, B. V., Tiessen, M., Björkman, M., Kragic, D., & König, P. (2021). Coordinating With a Robot Partner Affects Neural Processing Related to Action Monitoring. Frontiers in Neurorobotics, 15, 686010. (*shared first-author) Peter König, Danica Kragic, and Mårten Björkman : conceived the study. Artur Czeszumski, Anna Lisa Gert, Ashima Keshava, and Peter König: designed the study. Ali Ghadirzadeh and Mårten Björkman : programmed the tablet and the robot. Artur Czeszumski, Anna Lisa Gert, Ashima Keshava, Ali Ghadirzadeh, and Max Tiessen data collection. Anna Lisa Gert and Ashima Keshava: major data analysis. Artur Czeszumski, Anna Lisa Gert, and Ashima Keshava: initial draft of the manuscript. Artur Czeszumski , Anna Lisa Gert, Ashima Keshava, Ali Ghadirzadeh, Benedikt V. Ehinger, Mårten Björkman , and Peter König: revision and finalizing the manuscript. All authors contributed to the article and approved the submitted version. ◦Appendix D: Nolte, D., Vidal De Palol, M., Keshava, A., Carvajal, J. M., Gert, A.L., Von Butler, E., Komurluoglu, P., König, P. (2023). Combining EEG and eye-tracking in virtual reality: Obtaining fixation-onset event-related potentials and event-related spectral perturbations.Attention, Perception & Psychophysics. Debora Nolte designed the study, preprocessed and analyzed the data, and wrote the manuscript. Marc Vidal De Palol was involved in designing the experiment and contributed to discussing the results. Marc Vidal De Palol and John Madrid-Carvajal were involved in the aligning of time streams. John Madrid-Carvajal was involved in the data collection and preprocessing of the EEG data. Ashima Keshava was involved in designing the eyetracking algorithm. Anna Lisa Gert advised the EEG analysis and reviewed the manuscript. Eva-Marie von Butler and Pelin Kömürlüo˘ glu computed the time-frequency analysis and were involved in writing the manuscript. Pelin Kömürlüo˘ glu hand-labeled the data and worked on comparing the results to the algorithm classification. Peter König supervised the project and was involved in the data analysis and editing of the manuscript.
No other persons were involved in the substantive preparation of the present work. In particular, I have not made use of any paid services from intermediary or consulting agencies (doctoral advisors or other individuals). No one has received any financial or material compensation from me, either directly or indirectly, for work related to the content of the submitted dissertation. The dissertation has not previously been submitted, in whole or in part, in the same or a similar form to any other examination authority, either in Germany or abroad. Ort, Datum Unterschrift