Swiss Journal of Sociology 2025 sozciolog Schweizerische Zeitschrift für Soziologie Revue suisse de sociologie Swiss Journal of Sociology 51 2 Vol. 51 Issue 2, July 2025 Big Visual Data as a New Form of Knowledge – Potentials, Challenges, and Transformations / Big Visual Dataals neue Form des Wissens – Potenziale, Herausforderungen und Transformationen / Mégadonnées visuelles, une nouvelle forme de savoir – potentiels, défis et transformations Edited by Sebastian W. Hoggenmüller Sebastian W. Hoggenmüller Big Visual Data as a New Form of Knowledge. An Introduction to the Topic and Its Background [G] Ajit Singh Big Visual Data in Digital Infrastructure Planning. On the Processuality and Materiality ofSynthetic Planning Objects [G] Mina Godarzani-Bakhtiari From Flat Image to Spatialised Visual Analysis: and René Tuma Forensic Architecture and the Interweaving of Big Visual Data [G] Roland Meyer Operative Image Spaces. Navigating Virtual Museum Collections [E] Max Frischknecht Through the Eyes of the Machine: Exploring Historical Photo Collections WithConvolutional Neural Networks [E] Katrin Herms and Jörg Lehmann Seeing Like a Field? [E] Sebastian W. Hoggenmüller Metapictures as Research Tools. On the Contingency and Harald Klinke and Algorithmic Conditionality of Their Production [G]
Editors Roman Gibel (University of Zurich) Kenneth Horvath (Zurich University of Teacher Education) Stephanie Steinmetz (University of Lausanne) Núria Sánchez (University of Neuchâtel) Manuscripts and Editorial Correspondence Revue suisse de sociologie Faculté des sciences sociales et politiques Institut des sciences sociales Université de Lausanne Geopolis – Mouline CH-1015 Lausanne E-mail: [email protected] Subscription to the Swiss Journal of Sociology Seismo Press, Zeltweg 27, CH–8032 Zurich, tel. +41 (0)44 261 10 94 E-mail:
[email protected], http://www.seismoverlag.ch. Annual subscription (three issues) sFr. 120.– Individuals; sFr. 140.– Institutions; Overseas + sFr. 30.–
Schweizerische Zeitschrift für Soziologie Revue suisse de sociologie Swiss Journal of Sociology Vol. 51, Issue 2, July 2025 Big Visual Data as a New Form of Knowledge – Potentials, Challenges, and Transformations / BigVisual Data als neue Form des Wissens – Potenziale, Herausforderungen und Transformationen / Mégadonnées visuelles, une nouvelle forme de savoir – potentiels, défis et transformations Edited by Sebastian W. Hoggenmüller Inhalt / Sommaire / Contents 203 Big Visual Data als neue Form des Wissens. Eine Einführung in das Thema undseinen Hintergrund Big Visual Data as a New Form of Knowledge. An Introduction to the Topic and Its Background Mégadonnées visuelles, une nouvelle forme de savoir: introduction et contexte Sebastian W. Hoggenmüller 225 Big Visual Data in der digitalen Infrastrukturplanung. Zur Prozessualität undMaterialität synthetischer Planungsobjekte Big Visual Data in Digital Infrastructure Planning. On the Processuality and Materiality ofSynthetic Planning Objects Mégadonnées visuelles dans la planification digitale des infrastructures. Processus et matérialité des objets de planification synthétiques Ajit Singh 247 Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung von Big Visual Data From Flat Image to Spatialised Visual Analysis: Forensic Architecture and the Interweaving of Big Visual Data De l’image plate à l’analyse visuelle spatialisée: Forensic Architecture et l’imbrication desmégadonnées visuelles Mina Godarzani-Bakhtiari und René Tuma
273 Operative Image Spaces. Navigating Virtual Museum Collections Operative Bildräume. Zur Navigation in virtuellen Museumssammlungen Espaces d’images opératives. Naviguer dans les collections de musées virtuels Roland Meyer 291 Through the Eyes of the Machine: Exploring Historical Photo Collections WithConvolutional Neural Networks Durch die Augen der Maschine: Zur Untersuchung historischer Fotosammlungen mittels Convolutional Neural Networks À travers les yeux de la machine: exploration de collections de photos historiques à l’aide de Convolutional Neural Networks Max Frischknecht 317 Seeing Like a Field? Sehen wie ein Feld? Voir comme un champ? Katrin Herms and Jörg Lehmann 337 Metabilder als Forschungswerkzeuge. Zur Kontingenz und algorithmischen Bedingtheit ihrer Herstellung Metapictures as Research Tools. On the Contingency and Algorithmic Conditionality ofTheir Production Les métapictures en tant qu’outils de recherche: contingence et la conditionnalité algorithmique de leur production Sebastian W. Hoggenmüller und Harald Klinke
203 Swiss Journal of Sociology, 51 (2), 2025, 203–223 * Universität Luzern, Kulturund Sozialwissenschaftliche Fakultät, Soziologisches Seminar, CH-6002 Luzern, [email protected]. Big Visual Data als neue Form des Wissens. Eine Einführung in das Thema und seinen Hintergrund Sebastian W. Hoggenmüller* Zusammenfassung: Der Beitrag führt in das Thema und den Hintergrund des Sonderhefts Big Visual Data als neue Form des Wissens – Potenziale, Herausforderungen und Transformationen ein. Im Mittelpunkt steht dabei die Darstellung der beiden zentralen Forschungsinteressen, die das Heft programmatisch zusammenführt, um Big Visual Data aus einer dezidiert kritischen Perspektive zu erforschen: das Interesse an Big Data und das Interesse an Visualität. Abschließend werden die Struktur des Hefts sowie die Inhalte der Einzelbeiträge vorgestellt. Schlüsselwörter: Big Visual Data, Big Data, Visualität, Epistemologie, kritische visuelle Kompetenz Big Visual Data as a New Form of Knowledge. An Introduction to the Topic and Its Background Abstract: This article introduces the topic and background of the special issueBig Visual Data as a New Form of Knowledge – Potentials, Challenges, and Transformations. It foregrounds the two central research interests – big data and visuality – that the volume programmatically brings together to explore big visual data from a decidedly critical perspective. The article concludes with an overview of the structure of the volume and a summary of the individual contributions. Keywords: Big visual data, big data, visuality, epistemology, critical visual literacy Mégadonnées visuelles, une nouvelle forme de savoir: introduction et contexte Résumé: Cet article présente le thème et le contexte du numéro hors-série intitulé Mégadonnées visuelles, une nouvelle forme de savoir – potentiels, défis et transformations. Ce faisant, il éclaire les deux principaux axes de recherche réunis de manière programmatique dans ce numéro afin d’explorer les mégadonnées visuelles dans une perspective résolument critique: l’intérêt pour les mégadonnées d’une part, et celui pour la visualité d’autre part. L’article se clôt par un aperçu de la structure du numéro et des résumés des différentes contributions. Mots-clés: Mégadonnées visuelles, mégadonnées, visualité, épistémologie, compétence visuelle critique DOI 10.26034/cm.sjs.2025.7178 © 2025. This work is licensed under the Creative Commons Attribution-NonCommercialNoDerivatives 4.0 License. (CC BY-NC-ND 4.0)
204 Sebastian W. Hoggenmüller SJS 51 (2), 2025, 203–223 1 Big Visual Data – die neue Macht der Bilder? Riesige Mengen visueller Daten prägen zunehmend unser Wissen und beeinflussen, wie wir die Welt wahrnehmen: Auf Social Media etwa sind wir täglich einer Flut von visuellen Inhalten wie Fotos, KI-generierten Bildern, Livestreams und Reels ausgesetzt, die nicht nur die individuelle Aufmerksamkeit lenkt, sondern auch soziales Handeln bestimmt und öffentliche Diskurse beeinflusst (vgl. z. B. Schankweiler &Straub, 2023; Zulli & Zulli, 2022). In urbanen Zentren wiederum erfassen vernetzte Kamerasysteme fortlaufend Gesichter, überwachen Bewegungsmuster und kartieren Verkehrsströme, mit dem Anspruch, Mobilität zu steuern und sogenannte sicherheitsrelevante Ereignisse zu identifizieren (vgl. z. B. Butot et al., 2023; Monahan, 2018). Des Weiteren werden in der radiologischen Diagnostik große medizinische Bilddatensätze computergestützt analysiert, um Anomalien frühzeitig zu erkennen, was die Diagnosegenauigkeit verbessern und fundierte Behandlungsentscheidungen ermöglichen soll (vgl. z. B. Armato III, 2023; Lombi & Rossero, 2024). Und die optische Fernerkundung beeinflusst globale politische Entscheidungsprozesse, indem sie mit hochauflösenden Satellitenbildzeitreihen Umweltveränderungen sichtbar macht, die andernfalls nicht wahrnehmbar wären, und durch die Analyse von Infrastruktur und Landnutzung Indikatoren für Phänomene wie Armut oder kriegerische Aktivitäten liefert (vgl. z. B. Bouabid & Farah, 2024; Sako & Martinez, 2021; Uzhinskiy et al., 2018). Die Aufzählung der Beispiele ließe sich nahezu beliebig fortsetzen. Das vorliegende Sonderheft widmet sich dieser gegenwärtigen Konjunktur und wachsenden Bedeutung sehr großer digitaler visueller Datenmengen, die in verschiedenen Forschungsdisziplinen begrifflich unter Big Visual Data firmieren (vgl.z. B. in der Informatik, Elektrotechnik und Computer Vision Chen et al., 2016; Fang et al., 2017; Qin et al., 2015; in den Wirtschaftswissenschaften Giglio et al., 2020; in den Ingenieurswissenschaften Bhargava et al., 2018; in den Visual Studies Skarpelos, 2018).1 Im Zentrum steht dabei die Frage, welche Rolle Big Visual Data bei der Herstellung und Tradierung, Stabilisierung und Veränderung von gesellschaftlichem Wissen und sozialer Wirklichkeit spielen, und damit die Erkundung, inwiefern Big Visual Data eine neue Macht der Bilder (Sachs-Hombach, 1998) hervorbringen, worin diese besteht und wie sie sich entfaltet. Ausgehend von dieser Fragestellung setzt das Sonderheft zwei Analyseschwerpunkte: Einerseits stehen die Technologien, Methoden und infrastrukturellen Bedingungen im Fokus, durch die Big Visual Data erzeugt, gespeichert, integriert, 1 Ein weiterer geläufiger Begriff ist Big Image Data. Dieser findet Verwendung in Fachrichtungen wie der Digitalen Kunstgeschichte (vgl. z. B. Klinke, 2016), der Biologie(vgl. z. B. Smith et al., 2018) sowie erneut in der Informatik (vgl. z. B. Kanaparthi & Raju, 2022) und bezieht sich in der Regel auf explizit bildliche Daten wie insbesondere Fotografien, aber auch Gemälde, Drucke u. Ä., also auf visuelle Daten in einem engeren Sinne. Im Gegensatz dazu wird der Begriff Big Visual Data in diesem Sonderheft bewusst weiter gefasst, indem er unterschiedlichste visuelle Kommunikationsformen einbezieht und damit eine breitere Perspektive auf digitale visuelle Daten ermöglicht.
Big Visual Data als neue Form des Wissens. Eine Einführung in das Thema und seinen Hintergrund 205 SJS 51 (2), 2025, 203–223 verarbeitet, analysiert und visualisiert werden. Andererseits gilt das Interesse den konkreten sozialen Praktiken der Erzeugung, Interpretation und Nutzung von Big Visual Data in verschiedenen gesellschaftlichen Kontexten. Als Heuristik liegen diesen beiden Analyseschwerpunkten folgende Fragen zugrunde, die sich auf vier Kernaspekte beziehen: ›Epistemische Grundlagen von Big Visual Data: Welche Merkmale kennzeichnen Big Visual Data jenseits der bloßen Datenmenge? Worin liegt ihre spezifische kommunikative Qualität? Und welche persuasive Wirkung entfalten sie? › Soziotechnische Bedingungen von Big Visual Data: Wie beeinflussen situierte kommunikative Arrangements die Entstehung und das Verstehen von Big Visual Data? Und inwiefern führt insbesondere die Verflechtung sozialer und technischer Bedingungen zu neuen Formen der Wissensproduktion und epistemischer Autorität? › Stabilisierung oder Transformation von Machtverhältnissen durch Big Visual Data: Wer oder was entscheidet, wer Big Visual Data erfasst und zugänglich macht? Welche Rolle spielen dabei Akteure wie Plattformunternehmen und Regulierungsinstanzen? Und in welchem Verhältnis stehen datengetriebene Wissensökonomien zu anderen Wissensregimen? › (Forschungs-)Praktische Arbeit mit Big Visual Data: Welche Erkenntnisse eröffnet die Analyse von Big Visual Data? Welche Vorgehensweisen sind erforderlich, um diese Einsichten zu gewinnen? Und inwiefern verlangt der Umgang mit Big Visual Data nicht nur techn(olog)ische und methodische Innovationen, sondern – insbesondere in wissenschaftlichen Kontexten – auch eine grundlegende Reflexion über bestehende methodologische Prämissen sowie über theoretische Annahmen zu Daten, Information, Wissen und (Un-)Sichtbarkeit?2 Die Beiträge dieses Sonderhefts greifen diese Fragen aus unterschiedlichen Perspektiven auf und untersuchen kritisch die Potenziale und Herausforderungen von Big Visual Data sowie die Transformationen, die durch Big Visual Data hervorgerufen und mitgeprägt werden – und die wiederum auf Big Visual Data zurückwirken. Damit führt das Sonderheft zwei zentrale Forschungsinteressen zusammen, die in der sozialwissenschaftlichen Forschung im Allgemeinen und in der Soziologie im Besonderen üblicherweise getrennt voneinander behandelt werden, obwohl sie in anderen Disziplinen sowie interund transdisziplinär schon seit Längerem miteinander verknüpft und erforscht werden (vgl. allen voran die Arbeiten von Lev 2 Eine solche Diskussion über die Weiterentwicklung etablierter methodischer Ansätze und die Neubewertung methodologischer Grundannahmen lässt sich derzeit mit großer Dynamik in der qualitativen Sozialforschung im Zusammenhang mit KI-gestützten Analyseverfahren, insbesondere Machine-Learning-Modellen und Large Language Models (LLMs), beobachten (vgl. dazu in chronologischer Reihenfolge z. B. Christou, 2023; Şen et al., 2023; Eschrich & Sterman, 2024; Hitch, 2024; Krähnke et al., 2025; Nguyen-Trung & Nguyen, 2025).
206 Sebastian W. Hoggenmüller SJS 51 (2), 2025, 203–223 Manovich und seinem Cultural Analytics Lab, z. B. Manovich, 2017; 2020; zuletzt Manovich & Arielli, 2024). Gemeint sind das Interesse an Big Data und das Interesse an Visualität. Im Folgenden werden diese beiden Interessen kurz umrissen, zunächst jenes an Big Data (Abschnitt2), gefolgt von dem an Visualität (Abschnitt3). 2 Das Forschungsinteresse an Big Data Zum ersten der beiden Forschungsinteressen, dem Interesse an Big Data, ist in den letzten Jahren eine buchstäblich unüberschaubare Menge an Literatur erschienen (vgl. für Überblicksdarstellungen insbesondere Chen & Yu, 2018; Kaplan, 2015; Kitchin, 2017; speziell zur Entwicklung der Computational Social Science innerhalb der Soziologie Edelmann et al., 2020; für die Soziale Netzwerkanalyse Tindall et al., 2022; zur Datenanalyse in der sozialwissenschaftlichen Forschung William, 2024; allgemein zu Mustern, Trends und Lücken in der wissenschaftlichen Big-Data-Forschung Rossi et al., 2019; mit Fokus auf die letzten 15 Jahre Tosi et al., 2024). Aus dieser Literatur lässt sich nicht zuletzt entnehmen, dass das Phänomen Big Data äußerst unterschiedlich wahrgenommen und bewertet wird (wobei die Vielfalt der Stimmen insbesondere im frühen Diskurs deutlich hervortritt) – als neuartige Ressource, als heilbringende Revolution, als vorübergehender Hype, als ethische Herausforderung oder gar als Beginn des Untergangs von Theorie und wissenschaftlicher Methode (vgl.z. B. Anderson, 2008; Burrows & Savage, 2014; Diaz-Bone et al., 2020; Favaretto et al., 2019; Halford & Savage, 2017; Kitchin, 2014b; Miller, 2010; Noble, 2018; Schmitt, 2018; van Dijck, 2014; Wiegerling et al., 2018). Neben diesen Differenzen in der Bewertung von Big Data gibt es auch unterschiedliche Auffassungen über den Beginn und den Verlauf der exponentiellen Zunahme digitaler Daten, die in Big Data resultiert (vgl. z. B. Barnes, 2013; Diebold, 2012; Lünich, 2022). Zudem wird in diesem Zusammenhang diskutiert, inwiefern Big Data als eine Weiterentwicklung früherer Formen der Vervielfachung sozialer Daten verstanden werden kann. So ordnet beispielsweise Beer (2016) Big Data in die lange Geschichte der Sozialstatistik ein und setzt die heutige Zunahme digitaler sozialer Daten in Bezug zur „avalanche of printed numbers“ (Hacking, 1982, S.281) und zur „great explosion of numbers“ (Porter, 1986, S.11), die beide auf Entwicklungen im Zeitraum von 1820 bis etwa 1840 hinweisen (vgl. auch Ambrose, 2014): [T]he sense that we are being faced with a deluge of data about people is not something that is entirely new, in fact it has a long history. The type of data may have changed as might its analytics […] but the lineage is clear. There are, of course, features of the current data moment that are in some ways novel but it is still interesting to note that this idea of a scaling up of social data, the feeling that we are facing an unfathomable flow of social data, itself has a history. (Beer, 2016, S.2)
Big Visual Data als neue Form des Wissens. Eine Einführung in das Thema und seinen Hintergrund 207 SJS 51 (2), 2025, 203–223 Während Beer also in Bezug auf die gegenwärtige Masse an Daten historische Kontinuitäten hervorhebt, verweist er im angeführten Zitat zugleich auf einen anderen Aspekt, über den in der Literatur zu Big Data weitgehend Einigkeit herrscht: Die Art und Weise, wie wir im digitalen Zeitalter Daten sammeln, speichern, verarbeiten und analysieren bzw. interpretieren, hat sich grundlegend gewandelt und unterliegt weiterhin tiefgreifenden Veränderungen – nicht nur in der Wissenschaft (vgl.z. B.Boyd & Crawford, 2012; Kitchin, 2014a; Wolbring, 2020), sondern in nahezu allen gesellschaftlichen Bereichen (vgl. z. B. Kitchin, 2021; Mayer-Schönberger & Cukier, 2013; Sfetcu, 2023). In diesem Sinne machen Prietl und Houben (2018), die hier stellvertretend für viele andere Autor*innen genannt werden, auf entscheidende Brüche aufmerksam, die Big Data als etwas grundsätzlich Neues erscheinen lassen. In ihren Überlegungen zur Datafizierung des Sozialen identifizieren sie insgesamt sechs maßgebliche Veränderungen, die „die gegenwärtige Akkumulation von Daten und den aktuell beobachtbaren Umgang mit ihnen von historisch früheren Epochen unterscheide[n]“ (Prietl & Houben, 2018, S.9). Zu diesen Veränderungen gehören für sie unter anderem das Entstehen neuer sozialer Praktiken der Generierung und Verbreitung von Daten, die zunehmende Durchdringung aller Lebensbereiche durch datensammelnde digitale Technologien sowie das Aufkommen neuer Akteure wie Plattformunternehmen und Tech-Konzerne, die Staat und Kirche als maßgebliche Instanzen der Datensouveränität ablösen.3 Beispiele hierfür sind etwa Gesundheits-Apps und Wearables, die als Selbstvermessungstechnologien ununterbrochen individuelle Gesundheitsdaten aufzeichnen (vgl. z. B. Cappel, 2022; Gilmore, 2016), digitale Bezahlsysteme, die alltägliche Transaktionen ohne Bargeld ermöglichen und dabei Informationen über Zahlungsströme und Konsumpräferenzen erfassen (vgl. z. B. Bhuiyan et al., 2024; Mützel & Unternährer, 2024), Predictive-Policing-Technologien, die zur Vorhersage von Straftaten eingesetzt werden (vgl.z. B. Hälterlein, 2021; Miró-Llinares, 2020), sowie die algorithmische Steuerung von Inhalten in Sozialen Medien, die Echo-Kammern und Filterblasen begünstigt (vgl. z. B. Barberá, 2020; Palmieri, 2024). Ein weiterer zentraler Aspekt, der in vielen Forschungsarbeiten zu Big Data zu finden ist, ist der folgende: Big Data werden häufig anhand von drei grundlegenden Merkmalen beschrieben, die auf Doug Laney (2001) zurückgehen und als die drei Vs – volume, velocity und variety – bekannt geworden sind. Pointiert formuliert bezeichnet volume dabei die enormen Datenmengen, die durch digitale Prozesse kontinuierlich erzeugt werden. Velocity beschreibt die hohe Frequenz, mit der Daten generiert, übertragen und verarbeitet werden – von Near-Real-Time-Anwendungen bis hin zu Echtzeitanalysen. Variety bezieht sich auf die Heterogenität der Daten hinsichtlich Struktur, Quelle und Format. 3 Speziell zur Bedeutung des digitalen Kapitalismus für die computergestützte Sozialforschung vgl. jüngst die Laborstudie von Ak (2025).
214 Sebastian W. Hoggenmüller SJS 51 (2), 2025, 203–223 und generative KI zur Erzeugung neuer, synthetischer Bildwelten genutzt werden, setzt sich Meyer kritisch mit den ideologischen Implikationen dieser Transformation und den damit verbundenen Verfahren auseinander. Dabei argumentiert er, dass virtuelle Bildarchive oft eine imperialistische und kolonialistische Logik perpetuieren, und veranschaulicht dies anhand von Beispielen wie Google Arts & Culture sowie generativen KI-Modellen wie DALL-E. Als Gegenentwurf präsentiert er das Projekt Digital Benin, das eine nichteurozentrische Visualisierung von Museumsbeständen ermöglicht. Meyer zeigt, dass Digital Benin durch die Integration kontextualisierter Informationen und die Zusammenarbeit mit Expert*innen aus den Herkunftsländern der Artefakte eine vielschichtige Darstellung kultureller Objekte bietet, die deren historische und kulturelle Bedeutungen bewahrt. Über dieses konkrete Beispiel hinaus plädiert er für eine kritische Digitalisierungspolitik, die die bestehenden imperialen Strukturen überwindet. Max Frischknecht wiederum untersucht aus der Perspektive der Digital Humanities den Einsatz von Convolutional Neural Networks (CNNs) zur Clusterbildung historischer Fotosammlungen. Am Beispiel der Sammlung Ernst Brunner aus dem Archiv der Fachgesellschaft Empirische Kulturwissenschaft Schweiz (EKWS), die etwa 48 000 Negative aus den Jahren 1935 bis 1970 umfasst, geht er mithilfe der Software PixPlot der Frage nach, ob ein CNN zentrale Themen und Narrative der Sammlung erkennen kann. Sein Hauptinteresse gilt dabei nicht der Entdeckung neuer Aspekte, sondern der Evaluation, ob und inwiefern die algorithmisch durchgeführte Clusterbildung mit der auf traditionellem Weg erfolgten Werkanalyse übereinstimmt. Durch diesen Vergleich menschlicher und maschineller Ways of Seeing (Berger, 2008 [1972]) zeigt der Beitrag, wie stark die maschinelle Sehweise von Trainingsdaten und technischer Infrastruktur geprägt ist. So können nach Frischknecht Aspekte in den Vordergrund treten, die eher in den Trainingsdaten und ihrem Entstehungskontext als in der historischen Sammlung verankert sind – im vorliegenden Fall also im Trainingsdatensatz ImageNet, der vom Stanford Vision Lab entwickelt wurde. Darauf aufbauend führt Frischknecht aus, dass es einer erweiterten Zugänglichkeit und einer größeren Diversität von Trainingsdaten bedarf, um das Potenzial maschineller Analyseverfahren bei der Erforschung großer historischer Bildbestände besser auszuschöpfen und die damit verbundenen epistemologischen Implikationen umfassender zu verstehen. Ein grundlegendes Interesse an erkenntnistheoretischen Fragen im Kontext der wissenschaftlichen Forschung verfolgt auch der Beitrag von Katrin Herms und Jörg Lehmann, in dem sie im interdisziplinären Dialog zwischen Kultursoziologie, Netzwerkanalyse und den Digital Humanities die Feldtheorie Bourdieus mit aktueller Big-Visual-Data-Forschung verbinden. Im Mittelpunkt steht dabei die Frage, ob die Soziologie auf der Basis großer visueller Datensätze „wie ein Feld sehen“, also einen Blick aus dem Inneren eines Feldes richten kann. Zur Klärung dieser Frage diskutieren Herms und Lehmann unter systematischer Einbindung des Konzepts
Big Visual Data als neue Form des Wissens. Eine Einführung in das Thema und seinen Hintergrund 215 SJS 51 (2), 2025, 203–223 der Aufmerksamkeitsökonomie zwei aktuelle Forschungsansätze für große visuelle Datensätze aus dem Bereich der sozialen Netzwerkanalyse: Computer-Vision-Netzwerke und Bildähnlichkeitsanalysen. Anhand dieser Ansätze demonstrieren sie, wie maschinelle Lernverfahren Big Visual Data aggregieren und latente Beziehungen sichtbar machen, woraus sich ein erhebliches Potenzial für die soziologische Forschung ergibt. Gleichzeitig weisen sie auf die Grenzen dieser quantitativen Ansätze hin und warnen davor, kulturelle Komplexität auf formal messbare, relationale Netzwerke zu reduzieren. Kritisch weisen sie darüber hinaus auf die grundlegende Kluft zwischen Forschenden und großen Technologieunternehmen in Bezug auf Datenzugriff, ethische Standards und strategische Ziele hin. Diese Kluft, so halten Herms und Lehmann fest, hat weitreichende Konsequenzen – sowohl für die forschungsleitende Frage des Beitrags als auch für die Wissensproduktion insgesamt. Auch der abschließende Beitrag des Sonderhefts beschäftigt sich mit dem Bereich der wissenschaftlichen Forschung, richtet den Fokus jedoch auf den konkreten Umgang mit Big Visual Data in der Forschungspraxis: Sebastian W. Hoggenmüller und Harald Klinke untersuchen ein zentrales Forschungswerkzeug der computergestützten Analyse großer visueller Datenbestände – sogenannte Metabilder (auch bekannt als Image Plots). Basierend auf algorithmischen Verfahren machen Metabilder Muster und andere signifikante Zusammenhänge in den Daten sichtbar, die für das bloße Auge in der Regel nicht erkennbar sind und sich manuellen Analysen weitgehend entziehen. Ziel des Beitrags ist es, die üblicherweise verborgene Komplexität des Herstellungsprozesses von Metabildern offenzulegen, sodass besser verstehbar wird, wie das zustande kommt und funktioniert, was in der Erforschung von Big Visual Data als Instrument zur Erkenntnisgenerierung genutzt wird. Mit ihrem interdisziplinären Ansatz, der Visuelle Soziologie und Digitale Bildwissenschaft miteinander verschränkt, konzentrieren sich Hoggenmüller und Klinke insbesondere auf die Kontingenz des Herstellungsprozesses von Metabildern sowie auf deren algorithmische Bedingtheit. Dabei zeigen sie, dass Metabilder nicht als selbstevidente, objektive Darstellungen zu verstehen sind, sondern als Ergebnis komplexer mathematisch-statistischer Verfahren und kontextueller Entscheidungsprozesse, die ihren epistemischen Gehalt maßgeblich beeinflussen. Fernerhin erörtern sie die Herausforderungen einer kritischen Nutzung von Metabildern und betonen die Notwendigkeit, interdisziplinäre Ansätze zur kritischen Analyse von Big Visual Data weiterzuentwickeln. 5 Dank Mein Dank gilt allen Beitragenden des Sonderhefts für ihre engagierte Arbeit und die wertvollen Perspektiven, die sie eingebracht haben. Ebenso möchte ich den Gutachter*innen für ihr konstruktives Feedback danken, das die Beiträge maßgeblich
216 Sebastian W. Hoggenmüller SJS 51 (2), 2025, 203–223 bereichert hat. Darüber hinaus danke ich meinen Kolleg*innen für die verschiedenen Diskussionen, mit denen ich inspirierende Einsichten zum Thema gewinnen konnte– dieser Austausch war wesentlich für Form und Inhalt dieses Sonderhefts. Besonders bedanken möchte ich mich in diesem Zusammenhang bei Bettina Heintz, Charlotte Knorr, Helmut Grabner, Janna Muff, Johanna Wahl, Jürgen Raab, Katrin Herms, Sophia Cramer, Thomas Knüsel und Tobias Hodel. Und nicht zuletzt danke ich der Redaktion der Schweizerischen Zeitschrift für Soziologie sowie dem Seismo Verlag, insbesondere Katarzyna Czerwiec-Buczek, Kenneth Horvath und Marion Beetschen, für die durchweg angenehme und zuvorkommende Zusammenarbeit. 6 Literatur Ak, O. (2025). Platforms as Laboratories of the Social: How Digital Capitalism Matters for Computational Social Research in North America. Social Studies of Science, 1–21. Online-Vorabveröffentlichung. https://doi.org/10.1177/03063127251321826 Allan, A., & Tinkler, P. (2015). ‚Seeing‘ into the Past and ‚Looking‘ Forward to the Future: VisualMethods and Gender and Education Research. Gender and Education, 27(7), 791–811. https://doi.or g/10.1080/09540253.2015.1091919 Alloa, E. (2016). Iconic Turn: A Plea for Three Turns of the Screw. Culture, Theory and Critique, 57(2), 228–250. https://doi.org/10.1080/14735784.2015.1068127 Ambrose, M. L. (2014). From the Avalanche of Numbers to Big Data: A Comparative Historical Perspective on Data Protection in Transition. In K. O’Hara, M.-H. C. Nguyen, & P. Haynes (Hrsg.), Digital Enlightenment Yearbook 2014: Social Networks and Social Machines, Surveillance and Empowerment (S.25–48). IOS Press. https://doi.org/10.3233/978-1-61499-450-3-25 Anderson, C. (2008). The End of Theory: The Data Deluge Makes the Scientific Method Obsolete. Wired, https://www.wired.com/2008/06/pb-theory/ (letzter Zugriff am 14. Oktober 2024). Armato III, S. G. (2023). Computational Medicine in Radiology: Medical Images as Big Data. In JVA‘23: Proceedings of the 2023 IEEE John Vincent Atanasoff International Symposium on Modern Computing (S. 66–70). IEEE. https://doi.org/10.1109/JVA60410.2023.00021 Barberá, P. (2020). Social Media, Echo Chambers, and Political Polarization. In N. Persily & J. A. Tucker (Hrsg.), Social Media and Democracy (S. 34–55). Cambridge University Press. https://doi. org/10.1017/9781108890960.004 Barnes, T. J. (2013). Big Data, Little History. Dialogues in Human Geography, 3(3), 297–302. https:// doi.org/10.1177/2043820613514323 Beaulieu, A. (2002). Images Are Not the (Only) Truth: Brain Mapping, Visual Knowledge, and Iconoclasm. Science, Technology, & Human Values, 27(1), 53–86. https://doi.org/10.1177/016224390202700103 Beer, D. (2016). How Should We Do the History of Big Data? Big Data & Society, 3(1), 1–10 https:// doi.org/10.1177/2053951716646135 Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In FAccT ‘21: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (S. 610–623). Association for Computing Machinery. https://doi.org/10.1145/3442188.3445922 Bentkowska-Kafel, A., Cashen, T., & Gardiner, H. (Hrsg.). (2005). Digital Art History: A Subject in Transition. Intellect Books. Berger, J. (2008 [1972]). Ways of Seeing. Penguin Books.
Big Visual Data als neue Form des Wissens. Eine Einführung in das Thema und seinen Hintergrund 217 SJS 51 (2), 2025, 203–223 Berger, P. L., & Luckmann, T. (2010 [1966]). Die gesellschaftliche Konstruktion der Wirklichkeit: Eine Theorie der Wissenssoziologie. Fischer. Bhargava, M. G., Vidyullatha, P., Venkateswara Rao, P., & Sucharita, V. (2018). A Study on Potential of Big Visual Data Analytics in Construction Arena. International Journal of Engineering and Technology, 7(2.7), 652–656. https://doi.org/10.14419/ijet.v7i2.7.10916 Bhuiyan, M. R. I., Akter, M. S., & Islam, S. (2024). How Does Digital Payment Transform Society as a Cashless Society? An Empirical Study in the Developing Economy. Journal of Science and Technology Policy Management, 1–19. Online-Vorabveröffentlichung. https://doi.org/10.1108/ JSTPM-10-2023-0170 Bishop, C. (2018). Against Digital Art History. International Journal for Digital Art History, 3, 123–131. https://doi.org/10.11588/dah.2018.3.49915 Bleichmar, D., & Schwartz, V. (2019). Visual History: The Past in Pictures. Representations, 145(1), 1–31. https://doi.org/10.1525/rep.2019.145.1.1 Boehm, G. (2006 [1994]). Die Wiederkehr der Bilder. In G. Boehm (Hrsg.), Was ist ein Bild? (S.11–38). Wilhelm Fink. Boehm, G. (2007). Wie Bilder Sinn erzeugen: Die Macht des Zeigens. Berlin University Press. Boehm, G., & Mitchell, W. J. T. (2009). Pictorial versus Iconic Turn: Two Letters. Culture, Theory and Critique, 50(2–3), 103–121. https://doi.org/10.1080/14735780903240075 Bohnsack, R. (2024). Qualitative Bildanalyse. In H. Friese, M. Nolden, & M. Schreiter (Hrsg.), Handbuch Soziale Praktiken und Digitale Alltagswelten (S. 1–10). Springer VS. Online-Vorabveröffentlichung. https://doi.org/10.1007/978-3-658-08460-8_55-2 Bouabid, M., & Farah, M. (2024). GAZADeepDav: A High Resolution Geotagged Satellite Imagery Dataset for Analyzing War-Induced Damage. In IGARSS ‘24: 2024 IEEE International Geoscience and Remote Sensing Symposium (S. 8876–8879). IEEE. https://doi.org/10.1109/IGARSS53475.2024.10642306 Boxenbaum, E., Jones, C., Meyer, R. E., & Svejenova, S. (2018). Towards an Articulation of the Material and Visual Turn in Organization Studies. Organization Studies, 39(5–6), 597–616. https://doi. org/10.1177/0170840618772611 Boyd, D., & Crawford, K. (2012). Critical Questions for Big Data: Provocations for a Cultural, Technological, and Scholarly Phenomenon. Information, Communication & Society, 15(5), 662–679. https://doi.org/10.1080/1369118X.2012.678878 Burri, R. V. (2008). Doing Images: Zur Praxis medizinischer Bilder. transcript. https://doi.org/10.14361/ 9783839408872 Burri, R. V. (2009). Aktuelle Perspektiven soziologischer Bildforschung: Zum Visual Turn in der Soziologie. Soziologie, 38(1), 24–39. Burrows, R., & Savage, M. (2014). After the Crisis? Big Data and the Methodological Challenges of Empirical Sociology. Big Data & Society, 1(1), 1–6. https://doi.org/10.1177/2053951714540280 Butot, V., Jacobs, G., Bayerl, P. S., Amador, J., & Nabipour, P. (2023). Making Smart Things Strange Again: Using Walking as a Method for Studying Subjective Experiences of Smart City Surveillance. Surveillance & Society, 21(1), 61–82. https://doi.org/10.24908/ss.v21i1.15665 Cambrosio, A., Jacobi, D., & Keating, P. (2005). Arguing with Images: Pauling’s Theory of Antibody Formation. Representations, 89(1), 94–130. https://doi.org/10.1525/rep.2005.89.1.94 Cappel, V. (2022). Die Pluralität der digitalen Alltagsgesundheit: Das Aufkommen einer neuen Form der Gesundheitskoordination. In V. Cappel & K. E. Kappler (Hrsg.), Gesundheit – Konventionen– Digitalisierung: Eine politische Ökonomie der (digitalen) Transformationsprozesse von und um Gesundheit (S. 77–114). Springer VS. https://doi.org/10.1007/978-3-658-34306-4_3 Chen, C., Ren, Y., & Kuo, C.-C. J. (2016). Big Visual Data Analysis: Scene Classification and Geometric Labeling. Springer Nature. https://doi.org/10.1007/978-981-10-0631-9
218 Sebastian W. Hoggenmüller SJS 51 (2), 2025, 203–223 Chen, S.-H., & Yu, T. (2018). Big Data in Computational Social Sciences and Humanities: An Introduction. InS.-H. Chen (Hrsg.),Big Data in Computational Social Science and Humanities(S.1–25). Springer Nature. https://doi.org/10.1007/978-3-319-95465-3_1 Chmielecki, K. (2015). From Visual Culture to Visual Communication: The Pictorial and Iconic Turn in Contemporary Culture. Art Inquiry: Recherches sur les Arts, 17, 93–114. Christou, P. A. (2023). How to Use Artificial Intelligence (AI) as a Resource, Methodological and Analysis Tool in Qualitative Research? The Qualitative Report, 28(7), 1968–1980. https://doi. org/10.46743/2160-3715/2023.6406 Dang, S.-M. (2018). Digital Tools & Big Data: Zu gegenwärtigen Herausforderungen für die Filmund Medienwissenschaft am Beispiel der feministischen Filmgeschichtsschreibung. MEDIENwissenschaft: Rezensionen | Reviews, 35(2–3), 142–156. de Valle, M. K., Gallego-García, M., Williamson, P., & Wade, T. D. (2021). Social Media, Body Image, and the Question of Causation: Meta-Analyses of Experimental and Longitudinal Evidence. Body Image, 39, 276–292. https://doi.org/10.1016/j.bodyim.2021.10.001 Diaz-Bone, R., Horvath, K., & Cappel, V. (2020). Social Research in Times of Big Data: The Challenges of New Data Worlds and the Need for a Sociology of Social Research. Historical Social Research, 45(3), 314–341. https://doi.org/10.12759/hsr.45.2020.3.314-341 Diebold, F. X. (2012). On the Origin(s) and Development of the Term ‚Big Data‘. SSRN. PIER Working Paper 12–037. https://doi.org/10.2139/ssrn.2152421 D’Ignazio, C., & Klein, L. F. (2020). Data Feminism. MIT Press. https://doi.org/10.7551/mitpress/11805. 001.0001 Dodge, M., & Kitchin, R. (2005). Codes of Life: Identification Codes and the Machine-Readable World. Environment and Planning D: Society and Space, 23(6), 851–881. https://doi.org/10.1068/d378t Edelmann, A., Wolff, T., Montagne, D., & Bail, C. A. (2020). Computational Social Science and Sociology. Annual Review of Sociology, 46, 61–81. https://doi.org/10.1146/annurev-soc-121919-054621 Elkins, J. (2003). Visual Studies: A Skeptical Introduction. Routledge. Eschrich, J., & Sterman, S. (2024). A Framework for Discussing LLMs as Tools for Qualitative Analysis. arXiv. Preprint / Working Paper. https://doi.org/10.48550/arXiv.2407.11198 Fang, Z., Hwang, J.-N., Huo, X., Lee, H.-J., & Denzler, J. (2017). Emergent Techniques and Applications for Big Visual Data. International Journal of Digital Multimedia Broadcasting, 1, 1–2. https://doi. org/10.1155/2017/6468502 Favaretto, M., De Clercq, E., & Elger, B. S. (2019). Big Data and Discrimination: Perils, Promises and Solutions. A Systematic Review. Journal of Big Data, 6(12), 1–27. https://doi.org/10.1186/ s40537-019-0177-4 Fischer, M. D., & Zeitlyn, D. (2003). Visual Anthropology in the Digital Mirror: Computer-Assisted Visual Anthropology. The Virtual Institute of Mambila Studies, http://mambila.info/layers_nggwun. html (letzter Zugriff am 13. Januar 2024). Forsyth, D. A., & Ponce, J. (2003). Computer Vision: A Modern Approach. Prentice Hall. Galison, P. (2014). Visual STS. In A. Carusi, A. S. Hoel, T. Webmoor, & S. Woolgar (Hrsg.), Visualization in the Age of Computerization (S. 197–225). Routledge. https://doi.org/10.4324/9780203066973-10 Giglio, S., Pantano, E., Bilotta, E., & Melewar, T. C. (2020). Branding Luxury Hotels: Evidence from the Analysis of Consumers’ „Big“ Visual Data on TripAdvisor. Journal of Business Research, 119, 495–501. https://doi.org/10.1016/j.jbusres.2019.10.053 Gilmore, J. N. (2016). Everywear: The Quantified Self and Wearable Fitness Technologies. New Media & Society, 18(11), 2524–2539. https://doi.org/10.1177/1461444815588768 Glasze, G. (2014). Sozialwissenschaftliche Kartographie-, GISund Geoweb-Forschung. KN – Journal of Cartography and Geographic Information, 64(3), 123–129. https://doi.org/10.1007/BF03544141 Goffman, E. (1979). Gender advertisements. Harvard University Press.
Big Visual Data als neue Form des Wissens. Eine Einführung in das Thema und seinen Hintergrund 219 SJS 51 (2), 2025, 203–223 Goodwin, C. (1994). Professional Vision. American Anthropologist, 96(3), 606–633. https://doi. org/10.1525/aa.1994.96.3.02a00100 Grimshaw, A. D. (1982). Sound-Image Data Records for Research on Social Interaction: Some Questions and Answers. Sociological Methods & Research, 11(2), 121–144. https://doi.org/10.1177/ 0049124182011002002 Grohmann, R. (2025). Latin American Critical Data Studies. Big Data & Society, 12(2), 1–9. https:// doi.org/10.1177/20539517251330160 Hacking, I. (1982). Biopower and the avalanche of printed numbers. Humanities in Society, 5(1–2), 279–295. Halford, S., & Savage, M. (2017). Speaking Sociologically with Big Data: Symphonic Social Science and the Future for Big Data Research. Sociology, 51(6), 1132–1148. https://doi. org/10.1177/0038038517698639 Hälterlein, J. (2021). Epistemologies of Predictive Policing: Mathematical Social Science, Social Physics and Machine Learning. Big Data & Society, 8(1), 1–13. https://doi.org/10.1177/20539517211003118 Harley, J. B. (1989). Deconstructing the Map. Cartographica: The International Journal for Geographic Information and Geovisualization, 26(2), 1–20. https://doi.org/10.3138/E635-7827-1757-9T53 Heintz, B. (2010). Numerische Differenz: Überlegungen zu einer Soziologie des (quantitativen) Vergleichs. Zeitschrift für Soziologie, 39(3), 162–181. https://doi.org/10.1515/zfsoz-2010-0301 Heintz, B., & Huber, J. (Hrsg.). (2001). Mit dem Auge denken: Strategien der Sichtbarmachung in wissenschaftlichen und virtuellen Welten. Edition Voldemeer. Hentschel, K. (2014). Visual Cultures in Science and Technology: A Comparative History. Oxford University Press. https://doi.org/10.1093/acprof:oso/9780198717874.001.0001 Hitch, D. (2024). Artificial Intelligence Augmented Qualitative Analysis: The Way of the Future? Qualitative Health Research, 34(7), 595–606. https://doi.org/10.1177/10497323231217392 Hoggenmüller, S. W. (2020). Globalisierungsforschung als Bildforschung: Zur bildlichen Erzeugung globaler Beobachtungsordnungen und ihrer Analyse. In H. Bennani, M. Bühler, S. Cramer, &A.Glauser (Hrsg.), Global beobachten und vergleichen: Soziologische Analysen zur Weltgesellschaft (S. 435–472). Campus. https://doi.org/10.5281/zenodo.5060095 Hoggenmüller, S. W. (2022). Globalität sehen: Zur visuellen Konstruktion von „Welt“. Campus. https:// doi.org/10.12907/978-3-593-44235-8 Hoggenmüller, S. W., & Raab, J. (2022). Bilder. In N. Baur & J. Blasius (Hrsg.), Handbuch Methoden der empirischen Sozialforschung (S. 1581–1598). Springer VS. https://doi.org/10.1007/978-3-65837985-8_110 Imdahl, M. (2006 [1994]). Ikonik: Bilder und ihre Anschauung. In G. Boehm (Hrsg.), Was ist ein Bild? (S. 300–324). Wilhelm Fink. Ishwarappa, K., & Anuradha, J. (2015). A Brief Introduction on Big Data 5Vs Characteristics and Hadoop Technology. Procedia Computer Science, 48, 319–324. https://doi.org/10.1016/j.procs.2015.04.188 Jaton, F. (2017). We Get the Algorithms of Our Ground Truths: Designing Referential Databases in Digital Image Processing. Social Studies of Science, 47(6), 811–840. https://doi. org/10.1177/0306312717730428 Joo, J., & Steinert-Threlkeld, Z. C. (2018). Image as Data: Automated Visual Content Analysis for Political Science. arXiv. Preprint / Working Paper. https://doi.org/10.48550/arXiv.1810.01544 Kanaparthi, S. K., & Raju, U. S. N. (2022). Content-Based Image Retrieval on Big Image Data Using Local and Global Features. International Journal of Information Technology, 14(1), 49–68. https:// doi.org/10.1007/s41870-021-00806-8 Kaplan, F. (2015). A Map for Big Data Research in Digital Humanities. Frontiers in Digital Humanities, 2(1), 1–7. https://doi.org/10.3389/fdigh.2015.00001 Kauppert, M., & Leser, I. (Hrsg.). (2014). Hillarys Hand: Zur politischen Ikonographie der Gegenwart. transcript. https://doi.org/10.1515/transcript.9783839427491
220 Sebastian W. Hoggenmüller SJS 51 (2), 2025, 203–223 Khan, M. A.-u.-d., Uddin, M. F., & Gupta, N. (2014). Seven V’s of Big Data understanding Big Data to extract value. In Proceedings of the 2014 Zone 1 Conference of the American Society for Engineering Education (S. 1–5). IEEE. https://doi.org/10.1109/ASEEZone1.2014.6820689 Kitchin, R. (2014a). Big Data, New Epistemologies and Paradigm Shift. Big Data & Society, 1(1), 1–12. https://doi.org/10.1177/2053951714528481 Kitchin, R. (2014b). The Data Revolution: Big Data, Open Data, Data Infrastructures & Their Consequences. Sage. https://doi.org/10.4135/9781473909472 Kitchin, R. (2017). Big data: Hype or Revolution. In L. Sloan & A. Quan-Haase (Hrsg.), The SAGE Handbook of Social Media Research Methods (S. 27–39). Sage. https://doi.org/10.4135/9781473983847.n3 Kitchin, R. (2021). Data Lives: How Data Are Made and Shape Our World. Bristol University Press. https://doi.org/10.56687/9781529215649 Kitchin, R. (2025). Critical Data Studies: An A to Z Guide to Concepts and Methods. Polity Press. Kitchin, R., & McArdle, G. (2016). What Makes Big Data, Big Data? Exploring the Ontological Characteristics of 26 Datasets. Big Data & Society, 3(1), 1–10. https://doi.org/10.1177/2053951716631130 Klinke, H. (2016). Big Image Data within the Big Picture of Art History. International Journal for Digital Art History, 2, 14–37. https://doi.org/10.11588/dah.2016.2.33527 Knoblauch, H. (2017). Die kommunikative Konstruktion der Wirklichkeit. Springer VS. https://doi.org/10. 1007/978-3-658-15218-5 Knorr Cetina, K. (2001). Viskurse der Physik: Konsensbildung und visuelle Darstellung. In B. Heintz & J. Huber (Hrsg.), Mit dem Auge denken: Strategien der Sichtbarmachung in wissenschaftlichen und virtuellen Welten (S. 305–320). Edition Voldemeer. Kohle, H. (2013). Digitale Bildwissenschaft. Werner Hülsbusch. https://doi.org/10.5282/ubm/epub.25747 Krähnke, U., Pehl, T., & Dresing, T. (2025). Hybride Interpretation textbasierter Daten mit dialogisch integrierten LLMs: Zur Nutzung generativer KI in der qualitativen Forschung. GESIS. Preprint / Working Paper. https://nbn-resolving.org/urn:nbn:de:0168-ssoar-99389-7 Laney, D. (2001). 3D Data Management: Controlling Data Volume, Velocity, and Variety. META Group, Application Delivery Strategies, 949. Langer, S. K. (1987 [1942]). Philosophie auf neuem Wege: Das Symbol im Denken, im Ritus und in der Kunst. Fischer. Lombi, L., & Rossero, E. (2024). How Artificial Intelligence Is Reshaping the Autonomy and Boundary Work of Radiologists: A Qualitative Study. Sociology of Health & Illness, 46(2), 200–218. https:// doi.org/10.1111/1467-9566.13702 Long, D., & Magerko, B. (2020). What Is AI Literacy? Competencies and Design Considerations. In CHI ‘20: Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (S. 1–16). Association for Computing Machinery. https://doi.org/10.1145/3313831.3376727 Loukissas, Y. A. (2019). All Data Are Local: Thinking Critically in a Data-Driven Society. MIT Press. https://doi.org/10.7551/mitpress/11543.001.0001 Luccioni, A. S., Akiki, C., Mitchell, M., & Jernite, Y. (2023). Stable Bias: Analyzing Societal Representations in Diffusion Models. arXiv. Preprint / Working Paper. https://doi.org/10.48550/ arXiv.2303.11408 Lünich, M. (2022). Big Data und Wissen über Gesellschaft: Die Quantifizierung des Sozialen. In M. Lünich (Hrsg.), Der Glaube an Big Data: Eine Analyse gesellschaftlicher Überzeugungen von Erkenntnisund Nutzengewinnen aus digitalen Daten (S. 63–77). Springer VS. https://doi. org/10.1007/978-3-658-36368-0_5 Manovich, L. (2017). Instagram and Contemporary Image. http://manovich.net/index.php/projects/ instagram-and-contemporary-image (letzter Zugriff am 25. September 2024). Manovich, L. (2020). Cultural Analytics. MIT Press. https://doi.org/10.7551/mitpress/11214.001.0001 Manovich, L., & Arielli, E. (2024). Artificial Aesthetics: Generative AI, Art and Visual Media. https:// manovich.net/index.php/projects/artificial-aesthetics (letzter Zugriff am 05. Januar 2025).
Big Visual Data als neue Form des Wissens. Eine Einführung in das Thema und seinen Hintergrund 221 SJS 51 (2), 2025, 203–223 Mayer-Schönberger, V., & Cukier, K. (2013). Big Data: A Revolution That Will Transform How We Live, Work, and Think. Houghton Mifflin Harcourt. Mersch, D. (2006). Visuelle Argumente: Zur Rolle der Bilder in den Naturwissenschaften. In S. Maasen, T. Mayerhauser, & C. Renggli (Hrsg.), Bilder als Diskurse: Bilddiskurse (S. 95–116 ). Velbrück. Miller, H. J. (2010). The Data Avalanche Is Here: Shouldn’t We Be Digging? Journal of Regional Science, 50(1), 181–201. https://doi.org/10.1111/j.1467-9787.2009.00641.x Miró-Llinares, F. (2020). Predictive Policing: Utopia or Dystopia? On Attitudes Towards the Use of Big Data Algorithms for Law Enforcement. IDP: Revista de Internet, Derecho y Politica, 30, 1–18. https://doi.org/10.7238/idp.v0i30.3223 Mitchell, W. J. T. (1992). The Pictorial Turn. Artforum, 30(7), 89–94. Mitchell, W. J. T. (1994). Picture Theory: Essays on Verbal and Visual Representation. University of Chicago Press. Mitchell, W. J. T. (2007). Iconology: Image, Text, Ideology (10. Aufl.). University of Chicago Press. Mohn, B. E. (2023). Kamera-Ethnographie: Ethnographische Forschung im Modus des Zeigens. Programmatik und Praxis. transcript. https://doi.org/10.14361/9783839435311 Monahan, T. (2018). The Image of the Smart City: Surveillance Protocols and Social Inequality. In Y. Watanabe (Hrsg.), Handbook of Cultural Security (S. 210–226). Edward Elgar. https://doi. org/10.4337/9781786437747.00017 Mützel, S., & Unternährer, M. (2024). Digital Payments and Relational Embedding: Turning Relations into Data and Data into Relations. Big Data & Society, 11(3), 1–10. https://doi.org/10.1177/ 20539517241266432 Nguyen-Trung, K., & Nguyen, N. L. (2025). Narrative-Integrated Thematic Analysis (NITA): AI-Supported Theme Generation Without Coding. OSFPREPRINTS. Preprint / Working Paper. https:// doi.org/10.31219/osf.io/7zs9c_v1 Noble, S. U. (2018). Algorithms of Oppression: How Search Engines Reinforce Racism. New York University Press. https://doi.org/10.2307/j.ctt1pwt9w5 O’Neill, S. J., & Smith, N. (2014). Climate Change and Visual Imagery. WIREs Climate Change, 5(1), 73–87. https://doi.org/10.1002/wcc.249 Palmieri, E. (2024). Online Bubbles and Echo Chambers as Social Systems. Kybernetes, 54(4), 2457–2468. https://doi.org/10.1108/K-09-2023-1742 Paul, G. (2006). Visual History. Ein Studienbuch. Vandenhoeck & Ruprecht. Pauwels, L. (2021). Contemplating ‚Visual Studies‘ as an Emerging Transdisciplinary Endeavour. Visual Studies, 36(3), 211–214. https://doi.org/10.1080/1472586X.2021.1970326 Pink, S. (2011). Digital Visual Anthropology: Potentials and Challenges. In M. Banks & J. Ruby (Hrsg.), Made to Be Seen: Perspectives on the History of Visual Anthropology (S. 209–233). University of Chicago Press. Porter, T. M. (1986). The Rise of Statistical Thinking, 1820–1900. Princeton University Press. Prietl, B., & Houben, D. (2018). Einführung: Soziologische Perspektiven auf die Datafizierung der Gesellschaft. In D. Houben & B. Prietl (Hrsg.), Datengesellschaft: Einsichten in die Datafizierung des Sozialen (S. 18–32). transcript. https://doi.org/10.14361/9783839439579-001 Qin, H., Li, X., Yang, Z., & Shang, M. (2015). When Underwater Imagery Analysis Meets Deep Learning: A Solution at the Age of Big Visual Data. In OCEANS 2015 – MTS/IEEE Washington (S. 1–5). IEEE. https://doi.org/10.23919/OCEANS.2015.7404463 Raab, J. (2008). Visuelle Wissenssoziologie: Theoretische Konzeption und materiale Analysen. UVK. Rajagopal, A. (2011). Notes on Postcolonial Visual Culture. BioScope: South Asian Screen Studies, 2(1), 11–22. https://doi.org/10.1177/097492761000200103 Rossi, R., Hirama, K., & Franco, E. F. (2019). A Systematic Literature Map on Big Data. International Journal of Scientific Engineering and Science, 3(11), 25–32. https://doi.org/10.5281/zenodo.3570615
222 Sebastian W. Hoggenmüller SJS 51 (2), 2025, 203–223 Sachs-Hombach, K. (1998). Die Macht der Bilder. Zeitschrift für Ästhetik und Allgemeine Kunstwissenschaft, 43(2), 175–189. Sachs-Hombach, K. (Hrsg.). (2005). Bildwissenschaft: Disziplinen, Themen, Methoden. Suhrkamp. Sachs-Hombach, K. (2021). Das Bild als kommunikatives Medium: Elemente einer allgemeinen Bildwissenschaft (4. Aufl.). Herbert von Halem. https://doi.org/10.1453/2021_9783869625843 Sako, T., & Martinez, A. J. M. (2021). Seeing Poverty from Space: How Much Can It Be Tuned? arXiv. Preprint / Working Paper. https://doi.org/10.48550/arXiv.2107.14700 Schankweiler, K., & Straub, V. (2023). Bildproteste für die Freiheit im Iran: Die Memefication des Widerstands in den Sozialen Medien. 21: Inquiries into Art, History, and the Visual, 4(1), 97–110. https://doi.org/10.11588/xxi.2023.1.93820 Schmitt, M. (2018). Die Soziologie in Zeiten von Big Data. Angebote der Relationalen Soziologie. InD. Houben & B. Prietl (Hrsg.), Datengesellschaft: Einsichten in die Datafizierung des Sozialen (S.299–320). transcript. https://doi.org/10.1515/9783839439579-013 Schnettler, B. (2007). Auf dem Weg zu einer Soziologie visuellen Wissens. Sozialer Sinn, 8(2), 189–210. https://doi.org/10.1515/sosi-2007-0203 Şen, M., Sen, S. N., & Şahin, T. G. (2023). A New Era for Data Analysis in Qualitative Research: ChatGPT! Shanlax International Journal of Education, 11(1), 1–15. https://doi.org/10.34293/ education.v11iS1-Oct.6683 Sfetcu, N. (2023). Impact of Big Data Technology on Contemporary Society. IT & C, 3(1), 11–19. https://doi.org/10.58679/IT73078 Sittel, J. (2017). Digital Humanities in der Filmwissenschaft. MEDIENwissenschaft Rezensionen | Reviews, 34(4), 472–489. https://doi.org/10.17192/ep2017.4.7636 Skarpelos, Y. (2018). Big Visual Data in Social Sciences. In C. M. Stützer, M. Welker, & M. Egger (Hrsg.), Computational Social Science in the Age of Big Data: Concepts, Methodologies, Tools, and Applications (S. 235–265). Herbert von Halem. Smith, K., Piccinini, F., Balassa, T., Koos, K., Danka, T., Azizpour, H., & Horvath, P. (2018). Phenotypic Image Analysis Software Tools for Exploring and Understanding Big Image Data from Cell-Based Assays. Cell Systems, 6(6), 636–653. https://doi.org/10.1016/j.cels.2018.06.001 Sun, Z., Strang, K., & Li, R. (2018). Big Data with Ten Big Characteristics. In ICBDR ‘18: Proceedings of the 2nd International Conference on Big Data Research (S. 56–61). Association for Computing Machinery. https://doi.org/10.1145/3291801.3291822 Szeliski, R. (2022). Computer Vision: Algorithms and Applications. Springer Nature. https://doi. org/10.1007/978-3-030-34372-9 Tindall, D., McLevey, J., Koop-Monteiro, Y., & Graham, A. (2022). Big Data, Computational Social Science, and Other Recent Innovations in Social Network Analysis. Canadian Review of Sociology, 59(2), 271–288. https://doi.org/10.1111/cars.12377 Tosi, D., Kokaj, R., & Roccetti, M. (2024). 15 Years of Big Data: A Systematic Literature Review. Journal of Big Data 11(73), 1–39. https://doi.org/10.1186/s40537-024-00914-9 Uzhinskiy, A., Ososkov, G., Goncharov, P., & Frontsyeva, M. (2018). Combining Satellite Imagery and Machine Learning to Predict Atmospheric Heavy Metal Contamination. In GRID ‘18: Proceedings of the VIII International Conference Distributed Computing and Grid-technologies in Science and Education (S. 351–358). Indico. https://ceur-ws.org/Vol-2267/351-358-paper-67.pdf (letzter Zugriff am 12. August 2024). van Dijck, J. (2014). Datafication, Dataism and Dataveillance: Big Data between Scientific Paradigm and Ideology. Surveillance & Society, 12(2), 197–208. https://doi.org/10.24908/ss.v12i2.4776 Wang, W., Yang, Y., & Pan, Y. (2024). Visual Knowledge in the Big Model Era: Retrospect and Prospect. arXiv. Preprint/Working Paper. https://doi.org/10.48550/arXiv.2404.04308 Wiegerling, K., Nerurkar, M., & Wadephul, C. (2018). Ethische und anthropologische Aspekte der Anwendung von Big-Data-Technologien. In B. Kolany-Raiser, R. Heil, C. Orwat, & T. Hoeren
Big Visual Data als neue Form des Wissens. Eine Einführung in das Thema und seinen Hintergrund 223 SJS 51 (2), 2025, 203–223 (Hrsg.), Big Data und Gesellschaft: Eine multidisziplinäre Annäherung (S. 1–67). Springer VS. https://doi.org/10.1007/978-3-658-21665-8_1 Wilke, R. (2022). Wissenschaft kommuniziert: Eine wissenssoziologische Gattungsanalyse des akademischen Group-Talks am Beispiel der Computational Neuroscience. Springer VS. https://doi.org/10.1007/9783-658-36704-6 William, F. K. A. (2024). My Data Are Ready, How Do I Analyze Them: Navigating Data Analysis in Social Science Research. International Journal of Scientific Research and Management, 12(3), 1730–1741. https://doi.org/10.18535/ijsrm/v12i03.sh03 Wolbring, T. (2020). The Digital Revolution in the Social Sciences: Five Theses About Big Data and Other Recent Methodological Innovations from an Analytical Sociologist. In S. Maasen & J.-H.Passoth (Hrsg.), Soziologie des Digitalen – Digitale Soziologie? (S. 60–72). Nomos. https://doi. org/10.5771/9783845295008-60 Zulli, D., & Zulli, D. J. (2022). Extending the Internet Meme: Conceptualizing Technological Mimesis and Imitation Publics on the TikTok Platform. New Media & Society, 24(8), 1872–1890. https:// doi.org/10.1177/1461444820983603
230 Ajit Singh SJS 51 (2), 2025, 225–245 Metaprozess (vgl. Krotz, 2007; Knoblauch, 2013; Hepp et al., 2015) beschreibt die relationalen, nicht-linear zu lesenden Wechselwirkungen von historischem Medienwandel und kommunikativem (Wirk-)Handeln (vgl. Reichertz & Bettmann, 2018). Aufgrund eines weiteren Mediatisierungsschubs wird in jüngerer Zeit eine „tiefe Mediatisierung“ (Hepp, 2018, S. 35) diagnostiziert. Noch umfassender als zuvor wird damit die gesellschaftliche Konstruktion der Wirklichkeit medial vermittelt und kommunikatives Handeln durch die Einbindung in Plattformen und digitale Infrastrukturen vernetzt und datafiziert. Dergestalt wirkt sich die Mediatisierung nicht nur auf die Räumlichkeit, sondern auch auf die Zeitlichkeit kommunikativer Handlungen aus, „indem [es] die Wissensbestände in technologische Wissensträger auslagert, sie sozial entstrukturiert und in einer Dauerpräsenz verfügbar machen kann, die Zukunft und Vergangenheit verschmelzen lässt“ (Knoblauch, 2017, S. 341). Die Mediatisierungsperspektive rückt damit „[n]icht die Übermittlung (von Informationen), nicht Kommunikationskanäle, sondern die Vermittlung von Wahrnehmen, Denken und Handeln durch (im weiteren Sinne: technische) Medien“ (Pfadenhauer & Grenz, 2017, S. 5) in den Fokus. Überträgt man diese Gedanken auf die planerische Fokussierung auf digitale Objekte und Infrastrukturen, so gestaltet sich die zeitliche und soziale Ordnung kommunikativer Handlungen als eine intraund interaktive Synthese von körperlich-sinnlichen und technologischen Entitäten: sowohl als mediatisiertes Verhältnis zwischen Subjekten als auch zwischen Subjekten und interpretationsbedürftigen, prozessualen, digitalen Artefakten (3D-Modelle, Graphiken, Zahlen, Metadaten etc.). Digitales Planen verstehe ich in diesem Zusammenhang als einen spezifischen Modus mediatisierten kommunikativen Handelns, der menschliche Akteur*innen wie nichthumane Entitäten miteinander in ein relationales Verhältnis setzt und in zeitlicher, (digital-)materieller und sozialer Hinsicht auf die objektivierende Umsetzung von Planungsschritten gerichtet ist. Die visuelle und kommunikative Manifestation von synchronen und asynchronen Planungshandlungen und das in technische Datenträger eingeschriebene heterogene (Fach-)Wissen bildet sich zentral in der Objektivation von Koordinationsmodellen ab, die ich als synthetisches Objekt konzeptualisiere. Synthetisch meint weniger ein Verschwimmen von Akteursgrenzen, sondern eine sozio-technische Verschränkung zeichenhafter Sinnsetzungen, die einerseits im Zusammenspiel von Technologie und kommunikativem Handeln konstruiert werden, und die andererseits intersubjektiv und kommunikativ auslegungsbedürftig bleiben. Mit Knoblauch lässt sich dies als Indiz einer „Kommunikationsgesellschaft“ (2017, S. 329ff.) lesen, in der insbesondere die „Kommunikationsarbeit“ (1996, S. 344) ein wesentlicher Produktionsfaktor ist und Sinn menschlich und – in Erweiterung dazu – auch technisch und durch Algorithmen (re-)produziert wird.
Big Visual Data in der digitalen Infrastrukturplanung. Zur Prozessualität und Materialität … 231 SJS 51 (2), 2025, 225–245 3 Methodisches Vorgehen und Datengrundlage Meine hier zugrunde gelegten qualitativen Analysen stützen sich auf verschiedene Datensorten.3 Neben Felddokumenten (u. a. Webseiten, Strategiepapiere, Berichte aus der internen Unternehmenskommunikation) handelt es sich um Notizen aus ethnographischen Feldaufenthalten, Audio-Aufzeichnungen, visuelle Daten und vor allem elf qualitative Interviews (ca. 13 Std.). Im vorliegenden Text gehe ich (vorrangig) auf die Interviews zu einer Fallstudie in einem Betrieb ein, der BIM seit rund zehn Jahren projektbasiert erprobt. Das Unternehmen ist die Tochtergesellschaft eines Großkonzerns, der sich u. a. mit der Infrastrukturplanung von Mobilitätsund Transportwegen befasst. Darunter fallen Verkehrswege, aber auch andere Bauwerke wie Gebäude und Brücken. Die Tochtergesellschaft ist ein global agierendes Subunternehmen mit mehreren Tausend Mitarbeiter*innen an unterschiedlichen Standorten weltweit: Planen, Beraten und die Durchführung von Planungsprojekten, und damit ein umfassendes Projektmanagement, zählt zum Kerngeschäft, beginnend bei ersten Bedarfsanalysen bis hin zur Umsetzung auf der Baustelle. Zum Zeitpunkt der Erhebung (2022–2023) implementiert die Tochtergesellschaft im Rahmen eines Change Managements BIM in ihre Organisationsstrukturen und entwickelt damit nun die für den Betrieb notwendigen Standards, an denen sich die Mitarbeiter*innen in den jeweiligen BIM-gestützten Projekten orientieren sollen. Die leitfadengestützten qualitativen Interviews (vgl. Bogner et al., 2009; Helfferich, 2010) wurden sowohl mit Expert*innen auf der höheren Leitungsebene als auch mit Mitarbeiter*innen auf der mittleren Hierarchieebene durchgeführt: also mit Akteur*innen, die das Change Management und die Implementierung in der Tochtergesellschaft federführend verantworten, sowie mit BIM Manager*innen, BIM Koordinator*innen und Fachplaner*innen, die mit BIM in verschiedenen Projekten im Tagesgeschäft arbeiten und über eine Einschätzung zum praktischen Entwicklungsstand der Methode verfügen. Unsere offenen Fragen richteten sich erstens auf die Auswirkungen auf das Professionswissen, die infolge der digitalen und praktischen Neuerungen durch BIM in der Planung entstehen. Zweitens fragten wir nach dem Stellenwert, den BIM als kollaborative Planungsmethode in einzelnen Projekten einnimmt und wie es sich in der Praxis zeigt. Und drittens fragten wir nach dem Status des digitalen BIM Modells und dessen Bedeutung für die Planungskommunikation. Die aufgezeichneten Interviews wurden in Anlehnung an GAT2 (Selting et. al., 2009) transkribiert und mit Hilfe des Kodierparadigmas der Grounded Theory (Corbin & Strauss, 2007) analysiert. Unsere Kodierung zielt im ersten Schritt darauf, das Material aber auch die durch unsere Interviewfragen 3 Die zu diesem Zeitpunkt vorliegenden Daten habe ich gemeinsam mit Marie Marleen Heppner im Rahmen des von der DFG geförderten Forschungsprojektes Synthetische Planung – Digitale Mediatisierung von kollaborativer Kommunikationsarbeit und Veränderungen von Planungswissen (Laufzeit: 2022–2025) erhoben und ausgewertet. Die im Text aufgeführten Interviewzitate werden anonymisiert dargestellt.
232 Ajit Singh SJS 51 (2), 2025, 225–245 gesetzten Themen und Konzepte (wieder) aufzubrechen, um auf diesem Wege die Daten aus sich heraus (etwa durch in-vivo Codes) zu erschließen. Dabei richten wir fortwährend Fragen an das Material und auch an einzelne Codes, um spezifische Phänomene zu identifizieren. Ausgewählte, kürzere Passagen werden dann auch einer sequenziellen line-by-line Feinanalyse unterzogen. Die Analyse zielt damit nicht zwangsläufig auf die Bildung einer umfassenden Theorie, wie oftmals suggeriert wird, sondern auf die Entwicklung empirisch begründeter, theoretisierender Konzepte. 4 Empirische Befunde zur modellbasierten Planung mit BIM Die Darstellung der empirischen Ergebnisse erfolgt in einem Dreischritt: Zunächst wird die Genese des BIM-Modells beschrieben und rekonstruiert, welchen Status das Modell aus Sicht der Planer*innen hat (4.1). Danach richtet sich der Blick auf das visuelle und technische Professionswissen der Planer*innen, das sowohl für die Anwendung der BIM-Methode als auch für den Umgang mit Big Visual Data relevant ist (4.2). Schließlich werden die Verflechtungen von planerischem Handeln mit digitalen Planungsinfrastrukturen herausgearbeitet und wie diese zur zeitlichen und kommunikativen Handlungskoordination beitragen (4.3). 4.1 Von der Punktwolke zum 3D-Modell oder: Wo die „planerische Wahrheit“ liegt Die Arbeit am dreidimensionalen Modell spielt eine entscheidende Rolle für die planerische Integration großer heterogener Datensätze, denn das Modell bildet die Grundlage für alle weiteren Planungskommunikationen zwischen den verschiedenen Gewerken. Folglich ist es aufschlussreich, wie ein dreidimensionaler Entwurf konzipiert wird und welchen Stellenwert die visuelle Form aus Sicht der Planer*innen hat. Modelle entstehen selten aus dem Nichts – metaphorisch sprechen die interviewten Planer*innen vom „luftleeren Raum“ oder der „grünen Wiese“ –, sondern haben oftmals bereits einen realen Bezugspunkt. Damit verbunden sind einerseits planungsrechtliche Vorgaben und die Einteilung in Leistungsphasen gemäß der Honorarordnung für Architekt*innen und Ingenieur*innen (HOAI). Andererseits richten sich viele Flächenund Gebäudeplanungen auf bereits existierende räumliche und physisch-materielle (An-)Ordnungen von Objekten, die als Bestand gelten und den visuellen Ausgangspunkt der Planung bilden. Während etwa Flächennutzungspläne Bestandsräume optisch in 2D parzellieren und perspektivische Nutzungsmöglichkeiten der beplanten Fläche aufzeigen (vgl. Singh & Meißner, 2023, S. 241f.), wird der Bestand im Zuge der BIM-Methode über einen dreidimensionalen Laserscan (Ehm &Hesse, 2014) erfasst. Der Flächenraum wird, etwa über eine Drohne, digital bis ins kleinste Detail vermessen und dann in eine Punktwolke überführt, die die unüberschaubar große Menge von kleinsten Messpunkten in eine visuelle Darstellung übersetzt.
Big Visual Data in der digitalen Infrastrukturplanung. Zur Prozessualität und Materialität … 233 SJS 51 (2), 2025, 225–245 Die Punktwolke (in Verbindung mit Objektmodellen vgl. Abb. 1) bildet dann die digital-materielle Grundlage für die Erstellung des Planungsmodells, das weiterer teils automatisierter, teils manueller Anpassungen und Nachmodellierungen der Punktwolke bedarf: Früher, […] ja, man hat von der Vermessung so Lagepläne bekommen wo Linien waren oder vielleicht noch Querschnitte mit Höhen und DAS war halt schon eine große Veränderung, dass man halt plötzlich Punktwolken hatte, wo man ja das alles schon sieht […] wenn man jetzt rausgeht, so ähm Ortsbesichtigungen macht man ja auch und macht Bilder, ist ja meistens das, was man braucht, da hat man jetzt kein Foto von oder jetzt nicht irgendwie gefilmt. Und der Vorteil ist halt da in der BIM-Methode sehr hoch. Ich habe eine Punktwolke, die ist coloriert, da sind Bilder hinterlegt, ich sehe da alles. Ich habe so dreihundertsechzig Grad Panorama je nachdem, was halt gemacht wird. Und da hat man ne ganz andere Qualität von Daten […]. (BIM Koordinatorin) Abbildung 1 Mischung aus Punktwolke und Gebäudemodell (in der Cloudsoftware Cintoo) Quelle: Eigene Videoaufzeichnung [Screenshot, Singh].
234 Ajit Singh SJS 51 (2), 2025, 225–245 In dem Auszug hebt die befragte BIM-Koordinatorin zunächst den Unterschied zu den herkömmlichen Darstellungen hervor, die „früher“ in der Phase der Grundlagenermittlung auf zweidimensionalen Lageplänen („Linien“ und „Querschnitte mit Höhen“) basierten. Diese Art der geometrischen Darstellung vermittelt mit einem gewissen Verbindlichkeitsgrad Wissensstände über bauoder privatrechtliche Rahmenbedingungen (Erschließungen, Zuwege etc.). Mit BIM und der Erzeugung von Punktwolken können Planer*innen nunmehr ganz andere Quellen (in Form kleinster Messpunkte) und damit visuelle Qualitäten von Daten in die Grundlagenermittlung und die Modellierung einbeziehen. Die Punktwolke ist dabei nicht nur coloriert. Durch ihre Ausarbeitung zum Modell und die mögliche 360-Grad-Navigation im Modell erweitert sich die visuelle und räumliche Perspektive auf die zu beplanende Fläche. Die punktwolkenbasierte Erstellung eines Modells ist weder bloße Handwerkskunst noch technische Spielerei. Für die Planer*innen dient das Modell, wie in vielen Interviews betont wurde, als visuelle Quelle der „planerischen Wahrheit“: Das Modell ist eigentlich die planerische Wahrheit und daraus werden die 2D Pläne abgeleitet […]. Diese Modelldetaillierung oder Qualität ist natürlich in den verschiedenen Planungsgewerken unterschiedlich, weil wir erst überlegen müssen, wieviel 3D Modell brauche ich. Zum Beispiel, ich hatte vorhin die Oberleitung erwähnt ne, brauche ich jetzt nur den Oberleitungsmast, brauche ich auch die Fahrleitung, äh was brauche ich wirklich, um diese Informationen zu pflegen und auch gewisse wichtige Kollisionen durchführen zu können. (BIM Managerin) Was von der BIM Managerin als „planerische Wahrheit“ und von anderen Planer*innen auch als „Quelle der Wahrheit“ bezeichnet wird, beschreibt ein Prinzip des auf einer digitalen Infrastruktur aufsetzenden, organisationalen und netzwerkförmigen Handelns und kommt dort zur Geltung, wo Arbeitsphasen separiert ablaufen: Im Feld der BIM-basierten Planung erzeugen die Gewerke zunächst eigene Fachmodelle, die dann zu einem gewerkeübergreifenden Koordinationsmodell zusammengeführt werden. Die Modellierung unterliegt dabei Selektionsprozessen, die auf planerischen Relevanzsetzungen basieren, welche Details für wen in welcher Leistungsphase bedeutsam sind. Das Koordinationsmodell erfüllt aus Sicht der Planer*innen schließlich die Funktion einer „single source of truth“, in der (idealerweise) die notwendigen Datensätze („was brauche ich wirklich“) synthetisiert und, etwa im Zuge von Kollisionsprüfungen, miteinander abgeglichen werden, so dass alle (relevanten) Planungsbeteiligten auf der Grundlage eines vollständigen Datensatzes auf einem aktuellen Wissensstand sind. Das visuelle Modell ist folglich der materielle und zeitliche Referenzpunkt für die Planungskoordination. Veränderungen in der Planung sind sichtbare Manifestationen und (visuelle) Weiterentwicklungen im Modell, die für alle beteiligten Planer*innen
Big Visual Data in der digitalen Infrastrukturplanung. Zur Prozessualität und Materialität … 235 SJS 51 (2), 2025, 225–245 erkennbar sein sollten. Die digitale Materialität des Modells ist zugleich nicht losgelöst von den kommunikativen Handlungen des gem-/einsamen, kollaborativen Planens zu betrachten. Kommunikatives Planen und Modellieren amalgamieren im Prozess der synthetisierten Modellerstellung und resultieren im Ergebnis einer interferierenden Kopplung heterogener Daten, teilautomatisierter Technologien und händisch-verkörperter Prozeduren (etwa des Nachmodellierens). Die visuellen Daten haben damit unterschiedliche Entstehungsbedingungen, stammen in Teilen aus heterogenen Datenquellen und kennzeichnen sich auch durch eine fachliche Verschiedenartigkeit. Was als Big Visual Data in der Planung aufscheint, lässt sich im Wesentlichen als eine wandelbare und situativ nach planerischen Bedarfen abrufbare „Wahrheit“ beschreiben, die im synthetischen Objekt verankert ist. Zu klären ist nun, mit welchen (professionsspezifischen und verkörperten) Wissensbeständen und (intraaktiven) Sinnverknüpfungen Bedeutung in der Planung hergestellt wird. 4.2 Visuelles und technisches Professionswissen im Umgang mit digitalen Modellen Zweifelsohne treten alle Daten visuell (als Zeichen, Symbole, Bilder etc.) in Erscheinung. Entscheidend für die relationale Verwendbarkeit visueller Formen ist aber ihre interpretative Bedeutung für die jeweiligen Leistungsphasen der Planung. Die Diskrepanz zwischen verschiedenen visuellen Datentypen lässt sich anhand der aus den Interviews rekonstruierten Feldunterscheidung zwischen „Visualisierung“ und „technischem Modell“ verdeutlichen: Visualisierungen im Stile von Renderings (vgl. Mélix, 2022) konstruieren eine optische und ästhetische Vorstellung von einem Bauprojekt, gegebenenfalls mit „Himmel“ und „Schattenwurf“. In der konkreten Übersetzung in geometrische Anordnungen und Materialpassungen gibt dieser Visualisierungstyp im Gegensatz zum technischen Modell jedoch keine hinreichende Antwort für die konkrete planerische Umsetzung. Vermeintliche Genauigkeiten, die das Ergebnis konkreter Formdarstellungen sind, suggerieren den Stand einer planerischen Wirklichkeit, der in der Regel so nicht gegeben ist. Diese interpretierende und verstehende Einordnung von visuellen Formen und Daten bildet eine wichtige erworbene Handlungskompetenz von Planer*innen und verweist auf eine wissensbasierte, sinnlich-verkörperte Praxis des „professionellen Sehens“ (Goodwin, 1994, S. 626). Angewendet auf das Feld der Planung impliziert dies u. a., spezifische visuelle Wahrnehmungsmuster in planerisches Handeln und visuelle Kommunikationen selektiv zu transformieren. Die von mir kuratierte Darstellung eines Modells (Abb. 2) illustriert vier verschiedene Perspektiven auf einen Trassenabschnitt. Die visuellen Zuschnitte resultieren aus virtuellen Kamerafahrten, die es den Planer*innen ermöglichen, das Modell aus unterschiedlichen Blickwinkeln und mit unterschiedlicher Detailtiefe zu betrachten. BildA zeigt den geplanten Streckenabschnitt aus der Vogelperspektive, wobei die Informationstiefe der Messdaten (farblich differenziert) auf die Trasse
236 Ajit Singh SJS 51 (2), 2025, 225–245 (grau), Gebäude (rot) und die unmittelbare Umgebung (grüner Bereich) gerichtet ist (BildB). In den weiteren Bildern zoomt der interviewte Leiter des BIM Managements dann je nach Fokus spezifische Planobjekte wie einen Mast (BildC) an oder zeigt Simulationen eines Lichtraumprofils (BildD), das für die Bemessung von Abständen zu Gebäuden oder Fahrzeugen bereits in die frühen Phasen der Planung integriert wird. Wie voraussetzungsreich die Lesund Deutbarkeiten der visuellen Komponenten eines Modells sind, zeigt der folgende Interviewauszug mit dem Leiter des BIM Managements. In Abbildung3 ist erneut der bereits gezeigte Bahnhofsabschnitt zu sehen (Abb. 2). Der Schienenbereich erscheint in einem dunkleren Grau, davon abgesetzt in hellgrau ist ein Bahnsteig. Die in blau dargestellte Fläche zeigt das Gelände, auf dem sich die Objekte befinden. Anhand der Darstellung möchte uns der Leiter des BIM-Managements aufzeigen, wie sich unterschiedliche Modelle zueinander verhalten. Hier sieht man, ist das Geländemodell. Und der Bahnsteig hier, der ist jetzt von Hand modelliert nach ner Punktwolke [BILDA]. Das Geländemodell das steht halt so in zehn bis zwanzig Prozent der Fläche raus ähm und den Rest liegts da drunter [BILDB], und das heißt, dass das relativ genau ist. Also das äh ist ’n gutes Indiz [BILDC], dass die Daten, die aus völlig unterschiedlichen Quellen kommen, gut zueinander passen. (Leiter BIM-Management) Abbildung 2 Collage eines BIM Modells (Autodesk Large Model Viewer in Autodesk BIM360) Quelle: Eigene Videoaufzeichnung [Screenshot, Singh].
Big Visual Data in der digitalen Infrastrukturplanung. Zur Prozessualität und Materialität … 237 SJS 51 (2), 2025, 225–245 Für die ungeübten Betrachter*innen stellt sich das Problem, die verschiedenen Modellebenen in der visuellen Darstellung überhaupt zu erkennen: Nämlich, dass hier Messdaten zu Geländebestand und Gebäude mit unterschiedlich festgelegten Standards von verschiedenen Akteure*innen erstellt und zusammengeführt wurden. Auf der Grundlage des Modells machen Planer*innen technisch sichtbar, inwieweit unterschiedliche Datensätze miteinander kombinierbar sind. Letztlich gilt es zu bestimmen, ob die geplanten Objekte (Trassen, Gebäude etc.) zu der jeweiligen Geländebeschaffenheit (also zur lokalen Erdoberflächenstruktur) in einer Passung stehen. Bemerkenswert ist die angesteuerte immersive Perspektive des Bildausschnittes, die die Betrachtenden in das Modell hineinzoomt. Der visuelle Blickwinkel wie auch die sprachliche Erläuterung zu dem, was hier gezeigt werden soll (das Verhältnis von Bahnsteig und Geländemodell), illustrieren schließlich, dass Bedeutungen nicht nur aus dem Modell gewonnen werden („Das Geländemodell das steht halt so in zehn bis zwanzig Prozent der Fläche raus“). Als eine Form der inkorporierten Sehfertigkeit und der situierten Vermittlung einer Sehweise (in den Bildern A und B als von mir in Rot nachgezeichnete Markierungen des Mauskursors) wird professionsbezogenes visuelles Wissen an das Modell herangetragen, um die visuell in Erscheinung tretenden Datenbestände je nach Problemstellung innerhalb der Planung zu lokalisieren, interpretierend zu dekodieren („’n gutes Indiz“) und vor dem Hintergrund der relationalen Passung unterschiedlichen Objekte sinnhaft in die Planung einzuordnen. Dieses Erkennen von Referenzen und Objektbezügen, im Sinne einer Data Literacy, lässt sich zusätzAbbildung 3 Geländemodell (Autodesk Large Model Viewer in Autodesk BIM360) Quelle: Eigene Videoaufzeichnung [Screenshot, Singh].
238 Ajit Singh SJS 51 (2), 2025, 225–245 lich technisch unterstützen, indem Objekte coloriert und voneinander abgegrenzt werden können (BildD). Im praktischen Umgang mit BIM vermengt sich visuelles mit technischem Professionswissen, um die über die Zeit anwachsenden visuellen Datenund Informationsstände verstehend zu verarbeiten. Dabei geht es einerseits um die Ordnung von Daten, andererseits um die standardisierte Verfügbarmachung für den kollaborativen Prozess. Genau darin besteht ein zentrales Problem, weil in dem beschriebenen Unternehmen überwiegend in einzelnen Gewerken (oft umschrieben mit der Ethnokategorie des „Silos“) gearbeitet wird: […] die Daten tragen ja auch Informationen, haben Metadaten und sind dann natürlich über ein gut organisiertes Datenmanagement, was natürlich auch weiter noch entwickelt werden muss, dann auch verknüpft […]. Was die Vision ist, und es sind tatsächlich immer auch teilweise noch Datensilos, also selbst auch bei uns in der [Verkehrsnetz (VN) GmbH] gibt es ja auch verschiedene Bereiche, die sich in einem Projekt in einem Planungsprojekt auch […] mit dem Einsammeln von Daten beschäftigen, bleiben wir mal bei dem Thema Vermessungsdaten, Baugrunddaten, äh Bestandsdaten. Und auch da müssen wir intern irgendwie erst mal schauen, wie die auch gemeinsam verfügbar gemacht werden für die Projekte […]. (BIM Managerin) Die technische und gewerkeübergreifende Verfügbarmachung von Daten stellt eine komplexe Aufgabe dar („irgendwie erst mal schauen“). Wie eingangs erwähnt, handelt es sich um heterogene Daten, die visuelle wie nichtgrafische, alphanumerische oder geometrische Daten beinhalten und nicht zuletzt Wissensbestände, die auf der kollaborativen „Kommunikationsarbeit“ (Knoblauch, 1996, S. 344) zwischen den Gewerken beruhen. Big Visual Data lässt sich demnach nicht nur auf seine sichtbaren Qualitäten begrenzen, sondern ist in Verbindung zu ihrer referenzierten Informationstiefe zu interpretieren. Die Vernetzung und Selektion dieser großen Datensätze und dem, was als Wissen relevant wird und was nicht, ist zugleich ein kommunikativ zu lösendes Problem für die Herstellung einer datenbasierten Ordnung (strukturiert über „Metadaten“ vgl. Schapke et al., 2015, S. 209ff.) in der kollaborativen Planungspraxis. Für die termingerechte Realisierung von Planungsprozessen bedürfen daher alle planungsrelevanten Daten einer Strukturierung, die über die Nutzung von Datenmanagementsystemen und Infrastrukturen gewährleistet wird. 4.3 Vom Modell zur Infrastruktur: Standards, Zeitlichkeiten und Wirkzusammenhänge Die (gelingende) Planung mit Big Visual Data hängt an der Nutzung integrierender „Softwarelandschaften“, die es Planer*innen – im Zusammenhang mit der BIM-
Big Visual Data in der digitalen Infrastrukturplanung. Zur Prozessualität und Materialität … 239 SJS 51 (2), 2025, 225–245 Methode vor allem den BIM Manger*innen und die BIM Koordinator*innen– ermöglichen, Daten zusammenzuführen und zu gliedern. Das damit verbundene Problem des Datentransfers liegt allerdings in den digitalen Planungsinfrastrukturen begründet. Exemplarisch beschrieben wird eine „Hybrid Variante, wo ich halt einen Datenkreislauf habe und durch irgendwelche Nadelöhre muss, mit Einzeldaten und […] gerade in großen Datenpaketmengen wirkt das behindernd“ (Leiter Digitalisierungsabteilung). Sowohl auf Softwareebene, als auch für die Verbindung der sozialen Schnittstellen zwischen den Planungsgewerken spielt die unternehmensinterne und -übergreifende Einführung von Prozessstandardisierungen eine wichtige Rolle. Während Infrastrukturen das Problem technischer, zeitlicher und kommunikativer Abstimmungen lösen sollen, dienen Standards dazu, Übergänge, Vergleichbarkeiten und Vereinheitlichungen herzustellen, ein Phänomen, das Star und Ruhleder (1996) bei ihren Untersuchungen zu kollaborativen Infrastrukturen herausstellten. Mit Infrastrukturen4 bezeichnen sie in analytischer Hinsicht ein relationales Verhältnis von Technik, Arbeit und menschlichem Handeln, dessen Vieldeutigkeit sich (aus einer interaktionistischen Perspektive heraus argumentiert) erst in der praktischen Hervorbringung von Infrastrukturen begreifen lässt (Star & Ruhleder, 1996, S. 113f.). Die organisatorische Transformation von Technik denken Star und Ruhleder zweiseitig als Verbindungen und als infrastrukturelle Hindernisse, die kommunikativ, praktisch und situativ überwunden werden müssen. Und diese Herausforderung stellt sich auch für BIM und die dazugehörige Software: Warum ist Software halt momentan so wichtig oder eine integrierte Softwarelandschaft oder eine offene Softwarelandschaft so wichtig, weil die ganzen Austauschformate, die angeboten werden, […] ganz zu […] Anfang auch äh IFC genannt […] einfach noch nicht so weit sind, insbesondere für die Infrastruktur, um […] Datenaustausch möglich zu machen. Also das Übersetzerfischchen […] von Per Anhalter durch die Galaxis, das fehlt einfach ne. […] Also bleibt einem gar nichts anderes übrig, als eine Art Closed BIM zu verfolgen mit einer integrierten Softwarelandschaft. (Leiter Digitalisierungsabteilung) Die effiziente Nutzbarmachung der heterogenen Datentypen über Datenaustausch ist aus Sicht des Leiters der Digitalisierungsabteilung zunächst in technischer Hinsicht an die Verbindung von digitalen Schnittstellen gebunden. Das „Übersetzerfischchen“, auch bekannt als Babelfisch aus dem im Zitat anklingenden Roman von 4 In Anlehnung an Kropp et al. (2022, S. 244) ließe sich auch von Plattform sprechen: „BIM can be seen as a technology platform in that it entails a modular technological architecture composed of a core and a periphery that allows to create value by generating and harnessing an economy of scope. However, it also shows characteristics of a platform as a market that mediates transactions between planners or clients and producers of building components, such as doors, windows, walls, stairs, or others who can offer their products for sale through BIM software.“
ISBN 978-3-03777-302-4 270 Seiten 15.5 cm × 22.5 cm Fr. 38.– | Euro 38.– Während zu Beginn des Jahrhunderts 1,4 Millionen Ausländer:innen in der Schweiz lebten, ist ihre Zahl heute auf 2,2 Millionen angestiegen. Diese Zunahme geht mit einer starken Veränderung der sozio-professionellen und familiären Strukturen der ausländischen Wohnbe-völkerung einher. Dieses Buch zeigt die wirtschaftlichen und geopolitischen Faktoren auf, die diesem Wandel zugrunde liegen, und zeichnet die Entwicklung des gesetzlichen Rahmens in der Schweiz in den letzten zwei Jahrzehnten nach. Im Zuge der tiefgreifenden Veränderungen, die sowohl die Zulassungsregelungen als auch die Integrationsund Einbürgerungspolitik neu gestaltet haben, hat sich die Migrationslandschaft rasch gewandelt. Infolgedessen sind die soziodemografischen Merkmale der Zuwanderer, die Familiendynamik sowie das Integrationsund Einbürgerungsverhalten einer ständigen Transformation unterworfen. Gestützt auf statistische Originalquellen erstellen die Autor:innen eine Bestandesaufnahme der Arbeitsmigration in der Schweiz. Dabei gehen sie auf die drei grössten ausländischen Bevölkerungsgruppen ein: Italiener:innen, Deutsche und Portugies:innen. Auf der Grundlage dieser Ergebnisse skizzieren die Autor:innen die Prioritäten für die zukünftige Steuerung der Migrationspolitik. Somit stellt das Buch eine wichtige Grundlage für die aktuelle gesellschaftliche und politische Debatte in der Schweiz dar. Philippe Wanner ist Professor für Demografie an der Universität Genf, Schweizer Migrationsexperte und Stellvertretender Direktor des Nationalen Forschungsschwerpunkts über Migration und Mobilität «nccr – on the move». Rosita Fibbi ist ehemalige Dozentin für Soziologie der Migration an der Universität Lausanne und assoziierte Forscherin am Nationalen Forschungsschwerpunkt über Migration und Mobilität «nccr – on the move». Die Schweizer Migrationslandschaft im 21. Jahrhundert Seismo Verlag, Zürich und Genf www.seismoverlag.ch [email protected] Seismo Verlag Sozialwissenschaften und Gesellschaftsfragen Philippe Wanner, Rosita Fibbi (Hrsg.) Reihe Sozialer Zusammenhalt und kultureller Pluralismus
247 Swiss Journal of Sociology, 51 (2), 2025, 247–272 * Technische Universität Berlin, Institut für Soziologie, D-10587 Berlin, [email protected], [email protected]. Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung von Big Visual Data Mina Godarzani-Bakhtiari* und René Tuma* Zusammenfassung: Visuelle Dokumente von Gewaltereignissen sind oft umstritten und fragmentiert. Ihre Interpretation bedarf zusätzlicher Legitimation. Anhand des Fallbeispiels Hanau analysieren wir die Arbeit von Forensic Architecture (FA) im Kontext der Debatte um Big Visual Data. FA entwickelt Analyseverfahren und Darstellungsformen, um den Herausforderungen von Big Visual Data zu begegnen. Wir zeigen, wie FA audiovisuelle Daten als dynamische Referenzquellen behandelt und wie durch die Verschachtelung von Bildern ein neuer Blick im alltäglichen Feld der Analyse etabliert wird. Schlüsselwörter: Visuelle Soziologie, vernakulare Analysen, Forensic Architecture, digitale Raum-Modelle, Sehpraktiken From Flat Image to Spatialised Visual Analysis: Forensic Architecture and the Interweaving of Big Visual Data Abstract: Visual documents of violent events are often controversial and fragmented. Their interpretation requires additional legitimisation. Using the Hanau reconstruction as a case study, we analyse the work of Forensic Architecture (FA) in the context of the debate about Big Visual Data. FA develops analytic methods and forms of presentation to meet the challenges of Big Visual Data. We show how FA treats audiovisual data as dynamic reference sources and how the nesting of images establishes a new perspective in the vernacular field of analysis. Keywords: Visual sociology, vernacular analyses, Forensic Architecture, digital space models, visual practices De l’image plate à l’analyse visuelle spatialisée: Forensic Architecture et l’imbrication des mégadonnées visuelles Résumé: L’interprétation des documents visuels d’événements violents, souvent controversés et fragmentés, nécessite une légitimation supplémentaire. Sur la base du cas de Hanau, nous analysons le travail de Forensic Architecture (FA) dans le contexte du débat sur les mégadonnées visuelles. Pour répondre aux défis des mégadonnées visuelles, FA développe des méthodes d’analyse et des formes de représentation. Nous montrons comment FA traite les données audiovisuelles comme des sources de référence dynamiques et comment l’imbrication d’images permet d’établir un nouveau regard dans le champ d’analyse courant. Mots-clés: Sociologie visuelle, analyses vernaculaires, Forensic Architecture, modèles numériques d’espace, pratiques visuelles DOI 10.26034/cm.sjs.2025.6905 © 2025. This work is licensed under the Creative Commons Attribution-NonCommercialNoDerivatives 4.0 License. (CC BY-NC-ND 4.0)
248 Mina Godarzani-Bakhtiari und René Tuma SJS 51 (2), 2025, 247–272 1 Einleitung1 In diesem Beitrag untersuchen wir die Arbeit von Forensic Architecture (FA) als Beitrag zur Debatte um Big Visual Data. FA ist eine prominente Forschungsagentur, welche mit ihren Investigationen, die überwiegend auf visuellen Daten und raumanalytischen Methoden basieren, Vorwürfe von Vergehen und Verbrechen, die staatlichen Organisationen zugeschrieben werden, analysiert und öffentlich macht. Dabei bedienen sie sich eines avancierten wissenschaftlichen und forensischen Methodenapparates und einer theoretisch legitimierten Epistemologie, die sich aus politischen, raumund architekturtheoretischen und damit auch philosophischen und sozialwissenschaftlichen Debatten speist. Die Beiträge FAs, die wir als Teil eines Feldes vernakularer Analysen verstehen, die sich zwischen institutionellen Feldern bewegen und eigene kommunikative Formen der Analyse ausbilden, verweisen auf eine paradigmatische Herausforderung: Die Menge, Vielfalt und Komplexität audiovisueller Daten wirft Fragen für verschiedene Wissensfelder auf, und so stehen auch zivilgesellschaftliche Initiativen vor konkreten Handlungsproblemen – aber auch Chancen – im Umgang mit dem Phänomen Big Visual Data. Visuelle Abbilder sind durch Repräsentationsprobleme gekennzeichnet, strahlen nicht (mehr) unhinterfragte Objektivität aus und bedürfen zunehmend zusätzlicher Legitimation, um in ihrer argumentativen und evidenzbegründenden Verwendung kommunikative Kraft zu entfalten. Diese Themen sind nicht neu und werden an verschiedenen Stellen, z. B. in der Bildtheorie, und in der Medienwissenschaft, z. B. am Beispiel der „Smoking Gun“ in der Weltpolitik, seit langem verhandelt (Holert, 2004). Mit Big Visual Data spitzt sich die Problematik zu. Die omnipräsente Verfügbarkeit von Bildern ist eine praktische Ressource, aber auch ein Problem, besonders im Kontext des prekären Evidenzoder Indiziencharakters visueller Artefakte (vgl. auch Nohr, 2004). Damit verschiebt sich die Debatte um Deutungshoheit weg vom einzelnen Bild hin zu den Methoden der Analyse und den Formen ihrer Präsentation. In klassischen Vorgehensweisen der Forensik, und der Evidenzproduktion sind neben der Zeugenaussage für das Zum-Sprechen-Bringen von Bildern vor allem durch Professionen legitimierte und staatlich bestellte Sachverständige verantwortlich, die ihre Legitimität aus klassischen Professionen generieren (Schwartz, 2009; Milroy, 2017). In der Arena des öffentlichen Diskurses (siehe z.B. Seeliger & Sevignani, 2021) sind es zunehmend neue sich durch Methoden performativ legitimierende Akteur*innen, die Deutungswissen und visuelltechnische Expertise – also Methoden der forensischen Spurensuche und reflexive 1 Wir danken Sebastian W. Hoggenmüller, dem Herausgeber dieses Sonderhefts, den Gutachter*innen, Tom Berger, Frederike Brandt, Nele Dörl, Simon Egbert, Annika Haller, Lena Schober, Vivien Sommer und Talia Tuana Yücel für ihre Anregungen. Die Arbeit wurde von der Deutschen Forschungsgemeinschaft (DFG) gefördert – 502722049.
Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung … 249 SJS 51 (2), 2025, 247–272 Sehanleitungen – zur Verfügung stellen. Sie überschreiten die Grenzen etablierter Wissensfelder und bringen ihre Expertise in konkrete Anwendungsbereiche ein. Das fassen wir mit dem Begriff der Vernakularität. Expertise – und das möchten wir betonen – wird in spezifischen kommunikativen Formen reflexiv inszeniert, die damit insbesondere auch spezifische Kompetenzdarstellungskompetenzen zum Ausdruck bringen (vgl.Pfadenhauer & Dieringer, 2019).2 Anhand der investigativen Rekonstruktion des polizeilichen Umgangs mit dem rechtsextremen Anschlag in Hanau durch FA gehen wir der Frage nach, welche reflexiv-kommunikativen Formen der Analyse und ihrer Darstellung FA nutzt, um mit den Herausforderungen von Big Visual Data umzugehen. Wir argumentieren, dass das Aufkommen von FA (oder auch ähnlicher Initiativen wie Bellingcat und verschiedener journalistischer Formate) keinen Sonderfall darstellt, sondern eine paradigmatische Antwort auf ein politisches wie auch pragmatisches kommunikatives Problem der Gegenwart ist. Die Etablierung und Verbreitung dieser Analysemethoden und ihrer kommunikativen Formen bringen vereinzelte und prekär gewordene Bilder durch spezifische kommunikative Formen zum Sprechen oder vielmehr zum Zeigen. Die Besonderheit von FA liegt in der Relationierung von Bildern durch Raummodellierung und der reflexiven Offenlegung dieser Prozesse mittels Meta-Artfakte. Unsere Analyse knüpft an Konzepte wie Professional Vision (Goodwin, 1994) sowie an weitere feldüberschreitende, vernakulare Analyseformen an (vgl. Abschnitt2.4, Tuma, 2017). Anhand des Fallbeispiels rekonstruieren wir sechs verschiedene, ineinander greifende Methoden des Umgangs mit visuellen Artefakten, die FA anwendet: 1.Kontextualisierung; 2.Kontexturalisierung; 3.Verifikation; 4.Selektion und Extrahierung; 5.Vermessungen und 6.Synthese. Das Besondere an den Analysen von FA ist, dass sie bei der Analyse von audio/visuellen Daten über simulative Raumerweiterungen tiefe und breite Analysen der Referenzdaten vornehmen, mit denen die inhärente Begrenztheit der Daten kommunikativ überzeugend überwunden wird. Diese aufwendigen Analyseverfahren werden in visuellen MetaArtefakten (hier ein Video, aber auch Ausstellungen etc.) objektiviert, die bestrebt sind, nicht nur das Ergebnis, sondern reflexiv eben jene Praktiken der Herstellung mit-sichtbar zu machen. 2 Die Politisierung des Visuellen und Forensic Architecture In der Welt der mediatisierten Sichtbarkeit ist das visuelle Aufzeigen von Handlungen und Ereignissen eine wesentliche Strategie der Verhandlung des Politischen ( Thompson, 2005, 31). Insbesondere Gewaltereignisse sind in ihrer Produktion – wie auch der nachträglichen Vermittlung – Teil der Strategien der Beteiligten (Fujii et al., 2 Vgl. zur personalisierten Performanz von Expertise (Hill, 2022).
250 Mina Godarzani-Bakhtiari und René Tuma SJS 51 (2), 2025, 247–272 2021) und Teil der Konstitution von Events als historische Ereignisse (Wagner-Pacifici, 2017).3 Sie werden in den Nachrichten gezeigt, von Aktivist*innen aktiv präsentiert, teilweise gezielt produziert oder versteckt, abhängig vom jeweiligen Fall und der Positionierung der Beteiligten. Darüber hinaus erhalten auch beobachtende Dritte (vgl. Coenen & Tuma, 2022) durch die Allgegenwärtigkeit visueller Technologien eine niedrigschwellige Möglichkeit, sich an den Auseinandersetzungen zu beteiligen. Öffentliche Aufmerksamkeit erhalten vor allem Bilder sichtbarer Gewalt, die oft mit Smartphones und anderen Aufnahmegeräten aufgenommen werden und dann Teil des Gewaltdiskurses werden (Hoebel et al., 2022). Gleichzeitig werden visuelle Repräsentationen zunehmend aufgrund ihrer potenziellen digitalen Manipulierbarkeit in Frage gestellt. Daher gewinnt die Verifizierung, die reflexive Thematisierung von Repräsentationen und die Analyse von (visuellen) Dokumenten an Relevanz. Gewaltbilder sind also ein zentraler Gegenstand professioneller und vernakularer Bildanalysen. Verschiedene professionalisierte Akteur*innen unterziehen sie einer systematischen und detaillierten Analyse auf ihre Authentizität, Aussagekraft und Bedeutung. Dabei wird, ausgehend von der neuen (auch umstrittenen) Bildervielfalt die Fragwürdigkeit der Bilder selbst zum systematischen Teil der Diskurse um gewaltvolle Ereignisse. Dies zeigt, dass es bei der Verhandlung von Gewalt durch visuelle Daten nicht mehr nur um das Hervorheben spezifischer Unsichtbarkeiten geht, sondern um den Kampf um Deutungshoheit. Vor dem Hintergrund von Big Visual Data entstehen nicht nur Organisationen und professionelle Rollen, sondern insbesondere spezifische Methoden und vor allem kommunikative Formen, mit denen zunehmend auch nichtstaatliche Organisationen und Akteur*innen, z. B. im Rahmen von Open Source Intelligence (OSINT) auf die Prekarität der Bilder reagieren. Diese Umstrittenheit verschiebt die Gültigkeit der Bilder weg vom einzelnen Bild, hin zu ganzen Bilderserien, Videosammlungen und Archiven. Damit verlagert sich die Auseinandersetzung um die Legitimation von Deutungsmacht vom scheinbar einfachen sichtbaren Bild hin zu komplexen Methoden der Evidenzkonstruktion, die nun selbst legitimiert und präsentiert werden müssen (vgl. speziell zur Logik von Metabildern Hoggenmüller & Klinke in diesem Sonderheft). Die Frage, wie die visuellen Artefakte als Indizien oder Evidenzen interpretiert, wie sie als Spuren gelesen – genauer: konstruiert – werden (Krämer et al., 2007; Ginzburg, 2011; zur polizeilichen Spurenkonstruktion vgl. Reichertz, 1991; zu FA vgl. Harst, 2023), und wie und wo die Ergebnisse präsentiert werden, manifestiert sich als Diskursarena. Die Auseinandersetzung adressiert zunehmend methodische und technologische Aspekte der Bildinterpretation als zentrale Aspekte der Wirklichkeitskonstruktion. Somit werden die Formen der Analyse und die Auseinandersetzung zunehmend Teil der (Macht-)Diskurse. 3 Für Lindemann (2017, S.81) ist die diskursive Verhandlung zentraler Bestandteil der modernen Verfahrensordnung der Gewalt.
Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung … 251 SJS 51 (2), 2025, 247–272 Wir untersuchen im Folgenden die 2010 an der Londoner Goldsmith University gegründete Forschungsagentur FA, die sich der interpretativen Aushandlung von Gewalt auf der Basis vor allem visueller Daten widmet. Wir wählen sie aufgrund ihrer Pionierstellung und ihres Bekanntheitsgrades exemplarisch aus. FA’s interdisziplinäres Team führt im Auftrag von internationalen Staatsanwaltschaften und Umweltund Menschenrechtsinitiativen unabhängige Untersuchungen zu staatlicher oder durch Staaten verdeckter/unbearbeiteter Gewalt durch.4 Dabei verfolgt FA das Ziel, den forensischen Gaze umzukehren, was auch in der Selbstbezeichnung der Praxis als counter-forensics zum Ausdruck kommt. Gegenforensich geht sie insofern vor als sie den forensischen Blick auf staatliche Organisationen (Polizei oder Militär) selbst anwenden, welche eigentlich den forensischen Blick monopolisiert haben (Weizman, 2017, S. 9). Dabei positioniert sich FA als Vertreter*in der kämpferischen und aktivistischen Forschung, welche die Aufgabe verfolge, durch Materialund Medienanalysen neue Formen der Zeugenschaft zu schaffen (Bois et al., 2016, S. 121). Die Auswertung großer Mengen visueller Daten ist dabei zentral. FA rekonstruiert gewalttätige Ereignisse durch die Open-Source-getriebene Sichtung und Analyse vielfältiger Daten und durch die Anwendung neuer, vor allem raumbezogener Methoden. Den Anspruch der staatlichen Forensik mit eigens entwickelten epistemischen, methodologischen und methodischen Ansätzen entgegenzutreten leitet FA (Franke et al., 2014, S. 10) aus den gegebenen politischen Umständen ab: It is precisely because of the potential political agencies and the complexity of the emerging scientific-aesthetic-linguistic field of forensics that a new forensis must emerge to challenge the assumptions of received forensic practices. Diese Neuverhandlung ist bei FA in der Verwendung des Begriffs Forensic eingeschrieben, welcher als ästhetische Praxis umgedeutet, drei operational sites verbinden soll: Erstens das Feld, also den Tatort, an dem materielle Objekte Veränderungen der Umwelt registrieren würden; zweitens das Lab/Studio, in dem diese Veränderungen analytisch untersucht werden; und drittens das Forum, in dem die Ergebnisse narrativiert, diskutiert und Wahrheitsansprüche artikuliert werden (Weizman, 2017, S.94). Indem sich FA mit allen drei Sites auseinandersetzt, forciert sie durch die forensische Ästhetik die Zusammenführung von Ereignis und Öffentlichkeit. Damit zielt das Vorgehen von FA darauf ab, die im Zentrum stehenden Ereignisse zu politisieren, die kritische Debatte in der Öffentlichkeit zu fördern und zivilgesellschaftliche Kämpfe zu unterstützen. Dass FA eine methodische „Demokratisierung“ (Forensic Architecture o. J.) anstrebt, wie sie es nennt, zeigt sich darin, dass sie eigene Methoden und Software online zur Verfügung stellt. Wie viel Arbeit FA in die theoretische Fundierung und Legitimation des eigenen Vorgehens investiert, 4 In den meisten von der FA untersuchten Fällen spielt die Sinndimension der Akteur*innen keine Rolle. FA folgt einer Schuldbzw. Wahrheitsepistemologie.
252 Mina Godarzani-Bakhtiari und René Tuma SJS 51 (2), 2025, 247–272 wird an der der Vielzahl eigener Publikationen zur Epistemologie deutlich (Keenan et al., 2012; Franke et al., 2014; Weizman, 2017; Fuller & Weizman, 2021). Wir wollen weniger auf diese theoretisch-schriftliche Seite eingehen, sondern uns auf die in den Bildern eingebettete reflexive Legitimation fokussieren. Diese zeigt sich in den von FA produzierten Videos, die nach Abschluss der Untersuchungen die Analyseprozesse präsentieren. Sie können als „how to establish facts“ (Rothöhler, 2021, S. 155) gelesen werden. In den Sozialwissenschaften wird die Arbeit von FA als die Bearbeitung politischer Unsichtbarkeit verhandelt. Brown und Carrabine (2019, S.195) sehen in dieser den Versuch, Herrschaft abseits (auch visueller) öffentlicher Kontrolle zu kritisieren. Lee-Morrison (2015, S. 1) argumentiert, dass die Arbeit von FA als die Bekämpfung negativer Evidenz, also die Nicht-Existenz von Evidenz als Evidenz für politische Verschleierung, verstanden werden muss. Auch Gutiérrez (2022, S. 16) sieht in der Arbeit von FA die Bearbeitung datafizierter Unsichtbarkeiten. So kritisiert die Forschungsagentur über die Produktion eigener Daten staatliche und transnationale (z. B. EU) Kontrollmacht. Rothöhler (2021, S. 155) argumentiert ähnlich, wenn er hervorhebt, dass in der Praxis von FA jenes betrachtet wird, das staatliche Organisationen systembedingt ausblenden. Stuckey (2022, S. 66) bezeichnet die Arbeit von FA aufgrund der ihr eingeschriebenen Orientierung an staatlichen Praktiken (Forensik) und demokratischen Werten als hegemoniekritisch. Da FA darauf angewiesen ist, öffentlich zugängliche Daten, vor allem nicht verifizierte visuelle Daten, zu verwenden, klassifiziert Gutiérrez (2022) FA als eine Form des Datenaktivismus. Insbesondere ordnet sie FA als proaktiven Datenorganisationen ein, für die Daten Grundlage ihrer Arbeit und Mittel der Präsentation sind. In der Literatur wird die neue Form der Beweisführung von FA als Paradigmenwechsel diskutiert, in der das Materielle gegenüber der menschlichen Zeugenschaft eine neue Funktion übernimmt. Wie viele Autor*innen betonen, steht die Hinwendung zum Materiellen im Kern der Arbeit von FA (Kinstler, 2022, S. 329; Stuckey, 2022, S. 8). Dafür steht auch der von Weizman (2017, S. 67) formulierte Anspruch, durch die Analyse von Materialitäten Dinge zum Sprechen zu bringen. Der Umstand, dass FA in ihren Analysen Gebäude und materielle Infrastrukturen ins Zentrum stellt, wird unterschiedlich eingeordnet. Samuels (2013, S. 68) deutet das Vorgehen, bei dem Ereignisse über Tatorte rekonstruiert werden, als Prozess, bei dem es zur Einbindung disparater Daten in kohärente räumliche Narrative kommt. Mandolessi (2021, S. 628) argumentiert, dass „multifarious bits of data“ sinnvoll in Narrative assembliert werden, wobei die Konstituierung von digitalen Orten zentral sei. Laut Rothöhler (2021, S. 146) nutzt FA den Raum zur Verifizierung visueller Daten und zur Einbindung der Ergebnisse in eine verbindende Ereigniskonstruktion. Weizman (2017, S. 100) selbst konzeptualisiert die Evidenzbegrüdnung als die Produktion eines „architectural image complexes“. Er versteht darunter die Methode bildliche Beweise innerhalb eines räumlichen Modells zu positionieren, wodurch
Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung … 253 SJS 51 (2), 2025, 247–272 die lineare Trennung zwischen „images in before-and after montages“ (Weizman, 2017, S. 100) aufgebrochen und neue Form des Archivs entsteht.5 Naß (2021, S. 51) kritisiert, dass es sich bei den Bild-Raum-Figurationen um „eine komplexitätsreduzierende Methode der Vervollständigung eines sinngenerierenden Erzählraums“ und nicht um die Verräumlichung analytischen Denkens handelt. Diese Kritik ist auch an jene von Harst (2023, S. 40) anschlussfähig, welcher argumentiert, dass ein Widerspruch zwischen konstruktivistischer Epistemologie und Evidenzpraxis bestehe. Während in den epistemologischen Ausführungen und auch den methodisch geleiteten Ausführungen der Investigationen, der konstruktive Charakter hervorgehoben wird,6 wird, so Harst, letztlich Evidenz als visuell repräsentierbar und damit als in der Wirklichkeit selbst gegeben präsentiert, wodurch die Grenze zwischen Modell und Wirklichkeit verschwimmt. Wie umkämpft diese Deutung zwischen Analyse und Wirklichkeit sind, zeigt sich gegenwärtig in deutschsprachigen Debatten. Die Parteilichkeit von FA ist bereits durch ihre Positionierung auf Seiten der Marginalisierten gegeben und wird in aktuellen Debatten als Komplizenschaft mit politisch umstrittenen Gruppen problematisiert (Naß, 2021; 2024) oder als der Debattenkultur nicht förderlich angesehen (RWTH Aachen, 2024). 2.1 Fall: Forensic Architectures Investigation des Polizeinsatzes in Hanau Die Untersuchung des Polizeieinsatzes in Hanau von FA hat in der deutschen Öffentlichkeit große Aufmerksamkeit erregt. Am 19. Februar 2020 erschoss ein Rechtsterrorist in Hanau innerhalb von 12 Minuten 9 migrantisierte Personen. Als die Polizei fünf Stunden später das Wohnhaus des Täters stürmt, finden sie den Täter und die Mutter tot auf. Später wird bekannt, dass 13 der in der Tatnacht eingesetzten SEK-Beamt*innen einer rechtsextremen Chatgruppe angehören. Die Überlebenden und Angehörigen prangern das polizeiliche Versagen an. Sie kritisieren, dass staatliche Verantwortlichkeiten nicht erfüllt wurden (vor, während und nach der Tat) und fordern Aufklärung. Im selben Jahr noch wird FA von der Initiative 19. Februar beauftragt, den Polizeieinsatz in Hanau zu untersuchen. Der Fall ist für die jüngere deutsche Auseinandersetzung mit Rechtsterrorismus und staatlicher Verantwortung von besonderer Bedeutung. Wie Nobrega, Quent und Zipf (2021, S. 9) argumentieren existiert in Deutschland keine wissenschaftlich zufriedenstellende Datensammlung rechtsterroristischer Gewalt. Stattdessen belegt 5 Die visuelle Erfassung und die simulative digitale Rekonstruktion des Tatorts sind etablierte polizeiliche Verfahren (neben weiteren) in Deutschland. „Das ‚gedankliche Modell zum Ereignis‘ ist ein virtuell materialisierbares geworden, das [...]‚explorative‘ Handlungsressourcen“ (Rothöhler, 2021, S. 65f.) bereithält. Für FA ist das simulative Verfahren die wichtigste Ressource. 6 Harst (2023) sieht im Konstruktionscharakter der Arbeit von FA eine Veränderung des Indizienparadigmas. Wir argumentieren, dass gerade der Widerspruch zwischen Epistemologie und Praxis, den Harst benennt, darauf hinweist, dass sich das Paradigma nicht geändert hat, sich aber die kommunikativen Formen der Spurenkonstruktion geändert haben (siehe Reichertz, 1991).
254 Mina Godarzani-Bakhtiari und René Tuma SJS 51 (2), 2025, 247–272 der Nicht-Umgang mit rassistischem Terror durch „Behörden, Medien, Politik, Polizei, Justiz, Kultur, Stadtund Zivilgesellschaften sowie Wissenschaft“ (Nobrega et al., 2021, S. 17) rassistische Machtund Exklusionsverhältnisse. Der Anschlag in Hanau markiert einen Wendepunkt im offiziellen Umgang mit rassistischen Taten. In Folge wurde dessen erstmals in der in verschiedenen Bundesländern ein Rechtsterrorismus-Opferfond eingerichtet (Nobrega et al., 2021). Besonders die (post)migrantische zivilgesellschaftliche Organisierung durchlebte seit Hanau einen Strukturwandel, der zur Begründung eines translokalen Widerstandsund Unterstützer*innennetzwerkes führte (Stjepandić, 2022). Bis heute ist der anhaltende Kampf der Betroffenen, um eine lückenlose Aufklärung des Anschlags und des gescheiterten Polizeieinsatzes (Initiative 19. Februar 2021) ein entscheidender Bezugspunkt antirassistischer Arbeit in Deutschland. Diese Entwicklungen basieren maßgeblich auf der politischen Arbeit der Betroffenen. Auch die Ermittlungen von FA, die das Geschehen in der Tatnacht mit Fokus auf die polizeilichen Handlungen rekonstruierten, trugen dazu bei. FA nutze dabei Big Visual Data als Ressource, um das schwer nachzuweisende staatliche Versagen zu untersuchen und publizierte die Ergebnisse in zwei Videos, die die Datenanalyse transparent darstellen. 2.2 Big Visual Data als herausfordernde Ressource von Forensic Architecture Visuelle Daten sind zunehmend, und bei FA insbesonders, von einer doppelten Prekarität gekennzeichnet. Erstens repräsentieren diese perspektivisch gebundene Wirklichkeitsausschnitte. Es sind reduktionistische Repräsentation, bei denen das Aufgenommene als flache, zweidimensionale Wirklichkeit repräsentiert wird. Zweitens sind visuelle Daten keine „neutralen Registriermaschinen“ (Reichert, 2007, S. 45), welche die Wirklichkeitsausschnitte einfach wiedergeben. Vielmehr erfordert ihre Nutzbarmachung Interpretationen, deren Herleitung wiederum kommunikativ überzeugend dargestellt werden muss. In Bezug auf Big Visual Data kommt neben der Prekarität der einzelnen visuellen Daten hinzu, dass die Akteur*innen, die verschiedene, verifizierte und nicht verifizierte Daten verwenden, vor der Herausforderung stehen, analytisch nachvollziehbar nachzuweisen, dass die Daten miteinander in Beziehung stehen. Nur wenn die Daten als miteinander verknüpfbare Dokumente konstruiert werden, ist es überzeugend, dass die in ihnen herausgearbeiteten Informationen zu einer Informationscollage zusammengesetzt werden können. Ausgehend von diesen Überlegungen verstehen wir deshalb Big Visual Data als eine verstreute und zersplitterte Landschaft prekärer visueller Daten, die vielfältige kommunikative Herausforderungen bereithält. Bei der digitalen Ereignisrekonstruktion, wie FA sie z. B. im Fall Hanau vornimmt, handelt es sich um ein komplexes Unterfangen, bei dem das, was die Wirklichkeit repräsentieren soll, durch konstruktive Datenarbeit erst kommunikativ hergestellt, legitimiert und nachvollziehbar gemacht werden muss. Wir argumentieren, dass dabei spezifische
Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung … 255 SJS 51 (2), 2025, 247–272 kommunikative Praktiken, d. h. die Ausprägung jeweils angepasster, vernakularer Analysen, eine Schlüsselrolle spielen. 3 Vom Datum zu Praktiken des Sehens und Zeigens und zur Reflexivität Die Rekonstruktion sozialer Ereignisse anhand visueller Daten wird in verschiedenen Feldern verfolgt, die unterschiedlichen Erkenntnisinteressen, Epistemologien und auch Kommunikationsformen folgen. Übersetzungen bestehen systematisch zwischen Wissenschaft und angewandter Forensik, aber auch in das Feld der Kunst hinein. Reichert (2007) hat diese Kontinuität bereits für den Wissenschaftsfilm als Dispositiv herausgearbeitet. Die Bereitstellung technischer Werkzeuge zur Visualisierung kann als transversale Kultur bezeichnet werden und entwickelte sich z. B. bei Video in Ko-Evolution zwischen Technikentwicklung und Anwendung in verschiedenen Feldern (vgl. Tuma & Lettkemann, 2018). FA überschreitet die Feldgrenzen zwischen Wissenschaft, forensischer Ermittlung, Politik und Kunst explizit und gezielt. Sie greifen stark auf wissenschaftliche Methoden zurück, die sie anwenden und damit auch vernakular neu kontextualisieren. Das Erkenntnisinteresse, das schlussendlich auf Verantwortung und Schuld abzielt, stellt die Rekonstruktion von Ereignissen als raum-zeitliche Abläufe ins Zentrum der Analyse. Dabei verbleiben Parallelen auch zu sozialwissenschaftlichem Vorgehen, denn die genaue Ereignisrekonstruktion auf Basis visueller Daten ist auch Ansatzpunkt für interaktionsanalytische Forschungen u. a. in der Gewaltsoziologie (siehe Hoebel et al., 2022), wenn auch mit einem anderen– verstehend-erklärendem – Erkenntnisinteresse. Trotz der Parallelen tauchen die sozialwissenschaftlichen Perspektiven in den Analysen von FA, wie (überwiegend) auch in der klassischen Forensik, nicht auf. Wenden wir die Soziologie des Visuellen auf FA an, sind Perspektiven relevant, die das Sehen und Zeigen als Praktiken bzw. als (kommunikatives) Handeln begreifen, um die Sehgemeinschaften (Raab, 2008), in ethnomethodologischer Perspektive also members methods zu erfassen. Dabei geht es auch um Infrastrukturen des Sehens, des Abbildens und des Gesehenwerdens. In wissenssoziologischer Perspektive werden Varianten der Produktion und Distribution von visuellem Wissen thematisiert (Schnettler, 2007; Lucht et al., 2012). Eine Möglichkeit, sich diesen Fragen zu nähern, ist die reflexive Betrachtung der konkreten Handlungsoder Praxisformen. Hier bieten sich verschiedene Ansatzpunkte an, z.B. die reflexive Auseinandersetzung mit dem wissenschaftlich objektivierenden Blick und die reflexive Erforschung visueller Repräsentationen. Die Reflexion des Umgangs mit visuellen Daten wurde zunächst in der empirischen Wissenschaftsforschung entwickelt, die sich mit verschiedenen Formen der Visualisierung vor allem in den Naturwissenschaften auseinandergesetzt hat. In Lynch und Woolgars (1990) Representation in Scientific Practice finden sich einige wegweisende Studien, die sich mit dem Umgang mit und der Produktion von
262 Mina Godarzani-Bakhtiari und René Tuma SJS 51 (2), 2025, 247–272 4.3 Verifikation durch Relationierung Da die Forschungsagentur häufig nicht auf dieselbe Datengrundlage zurückgreifen kann, wie z. B. staatliche Behörden, werden in den Untersuchungen nicht verifizierte Daten herangezogen, die erst durch FA verifiziert, damit objektiviert werden. Anders als bei verifizierten Daten wird von FA bei unverifizierten Daten der Ursprung der Daten nicht benannt.10 Stattdessen steht deren Verifizierung durch die Methode des Cross-Referencing (Weizman, 2017, S. 58) im Zentrum. Cross-Referencing meint die Überprüfung einer Quelle durch das Aufspüren einer anderen, von der ersten unabhängigen Quelle, welche die gleiche Information aus einer anderen Perspektive wiedergibt. Anders als bei gängigen journalistischen Tätigkeiten, bei denen auch die Technik des Cross-Referencing zur Überprüfung unverifizierter Aussagen (textuelle Daten) angewendet wird, setzt FA die Methode visuell ein. Dabei spielt Raum bzw. Architektur eine entscheidende Rolle, wie Weizman (2017, S.132) selbst schreibt: „We use architecture [...] to create ,evidence assemblages‘ that locate these elements in space and study the time/space relations between them“. Über die raumanalytische Methode der Breitenund Tiefenanalyse visueller Daten werden Quellen zueinander in Verhältnis gesetzt, um die in den Daten vorhandenen Informationen als bestätigte und damit aussagekräftige Informationen herzustellen. Damit kommt es zur Verifikation über die Kontexturalisierung der Daten zueinander (Relationierung), d. h.der Verifikation durch die kommunikative Herstellung und Verknüpfung materieller und räumlicher Bezüge über die Grenzen der verschiedenen visueller Daten hinweg. Ein Beispiel für die Verifizierung eines Gesamtdatums sind die Abbildungen 4 und 5. Um die Frage zu beantworten, wann die Polizei vor dem Wohnhaus des Täters eingetroffen ist, zieht FA zunächst ein Video einer Überwachungskamera (mit Zeitstempel) aus der Nachbarschaft heran, welches in schwarz-weis rotierendes Licht zwischen den Häusern zeigt (Abb. 4). Dieses flackernde Licht wird von FA als Blaulicht interpretiert. Um das Datum als objektiven Informationsspeicher zu konstruieren, wird ein zweites Video herangezogen (Abb.5). Es zeigt das Wohnhaus des Täters, von einer Privatperson mit einem Smartphone aus dem gegenüberliegenden Haus aufgenommen. Zu sehen ist, in Farbe, wie zu einer bestimmten Zeit die Reihenhäuser durch ein rotierendes Blaulicht angestrahlt werden. Um zu überprüfen, ob es sich um dasselbe rotierende Licht in beiden Videos handelt, wird der Raum ausgehend von den beiden Videos jeweils in die Breite und Tiefe simulativ erweitert (siehe Bilderrahmen Abb. 4 und 5). Außerdem wird in der Sequenz von einem Video fließend über das Raum-Modell eine Perspektivverschiebung vorgenommen, die mit der Perspektive des anderen Videos endet. Überzeugend werden durch die erweiternde Raumanalyse beide Videos als verschieden perspektivistisch definierte, jedoch einer kongruenten spatio-temporal-order entstammende Repräsentationen 10 Naß (2021, S.52) sieht hierin eine parteiisch agierende Quellenarbeit. In ihrer Argumentation verkennt sie u. E., dass FA den Prozess der Kontextualisierung durch den der Kontexturalisierung ersetzt, wodurch eine andere Form kommunikativer Objektivierung vorgenommen wird.
Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung … 263 SJS 51 (2), 2025, 247–272 Abbildung 4 Überwachungskamera nimmt rotierendes Licht in der Straße auf Quelle: Forensic Architecture (2022b, 25. April ), https://www.youtube.com/watch?v=N7H5fhokpLU (c) Forensis, 2022 [Screenshot bei 27:44 min durch Godarzani-Bakthiari und Tuma]. Abbildung 5 Smartphone Aufnahme filmt Polizeieinsatz aus gegenüberliegendem Haus Quelle: Forensic Architecture (2022b, 25. April ), https://www.youtube.com/watch?v=N7H5fhokpLU (c) Forensis, 2022 [Screenshot bei 28:01 min durch Godarzani-Bakthiari und Tuma].
264 Mina Godarzani-Bakhtiari und René Tuma SJS 51 (2), 2025, 247–272 hergestellt. Durch die visuelle Raumanalyse, die wir als multiperspektivische Tiefenund Breitenanalyse bezeichnen, bestätigt FA, dass die Polizei zu einem bestimmten Zeitpunkt an einem bestimmten Ort anwesend gewesen ist. Um dieses analytische Vorgehen visuell nachzuvollziehen, haben wir eine Visualisierung der räumlichen Analyse (Abb. 6) erstellt. 4.4 Selektion und Extrahierung von Informationen Die Analyse von Ereignisdaten, mit dem Ziel relevante Informationen über das Ereignis herauszuarbeiten, wird von FA getrennt in einzelnen Daten vollzogen. Dabei repräsentiert FA das Datum (z. B. eine Smartphoneaufnahme) und macht gleichzeitig die Interpretation des Datums visuell manifest. Das, was von FA als relevant erachtet wird, wird mit der Technik der Hervorhebung (Highlighting bei Goodwin, 1994) sichtbar gemacht: Es wird unterstrichen, umkreist, farblich abgehoben oder vergrößert. Dadurch kommt es zur Selektion und Extrahierung relevanter Informationen. Das Beispiel der Analyse eines Polizeihelikoptervideos (Video2) veranschaulicht den Prozess. Der Helikopter filmt aus dem Luftraum den Tatraum. FA zieht dieses Video heran, um den Standort der Polizei zu einem bestimmten Zeitpunkt zu rekonstruieren. Wie sie die Analyse und deren Ergebnisse direkt im Material manifest machen, zeigt die Abbildungen7. Auf Abbildung7 (oben) wird das Video Abbildung 6 Unsere Visualisierung des Ergebnisses der weiten Analyse des Raumes durch die Videos Quelle: Eigene Visualisierung [Godarzani-Bakthiari und Tuma].
Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung … 265 SJS 51 (2), 2025, 247–272 des Helikopters zunächst eingeblendet und vor dem Hintergrund des Raum-Modells positioniert. Damit wird die flache und begrenzte 2D-Repräsentation über seine Begrenztheit hinaus eingebettet und kontexturalisiert. Gleichzeitig wird unterhalb des Bildes eine Zeitleiste eingeblendet, die das Video zeitlich verortet. Anschließend (Abb. 7, unteres Bild) wird eine bestimmte visuelle Information, in diesem Fall die Position eines (Polizei-)Fahrzeuges hervorgehoben und damit kommunikativ als relevant markiert und visuell manifestiert. Abbildung 7 Selektion und Manifestierung relevanter Informationen bei der Analyse eines Helikoptervideos Quelle: Forensic Architecture (2022b, 25. April ), https://www.youtube.com/watch?v=N7H5fhokpLU (c) Forensis, 2022 [Screenshot oben bei 20:24 min; unten bei 20:35 min durch Godarzani-Bakthiari und Tuma].
266 Mina Godarzani-Bakhtiari und René Tuma SJS 51 (2), 2025, 247–272 4.5 Vermessung von Informationen in/aus Daten FA nutzt visuelle Daten auch um sekundäre Analysen zu betreiben, in deren Ergebnis sie neue quantifizierte Daten produziert. Ein Beispiel ist die Bewegungsanalyse der Subjekte in der Arena Bar. Nachdem der Handlungsverlauf in der Bar über das Cross-Referencing und der Synchronisation der Repräsentationen verschiedener Überwachungskameras rekonstruiert wurde, misst die Forschungsagentur die Geschwindigkeit der Personen im Zeitverlauf. Auf Abbildung 8 sehen wir die Bewegungsanalyse in der Arena Bar. Neben dem Grundriss der Bar, auf welchem die Bewegung der Subjekte räumlich nachvollzogen wird, ist auf der rechten Seite ein Diagramm eingeblendet, in welchem die Geschwindigkeit pro Meter eingezeichnet wird. Da beides die Bewegung des Einzelnen und der Geschwindigkeitsverlauf im Diagramm gleichzeitig von FA eingezeichnet wird, ist visuell nachvollziehbar, in welchem Verhältnis die Geschwindigkeit zu den Raumattributen steht. Diese Art von Daten, die von FA mit Rückgriff auf andere Daten produziert werden, erhalten besonders durch die diagrammatische Darstellung einen objektivierten Status. 4.6 Synthese Für die Begründung von Evidenz ist der letzte Schritt der Analyse, die Synthese, entscheidend. Nun werden von FA verschiedene Informationen, die aus visuellen Daten extrahiert wurden, auf dem Raum-Modell positioniert. Dieser Prozess kann Abbildung 8 Visualisierung der Bewegungen der Subjekte im Raum Quelle: Forensic Architecture (2022a, 14. März), https://www.youtube.com/watch?v=gwEMMI_zGas (c) Forensis, 2022 [Screenshot bei 6:51 min durch Godarzani-Bakthiari und Tuma].
Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung … 267 SJS 51 (2), 2025, 247–272 im Sinne Löws (2001) als Spacing und Synthese verstanden werden. Indem die Platzierung gleichzeitig als Relationalisierung verschiedener Informationen zueinander funktioniert (Platzierung im Verhältnis zu anderen Platzierungen) wird von FA ein Ereignisraum geschaffen, welcher eine verknüpfende Betrachtung, die Synthese, ermöglicht. Die Relationierung fungiert hier gleichsam als räumliche Plausibilisierung der Informationen im Raum. Im Zentrum stehen dabei die Fragen, fügen sich die Informationen logisch in einen größeren Sinnzusammenhang ein oder treten bei dem spatio-temporal-ordering der Informationen Konflikte auf. Dieser komplexe Vorgang lässt sich nachvollziehbar an dem Beispiel der Synthese aus Video2 verdeutlichen. Das Still auf Abbildung9 entstammt einem Videoausschnitt, bei dem FA der Frage nachgeht, inwiefern die Polizei die Schüsse, mit denen der Täter seine Mutter und sich selbst erschoss, hätte hören müssen. Auf der Abbildung 9 symbolisiert die farbliche Einfärbung des Modells die zuvor gemessene Schallausbreitung eines Schusses im Wohnhaus des Täters. Die Schallreichweite wird im Modell ins Verhältnis zu den Positionen der Polizei, welche unter anderem durch das Helikoptervideo rekonstruiert wurden, gesetzt. Ein Ergebnis der Relationierung sind die Dezibelangaben bei den Polizeiautos, welche angeben, wie hoch die Schusslautstärke bei den Positionen der Fahrzeuge gewesen sein müsste. FA weist räumlich durch diese Visualisierung nach, dass die Polizeibeamten mit mittlerer bis hoher Wahrscheinlichkeit die Schüsse gehört haben müssten. Indem zuvor das Modell und die extrahierten Informationen von FA als objektiv Abbildung 9 Verhältnis des Standorts Polizei und der Schallausbreitung des Schusses wird räumlich durch das Modell plausibilisiert Quelle: Forensic Architecture (2022b, 25. April ), https://www.youtube.com/watch?v=N7H5fhokpLU (c) Forensis, 2022 [Screenshot bei 22:58 min durch Godarzani-Bakthiari und Tuma].
268 Mina Godarzani-Bakhtiari und René Tuma SJS 51 (2), 2025, 247–272 hergeleitete Ergebnisse der Analysen dargestellt werden, deren Wirkkraft performativ visuell hergestellt wurden, kommt es zur Begründung von Evidenz anhand räumlicher Rationalisierungslogik. So wird Evidenz entlang eines räumlichen Objektivitätsverständnisses (spatial objectivity) konstruiert (Godarzani-Bakhtiari, 2024). 5 Fazit Wir haben die Frage gestellt, mit welcher reflexiv-kommunikativen Analyseform FA den Herausforderungen von Big Visual Data begegnet. FA untersucht staatliches Handeln mittels Big Visual Data und nutzt neue kommunikative Formen der Kritik und öffentlichen Diskursivierung, die sich durch eine reflexive Darstellung der eigenen Analyseverfahren auszeichnen. Dabei werden die Herausforderungen systematisch durch die Anwendung vor allem raumbezogener Verfahren adressiert. Wir konnten verschiedene analytische Verfahren identifizieren, die unterschiedliche Funktionen im Prozess der Evidenzkonstruktion übernehmen: Kontextualisierung, Kontexturalisierung, Verifikation durch Relationierung, Selektion und Extraktion von Informationen, (sekundäre) Messung und Synthese. Die Daten durchlaufen einen reflexiv offen gelegten Geneseprozess, in dem Informationen herausgefiltert und innerhalb eines räumlichen Modells zueinander in Beziehung gesetzt werden. In den Videos (wie auch in analogen Rauminstallationen in international renommierten Ausstellungsorten) werden die Daten und ihre Analysen nach und nach miteinander verschränkt, wodurch die empirische Visualisierung zweiter Ordnung ausgestellt und damit im Artefakt nachvollziehbar wird. Visuelle Daten werden nicht als geschlossene statische Informationsträger behandelt, sondern als Referenzquellen, über deren raumzeitlich gebundene visuelle Darstellungsrahmen hinaus Analysen durchgeführt werden. Spuren werden unter Rückgriff auf zeitliche und räumliche Modelle als komplexe Verweisungszusammenhänge digital (re)konstruiert. Flache 2D-Repräsentationen werden so in tiefe11 und weite virtuelle Repräsentationen verwandelt, bei denen das ursprüngliche Datum nur noch einen Teil neben weiteren modellierten Teilen darstellt. So werden zweidimensionale visuelle Zusammenhänge in dreidimensionale Beziehungen aufgespannt und expandiert und über ein vereinendes synthetisierendes Raum-Modell aneinandergebunden. Nach McIlvenny (2018, S. 2), der ein ähnliches Verfahren methodisch für die sozialwissenschaftliche Videoanalyse nutzt, lässt sich diese neue Praxis des Umgangs mit visuellen Dokumenten als Virtualisierung audio-visueller Daten verstehen. Diese methodische Entwicklung steht nach ihm für einen „scenographic turn“. 11 Hier bestehen auch Gemeinsamkeiten zum Deep Mapping (Bodenhamer et al., 2015).
Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung … 269 SJS 51 (2), 2025, 247–272 Bei FA mündet die Analysearbeit in der Verschachtelung visueller Fragmente in ein Meta-Artefakt. Als Metaartefakte verstehen wir Dokumente,12 die alle drei Ebenen des Erkenntnisprozesses (Datensampling, Datenanalyse, Ergebnisdarstellung) beinhalten und diese als differente, jedoch aufeinander aufbauende Ebenen sichtbar machen. Sie präsentieren selbst eine Narration über die eigene trans-sequentielle Entstehungsgeschichte. Somit wird kommunikationsmächtige Evidenz über das Aufzeigen eines Pfades durch diese verschachtelte Welt der Bilder hergestellt. Während durch das Aufzeigen einerseits dessen konstruktiver Akt offengelegt wird, sorgt die visuell verankerte Manifestierung der Analyse und die Evidenzkonstruktion entlang eines spatial objectivity Verständnisses andererseits, für die Stabilisierung der dabei begründeten Evidenz. Die kommunikative Macht ergibt sich aus der Verschränkung von Offenheit (Sichtbarmachung der Konstruktionsarbeit) und (konstruierter) Gesetzmäßigkeit (spatial objectivity). FA schließt an wissenschaftliche Praktiken an, beruft sich auf etablierte Expertise und knüpft an künstlerische und populäre Mediengattungen an.13 Die reflexive Offenlegung dieser Vernakularität ist vor dem Hintergrund einer fragmentierten Öffentlichkeit (Gates, 2024) und der Umkämpftheit der Deutungshoheit einzuordnen. Außerdem ist die Positionierung von FA an der Schnittstelle verschiedener Felder und der explizite Beitrag zum Politischen, den FA leisten möchte, wesentlich für ihre Sehpraktiken. Ihre visuell-narrativen Präsentationen zielen darauf ab, die Perspektiven der Betroffenen in der Öffentlichkeit zu stärken und trotz aller GegenEpistemologie dennoch im Kern positivistisch zu legitimieren. Die Sichtbarmachung der Analyse-Praktiken in den FA-Präsentationen/Inszenierungen (siehe nochmal Naß, 2021) dient einer doppelten Verschiebung: Das Urteilen wird einerseits vom Gerichtssaal in die Öffentlichkeit verlegt, andererseits von den Ergebnissen zu den Methoden verschoben. Es sind die sich als objektiv darstellenden Verfahren der Analysen, die überzeugen sollen. Meta-Artefakte fordern so Deutungshoheit und eine Politisierung der Öffentlichkeit ein. Dass FAs Ansatz eine Lücke im öffentlichen Diskurs füllt, wird an der Übernahme ihrer Methoden im klassischen Investigativjournalismus deutlich. Damit entwickeln sich forensischer Journalismus und media evidence allmählich zu einem standardisierten, populären Format (siehe Gates, 2020). 6 Literatur Bodenhamer, D. J., Corrigan, J., & Harris, T. M. (Hrsg.). (2015). Deep Maps and Spatial Narratives. Indiana University. Bois, Y.-A., Feher, M., Foster, H., & Weizman, E. (2016). On Forensic Architecture: A Conversation with Eyal Weizman. October Magazine, 156, 116–140. 12 Vgl. Gutiérrez (2021) Begriff der Meta Documentaries. 13 Form-Erwartungen bestehen zur Populärkultur (siehe Englert & Reichertz, 2016).
270 Mina Godarzani-Bakhtiari und René Tuma SJS 51 (2), 2025, 247–272 Brown, M., & Carrabine, E. (2019). The Critical Foundations of Visual Criminology: The State, Crisis, and the Sensory. Critical Criminology, 27(1), 191–205. Coenen, E., & Tuma, R. (2022). Contextural and Contextual – Introducing a Heuristic of Third Parties in Sequences of Violence. Historical Social Research, 47(1), 200–224. Daston, L. (2007). Objectivity. MIT. Englert, C. J., & Reichertz, J. (2016). CSI – Rechtsmedizin – Mitternachtsforensik. Springer. Forensic Architecture (o. J.). Open Source Software. Forensic Architecture. https://forensic-architecture. org/subdomain/oss Forensic Architecture (2022a, 14. März). Hanau-Anschlag: Der Notausgang [Video]. YouTube. https:// www.youtube.com/watch?v=gwEMMI_zGas Forensic Architecture (2022b, 25. April). Rassistischer Terror-Anschlag in Hanau: Der Polizeieinsatz [Video]. YouTube. https://www.youtube.com/watch?v=N7H5fhokpLU Franke, A., Weizman, E., & HKW (Hrsg.). (2014). Forensis: The architecture of public truth. Exhibition „Forensis“. Sternberg. Fujii, L. A., Finnemore, M., & Wood, E. J. (2021). Show time: The logic and power of violent display. Cornell University. Fuller, M., & Weizman, E. (2021). Investigative Aesthetics: Conflicts and commons in the politics of truth. Verso. Gates, K. (2013). The cultural labor of surveillance: Video forensics, computational objectivity, and the production of visual evidence. Social Semiotics, 23(2), 242–260. Gates, K. (2020). Media Evidence and Forensic Journalism. Surveillance & Society, 18, 403–408. Gates, K. (2024). Day of Rage: Forensic journalism and the US Capitol riot. Media, Culture & Society, 46(1), 78–93. Ginzburg, C. (2011). Spurensicherung: Die Wissenschaft auf der Suche nach sich selbst (G. Bonz & K. F.Hauber, Übers.). Klaus Wagenbach. Godarzani-Bakhtiari, M. (im Review). Die Analyse des Analysierens visueller Artefakte: Vergleichende Video-Artefakt-Analyse der Evidenzkonstruktion von Forensic Architecture. Forum Qualitative Sozialforschung. Godarzani-Bakhtiari, M. (2024). Gegenöffentliche Problematisierung polizeilicher Nekropolitik: Forensic Architecture’s Investigation des Polizeieinsatzes in Hanau. sub\urban. Zeitschrift für kritische Stadtforschung 12(2),13–42. Goodwin, C. (1994). Professional Vision. American Anthropologist, 96(3), 606–633. Gutiérrez, M. (2021). Data activism and meta-documentary in six films by Forensic Architecture. Studies in Documentary Film, 1–21. Gutiérrez, M. (2022). Documenting the Invisible: How Data Activism Fills Visual Gaps. Papeles de Identidad, 2, 1–21. Harst, J. (2023). Virtuelle Investigationen. Transformationen des Indizienparadigmas zwischen Sherlock Holmes und Forensic Architecture. In G. Lisa & A. Simonis (Hrsg.), Medienkomparatistik: 4 (2022) (S. 23–44). Aisthesis. Heßler, M., & Mersch, D. (2009). Logik des Bildlichen: Zur Kritik der ikonischen Vernunft . Transcript. Hill, M. B. (2022). The New Art of Old Public Science Communication: The Science Slam. Routledge. Hoebel, T., Tuma, R., & Reichertz, J. (2022). Visibilities of Violence, Special Issue Historical Social Research 47/1. H-Soz-Kult. http://www.hsozkult.de/searching/id/z6ann-130243 Hoggenmüller, S. W., Klinke, H. (2025). Metabilder als Forschungswerkzeuge: Zur Kontingenz und algorithmischen Bedingtheit ihrer Herstellung. Schweizerische Zeitschrift für Soziologie, 51(2), Special Issue hrsg. von S. W. Hoggenmüller, Big Visual Data als neue Form des Wissens: Potenziale, Herausforderungen und Transformationen.
Vom flachen Bild zur verräumlichten visuellen Analyse: Forensic Architecture und die Verschachtelung … 271 SJS 51 (2), 2025, 247–272 Holert, T. (2004). Smoking Gun. Über den Forensic Turn der Weltpolitik. In R. F. Nohr (Hrsg.), Evidenz… das sieht man doch! LIT. Hüppauf, B., & Weingart, P. (Hrsg.). (2009). Frosch und Frankenstein: Bilder als Medium der Popularisierung von Wissenschaft. Transcript. Initiative 19 Februar (2021). Ein Jahr nach dem 19. Februar in Hanau: Die Kette behördlichen Versagens vor dem rassistischen Terroranschlag, in der Tatnacht und in den Monaten danach [Bericht]. https://19feb-hanau.org/wp-content/uploads/2021/02/Kette-des-Versagens-17-02-2021.pdf Kammerer, D. (2020). Qualitative Verfahren der Filmanalyse. In M. Hagener & V. Pantenburg (Hrsg.), Handbuch Filmanalyse (S. 385–397). Springer. Keenan, T., Weizman, E., & Steyerl, H. (2012). Mengele’s skull: The advent of forensic aesthetics. Sternberg. https://research.gold.ac.uk/id/eprint/9292 Kinstler, L. (2022). Situated Testimony: Forensic Architecture’s Memory Objects. Space and Culture, 25(2), 327–330. Knoblauch, H., Janz, A., & Schröder, D. J. (2021). Kontrollzentralen und die Polykontexturalisierung von Räumen. In M. Löw, V. Sayman, J. Schwerer, & H. Wolf (Hrsg.), Am Ende der Globalisierung (S. 157–182). Transcript. Krämer, S., Kogge, W., & Grube, G. (Hrsg.). (2007). Spur: Spurenlesen als Orientierungstechnik und Wissenskunst. Suhrkamp. Lee-Morrison, L. (2015). The Forensic Architecture Project: Virtual imagery as evidence in the contemporary context of the war on terror. https://lucris.lub.lu.se/ws/portalfiles/portal/6367379/7862140.pdf Lindemann, G. (2017). Verfahrensordnungen der Gewalt. Zeitschrift für Rechtssoziologie, 37(1), 57–87. Löw, M. (2000). Raumsoziologie (6. Aufl.). Suhrkamp. Lucht, P., Schmidt, L.-M., & Tuma, R. (mit Soeffner, H.-G., Hitzler, R., Knoblauch, H., & Reichertz,J.). (2012). Visuelles Wissen und Bilder des Sozialen. Aktuelle Entwicklungen in der Soziologie des Visuellen. VS. Lynch, M., & Woolgar, S. (1990). Representation in Scientific Practice. Kluwer. Mandolessi, S. (2021). Challenging the placeless imaginary in digital memories: The performation of place in the work of Forensic Architecture. Memory Studies, 14(3), 622–633. McIlvenny, P. (2018). Inhabiting spatial video and audio data: Towards a scenographic turn in the analysis of social interaction. Social Interaction. Video-Based Studies of Human Sociality, 2(1). Meier zu Verl, C., & Tuma, R. (2021). Video Analysis and Ethnographic Knowledge: An Empirical Study of Video Analysis Practices. Journal of Contemporary Ethnography, 50(1), 120–144. Merleau-Ponty, M. (2004). The world of perception. Routledge. Milroy, C. M. (2017). A Brief History of the Expert Witness. Academic Forensic Pathology, 7(4): 516–526. Naß, M. A. (2021). Bilder von Überwachung oder Überwachungsbilder? Zur Ästhetik des Kritisierten als Ästhetik der Kritik bei Hito Steyerl und Forensic Architecture. Naß, M. A. (2024, Januar 3). Kritik an Forensic Architecture: Zweifelhafte Beweisbilder. Die Tageszeitung: taz. https://taz.de/!5983353/ Nassehi, A. (2019). Muster: Theorie der digitalen Gesellschaft (2. Auflage). C. H. Beck. Nohr, R. F. (2004). Evidenz – Das sieht man doch! LIT. Peltzer, A. & Keppler, A. (2015). Die soziologische Film-und Fernsehanalyse: EineEinführung. De Gruyter. Peltzer, A., & Sommer, V. (2020). Stolpersteine digitaler Erinnerungskulturen. Eine komparative Analyse digitaler Zeitzeugenvideos über den Holocaust. In Medien + Erziehung (Bd. 64, Nummer 6, S. 74–86). Pfadenhauer, M., & Dieringer, V. (2019). Professionalität als institutionalisierte Kompetenzdarstellungskompetenz. In C. Schnell & M. Pfadenhauer (Hrsg.), Handbuch Professionssoziologie (S. 1–21). Springer.
278 Roland Meyer SJS 51 (2), 2025, 273–289 2 Models of Latent Space Operative image spaces that assemble seemingly weightless and placeless museum artifacts in the form of decontextualized image data have become almost a standard interface for visualizing extensive collections. One current example of this trend is the bauhaus infinity archive, an interactive installation that promises virtual access to around 15 000 collection objects during the temporary closure of the Bauhaus Archive Berlin. Again, this vast collection is visualized as a galaxy of digital images floating in an endless black universe: an explorable, navigable, immersive three-dimensional image space made up of seemingly immaterial objects, waiting to be sorted and rearranged into clusters following the user’s commands. Designed by the renowned Berlin Art+Com studios, the installation allows users to navigate the collection by drawing lines on a pad or selecting colors from a menu, making visible new, supposedly before unseen connections between collection objects based on pattern recognition. As the designers explain in an interview, the precondition for this is a specific form of virtual spatialization that goes beyond the mere interface: The images are initially vectorised using a convolutional neural network, i.e. translated into sequenced group of numbers – a so-called vector. After the images are vectorised, an algorithm called UMAP processes the dataset. This ensures that each vector, and with it, each picture is assigned a position in three-dimensional space. The result is a spatial depiction of the images, arranged in visually similar groups which the visitors can experience live in the bauhaus infinity archive. (Brafa, 2022) As this statement makes clear, the spatial visualization the users explore via the interface is a three-dimensional representation of the high-dimensional vector space by which these images are internally processed. Virtual spatialization is thus not merely a form of representation of big visual data intended for human eyes but also lies at the conceptual core of how contemporary forms of machine learning and pattern recognition make similarities within large data sets operative. When deep learning algorithms are trained on vast quantities of digital objects such as images, the features abstracted from these objects are encoded in a so-called latent space, amultidimensional vector space in which similarities between two images, be it in form, style, color, or any other aspect, are represented as quantifiable proximities (Somaini, 2023, p. 77). While such latent spaces themselves are abstract, purely mathematical, multi-dimensional, and therefore not only invisible but ultimately impossible to visualize, three-dimensional interface visualizations such as the bauhaus infinity archive function as models of latent space as a symbolic form. Radically reduced
Operative Image Spaces. Navigating Virtual Museum Collections 279 SJS 51 (2), 2025, 273–289 in their dimensions and made accessible to the human eye in ultimately diagrammatic form (Hunger, 2023), such visualizations of mathematical relationships and statistical distributions as spatial patterns nevertheless convey essential aspects of latent spaces: homogeneity, quantifiability, and continuity. Firstly, by staging the virtual image archive as a homogeneous space of universal comparison, in which all differences of media, genre, dimension, format, and cultural context are erased, these operative image spaces reflect the technical requirements of machine learning algorithms, which reduce all realized objects to a matrix of pixels, ultimately a series of numbers indicating color values. Converted into a table of discrete values, each digital image can be described as a vector in high-dimensional coordinate space, and its relative position in this space provides information about its relationship to other image vectors. Therefore, and secondly, such relationships between digital images, be it formal similarities, or, at least in some cases, iconographic references, can also be represented spatially as quantifiable proximities and distances. The closer two images appear in these spaces, at least in a certain dimension, the more similar they are said to be. Similarity, once an elusive category, thus seems to become measurable (see Hoggenmüller and Klinke in this special issue for more details). Thirdly and finally, these spaces are not only discretely addressable but also designed to be (almost) continuously navigable – from one image to another, there is always a path to follow, and each image is connected to every other image by a chain of similarities. Figure 2 Mario Klingemann and Simon Doury (Google Cultural Institute), X Degrees of Separation, 2017 (Screenshot) Source: artsexperiments.withgoogle.com/xdegrees/.
280 Roland Meyer SJS 51 (2), 2025, 273–289 The idea that every existing image is just one link in a chain of similarities is maybe best illustrated by X Degrees of Separation (2016), an experimental collection interface designed by artist Mario Klingemann in collaboration with the Google Cultural Institute. Its stated aim was to playfully explore similarities in a collection of over 250 000 image data objects. From each object, a path of visual similarities to every other object was to be found or constructed. All these paths, it is suggested, coexist in a common “art space”.2 In contrast to the previous examples, this space is not visualized as a three-dimensional, perspective space but nevertheless forms the conceptual basis of the entire undertaking. Imagining a path leading from one object to the next only becomes plausible by spatializing similarities and differences between entirely different and physically unrelated objects. However, the supposed similarities that are traced here, for example, between a bronze sculpture and awatercolor drawing (fig.2), are primarily those between digital image data, not between the actual objects themselves. Thus, a light background can sometimes be part of an artistic concept, as in the case of a watercolor drawing, in other cases it can simply be an arbitrary feature of the standardized photographic format of museum collection documentation. Similarity, abstracted from any context and reduced to a mere statistical proximity between flattened, standardized, pre-formatted digital representations, threatens to become an almost meaningless category (Wasielewski, 2023). Nevertheless, it is also becoming a productive category, as latent spaces are not only used to compare, sort and classify big visual data using discriminative AI such as pattern recognition algorithms, but also form the core of what is now known as generative AI. 3 Generative Spaces Ultimately, the idea that all images coexist in a homogeneous, quantifiable, and continuously navigable space of universal comparison also blurs the difference between the actual and the virtual. If every possible image occupies a specific position in latent space and there are always countless other images to be found between two actual images, what seems more tempting than trying to visualize these latent, potential, only virtually existing images? This was the idea behind GenStudio, an experimental interface launched in 2019 by the Metropolitan Museum in collaboration with Microsoft and MIT. This interface goes beyond simply navigating existing collections. It uses an early form of generative AI to create purely synthetic images from the collection data that do not resemble any pre-existing artifacts. However, this synthesis is understood as an exploration of a new, previously unexplored space: “Based on given artworks from the Met’s Open Access collection, a Generative Ad2 https://artsexperiments.withgoogle.com/xdegrees/ (19. 12. 2024).
Operative Image Spaces. Navigating Virtual Museum Collections 281 SJS 51 (2), 2025, 273–289 versarial Network (GAN) allows you to explore and visualize the spaces in between those pieces” (Fenstermaker, 2019). However, this space in between is not a space between physical objects in the collection, for example, between different historical teapots (fig.3), but an imaginary space of mere statistical possibilities. The old universal museum’s imperial claim to all-encompassing representation thus becomes atechnical utopia, the empirical space of the collectible expands into astatistical space of endless possibilities, and digital representations of collection artifacts become aresource for generating ever-new variants of images. Since OpenAI’s Dall-E 2 in 2022, a wave of new generative AI models for image and even video synthesis has emerged, making GANs like the one used in the example above look old-fashioned by comparison (Wilde, 2023). While GANs have typically been trained on limited databases of thousands or tens of thousands of pre-selected images, so-called foundation models such as Dall-E, Stable Diffusion, or Midjourney are trained on billions of image-text pairs harvested from all over the web. Moreover, while GANs only reproduce and synthesize recurring visual patterns found in the training data, these models learn relationships between images and their surrounding text to transform written prompts into visual images. Despite these and other fundamental differences, all these forms of generative AI are based Figure 3 Metropolitan Museum and Microsoft, GenStudio, 2019 Source: https://microsoft.github.io/GenStudio/.
282 Roland Meyer SJS 51 (2), 2025, 273–289 on similar conceptual premises: the spatialization of similarities and the creation of a homogenized, quantifiable, and continuously navigable operative space of possible images, in which all images, regardless of format, style, origin, and materiality, virtually coexist. From the perspective of these models, every image they are able generate – which is, of course, not every possible image, as these models are highly biased and ultimately limited by the boundaries of their training data – already exists as a potential image within these latent spaces, as do all the images, albeit in acompressed and abstracted form, with which they have been trained. In other words, for these models, any imaginable image – and again, the realm of the imaginable may seem endless but is, in fact, limited, incomplete, and distorted – is just one more or less probable variant in an endless chain of variations (Meyer, 2023). In fact, variations inspired by the original was one of the first features announced when Dall-E went public in 2022. On its website, Open AI showed a series of variations of George Seurat’s famous pointillist painting Un dimanche après-midi à l’Île de la Grande Jatte (1884–86) as a demonstration (fig.4). These pictures are not simply collages or remixes. Rather, they are interpolations in which the virtual image archive of existing images is used as a source of data points and machine learning is supposed to fill the gaps between them. Such AI-generated variations Figure 4 OpenAI, Dall-E 2, 2023 Source: https://openai.com/dall-e-2.
Operative Image Spaces. Navigating Virtual Museum Collections 283 SJS 51 (2), 2025, 273–289 are already used by museums as a form of marketing. In 2023, the Vienna Tourist Board presented an advertising campaign entitled UnArtificial Art, which featured AI variations of famous artworks by Gustav Klimt, Egon Schiele, and others, all now turned into cat content in their respective styles (fig.5). As they stated on their website, “AI mines vast repositories of existing artworks for data before replicating their substance and style. So, you could say it was era-defining artists like Klimt (ahuge cat fan, by the way) and Schiele that made AI artworks possible in the first place” (Vienna Tourist Board, 2023). The “value of the archive” (Meyer, 2023) is fundamentally redefined here – the art of the past becomes a resource of styles to be mined and fuel the production of ever new variants. But before Klimt and Schiele could “teach artificial intelligence a thing or two” (Vienna Tourist Board, 2023), their paintings first had to be digitally reproduced and converted into training data, transformed from individual masterpieces into vectors and data points in a huge latent space of billions of images. In these latent spaces, virtual and actual Klimts or Schieles, what they actually painted and what they could have painted potentially, coexist as equally possible variations of patterns, and what makes them images in the style of Klimt or Schiele is that their relative proximity in the latent space. As already stated, the invisible, multidimensional latent spaces of generative AI should not be confused with the three-dimensional galaxies of images spaces visualized in interfaces such as bauhaus infinity archive. But despite their differences in complexFigure 5 Campaign UnArtificial Art, 2023 © ViennaTouristBoard Source: https://b2b.wien.info/de/see-the-art-behind-ai-art-klimt-sujet-451836?view=asDownload.
284 Roland Meyer SJS 51 (2), 2025, 273–289 ity and function, all examples discussed so far ultimately share the same operative imaginary of seemingly unlimited access and control. Thus, it is no wonder that in a video Open AI produced to explain how Dall-E 2 was trained, they used almost exactly the same kind of imagery: a boundless black galaxy of free-floating images arranged into clusters and forming networks of relations.3 Models such as Dall-E, Midjourney and Stable Diffusion are a manifestation of a very specific, contemporary understanding of virtual image archives as both navigable spaces and exploitable resources. In this respect, they are more than just another tool for image production. Rather, they are the medium through which we negotiate what it possibly means to produce new images when almost every conceivable future image already seems to exist as a statistical possibility in a latent image space spanned by images of the past. In some ways, image generation by generative AI is indistinguishable from image search. When you enter a prompt into Dall-E, Midjourney, or Stable Diffusion, the software treats it less as an instruction to be executed and more as a search command that guides the model to a particular result– not unlike searching a database or catalogue, although you are not searching acollection of pre-existing images, but a latent space of possible images (Meyer, 2023). This latent space of possible images is, however, completely defined and determined by images already existing: the billions of training images scraped from the web and used for training these models. The underlying archival fantasy of generative AI is that there is no outside of the archive: Everything can be created, can be interpolated from what is already stored and made accessible. This totalizing fantasy of an archive without outside, in some way or the other, connects all examples mentioned so far, from Sood’s Cultural Big Bang to the generative spaces of today’s AI models. It builds on and ties in with a second fantasy: that the virtual collection objects do not represent physical objects in specific institutions with their own concrete history but stand for themselves as the main object of interest. Only as data objects can all these diverse pictures and artifacts become the object of operations of comparing, ordering, interpolating, and synthesizing – operations that would be impossible with physical collections. Far from being a mere double of physical collections, a deficient copy, or mere add-on, virtual image archives have become, as big visual data, a valuable resource to be mined, mobilized, and monetized (Allert & Richter, 2018). 4 Beyond Extraction With the progressive transformation of virtual image archives into an exploitable data resource, image operations tend to focus less and less on the individual image 3 https://openai.com/dall-e-2 (19. 12. 2024).
Operative Image Spaces. Navigating Virtual Museum Collections 285 SJS 51 (2), 2025, 273–289 and more and more on the modulation of visual patterns extracted from big visual data. Adrian MacKenzie and Anna Munster (2019) have described this new visual regime as “platform seeing”, a form of distributed visuality that emerges from the mass acquisition, accumulation, and operationalization of “image ensembles” through online digital platforms. Experimental museum interfaces serve as a playful introduction to this explorative and exploitative form of access to the virtual image archives aggregated by actors such as Google or Microsoft. Visualized as floating image populations in infinite space, images of the past become a seemingly natural resource that can be appropriated, varied, and transformed at will. Such an operative imaginary of a statistically controllable and sovereignly explorable space of all possible images is by no means harmless. Rather, in these interfaces, ahighly ideological dream of overview and control manifests itself, perpetuating the imperial, colonial, and extractivist logic that has already driven the emergence of Western museum collections. Since the 1980s, Tony Bennett (1988) and many other representatives of critical museology have analyzed how museums, as part of a larger “exhibitionary complex”, establish a particular order of visibility, an order in which the world in its entirety is metonymically made present and subjected to a classifying gaze through isolated objects torn from their context of origin and production. Museums, as Ariella Aisha Azoulay (2019, p. 109) has put it, are “worldless depositories” – they destroy the living networks of relationships in which cultural objects were once integrated, reduce them to their collectability and displayability, and replace complex and diverse cultural practices with standardized bureaucratic procedures that are equally applicable to any and all objects (Azoulay, 2019, p. 96). Perhaps there is no better image for these worldless depositories than the endless galaxies of free-floating image clusters offered by Google, Microsoft, and OpenAI: placeless spaces that can be navigated by a disembodied gaze, digital universal museums in the age of data extractivism. As artist Nora Al-Badri (2021) reminds us, “We live in a post-digital world as much as a post-colonial one”, and both perspectives cannot be separated. Thus, regarding the examples presented in this essay, the question arises: What could be possible alternatives to their imperialist, ultimately neocolonial logic? Are there alternative spaces that make virtual museum collections navigable without imagining them as exploitable resources? One potential model could be found in Digital Benin, an online project launched in 2022 (Agbontaen-Eghafona et. al., 2022). On the surface, Digital Benin looks like a straightforward digital online catalog: via the website digitalbenin.org information on more than 5 000 objects from 131 museums is available for the first time in a common database (fig.6). And thus, for the first time, the full extent of the looting becomes visible, which the often-used term Benin bronzes tends to obscure. Clicking through the catalog, one quickly comes across hundreds of musical instruments, spoons and combs, boxes, containers, and other household objects, in addition to the world-famous bronze heads, relief plates, and
286 Roland Meyer SJS 51 (2), 2025, 273–289 ivory masks. Bringing together information on all these objects, which until now had been difficult or almost impossible to access, was the focus of the project funded by the Ernst von Siemens Foundation, on which a fourteen-member international project team, supplemented by five scientific advisors in Nigeria, Kenya, and the United States, worked for two years. The desire for such a cross-collection overview is decades old (Savoy, 2021, p. 150). Still, the fact that over one hundred museums and institutions from twenty countries cooperated and shared their data would only have been conceivable after the current restitution debate. But Digital Benin is much more than just a cataloging project; it is perhaps the most ambitious attempt to date to think about virtual collections in an explicitly nonEurocentric, or in this case consciously “Edo-centric” way (Agbontaen-Eghafona et al., 2022). In addition to the catalog, the website offers seven additional sections called spaces, which go far beyond the usual logic of museum databases. The space “Ẹyo Otọ”, for example, groups the objects along categories that correspond to their original Edo designations. Here, one can not only hear the names of the various object categories read aloud in the language of the Kingdom of Benin, one learns, above all, something about the concrete ways in which the artifacts were used – acontextual knowledge that had been lost with the looting and musealization of the artifacts. In Figure 6 Digital Benin, 2023 (starting page) Source: https://digitalbenin.org.
Operative Image Spaces. Navigating Virtual Museum Collections 287 SJS 51 (2), 2025, 273–289 order to reconstruct this knowledge, the project team not only conducted archival research in Nigeria, but also spoke with a variety of Nigerian experts, curators, historians, and linguists, as well as with craftsmen and artists who continue to produce and use similar objects today. So instead of making publicly available only the incomplete object data that the apparatus of the Western museum deemed worthy of recording, Digital Benin lays the foundations for a new, polyphonic, networked, and living knowledge of these objects and their cultural references. While projects like Google Arts & Culture stage an operative imaginary as aspectacle of automated access to resources, Digital Benin uses modest technical means to show an alternative way of visualizing virtual museum collections: Instead of projecting isolated data points into a virtual space devoid of context and history, it opens up a multitude of situated and contextualized spaces for interpretation. And instead of nourishing the idea of a totalizing, all-encompassing virtual archive that is seemingly beyond all spatial and temporal limitations and detached from its history of origin, as was the focus of all the examples mentioned so far, Digital Benin offers access to historically located collections and the stories hidden within them. Rather than presenting us with an archive without an outside, in which what has already been stored, collected and classified marks the horizon of what can be represented, it strives to map the diverse, dynamic, and constantly growing networks of relationships that connect archived objects with historical events, physical places, and living practices. When thinking about alternatives to the prevalent representations of big visual data, we have to acknowledge how deeply our operative imaginaries of how to handle, access, and navigate virtual collections owe to the specific presuppositions of Western image cultures. That it is possible to imagine that highly diverse museum objects share the same homogenized virtual art space is not least due to the standardized image format that is typical for the photographic recording of museum collection objects: physicality is reduced to a surface, materiality becomes a visual texture, and differences in dimensions, formats, and media disappear. The mass digitization of museum artefacts thus ultimately reduces the diversity of cultural heritage to a set of visual data, a two-dimensional pixel matrix that can be calculated with, and thus establishes a Eurocentric understanding of images as the basis for the supposedly universal comparison of visual similarities (Schröter, 2022). In order to think beyond operative image spaces, therefore, we need a politics of digitization that does not simply extend the imperial and colonial logic of the universal museum to virtual space but radically breaks with its underlying ideological premises and operative imaginaries. Instead of sustaining the illusion of a universal, free-floating, disembodied gaze, we need to build interfaces suited to specific needs and interests, reflecting the diversity of subject positions and personal as well as collective histories. Instead of imagining new, seemingly neutral spaces of universal comparison, we need situated, specific and diverse spaces in which we can confront
294 Max Frischknecht SJS 51 (2), 2025, 291–315 data (Savage & Burrows, 2007; Burrows & Savage, 2014; Frade, 2016). In the digital humanities, advocates argue that big data analysis facilitates and augments research (Manovich, 2011) while critics describe the results as superficial and reductionist (Kitchin, 2014, p. 142 with reference to Trumpener 2009). Either way, big (visual) data holds the potential to reframe the epistemology of social science and humanities and the chosen methodological approaches must be thought through accordingly (Kitchin, 2014). This article attempts to make a humble contribution to these methodological discussions through an empirical study that outlines a possible humanmachine collaboration. It draws on a multidisciplinary approach combining concepts from digital humanities, media theory, science, and technology studies (STS), and critical data studies to explore how a “machine way of seeing” (Cox, 2022, p. 103) is shaped by its underlying infrastructure, and how this infrastructure, in turn, is shaped by sociotechnical imaginaries. In doing so, it provides a critical framework for how knowledge is created through big visual data and shaped by technology. 2.1 Machine Ways of Seeing In his seminal essay Ways of Machine Seeing as a Problem of Invisual Literacy Geoff Cox (2022) outlines the characteristics of how machines see in reference to John Berger’s Ways of Seeing (Berger, 1972). Seeing and naming things is a question of literacy. Literacy can be defined as “competence or knowledge of practices that allow users to maintain and build social imaginaries”, it is the ability to “read, write and program” (Cox, 2022, p. 105–106). It is a form of power and authority when certain ways of describing and naming things are held over others. This is the case with ImageNet and how it manifests what PixPlot can see through a defined set of images and categories. Building on Geoff Cox we can identify three aspects of a machine way of seeing. 1)Machines don’t see, they relate. Seeing for a machine is no longer singular or indexical, but rather distributed and multimodal (Cox, 2022, p. 110). At one point the image is digitized or created, at another it is given a label, at a third, it is viewed. The meaning that an image has for a machine is not based on its indexicality but rather on its relation to its categories and other images and their categories. 2)Machines don’t see, they read. Machines don’t see an image, but rather read it according to the model of the world they know (Cox, 2022, p. 108). What lies outside of this model can’t be recognized. This also applies in a technical sense as an algorithm reads an image pixel by pixel to interpret the correlation of a pixel with its neighbors. 3)Machines don’t see, they calculate probabilities. With reference to Crawford and Paglen (2019), Cox argues that seeing for a machine is a calculative practice, where the algorithm calculates the probability that, for example, the image shows a totem pole rather than a cactus. These models of probability are “built upon inherent human prejudices related to class, gender, and race” (Cox, 2022, p. 109; with reference to Crawford & Paglen, 2019). In summary, a machine way of seeing can be described
Through the Eyes of the Machine: Exploring Historical Photo Collections … 295 SJS 51 (2), 2025, 291–315 as based on relational perception, algorithmic reading, and probabilistic interpretation. To understand how the machine sees the world is to understand how the humans understand the machine and how they see and teach the machine to see the world (see also Hoggenmüller & Klinke in this Special Issue). 2.2 Data as Infrastructure Infrastructure study broadens our view of the various components that are at play when a machine sees. PixPlot and its CNN are built upon existing information infrastructure, specifically cyberinfrastructure. Cyberinfrastructures are “those layers that sit between base technology (a computer science concern) and discipline-specific science” (Bowker et al., 2010, p. 100). This applies to PixPlot as it builds on existing base technologies such as Tensorflow4 or Keras5 and was specifically developed for ahumanities context. Susan Leigh Star introduced us to the idea that infrastructures are relational to social practices and knowledge (Star, 1999). Bowker et al. (2010, p. 102) further proposed to investigate cyberinfrastructures as a set of distributed activities along a technical/social and a local/global axis: “The key question is not whether a problem is a ‘social’ problem or a ‘technical’ one. […] The question is whether we choose, for any given problem, a primarily social or a technical solution, or some combination.” If the CNNs training data doesn’t recognize certain aspects of the collection Brunner, we could define it as a technical problem. But at the same time, we can frame it as a social problem if we ask why certain motives appear in ImageNet and others do not. 2.3 Sociotechnical Imaginaries Technical systems are inseparable from the social contexts from which they emerge and operate. The exploration of infrastructures such as ImageNet leads us to the sociotechnical imaginaries (Jasanoff, 2015) embedded within such technological systems, influencing the development of computer vision and their applications and implications to explore big visual data. I understand a CNN and its training data as collectively held and institutionally stabilized reflections of imagined forms of social life and order, which, in the context of this study, are used to explore yet another imagined form of social life and order embedded within the collection Brunner. The concept of sociotechnical imaginaries allows us to recognize and investigate the multilayered levels of meaning embedded in using CNNs for historical big visual data exploration. It highlights the intertwined nature of technical systems and social contexts. The CNNs way of seeing is built upon an archive (ImageNet) and is used 4 Tensorflow is an open-source software library for machine learning applications developed and maintained by Google, https://www.tensorflow.org/ (14.6.2024). 5 Keras is an open-source deep learning library written in the Python programming language, https://keras.io/ (14. 6. 2024).
296 Max Frischknecht SJS 51 (2), 2025, 291–315 to explore an archive (the collection Brunner). Both of these imagined forms of social life, lead, as I will try to show, to a potential clash of meanings. 3 Object of Study, Data, and Methods This section sets out the methodological approach of this study. 3.1)Examines the human way of seeing the collection through a literature-based historical recontextualization of the photographer Ernst Brunner and his work. 3.2)Briefly introduces the digitization of the collection and presents the Brunner data set that was clustered with PixPlot. 3.3)Explains in detail the underlying mechanics of the PixPlot application and the ImageNet data set that was used to train the CNN. 3.4)Describes the analytical approach for the interpretation of the clusters through a combination of close and distant reading. 3.1 The Collection Ernst Brunner Ernst Brunner (1901–1979) attended a carpentry apprenticeship in his father’s company in Mettmenstetten, Switzerland. After two semesters at the carpentry college in Nürnberg and studying interior design at the Kunstgewerbeschule Zurich, Brunner moved to Lucerne and worked as an interior designer. During the great depression (1929–1939) Brunner lost his job and attended a public employment program where he worked on an inventory of historical monuments. He taught himself photography autodidactically and presented his first pictures to Zurich Publisher Regina around 1936. He quickly began photographing for magazines such as Das Schweizer Heim and Die Schweizer Familie. Starting in the 1940s, photographs were repeatedly published in the Swiss fine art magazine Du (Lüthi & Frei, 2024). In 1955 Brunner was part of the famous exhibition Family of Man by Edward Steichen at MoMA New York (Steiger, 1998). From the mid-fifties till his death in 1979 Ernst Brunner shifted focus and became part of the Aktion Bauernhausforschung in der Schweiz (Farmhouse Research Campaign in Switzerland) initiated by the CAS (2023c) between 1919 and 1960. He documented the distinct architecture of farmhouses in Lucerne and published a corresponding book in 1977 (Brunner, 1977).6 It was not until the 1990s that Brunner’s photographic work was rediscovered by a broader public (Lüthi &Frei, 2024). Most notably through Peter Pfrunder’s monograph Ernst Brunner: Photographien, 1937–1962 (1995) and an accompanying travelling exhibition (Verlorene Welten. Ernst Brunner Photographien 1937–1962). It prominently features Brunner’s work to illustrate the historical timber industry, farming, milling, or soil cultivation. 6 Brunner’s photographic material on farmhouses in the canton of Lucerne is not part of the CAS collection. The many farmhouses from other cantons found in the CAS collection resemble Brunner’s general interest and are not directly part of this research.
Through the Eyes of the Machine: Exploring Historical Photo Collections … 297 SJS 51 (2), 2025, 291–315 Till today, the monograph has decisively shaped the perception of Brunner as the photographer who documented the vanishing rural world (Özvegyi, 2020, p. 26). While Brunner’s work has been published in various magazines, the available academic literature on Ernst Brunner is yet very limited. To the best of my knowledge, only two articles contextualize his work so far with a focus on his oeuvre (Steiger, 1998; Özvegyi, 2020). At the moment, the first dissertation on Brunner’s collection is being written at the University of Basel (Lüthi, 2024). Although a large body of Brunner’s work is concerned with agriculture and craftsmanship documenting the everyday lives of farmers in rural Switzerland, the academic literature also highlights the diversity of the collection, including photographs about city life, industry, construction projects, or military service (Özvegyi, 2020, p. 28). Due to Brunner’s serial approach and rigid cataloguing, his work has further been described as “systematic” and “intended as objective documentations” (Steiger, 1998, pp. 26, 36). Brunner Figure 1 First Part of a Longer Photo Series on the Production of Charcoal (SGV_12N_04301 to SGV_12N_04330) Source: Collection Ernst Brunner, photo archive of Cultural Anthropology Switzerland, https://archiv.sgv-sstp.ch.
298 Max Frischknecht SJS 51 (2), 2025, 291–315 created an extensive series on work processes, for example, timber processing or charcoal making (cf. Fig. 1). The latter is a prominent example of his photographic work that aimed at visually preserving knowledge that could potentially be lost. This effort to preserve can also be recognized in his involvement in the Swiss farmhouse research movement. Ernst Brunner’s work must further be understood in the political context of its time. As Steiger illustrates, Brunner’s photographs contributed to the construction of a public image of Switzerland as a country of strong and free people living in an alpine landscape during World War II (Steiger, 1998, p. 33). Brunner’s images were removed from their serial context and shown as collages of national unity in popular Swiss magazines such as Das Schweizer Heim. In the succeeding decades the photographs were published several times and generally in a way “which emphasized their formal and artistic character rather than their documentary purpose” to support the construction of a national myth of “wise, but seemingly uncomplicated farmers” that perform “real” work (Steiger, 1998, p. 47). During the war, Switzerland became politically and militarily isolated causing the desire to ensure one’s own identity. Figure 2 Ernst Brunners Multifaceted Portrayal of Swiss Soldier During World War II (From Left to Right, Top to Bottom: SGV_12N_03504, SGV_12N_03563, SGV_12N_05317, SGV_12N_04682, SGV_12N_20301, SGV_12N_04603) Source: Collection Ernst Brunner, photo archive of Cultural Anthropology Switzerland, https://archiv.sgv-sstp.ch.
Through the Eyes of the Machine: Exploring Historical Photo Collections … 299 SJS 51 (2), 2025, 291–315 Steiger and Özvegyi both point out that Brunner knew how to create photographs that could be sold to magazines in the context of WWII. Ernst Brunner’s depictions of farmers as free, independent, and hard-working people and of brave heroic soldiers were appreciated visual material in the effort for an intellectual national defence (Geistige Landesverteidigung). However, it would be short-sighted to impute a propagandistic intention to Brunner’s work. As Özvegyi (2020, p. 41) shows, Brunner’s photographs of the Swiss military not only included portrayals of soldiers fit for service, but also sober recordings of their daily lives (cf.Fig.2). To summarize, the way of seeing the collection driven by public discourse, magazines, and exhibitions, focuses on the depiction of rural life and craftsmanship. The academic perspective complements this view by highlighting the documentary and rigid photographic approach, the complexity and diversity of the collection and its specific historical context. For example, its use for the construction of a nation’s image of brave soldiers and independent farmers during WWII while providing a far more nuanced view. These observations will guide the following examination of the PixPlot clusters. Will the machine see the same? 3.2 Digitization of the Collection and Dataset After the death of Ernst Brunner in 1979 the collection was handed over to the CAS photo archive which is responsible for its archiving and digitization. The physical collection contains approximately 48 000 black-and-white negatives in medium format, 20 000 prints on index cards organized in a corresponding file system and additional material such as handwritten indexes, several hundred historical prints and specimen copies of published photos. Between 2014 and 2018 CAS conserved, restored, digitized, and, for the most part, also indexed the black-and-white negatives within a larger digitalization effort. Since 2021 and in the context of the SNSF research project Participatory Knowledge Practices in Analogue and Digital Image Archives (PIA, 2023a)7 this is also being done for the additional material. The collection is further transferred into a new data model and base and an extended cataloguing is carried out. These efforts are being made not least to make the collection more accessible and understandable. As shown, the collection has so far received little academic attention despite its size and significance. The Brunner data set used for clustering in this article has been created by accessing the PIA metadata API (PIA, 2023b) and IIIF API (PIA, 2023c). A Python script collected ID (e. g.SGV_12N_20301) and image title (e. g. “Soldaten beim Sport”) for each available digital object in the API. At the time of the collection, the script collected a total of 47 837 ID and title pairs which were saved in a CSV file. A second Python script was used to download each image as a JPG file based on its 7 SNSF Grant Number 193788, cf. https://data.snf.ch/grants/grant/193788 (14. 6. 2024).
300 Max Frischknecht SJS 51 (2), 2025, 291–315 ID from the PIA image server. Interestingly, this produced a total of 47 020 image files, 817 files less than there are objects in the metadata API. One example of such a missing image file is SGV_12N_27618 (“Häuser auf einer Alp”): While the physical negative exists in the archive and the metadata API returns information on the object, the image server responds with an internal error that the file is temporarily not available. These interruptions in file availability are due to the complexity of the infrastructure and the fact that development is ongoing. They beautifully echo Susan Leigh Stars statement that infrastructure becomes visible upon breakdown (Star, 1999). The historical photographic collection, which we perceive as a stable entity, develops a certain dynamic in the digital. The PixPlot clusters do not represent the collection in itself, but the collection in a specific state at a specific point in time. As PixPlot requires an image file to see, the 817 digital objects with no available image file were excluded from the data set and a total of 47 020 images (98.3%) were used for clustering (cf. Table 1). 3.3 PixPlot and Convolutional Neural Networks (CNN) PixPlot is a free, open-source application developed by Yale’s Digital Humanities Lab in 2017 (see also Herms & Lehmann in this Special Issue). It has been used to cluster big visual data sets such as the collections of the Yale Center for British Art (Duhaime, 2017) or the Harvard Art Museum (Rodighiero et al., 2022). Cyberinfrastructures like PixPlot are not stand-alone software packages but built upon (and dependent on) existing code libraries.8 PixPlots convolutional neural network (CNN) was pre-trained on ImageNet (Stanford Vision Lab, 2011) to detect key features from images, such as edges, shapes, textures, and patterns, and calculate the probability that the image shows a certain object such as a “tree” or “house”. Based on these identified characteristics PixPlot clusters images according to their visual similarity. For each image, a featurization space of 2 048 dimensions is created, meaning that the CNN produces a vector with 2 048 values, each of which corresponds to a specific feature of the image. A vector is basically a list of numbers 8 For a complete list of libraries see the PixPlot code repository on Github (Duhaime, 2017). Table 1 Construction of the Image Set for Clustering Type Count Available Objects in the Metadata API 47 837 Available Files on the Image Server 47 020 Unavailable Files on the Image Server 817 Total Files used in PixPlot to cluster 47 020 Source: Frischknecht, November 2023.
Through the Eyes of the Machine: Exploring Historical Photo Collections … 301 SJS 51 (2), 2025, 291–315 that serve as a kind of coordinate that situates an image and its features in relation to other images. The featurization space holds the information on how similar an image is to another one. These features are not necessarily visible or understandable for humans but rather created by the algorithm through an iterative process of guess and check. Finally, to plot the images in two-dimensional space (on an X and Y axis) the 2 048 image features need to be reduced to two. This is achieved using a dimensionality reduction algorithm, specifically UMAP (Uniform Manifold Approximation and Projection) (McInnes et al., 2020), that aims at reducing the dimensions while retaining as much of the relevant information as possible (see also Hoggenmüller & Klinke in this Special Issue). To understand the in PixPlot embedded sociotechnical imaginaries we need to analyze its training data set ImageNet. The central characteristics of PixPlot’s way of seeing – relational perception, algorithmic reading, and probabilistic interpretation– are essentially derived from this training data. ImageNet, originally created for visual object recognition, was one of the first widely available large-scale image data sets and has been central for the advancement of computer vision and deep learning research. It was developed at Stanford Vision Lab and first presented in 2009 at the IEEE Conference on Computer Vision and Pattern Recognition (Deng et al., 2009). Each year between 2010 and 2017 the data set and its accuracy have been developed further through the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) (Russakovsky et al., 2015). The data set contains 1 281 167 training images (to learn visual features), 50 000 validation images (to validate how well these features can be generalized), 100 000 test images (to test what has been learned on unknown data) and 1 000 object classes (that specify the labels such as “tree” or “house”). The images were collected from the internet through automated and manual searches. It remains unclear from which particular sources the images come but it seems likely that they are the results of well-known search engines such as Google or Yahoo on the one hand, and popular image websites such as Flickr on the other (cf. Dengetal., 2009; Russakovsky et al., 2015). The images are largely derived from North American amateur photography (Cox, 2022) and were labelled by precarious workers through Amazon’s Mechanical Turk (Crawford, 2022). The classes to label the images are based on WordNet, a lexical database of nouns, verbs, adjectives, and adverbs that are grouped into sets of synonyms (Princeton University, 2010). In the subsequent analysis of this article, the openly available ImageNet data set on Kaggle is used for comparison and interpretation of the clustering results (Kaggle, 2020). As we know, each data set comes with inherent bias. Generally, we can examine three levels of bias that can vary significantly across different data sets (Tommasi etal., 2017). Knowledge of these biases will support our assessment of the embedded sociotechnical imaginaries and whether a certain behavior of PixPlot can be framed as a technical or social problem. The first form of bias, the capture bias, relates to the distinct features of the images, such as angle or lightning. ImageNet-trained CNNs
302 Max Frischknecht SJS 51 (2), 2025, 291–315 tend to have a bias towards identifying images based on texture rather than shape, contrary to humans where the case is the opposite (Geirhos et al., 2018). The ImageNet data set holds primarily images taken in well-lighted situations contributing to its texture-bias (Hermann et al., 2020). In contrast, humans are trained to see in diverse lighting conditions which is why the recognition of shapes takes on greater importance. The second form of bias, label bias, relates to the data sets visual semantic categories. According to Yang et al. (2020), WordNet (the lexical database used to create the ImageNet categories) includes words that are offensive in terms of sexuality or race, sensitive terms that can be offensive in a specific context, and terms that are hardly applicable to the description of images (e. g.“vegetarian”). While some of these words have been removed from the ImageNet data set, Yang et al. show that many slipped through the filtering process. Further, it is worth considering the language difference that results from the temporal difference of ImageNet and the collection Ernst Brunner. Lastly, the negative bias relates to the limits of the available images and categories and their representation of the world. A negative bias is challenging to address because changing the form of representation doesn’t necessarily lead to a broader or more inclusive representation (Tommasi et al., 2017). Addressing the negative bias would lead to a more extensive data set, but it can never be reduced entirely. The negative bias can be understood as the periphery of the machine’s eye. 3.4 Analysis Method: Distant and Close Reading For the following analysis of the PixPlot clusters a combination of distant and close reading is proposed (Jockers, 2013; Moretti, 2016). Distant reading is understood as the “not reading” of (or not looking at) the collection photograph by photograph but rather from afar “to focus on units that are much smaller or much larger” (Moretti, 2016, p. 50).9 4.1)Describes the clusters from afar, not the individual photographs are of importance, but rather groups of photographs and how the clusters relate to each other. 4.2)Proposes a close reading of specific groups of photographs to identify what the machine might see, how this relates to the ImageNet data set and how it aligns or differs from the human perspective identified previously under section3.1. 4 Analysis and Results 4.1 Cluster Analysis: A Panorama of Brunner’s Work PixPlot created and numbered ten visual clusters from the collection Brunner that are accessible via side navigation in the application. Figure3 shows these clusters and 9 While Moretti’s distant reading approach was originally developed for literature studies it has been adopted to analyze various media, including images (cf. Arnold & Tilton, 2019).
Through the Eyes of the Machine: Exploring Historical Photo Collections … 303 SJS 51 (2), 2025, 291–315 visualizes how they overlap. Table2 provides a brief description of each cluster starting at the top left corner with cluster3 and then continuing in a clockwise rotation. The ten clusters provide a distant view of the collection and highlight the large diversity of Ernst Brunner’s work. Equipped with such an overview, the next section looks at three concrete examples to examine how human and machine ways of seeing might align, differ, or complement. Figure 3 A Rough Outline of the Ten Clusters Produced by Pixplot Source: Collection Ernst Brunner clusters, local instance of the PixPlot application [Screenshot, red outlines added, Frischknecht].
310 Max Frischknecht SJS 51 (2), 2025, 291–315 how most farmers appear in conjunction with distinct objects, especially long tools like scythes, shovels, or pitchforks. A look into ImageNet reveals that there are not many categories for specific agricultural tools nor does the category “farmer” exist. The categories “parallel bars, bars” and “horizontal bar, high bar” are probably the closest thing to a pitchfork. Searching for “parallel bars” in the ImageNet data set reveals images of athletes performing high jump (cf. Fig. 8). Comparing them with the photographs in Figure7 lets us suspect why PixPlot clustered the way it did. The machine doesn’t see “wise, but seemingly uncomplicated farmers” doing “real work” (Steiger 1998, 47). The machine’s perception is relational and determined by the distinctive horizontal and vertical objects. Looking at the second historical narrative of the heroic soldier and national defense, PixPlot seems to have a clearer vision. The algorithm clusters photographs with greater visual variance than in the previous example on farmers. Soldiers are shown alone, in groups of different sizes and with different equipment (Fig. 9). Looking at the images makes clear that soldiers introduce more visual consistency due to their helmets and uniforms. Interestingly, in some cases, PixPlot places Soldiers in close proximity to military equipment such as tanks or anti-aircraft guns, which differ greatly visually. This alleged contextual knowledge shows how relational Figure 8 Selection of Imagenet Images Labeled With “Horizontal Bars, Bars” Source: Navigu, ImageNet Dataset Explorer, https://navigu.net/#imagenet [Screenshot, Frischknecht].
Through the Eyes of the Machine: Exploring Historical Photo Collections … 311 SJS 51 (2), 2025, 291–315 perception, algorithmic reading, and probabilistic interpretation influence each other when concepts often co-occur in the training data. It must be noticed that ImageNet is highly militarized containing many categories such as “rifle”, “assault rifle”, “tank, armored combat vehicle”, “warplane”, or “cannon”.10 Again, with a focus on objects as the data set includes “military uniform” but not “soldier”. This militarized view overlooks the versatility of Brunner’s work and his portrayal of soldiers as playful and vulnerable humans playing soccer or sleeping (cf. Fig. 2). On the contrary, the machine falls into an almost propagandistic mode. 10 Here it would be interesting to examine more closely how this speaks for North American society, from which large parts of the ImageNet images originate. Figure 9 Clustering of Soldiers in Diverse Situations Source: Ernst Brunner clusters, local instance of the PixPlot application [Screenshot, Frischknecht].
312 Max Frischknecht SJS 51 (2), 2025, 291–315 5 Conclusion The article started with the assumption that visual similarity can be fruitful for the exploration and examination of large collections if the CNN clusters the photographs along the central topics and narratives inherent in the collection. The study juxtaposed a human perspective on the collection derived from literature with amachine perspective through PixPlot and its underlying infrastructure, particularly ImageNet. The main focus of the study was the examination of the epistemological implications of such a human-machine interpretation interplay, and it was conducted along four research questions: 1) Can a CNN recognize the central (visual) topics and narratives of a historical photographic collection? 2)To what extent does this machine way of seeing align or differ from a human perspective on the collection? 3)Does this difference allow for interesting modes of human-machine collaboration? 4)And what epistemological implications arise from such a collaboration? The central topics and narratives prevalent in public discourse, magazines, and exhibitions, are Brunner’s depiction of rural life and craftsmanship. Academic perspectives complement this view by highlighting the collection’s complexity, diversity, historical context, and the documentary and rigid photographic approach. The CNN recognized some, but not all, of these central topics and narratives. The clusters make the important role of rural life and craftsmanship comprehensible while simultaneously showing the impressive size and diversity of the collection. This is a result mainly due to the scalability of the computational approach. The CNNs bias towards textures aligns well with Brunner’s documentary approach, in particular the extensive photo series, such as the one on coal making, are clustered together, at least for the most part. In relation to historical context, matters get more complicated. A narrative such as the free, independent farmer becomes partly comprehensible while the soldiers are mostly shown in training situations lacking Brunner’s humanistic portrayal of their everyday lives. Besides accessibility resulting from scalability, I see the potential of an interpretative collaboration above all in the fact that the machine view emphasizes formal aspects of Brunner’s photographic language. Building on the texture bias, a further specialized CNN could, for example, support questions in art history regarding image composition. On the contrary, the collaboration must be viewed critically when the machine’s view serves as the basis for the development of content-related questions, for example, to inform potential novel research directions. Here, the sociotechnical imaginaries embedded in cyberinfrastructure turn out to be a kind of epistemological trojan horse. Hidden behind layers of interface, software and code lies the ImageNet data set with its own classification of the world from which the machine’s way of seeing emerges. This has epistemological consequences insofar as that the CNN has the ability to shift our attention towards certain narratives and imaginaries that we might assume arise from the collection itself but actually originate
Through the Eyes of the Machine: Exploring Historical Photo Collections … 313 SJS 51 (2), 2025, 291–315 in the infrastructure. A clash of meaning arises between a historical photo collection and the CNNs training data set. In particular, the different time dimensions of the collection and the training data create hidden temporal references (examining these in more detail would be an exciting undertaking in itself). To assess the origin of phenomena that one might observe in the collection through the CNNs clustering, it is necessary to cross-reference the machine ways of seeing with its original ontology ImageNet. Here, the study showed that infrastructural and critical data studies prove to be a suitable tool to critically interpret, and where necessary, recontextualize information gained from machine-learning based big visual data exploration. In conclusion, it can be stated that an improved human-machine interpretation interplay depends on training data sets tailored to the human’s interpretation interest. It would be interesting to see how a CNN clusters the collection Brunner if trained with historical data including categories derived from description of work processes or oral history. To further promote the potential of machine-learning-based approaches for big visual data analysis, more diverse and thematically specific data sets must become publicly accessible. Inspired by human ways of seeing we should train many ways of machine seeing instead of a few trying to depict the whole world. 6 References Arnold, T., & Tilton, L. (2019). Distant Viewing: Analyzing Large Visual Corpora. Digital Scholarship in the Humanities, 34(1), i3–i16. https://doi.org/10.1093/llc/fqz013 Berger, J. (1972). Ways of seeing: Based on the BBC television series. Penguin Books. Bowker, G. C., Baker, K., Millerand, F., & Ribes, D. (2010). Toward Information Infrastructure Studies: Ways of Knowing in a Networked Environment. In J. Hunsinger, L. Klastrup, & M. Allen (Eds.), International Handbook of Internet Research (pp. 97–117 ). Springer Netherlands. https:// doi.org/10.1007/978-1-4020-9789-8_5 Brunner, E. (1977). Die Bauernhäuser im Kanton Luzern. Empirische Kulturwissenschaft Schweiz. Burrows, R., & Savage, M. (2014). After the crisis? Big data and the methodological challenges of empirical sociology. Big Data & Society, 1(1), 1–6. https://doi.org/10.1177/2053951714540280 CAS (Cultural Anthropology Switzerland). (2023a). Das Fotoarchiv der EKWS (Empirische Kulturwissenschaft Schweiz). https://archiv.sgv-sstp.ch/ (accessed March 12, 2023). CAS (Cultural Anthropology Switzerland). (2023b). SGV_12 Ernst Brunner. https://archiv.sgv-sstp.ch/ collection/sgv_12/all/1 (accessed January 30, 2023). CAS (Cultural Anthropology Switzerland). (2023c). Die Bauernhäuser der Schweiz. https://www.volkskunde.ch/sgv/publikationen/reihen/die-bauernhaeuser-der-schweiz/ (accessed January 30, 2023). Cox, G. (2022). Ways of machine seeing as a problem of invisual literacy. In A. Dewdney & K. Sluis (Eds.), The networked image in post-digital culture (pp. 101–113). Routledge. Crawford, K. (2021). Atlas of AI: Power, politics, and the planetary costs of artificial intelligence. Yale University Press. Crawford, K., & Paglen, T. (2019). Excavating Ai: The Politics of Images in Machine Learning Training Sets. https://excavating.ai/
314 Max Frischknecht SJS 51 (2), 2025, 291–315 Deng, J., Dong, W., Socher, R., Li, L., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition (pp. 248–255). https://doi.org/10.1109/CVPR.2009.5206848 Duhaime, D. (2017). PixPlot. https://github.com/YaleDHLab/pix-plot (accessed February 30, 2023). Frade, C. (2016). Social theory and the politics of big data and method. Sociology, 50(5), 863–877. https://doi.org/10.1177/0038038515614186 Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., & Brendel, W. (2018). ImageNettrained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. arXiv. https://arxiv.org/abs/1811.12231 (accessed June 14, 2024). Hermann, K. L., Chen, T., & Kornblith, S. (2020). The origins and prevalence of texture bias in convolutional neural networks. Proceedings of the 34th International Conference on Neural Information Processing Systems, 19000–19015. Herms, K., Lehmann, J. (2025). Seeing Like a Field? Schweizerische Zeitschrift für Soziologie, 51(2), Special Issue hrsg. von S. W. Hoggenmüller, Big Visual Data als neue Form des Wissens: Potenziale, Herausforderungen und Transformationen. Hoggenmüller, S. W., Klinke, H. (2025). Metabilder als Forschungswerkzeuge: Zur Kontingenz und algorithmischen Bedingtheit ihrer Herstellung. Schweizerische Zeitschrift für Soziologie, 51(2), Special Issue hrsg. von S. W. Hoggenmüller, Big Visual Data als neue Form des Wissens: Potenziale, Herausforderungen und Transformationen. Jasanoff, S. (2015). One. Future Imperfect: Science, Technology, and the Imaginations of Modernity. In Dreamscapes of Modernity (pp. 1–33). University of Chicago Press. https://doi.org/10.7208/ chicago/9780226276663.001.0001 Jockers, M. L. (2013). Macroanalysis: Digital methods and literary history. University of Illinois Press. Kaggle. (2020). ImageNet object localization challenge. https://kaggle.com/competitions/imagenet-object-localization-challenge (accessed December 11, 2023). Kitchin, R. (2014). The Data Revolution: Big Data, Open Data, Data Infrastructures & Their Consequences. SAGE Publications Ltd. https://doi.org/10.4135/9781473909472 Lüthi, F. (2024). Partizipative Wissenspraktiken in analogen und digitalen Bildarchiven am Beispiel der Sammlung Ernst Brunner (SGV). https://universe.unibas.ch/projects-collaborations/9088 (accessed March 12, 2024). Lüthi, F., & Frei, F. (2024). Objektbiografien im interdisziplinären Fokus. Ein Werkstattbericht aus der Sammlung Ernst Brunner. Lecture, Seminar für Kulturwissenschaft und Europäische Ethnologie, Universität Basel. Manovich, L. (2011). Trending: The Promises and the Challenges of Big Social Data. In M. K. Gold (Ed.), Debates in the Digital Humanities (pp. 460–475). University of Minnesota Press. https:// doi.org/10.5749/minnesota/9780816677948.003.0047 McInnes, L., Healy, J., & Melville, J. (2020). UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv. https://doi.org/10.48550/arXiv.1802.03426 Moretti, F. (2016). Distant Reading. Konstanz University Press. Özvegyi, A. (2020). Von der heroisierten Inszenierung zur ernüchterten Darstellung? Fotografien von Ernst Brunner aus seiner Militärdienstzeit bei der Fliegerabwehrbatterie 311. Schweizerisches Archiv für Volkskunde/ Archives suisses des traditions populaires, 116(2), 25–46. https://doi.org/10.5169/ SEALS-913987 PIA (Participatory Image Archives). (2023a). Participatory Image Archives. https://about.participatory-archives.ch/ (accessed January 30, 2023). PIA (Participatory Image Archives). (2023b). PIA Metadata API. https://opendata.swiss/de/dataset/ pia-metadata-api (accessed December 23, 2023).
Through the Eyes of the Machine: Exploring Historical Photo Collections … 315 SJS 51 (2), 2025, 291–315 PIA (Participatory Image Archives). (2023c). PIA IIIF API. https://opendata.swiss/de/dataset/pia-iiif-api (accessed December 23, 2023). Pfrunder, P. (1995). Ernst Brunner: Photographien, 1937-1962. Schweizerische Gesellschaft für Volkskunde. Princeton University. (2010). WordNet: A lexical database for English. https://wordnet.princeton.edu (accessed April 30, 2023). Rodighiero, D., Derry, L., Duhaime, D., Kruguer, J., Mueller, M. C., Pietsch, C., Schnapp, J. T., Steward, J., & metaLAB. (2022). Surprise machines: Revealing Harvard Art Museums’ image collection. Information Design Journal, 27(1), 21–34. https://doi.org/10.1075/idj.22013.rod Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., & Fei-Fei, L. (2015). ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision, 115(3), 211–252. https://doi.org/10.1007/ s11263-015-0816-y Savage, M., & Burrows, R. (2007). The Coming Crisis of Empirical Sociology. Sociology, 41(5), 885–899. https://doi.org/10.1177/0038038507080443 Stanford Vision Lab. (2011). ImageNet. https://image-net.org/index.php (accessed May 17, 2023). Star, S. L. (1999). The Ethnography of Infrastructure. American Behavioral Scientist, 43(3), 377–391. https://doi.org/10.1177/00027649921955326 Steiger, R. (1998). On the uses of documentary: The photography of Ernst Brunner. Visual Sociology, 13(1), 25–47. https://doi.org/10.1080/14725869808583785 Stevenson, J. (2008). The Online Archivist: A Positive Approach to the Digital Information Age. In L. Craven (Ed.), What are archives? Cultural and theoretical perspectives a reader (pp. 89–108). Ashgate. Terras, M. M. (2011). The Rise of Digitization. In R. Rikowski (Ed.), Digitisation Perspectives (pp. 3–20). SensePublishers. https://doi.org/10.1007/978-94-6091-299-3_1 Tommasi, T., Patricia, N., Caputo, B., & Tuytelaars, T. (2017). A Deeper Look at Dataset Bias. In G.Csurka (Ed.), Domain Adaptation in Computer Vision Applications (pp. 37–55). Springer International Publishing. https://doi.org/10.1007/978-3-319-58347-1_2 Trumpener, K. (2009). Critical Response I. Paratext and Genre System: A Response to Franco Moretti. Critical Inquiry, 36(1), 159–171. https://doi.org/10.1086/606126 Yale Digital Humanities Lab. (2017). Yale DHLab – PixPlot. https://dhlab.yale.edu/projects/pixplot/ (accessed May 13, 2023). Yang, K., Qinami, K., Fei-Fei, L., Deng, J., & Russakovsky, O. (2020). Towards fairer datasets: Filtering and balancing the distribution of the people subtree in the ImageNet hierarchy. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 547–558. https://doi. org/10.1145/3351095.3375709
Marc Szydlik ISBN 978-3-03777-303-1 200 pages 15.5 cm × 22.5 cm Fr. 34.– | Euro 34.– What do adults say about their parents? What emotions do daughters and sons have when it comes to their mothers and fathers? What stories do they tell? This book offers personal firsthand thoughts on family situations and histories. Daughters and sons express and explain their connections with their parents from early childhood across the whole life course. They talk about cohesion, ambivalence, conflict and distance. They report love and hate, eternal bonds and painful separations. The statements address both relationships with living parents and past ties to mothers and fathers who have passed away. This is the fourth book of the SwissGen project, arepresentative survey of intergenerational relations in Switzerland. The analysis volumes offer key findings and examine central generational issues in depth (“Generationen zwischen Konflikt und Zusammenhalt” / “Generations between Conflict and Cohesion”). The data volume provides general information on the research project and basic quantitative results in form of summarised tables (“Relations with Parents: Questions and Results”). The book at hand is the qualitative complement to the analysis volumes. It offers over 1 500 statements of adults in their own words. The study was conducted under the direction of Marc Szydlik at the Department of Sociology at the University of Zurich. Relations with Parents Statements Seismo Press, Zurich und Geneva www.seismopress.ch [email protected] Seismo Press Social sciences and social issues
317 Swiss Journal of Sociology, 51 (2), 2025, 317–336 * Université de Lausanne, Institut des sciences sociales, CH-1015 Lausanne, [email protected]. ** Staatsbibliothek zu Berlin Preußischer Kulturbesitz, D-10772 Berlin, [email protected]. Seeing Like a Field? Katrin Herms* and Jörg Lehmann** Abstract: This article introduces the notion of seeing like a field – the view from the inside of a Bourdieusian field in which agents and their social positions are located. Potentials and fallacies of concretizing field theory with social network analysis (SNA) are evaluated on the basis of two examples of current research approaches to big visual data: computer vision networks and image similarity analysis. Keywords: Field analysis, computer vision networks, image similarity analysis, machine learning, predictive analytics Voir comme un champ? Résumé: Cet article introduit la notion de voir comme un champ – la vue de l’intérieur d’un champ dans lequel les agents et leurs positions sociales sont situés. Le potentiel et les pièges d’une concrétisation de la théorie des champs par l’analyse des réseaux sociaux (ou Social Network Analysis – SNA en anglais) sont discutés au travers de deux exemples d’approches de recherche actuelles par l’analyse du big visual data: les réseaux de vision par ordinateur et l’analyse de similarité des images. Mots-clés: Analyse de champ, réseaux de vision par ordinateur, analyse de similarité des images, apprentissage automatique, analyse prédictive Sehen wie ein Feld? Zusammenfassung: In diesem Artikel wird das Konzept Sehen wie ein Feld eingeführt – der Blick aus dem Inneren eines Bourdieu’schen Felds, in dem sich Akteurinnen und ihre sozialen Positionen befinden. Potenziale und Fallstricke einer Konkretisierung der Feldtheorie durch soziale Netzwerkanalyse (SNA) werden anhand von zwei Beispielen aktueller Forschungsansätze für große visuelle Datensätze diskutiert: Computer-Vision-Netzwerke und Bildähnlichkeitsanalysen. Schlüsselwörter: Feldanalyse, Computer Vision Netzwerke, Bildähnlichkeitsanalyse, maschinelles Lernen, prädiktive Analytik DOI 10.26034/cm.sjs.2025.6906 © 2025. This work is licensed under the Creative Commons Attribution-NonCommercialNoDerivatives 4.0 License. (CC BY-NC-ND 4.0)
318 Katrin Herms and Jörg Lehmann SJS 51 (2), 2025, 317–336 1 Introduction1 Internet communication has become omnipresent. Images, texts, and videos retrieved in social media offer insights into everyday situations like cooking or spending time with family, friends, and pets. Digital stories are characterised by increasing imagery, as “people are three times more likely to engage with tweets that include visual content” (Alton, 2024). Users build social ties with other platform users when following, liking, commenting, or sharing their images. Since more and more people are posting pictures from their daily lives online, characteristics of circulating images, their production and reception in a social media environment form new fields of research (Ma & Fan, 2022). The social network Instagram is aplatform that combines images and videos with microblogging. With more than 2.4 billion users since its launch in 2010, out of which 500 million users access the platform each day and who created over 990 million daily photo sequences that were published in 2023 as “stories” (Demandsage Instagram Statistics, 2024), Instagram can be seen to hold big data. It is an archive of everyday digital practices. Based on the available data (i.e. visual content and trace data including socio-demographic variables), the service offers a source for linking observations of aesthetic preferences with an analysis of users’ characteristics and properties. Under the condition of data access, analysis enables to group users into categories of taste, explore their networks, and make statements about their lifestyles. Such an analysis can be done with economic interest or with sociological ambitions. Whilst sociologists study social dynamics and consequences of media use, ideally by following ethical standards of data protection, for instance to avoid user profiling, big tech platforms primarily long for profit maximisation based on trace data. Instagram is but one example of how the availability of big data – and not only big visual data, but also data on the users – changes the view of the big tech corporations onto their users. For them, it is now possible to See like a market, to cite Fourcade and Healy (2017), which indicates the ability to calculate, stratify, and apprehend user behaviour based on consumer tracking and measurements that are applied on massive collections of trace data: “As new techniques allow for the matching and merging of data from different sources, the results crystallise – for the individuals classified – into what looks like asupercharged form of capital” (Fourcade & Healy, 2017, p. 10). To give an example, Meta as the parent company of Instagram is able to combine images with further data sources from Facebook and existing telephone numbers from WhatsApp as well as the respective real names. The result is a precise set of digital records, a data double (Bouk, 2017) of all users active 1 The contributors would like to thank the anonymous peer reviewers for their valuable feedback and their guidance and support, as well as the editor of this volume, Sebastian W. Hoggenmüller, for his patience, perseverance and all his effort in bringing this article to publication. The contributors would also like to cordially thank Janna Joceli Omena, Elena Pilipets, Beatrice Gobbo, and Jason Chason for their kindness to provide illustrations for being discussed in this article. The first three illustrations (Figure 1.1–1.3) below stem from their article in the journal Diseña, published under the CC-BY-SA license.
Seeing Like a Field? 319 SJS 51 (2), 2025, 317–336 on these platforms. Historically, capitalist markets and bureaucratic organisations in the service of nation-states have always strived to collect data about consumers and citizens. These data serve to apply rules and measures, to implement classificatory schemes and analyse consumption patterns and scores in order to make the real world legible to companies and states. However, in the 21st Century analytic capacities supported by computing power go way further. Predictive analytics – the identification of preferential consumption patterns – enables the anticipation of new needs, desires, and trends in real time during their formation. As big tech companies are able to distinguish consumers along categories such as riskiness or worth, such new stratifying technologies also have the power to discriminate in a way that was yet unimaginable in the 20thCentury (Fourcade & Healy, 2017, p. 24). At the same time, big (visual) data are available for scientific purposes. Social media data are used in computational social science projects to test, among others, machine learning applications and to optimise models that are supposed to predict social behaviour. A risk of this development is the creeping economisation of digital sociology if researchers uncritically adopt and reproduce economic categories of user stratification. With the increase in publicly accessible trace data, another challenge is to choose meaningful heuristics for the interpretation of big (visual) data. A classic epistemological framework for the analysis of everyday social practices was developed by the French sociologist Pierre Bourdieu. He was the first to ask how everyday situations are regulated without people consciously following predetermined rules. Inspired by literary descriptions and based on his empirical observations of French cultural idiosyncrasies in living, eating, dressing, and receiving art, Bourdieu unfolds his analysis in la Distinction (Bourdieu, 1979). Several researchers argued later on that social network analysis (SNA) would enable a concrete method of application for Bourdieu’s field theory. While Bourdieu was amongst the first sociologists to employ correspondence analysis (Bourdieu, 1979, pp. 296, 596, 622), Wouter de Nooy (2003) proposed to transfer the contingency tables on which Bourdieusian correspondence analysis rests into adjacency matrices used for social network analysis. A more recent proposal is given by Stefan Bernhard (2008) who argues that the theoretical strengths of field theory should be combined with the empirical possibilities of network analysis in order to overcome the weaknesses of both approaches. In light of the growing possibilities for sociometry based on trace data, more and more quantitative and structural approaches have been developed to combine field theory and SNA. For example, Serino et al. (2017) suggested a practical implementation to combine blockmodeling and multidimensional data analysis in a case study on theatre industry as a field of cultural production. Nevertheless, concrete suggestions to use field theory for analysing the phenomenon of pictorial digital forms of expression with big visual data have not yet been developed. Therefore, and in allusion to Fourcade and Healy’s critical study Seeing like a market (2017), we ask: Is it possible for sociology resting on big visual data to See like a field?
326 Katrin Herms and Jörg Lehmann SJS 51 (2), 2025, 317–336 references from the image (here “mosquito”) to entities such as persons, places, and things (here “zika virus”, “microcephaly”, “pregnancy”, “ministry of health” etc). In this way, the Knowledge Graph goes beyond the content of the image itself and provides a structure to the network, where images and web entities serve as nodes and the occurrence of web entities (or other textual descriptions) in relation to images serve as edges. The dichotomy between certain web entities (here disease versus health) may lead to the interpretation of a public controversy. As a methodological extension with respect to attention economy, we propose to combine this network approach with a qualitative analysis of image-text annotations to decide and evaluate what the entities refer to in specific situations. The third kind of network rests on the images in combination with the URLs on which the images have been published. Similar visual contents are suggested by the powerful Google Image Search. These image-domain networks (c) allow to explore and analyse the web origin of those images, whereby images and domains serve as nodes and the occurrence of domains in relation to images serve as edges. In this way, similar or identical images can be traced back to scientific or social communities, or influential agents in terms of distributing images across the web can be identified. In a way, this is a double-edged sword, since the provision of URLs reflects Google’s opaque ranking algorithm and the particular order in which the web content is delivered to users. Omena et al.’s method can be extended by an analysis of concrete and situated user interactions on the detected websites to better understand their digital sociability. It is also possible to combine the presented computer vision Figure 1.2 Example of Image-Web Entities. Nodes Reveal Which Contexts orKnowledge Graph Entities Are Most Frequently Associated With an Image in Question (Here a Mosquito) pink temple fashion accessory knitting teddy bear smile hug facial hair cap beanie hat sun hat glitter vision care eyewear snapshot interaction headgear sunglasses clown skin beauty cosmetics people day toddler sibling lady doll person hair accessory headpiece child cool fun play t shirt standing textile sleeve baby products clothing woman plaid family socialite lap polka dot furniture costume outdoor play equipment baby carriage shirt playground slide educational toy journalist daughter father infant mother boy man jewellery aggression headband black hair shoulder joint sitting peach photo shoot scarf bathing nightwear lace washing wrist elbow therapy footwear shoe patient jeans trousers ankle dress sports uniform gown boxing glove denim leggings blouse rave pocket bowling pin sweater bean bag pajamas girl muscle grandparent human body erotic literature knit cap bonnet outerwear �ooring linens bed sheet kindergarten �oor dance dress in�atable health care room sportswear bedding sock car seat fashion tartan dentist sneakers top baby �oat toy block couch club pillow nursery bowling ball jacket dress shirt ball tablecloth pumpkin halloween coat ballet tutu superman nap mat image-label (1) pink temple fashion accessory knitting teddy bear smile hug facial hair cap beanie hat sun hat glitter vision care eyewear snapshot interaction headgear sunglasses clown skin beauty cosmetics people day toddler sibling lady doll person hair accessory headpiece child cool fun play t shirt standing textile sleeve baby products clothing woman plaid family socialite lap polka dot furniture costume outdoor play equipment baby carriage shirt playground slide educational toy journalist daughter father infant mother boy man jewellery aggression headband black hair shoulder joint sitting peach photo shoot scarf bathing nightwear lace washing wrist elbow therapy footwear shoe patient jeans trousers ankle dress sports uniform gown boxing glove denim leggings blouse rave pocket bowling pin sweater bean bag pajamas girl muscle grandparent human body erotic literature knit cap bonnet outerwear �ooring linens bed sheet kindergarten �oor dance dress in�atable health care room sportswear bedding sock car seat fashion tartan dentist sneakers top baby �oat toy block couch club pillow nursery bowling ball jacket dress shirt ball tablecloth pumpkin halloween coat ballet tutu superman nap mat image-label (1)image-label (1) 2016 summer olympics abortion abortion law academic journal aerosol spray agãªncia brasil agãªncia nacional de saãºde suplementar aids all india pre medical test allergy anesthesia ankyloglossia aortic aneurysm artery arthropod aspm asymptomatic atherosclerosis austria autism speaks autism-europe avocado axon babycenter bachelor of medicine and bachelor of surgery bacteria band fm bee sting belã©m biology birth control birth defect biting bleeding blighted ovum blood blood test blood–brain barrier body ache boqueirã£o brain brain damage brain tumor bronchitis brumado bug zapper calci�cation campina grande cancer immunotherapy cantoria caraãºbas carbon–carbon bond cardiology cardiovascular disease carmen lãºcia carpal tunnel carãºpano castiglione della pescaia cearã¡ celina turchi cell cell division cerebrospinal �uid cerebrum charlotte denman lozier chemical bond chickenpox children's hospital chordata chromosome chromosome 3 citronella oil city of guaxupe civil police cleaning climate change clinical hospital teacher clã©riston andrade cochlear implant college of war collegium coluna community health center companhia de ãgua e esgoto da paraã-ba complication compounds of carbon congenital abnormality container continent cooking oil correios cri du chat cross-functional team curve cã¡ceres deformity delayed milestone dermatologist dermatology development of the nervous system diabetes mellitus diagnosis of hiv/aids diagnostic test diarrhea digg reader digital signal director disease disease management double bond dressing drug drug test drugstore early childhood intervention eastern region eduardo amorim egg cell el salvador electron microscope epileptic spasms erythema escola de formaã§ã£o de o�ciais da marinha mercante estado de minas evandro chagas evandro chagas institute exame nacional do ensino mã©dio extreme poverty eye injury fabio rodrigues pozzebom facipe - unit caxangã¡ family medicine february 8 federal district federal highway police federal institute of pernambuco federal university of bahia federal university of goiã¡s federal university of rio de janeiro federal university of rio grande do sul federal university of santa catarina federal university of sã£o paulo fetus �fth disease �rst aid kit �ushing foreign exchange market franco da rocha fred taylor frederick banting g1 genetic disorder gland glaucoma foundation global health guaxupã© gynaecology hayner headache health system hearing loss henry k. beecher herpes simplex histamine hiv honduras hong kong hospital das clã-nicas da universidade de sã£o paulo hydra hydrocortisone hypertension hypothesis icaridin igor padilla immune system incident infertility infestation in�ammation in�uenza vaccine infographic insect repellent integrative medicine internal medicine intrauterine growth restriction ipub - instituto de psquiatria da ufrj irritation itch jama january jaundice john langdon down joã£o �lgueiras lima juazeiro juca ferreira juscelino kubitschek kaiser permanente keratitis kerosene laboratory ld&b insurance and �nancial services legislative assembly of paraã-ba lesion life tenure lingual frenectomy macapã¡ magnetic resonance imaging major depressive disorder malnutrition manual therapy marcos espinal mata south campus mato grosso medical college medical glove medical technologist medicine meningitis meningococcal disease michael greger microbiology microcephaly microorganism microscope ministry of culture ministry of education ministry of health ministry of social development and �ght against hunger mogi das cruzes molecule mosquito net movember myiasis myofascial trigger point national council for scienti�c and technological development national electrical code national eligibility and entrance test national �re protection association national sanitary surveillance agency nature microbiology neonatal acne nervous system neuroglia neurological disorder neurosurgery niterã³i non-communicable disease nutrient obstetrics and gynaecology o�! oswaldo cruz oswaldo cruz foundation pain palmares pamphlet panama paranã¡ parietal lobe parkinson's disease patient pelourinho pericardial e�usion perimeter permethrin pertussis peter the great pharmaceutical drug pharmacy phenylketonuria phylum physical examination piedade plastic surgery polyclinic lessa de andrade polymicrogyria pre-eclampsia precedent pregnancy presbyterian college xv de novembro primary healthcare prognosis programa saãºde da famã-lia prostate cancer prosthesis psychiatry punta ala pyriproxyfen raul henry referral reform rehabilitation rehabilitation hospital reproductive health residency risk factor rna rodrigo rollemberg rsw software rubella rã¡dio cidade 99.1 saliva salivary gland santo domingo sarah network of rehabilitation hospitals sarvey insurance school of medicine scientist secondary education sertã£o sexual intercourse sexual partner signal sinusitis sistema ãšnico de saãºde spinal cord sprain state university of alagoas state university of feira de santana stem cell stevens rehen stylianos antonarakis sucre surgeon surubim swelling symmetry in biology syphilis teaching hospital teixeira de freitas tejuã§uoca televisiã³n nacional de chile test strategy time 100 tissue tobacco 21 toothpaste toxicology toxoplasmosis transatlantic trade and investment partnership translational research translational research institute trisomy tsgi inc. tvn ultrasound unicef australia united nations university centre of joã£o pessoa urinary tract infection urology vertically transmitted infection vigabatrin vitamin vitamin d vitã³ria da conquista vomiting water storage water tank week 30 of pregnancy wilson's disease women's health women's rights wound wound healing ya yolani batres zika virus ㉠o tchan! health image-web entities (2) cdninstagram.com content1.jdmagicbox.com dab.saude.gov.br pbs.twimg.com petitebox.com.br pm1.narvii.com saude-rioclaro.org.br scontent-yyz1-1.cdninstagram.com sobep.org.br u.saude.gov.br wotspeak.ru www.gammacenter.net www.hebiatriabatistela.com.br www.megaimagem.com.br www.nixorclinik.ru www.nursing.com.br www.rioclaro.sp.gov.br www.saude.rc.sp.gov.br image-domain (3) Janna Joceli omena elena PiliPets Beatrice GoBBo Jason chao The poTenTials of GooGle Vision api-based neTworks To sTudy naTiVely diGiTal imaGes diseña 19 auG 2021 arTiCle.1 9 Source: Omena et al., 2021, p. 9.
Seeing Like a Field? 327 SJS 51 (2), 2025, 317–336 networks with each other (e.g. web-entities and image-domains). The potential for sociological research is enormous: Not only can networks of images available on the internet be explored at scale – up to tens of thousands of images are thinkable –, but image-web entity networks allow for the analysis of the circulation of the recognised entities and can be blended into social networks, whereas image-domain networks open up the possibility to identify domain-specific social communities gathered around the use of the images which have been selected by the researchers analysing such networks. With respect to image-label networks, adependency on the data provided by the Google Image API has certainly to be noted. However, at least with respect to photographic material, alternative means to annotate images are currently available: Even without the labels attached to the images from the internet it is possible to first perform object detection, object localization, and object qualification by applying open source algorithms such as InceptionV3, adeep learning model capable to autonomously detect objects like persons, horses, or a car. In a second step it is possible to extract and represent semantic elements by establishing a code system which serves the needs of the researchers working on this material (Arnold & Tilton, 2019). The latter approach does not depend on the labels attributed by the users of the internet, but allows for a culturally and socially constructed code system that has to be established by the researchers according to their needs. Consequently, the analysis of the images themselves and the objects contained therein can be conducted according to the premise of a sociologist. In sum, the analysis of the three types of networks provides the basis for a transfer into social networks and thus into a modified field analysis in the Bourdieusian sense. Besides, the point of Figure 1.3 Example of an Image-Domain Network. Nodes Display Web Domains Which Host the Chosen Images pink temple fashion accessory knitting teddy bear smile hug facial hair cap beanie hat sun hat glitter vision care eyewear snapshot interaction headgear sunglasses clown skin beauty cosmetics people day toddler sibling lady doll person hair accessory headpiece child cool fun play t shirt standing textile sleeve baby products clothing woman plaid family socialite lap polka dot furniture costume outdoor play equipment baby carriage shirt playground slide educational toy journalist daughter father infant mother boy man jewellery aggression headband black hair shoulder joint sitting peach photo shoot scarf bathing nightwear lace washing wrist elbow therapy footwear shoe patient jeans trousers ankle dress sports uniform gown boxing glove denim leggings blouse rave pocket bowling pin sweater bean bag pajamas girl muscle grandparent human body erotic literature knit cap bonnet outerwear �ooring linens bed sheet kindergarten �oor dance dress in�atable health care room sportswear bedding sock car seat fashion tartan dentist sneakers top baby �oat toy block couch club pillow nursery bowling ball jacket dress shirt ball tablecloth pumpkin halloween coat ballet tutu superman nap mat image-label (1) pink temple fashion accessory knitting teddy bear smile hug facial hair cap beanie hat sun hat glitter vision care eyewear snapshot interaction headgear sunglasses clown skin beauty cosmetics people day toddler sibling lady doll person hair accessory headpiece child cool fun play t shirt standing textile sleeve baby products clothing woman plaid family socialite lap polka dot furniture costume outdoor play equipment baby carriage shirt playground slide educational toy journalist daughter father infant mother boy man jewellery aggression headband black hair shoulder joint sitting peach photo shoot scarf bathing nightwear lace washing wrist elbow therapy footwear shoe patient jeans trousers ankle dress sports uniform gown boxing glove denim leggings blouse rave pocket bowling pin sweater bean bag pajamas girl muscle grandparent human body erotic literature knit cap bonnet outerwear �ooring linens bed sheet kindergarten �oor dance dress in�atable health care room sportswear bedding sock car seat fashion tartan dentist sneakers top baby �oat toy block couch club pillow nursery bowling ball jacket dress shirt ball tablecloth pumpkin halloween coat ballet tutu superman nap mat image-label (1)image-label (1) 2016 summer olympics abortion abortion law academic journal aerosol spray agãªncia brasil agãªncia nacional de saãºde suplementar aids all india pre medical test allergy anesthesia ankyloglossia aortic aneurysm artery arthropod aspm asymptomatic atherosclerosis austria autism speaks autism-europe avocado axon babycenter bachelor of medicine and bachelor of surgery bacteria band fm bee sting belã©m biology birth control birth defect biting bleeding blighted ovum blood blood test blood–brain barrier body ache boqueirã£o brain brain damage brain tumor bronchitis brumado bug zapper calci�cation campina grande cancer immunotherapy cantoria caraãºbas carbon–carbon bond cardiology cardiovascular disease carmen lãºcia carpal tunnel carãºpano castiglione della pescaia cearã¡ celina turchi cell cell division cerebrospinal �uid cerebrum charlotte denman lozier chemical bond chickenpox children's hospital chordata chromosome chromosome 3 citronella oil city of guaxupe civil police cleaning climate change clinical hospital teacher clã©riston andrade cochlear implant college of war collegium coluna community health center companhia de ãgua e esgoto da paraã-ba complication compounds of carbon congenital abnormality container continent cooking oil correios cri du chat cross-functional team curve cã¡ceres deformity delayed milestone dermatologist dermatology development of the nervous system diabetes mellitus diagnosis of hiv/aids diagnostic test diarrhea digg reader digital signal director disease disease management double bond dressing drug drug test drugstore early childhood intervention eastern region eduardo amorim egg cell el salvador electron microscope epileptic spasms erythema escola de formaã§ã£o de o�ciais da marinha mercante estado de minas evandro chagas evandro chagas institute exame nacional do ensino mã©dio extreme poverty eye injury fabio rodrigues pozzebom facipe - unit caxangã¡ family medicine february 8 federal district federal highway police federal institute of pernambuco federal university of bahia federal university of goiã¡s federal university of rio de janeiro federal university of rio grande do sul federal university of santa catarina federal university of sã£o paulo fetus �fth disease �rst aid kit �ushing foreign exchange market franco da rocha fred taylor frederick banting g1 genetic disorder gland glaucoma foundation global health guaxupã© gynaecology hayner headache health system hearing loss henry k. beecher herpes simplex histamine hiv honduras hong kong hospital das clã-nicas da universidade de sã£o paulo hydra hydrocortisone hypertension hypothesis icaridin igor padilla immune system incident infertility infestation in�ammation in�uenza vaccine infographic insect repellent integrative medicine internal medicine intrauterine growth restriction ipub - instituto de psquiatria da ufrj irritation itch jama january jaundice john langdon down joã£o �lgueiras lima juazeiro juca ferreira juscelino kubitschek kaiser permanente keratitis kerosene laboratory ld&b insurance and �nancial services legislative assembly of paraã-ba lesion life tenure lingual frenectomy macapã¡ magnetic resonance imaging major depressive disorder malnutrition manual therapy marcos espinal mata south campus mato grosso medical college medical glove medical technologist medicine meningitis meningococcal disease michael greger microbiology microcephaly microorganism microscope ministry of culture ministry of education ministry of health ministry of social development and �ght against hunger mogi das cruzes molecule mosquito net movember myiasis myofascial trigger point national council for scienti�c and technological development national electrical code national eligibility and entrance test national �re protection association national sanitary surveillance agency nature microbiology neonatal acne nervous system neuroglia neurological disorder neurosurgery niterã³i non-communicable disease nutrient obstetrics and gynaecology o�! oswaldo cruz oswaldo cruz foundation pain palmares pamphlet panama paranã¡ parietal lobe parkinson's disease patient pelourinho pericardial e�usion perimeter permethrin pertussis peter the great pharmaceutical drug pharmacy phenylketonuria phylum physical examination piedade plastic surgery polyclinic lessa de andrade polymicrogyria pre-eclampsia precedent pregnancy presbyterian college xv de novembro primary healthcare prognosis programa saãºde da famã-lia prostate cancer prosthesis psychiatry punta ala pyriproxyfen raul henry referral reform rehabilitation rehabilitation hospital reproductive health residency risk factor rna rodrigo rollemberg rsw software rubella rã¡dio cidade 99.1 saliva salivary gland santo domingo sarah network of rehabilitation hospitals sarvey insurance school of medicine scientist secondary education sertã£o sexual intercourse sexual partner signal sinusitis sistema ãšnico de saãºde spinal cord sprain state university of alagoas state university of feira de santana stem cell stevens rehen stylianos antonarakis sucre surgeon surubim swelling symmetry in biology syphilis teaching hospital teixeira de freitas tejuã§uoca televisiã³n nacional de chile test strategy time 100 tissue tobacco 21 toothpaste toxicology toxoplasmosis transatlantic trade and investment partnership translational research translational research institute trisomy tsgi inc. tvn ultrasound unicef australia united nations university centre of joã£o pessoa urinary tract infection urology vertically transmitted infection vigabatrin vitamin vitamin d vitã³ria da conquista vomiting water storage water tank week 30 of pregnancy wilson's disease women's health women's rights wound wound healing ya yolani batres zika virus ㉠o tchan! health image-web entities (2) cdninstagram.com content1.jdmagicbox.com dab.saude.gov.br pbs.twimg.com petitebox.com.br pm1.narvii.com saude-rioclaro.org.br scontent-yyz1-1.cdninstagram.com sobep.org.br u.saude.gov.br wotspeak.ru www.gammacenter.net www.hebiatriabatistela.com.br www.megaimagem.com.br www.nixorclinik.ru www.nursing.com.br www.rioclaro.sp.gov.br www.saude.rc.sp.gov.br image-domain (3) Janna Joceli omena elena PiliPets Beatrice GoBBo Jason chao The poTenTials of GooGle Vision api-based neTworks To sTudy naTiVely diGiTal imaGes diseña 19 auG 2021 arTiCle.1 9 Source: Omena et al., 2021, p. 9.
328 Katrin Herms and Jörg Lehmann SJS 51 (2), 2025, 317–336 independent data collection with regard to photographic material should not be underestimated: While the knowledge necessary to use Google’s Image API has been extensively documented (Omena & Currie, 2022a; 2022b) and the methodology to construct the three types of networks has been thoroughly described (Omena et al., 2021), data collection and therefore all subsequent steps depend on Google’s Knowledge Graph and/or on Google’s ranking systems and search capabilities: “researchers must understand that what they are seeing includes the layered structure of online connectivity through Google’s eyes” (Omena et al., 2021, p. 20). However, currently available sophisticated methodologies using existing machine learning applications facilitate data collection independent from data provision by a big tech company wherever this deems suitable to the researcher (Smits & Wevers 2023). 3.2 Image Similarity Analysis on Digitised Cultural Heritage Images The second example presented here does not necessarily rely on images to be found on the internet, even though a sociological perspective on contemporary societies might be a priority. If the emphasis of analysis is more on a historical perspective and focuses on longitudinal evolution of e.g., developments in the field of art, digitised visual cultural heritage assets might be preferred. Similar to the approach described in the section above, such material can equally be analysed with computer vision methods; however, the methodological approach is quite different. Digital cultural heritage datasets usually contain rich metadata; in the case of images from art history these metadata contain information on the artist, the year when the work of art was created, the place of creation or publication as well as further information like the artistic movement to which the artist belonged or stylistic or technical features characterising this specific work of art. In computer vision, image-text combinations are termed multimodal and allow for the deployment of deep learning models (Smits & Wevers, 2023). Such an image dataset can be analysed on apixel level using aconvolutional neural network. The computation of the location of each image in ahigh-dimensional vector space is subsequently possible, since the number of pixels with identical or similar colour intensities as well as their characteristic combinations within an image allows for the transformation of the image files into vectors consisting of numbers. This calculation also provides the basis for the determination of image similarities and the application of clustering algorithms (cf.on the logic of metapictures Hoggenmüller & Klinke in this Special Issue). While this approach resembles the sociological grouping of individuals on the basis of features characterising them, as it may be employed for example in lifestyle analysis, the methodology of creating a high-dimensional vector space for big data analysis has been developed in the field of natural language processing (NLP), where the number of word occurrences in a corpus of documents form the strictly quantitative basis for arranging
Seeing Like a Field? 329 SJS 51 (2), 2025, 317–336 each document in a vector space (Turney & Pantel, 2010). With the advent of neural networks, this method was implemented in computer vision applications. It not only has the benefit of computationally determining the position of each image in the high-dimensional vector space, but also in a visual space in which the clustered images can be explored as if they were located on a twoor three-dimensional Cartesian coordinate system. To achieve this, the numbers of each vector are computed using a technique of dimensionality reduction – from a high-dimensional vector space into a two or three dimensional one – using an algorithm called UMAP (McInnes et al., 2018; specifically for t-SNE see again Hoggenmüller & Klinke in this Special Issue). PixPlot is one of those computer vision applications. It uses the Inception deep neural network for analysis of image similarities, clustering and visualisation of large image datasets. It has been developed by the Yale Digital Humanities Lab. If metadata associated with an image are available, it is possible to display them in an interactive visualisation. A couple of research projects, especially from the digital humanities, have employed PixPlot in order to visualise, explore, and analyse large visual datasets (see Frischknecht in this Special Issue as an example). Furthermore, several cultural heritage institutions have used the technology developed by Yale (see e.g. the implementations by the National Library of Denmark https://labs.statsbiblioteket. dk/pixplot/ as well as the National Gallery of Denmark https://pixplot.smk.dk/). The Surprise Machines project has visualised more than 200 000 images from the Harvard Art Museums usually inaccessible to visitors of the museum. While each of these digital objects are publicly accessible online, a visualisation of the extensive image collection was presented as an exhibition on the premises of the museum (Rodighiero et al., 2022). The title of the exhibition and of the research project is revealing: Surprise Machines refers to the often unexpected and unpredictable outcomes of machine learning applications, thereby generating surprises in the individuals exploring the visualisation of the vector space model. PixPlot has also been used to provide a panoramic overview of more than 1 000 digital images which have been created as illustrations of Dante Alighieri’s Divina Comedia since the publication of this narrative poem in the 13th Century. Images are arranged in clusters based on their similarities and presented alongside the Comedy’s structure and plot; metadata can as well be browsed by the users on the project’s website at https://divinecomedy. digital/#/ (Bonera & Bardazzi, 2022). While these projects refer to similar ‘clustering’ approaches in art history like Aby Warburg’s Atlas Mnemosyne (Warburg, 2020), the difference between the approach of an art historian with her rich contextual knowledge and the computational determination of image similarities is obvious. In contrast, the Poscatálogo (or Postcatalog in English) project took a different approach. The starting point here was Alfred H. Barr’s famous diagram representing the most important artistic movements of the first decades of the 20th Century,
330 Katrin Herms and Jörg Lehmann SJS 51 (2), 2025, 317–336 which charts formal influences and thus interprets the evolution and genealogy of art of modernity. This diagram was transformed into a high-dimensional vector space using the metadata accompanying the image dataset, especially the dates when the works of art were created. In a second step, an Inception convolutional neural network (CNN) was used to group the dataset of about 2 000 images into clusters and map them onto the diagram (Rodríguez Ortega et al., 2021). The resulting visualisation combines hand-curated metadata and computed clusters, thus allowing for an immersive exploration of the grouped images and their various relationships. The two-dimensional display of these grouped images enables an analysis of the development of the field over time and supports – in sociological terms – the exploration of the tensions between and sequence of “orthodox” and “heretical” artists (Bourdieu, 1979, p. 257). However, the exploration of structuring oppositions – such as cubism vs. surrealism – does not unfold in a straightforward way, because the images are clustered on the basis of their similarity and not according to their affiliation to a modernist movement, the painting style of which might have changed over time. This insight is put in a nutshell by using the term Poscatálogo– it implies the detachment from the traditional, man-made cataloguing systems and the interpretations provided by researchers. By contrast, it inaugurates visible and physically perceptible approaches to large image datasets grouped according to their computed similarities. Figure 2.1 Barr X Inception CNN Visualisation of Clustered Image Data in a Vector Space Model, Which Can Be Understood as a Field inBourdieusian Terms, Here a Full View of the Visualisation. Source: Available online at https://digital-narratives.versae.es/ [Screenshot, Herms and Lehmann].
Seeing Like a Field? 331 SJS 51 (2), 2025, 317–336 Both figures illustrate the potential of image similarity approaches for arranging large amounts of images from art history, if they are combined with metadata. In the first figure, the elder modernist movements (e.g. impressionism) can be found on the left, while younger movements (e.g. abstract expressionism) are to be found on the right. The second figure shows how the clustering algorithm works, since it arranges similar images into one group (e.g. cubist paintings). However, not all cubist paintings are accumulated, since motifs and the colours used may vary. Compared to traditional cultural sociological approaches, the potential of image similarity analyses to scale up from a narrow, qualitative approach to big visual data is particularly evident here. They could be further elaborated if the neural networks would be trained on traits like brushwork, stroke weight, composition or else. Afurther extension would be realised by including the likes or other kind of feedback which the used images received by retrieving such attention economy data via the Pinterest REST API, thus providing the clustering algorithm with further data. At the same time and as it is often the case, the digital methodologies used in the named projects bring man-made epistemologies as cultural constructs and traditional representations of art history to the surface. In this way, they question Figure 2.2 The Same Visualisation as in Figure 2.1, With a Zoom into One of the Clusters, in This Case Cubist Paintings. Source: Available online at https://digital-narratives.versae.es/ [Screenshot, Herms and Lehmann].
332 Katrin Herms and Jörg Lehmann SJS 51 (2), 2025, 317–336 customary classification systems and genealogical narratives, conceptions of creativity, originality, and influence with a long pedigree in art history, or the determination of similarity and difference as they have been established and systematically elaborated by art historians. The digital turn in art history might therefore unfold its disruptive potential in a way which may be perceived as unsettling by scientific researchers – the promise of advancing human knowledge by using computational approaches might be accompanied by the shattering of long-established art historical or sociological methodologies, a collateral damage that is not always welcomed if it contests the foundation of art history as a discipline (Rodríguez Ortega, 2019). However, the example of Poscatálogo shows that the clustering of similar images presents a productive provocation to art history oriented towards Bourdieu, because it would have to clearly identify what exactly “orthodoxy” and “heresy” (Bourdieu, 1979, p. 257) mean in terms of content. The metadata collected by cultural heritage institutions use categories that were established beforehand; it would be possible to group the images according to these descriptive labels, for example to visually arrange Cubism as opposed to Surrealism. Meanwhile, this would only result in avisualisation of an interpretation established beforehand. Image similarity clustering, by contrast, reveals new visual contexts, challenges inherited interpretations and concretises the visual rules established in the field. It thus supports the analysis of big visual data as a field in the Bourdieusian sense. 4 Conclusion The two presented examples have underlined the significant potential of analytic possibilities for sociological analyses resting on big visual data. The first case study exemplifies how image-label networks can be created, answering the question What does the users’ visual attention focus on? With respect to image-web entity networks, it has to be taken into account that those entities are added by Google and linked to the Google Knowledge Graph. As such, they enable the identification of various perspectives onto related topics, revealing image-related debates from a Google Knowledge Graph perspective. Image-domain networks, finally, enable the exploration of the socio-technical origin of the circulating images, sometimes pointing to domain-specific audiences gathered around the use of specific images, thus answering the question Which social groups participate in the visual debate? This methodology can be understood as an alternative to the contingency tables used by Bourdieu to analyse aesthetic preferences: The size of nodes in all three types of presented computer vision networks reflects different levels of image visibility, be it through image-labels, web-entities or image-domains. Visual attention can thus be interpreted as a resource that makes circulating images gain influence. Therefore, we propose that attention economy may be used as a helpful concept to interpret computer vision networks from a Bourdieusian perspective.
Seeing Like a Field? 333 SJS 51 (2), 2025, 317–336 The second case study does not take images drawn from the web into focus, but image collections from cultural heritage institutions alongside with metadata curated by cultural heritage practitioners. The pixelwise analysis of the image allows, alongside with the available metadata, a clustering and visual arrangement of large amounts of images (tens or hundreds of thousands). Subsequently, an analysis in terms of content of the image clusters and the resulting visual arrangement can be performed, thus identifying the position of the groups of images in a virtual space. Coming back to the main question of our contribution – is it possible for sociology resting on big visual data to See like a field?, we highlight that the methodologies presented above need to be complemented by qualitative, contextualised, and interpretative approaches. In this way, both examples open the door widely for the analysis of big visual data as a field in a Bourdieusian sense: The network analyses do not only enable interpretation about what these visual data are about, but also which social groups are engaged in the visual debates. Furthermore, if available attention economy data such as likes, emojis, or reposts are integrated as well, aesthetic preferences of the users can be analysed. Cultural styles around circulating images will then materialise in the process of analysis. Image similarity clustering draws the attention of the researcher to the content of the images themselves, thereby enabling an interpretation of who would constitute orthodox and heretic groups of image creators in the given image sample, as well as fostering an interpretation of what the visual exchange is really about. Taken together, these methodological approaches allow to a large extent for an interpretation of big visual data as Bourdieusian fields. However, there are limitations to what can analytically be achieved with an approach that tries to see like a field. These limitations result from the notable gaps between researchers and big tech companies in terms of data access, ethical standards and strategic goals of big data analysis. While companies long for prediction of consumers’ behaviour, researchers analyse trace data with the aim to better understand social dynamics in a platform context. By contrast, the asset of behavioural data and derived analytic products like identified patterns of aesthetic consumption and production as well as preferences of taste and social relationships are what enable big tech companies to match user dispositions and advertisements and thus to generate profits out of data analysis. This process of “commodifying people’s behaviours” (Fourcade & Healy, 2017, p. 16) has been analysed by Shoshana Zuboff (2019) in her study The Age of Surveillance Capitalism. The economist analyses digital capitalism as a novel market form and develops concepts like “behavioural surplus” (Zuboff, 2019, pp. 63–97), the “epistemic coup” (i.e. the claim of ownership of knowledge in society by tech corporations, Zuboff, 2019, pp. 176–195, 495–595) and “instrumentarian power” (Zuboff, 2019, pp. 351–444). The first one is relevant in the context discussed here. As Zuboff explains, the identification of behavioural patterns allows for predictive analytics: “These machine intelligence operations convert raw material into the firm’s highly profitable algorithmic products designed to predict the behaviour of its users” (Zuboff, 2019, p. 65). In combination with
334 Katrin Herms and Jörg Lehmann SJS 51 (2), 2025, 317–336 the trace data on language, cultural assets, traditions, customs, and preferences in terms of taste available only to big tech companies, such analytic products enable even the prediction of cultural production and consumption. This is a central point in Zuboff’s argumentative framework, since it is exactly here that surveillance capitalism transcends the market logic as it was conceived so far, namely as an opaque mechanism where supply and demand are balanced: “Surveillance capitalism thus replaces mystery with certainty as it substitutes rendition, behavioural modification, and prediction for the old ‘unsurveyable pattern’. This is a fundamental reversal of the classic ideal of the ‘market’ as intrinsically unknowable” (Zuboff, 2019, p. 497). The availability of complete trace data at the hands of big tech corporations thus has significant consequences for the division of knowledge production especially in Western societies. Predictive analytics fills the blank in Bourdieusian field theory which we marked above, but it cannot be performed on the basis of attention economy data alone. If the power of disposition over complete trace data is decisive for the analysis of big visual data, sociological research endeavours are to be found on the dominated side of the division of knowledge production that has taken place since the advent of the Internet. For scientific research, it may be extremely attractive to explore cultural creativity as patterns of behaviour and to predict cultural production, but it has to be admitted that those analytic products rely on trace data which are beyond the reach of scientific knowledge. Big tech corporations will not make such data available since such patterns of consumption are eminently exploitable and thus serve the maximisation of profit they are looking for. As long as researchers don’t have full access to trace data, it will not be possible for sociology to see like afield, to the same extent as platform companies see like a market. 5 References Ahnert, R., Ahnert, S. E., Coleman, C. N., & Weingart, S. B. (2020). The Network Turn: Changing Perspectives in the Humanities (1st ed.). Cambridge University Press. https://doi.org/10.1017/9781108866804 Alton, Liz (2024, April 4). 7 tips for creating engaging content every day. https://web.archive.org/ web/20240404010929/https://business.twitter.com/en/blog/7-tips-creating-engaging-contentevery-day.html Arnold, T., & Tilton, L. (2019). Distant viewing: Analyzing large visual corpora. Digital Scholarship in the Humanities, 34(Supplement_1), i3–i16. https://doi.org/10.1093/llc/fqz013 Barabási, A.-L. (2009). Linked. How everything is connected to everything else and what it means for business, science, and everyday life ([18. edition] Plume print; authorized reprint of a hardcover ed. publ. by Perseus Publ.). Plume. Bernhard, S. (2008). Netzwerkanalyse und Feldtheorie. Grundriss einer Integration im Rahmen von Bourdieus Sozialtheorie. In C. Stegbauer (Ed.), Netzwerkanalyse und Netzwerktheorie: Ein neues Paradigma in den Sozialwissenschaften (pp. 121–130). VS Verlag für Sozialwissenschaften. https:// doi.org/10.1007/978-3-531-91107-6_8 Bonera, M., & Bardazzi, A. (2022). Data Visualization as a Tool to Experience the Legacy of Dante’s Divine Comedy and its Influence on the Cultural Heritage. Bibliotheca Dantesca: Journal of Dante Studies, 5(1), 288–297.
Seeing Like a Field? 335 SJS 51 (2), 2025, 317–336 Bouk, D. (2017). The History and Political Economy of Personal Data over the Last Two Centuries in Three Acts. Osiris, 32(1), 85–106. https://doi.org/10.1086/693400 Bourdieu, P. (1971). Une interprétation de la théorie de la religion selon Max Weber. European Journal of Sociology / Archives Européennes de Sociologie / Europäisches Archiv Für Soziologie, 12(1), 3–21. Bourdieu, P. (1979). La distinction. Critique sociale du jugement. Les Éd. de Minuit. Bourdieu, P. (1992). Les Règles de l’Art. Genèse et Structure du Champ Littéraire. Éd. du Seuil. Centola, D. (2011). An Experimental Study of Homophily in the Adoption of Health Behavior. Science, 334(6060), 1269–1272. https://doi.org/10.1126/science.1207055 Cha, M., Haddadi, H., Benevenuto, F., & Gummadi, K. (2010). Measuring User Influence in Twitter: The Million Follower Fallacy. Proceedings of the International AAAI Conference on Web and Social Media, 4(1), 10–17. https://doi.org/10.1609/icwsm.v4i1.14033 de Nooy, W. (2003). Fields and networks: Correspondence analysis and social network analysis in the framework of field theory. Poetics, 31(5), 305–327. https://doi.org/10.1016/S0304-422X(03)00035-4 Demandsage Instagram Statistics. (2024). https://www.demandsage.com/instagram-statistics/ Escofier, J.-P. (2024). Petite histoire des mathématiques. Dunod. Fourcade, M., & Healy, K. (2017). Seeing like a market. Socio-Economic Review, 15(1), 9–29. https:// doi.org/10.1093/ser/mww033 Fraiberger, S. P., Sinatra, R., Resch, M., Riedl, C., & Barabási, A.-L. (2018). Quantifying reputation and success in art. Science, 362(6416), 825–829. https://doi.org/10.1126/science.aau7224 Franck, G. (1998). Ökonomie der Aufmerksamkeit: Ein Entwurf. Hanser. Frischknecht, M. (2025). Through the Eyes of the Machine: Exploring Historical Photo Collections with Convolutional Neural Networks. Schweizerische Zeitschrift für Soziologie, 51(2), Special Issue hrsg. von S. W. Hoggenmüller, Big Visual Data als neue Form des Wissens: Potenziale, Herausforderungen und Transformationen. Hennig, M., & Kohl, S. (2012). Fundierung der Netzwerkperspektive durch die Habitusund Feldtheorie von Pierre Bourdieu. In M. Hennig & C. Stegbauer (Eds.), Die Integration von Theorie und Methode in der Netzwerkforschung (pp. 13–32). VS Verlag für Sozialwissenschaften. https:// doi.org/10.1007/978-3-531-93464-8_2 Hoggenmüller, S. W., Klinke, H. (2025). Metabilder als Forschungswerkzeuge: Zur Kontingenz und algorithmischen Bedingtheit ihrer Herstellung. Schweizerische Zeitschrift für Soziologie, 51(2), Special Issue hrsg. von S. W. Hoggenmüller, Big Visual Data als neue Form des Wissens: Potenziale, Herausforderungen und Transformationen. Janning, F. (1991). Pierre Bourdieus Theorie der Praxis. Analyse und Kritik der konzeptionellen Grundlegung einer praxeologischen Soziologie. (Vol. 105). Westdeutscher Verlag. Ma, X., & Fan, X. (2022). A review of the studies on social media images from the perspective of information interaction. Data and Information Management, 6(1), 100004. https://doi.org/10.1016/j. dim.2022.100004 McInnes, L., Healy, J., Saul, N., & Großberger, L. (2018). UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software, 3(29), 861. https://doi.org/10.21105/joss.00861 Omena, J. J., & Currie, M. (2022a). Collecting Data Using APIs Part 1: How to Understand APIs and Navigating API Documentation [How-to Guide]. SAGE Research Methods: Doing Research Online. https://doi.org/10.4135/9781529611441 Omena, J. J., & Currie, M. (2022b). Collecting Data Using APIs Part 2: How to Communicate With and Access APIs [How-to Guide]. SAGE Research Methods: Doing Research Online. https://doi. org/10.4135/9781529611458 Omena, J. J., Elena, P., Gobbo, B., & Jason, C. (2021). The Potentials of Google Vision API-based Networks to Study Natively Digital Images. Diseña, 19, Article 1, 1–25. https://doi.org/10.7764/ disena.19.Article.1
342 Sebastian W. Hoggenmüller und Harald Klinke SJS 51 (2), 2025, 337–359 nante Farbe, Kontrast) andererseits können die Bilddaten sodann auf der Suche nach Ordnungsmustern variabel sortiert werden, was potenziell zu einer Vielzahl und Vielfalt an Metabildern führen kann.6 Ausschlaggebend für die Erstellung des hier gezeigten Beispiels war das spezifische Interesse an Kamerabildern aus Asien in einer frühen Phase des Forschungsprozesses, insbesondere die Frage, inwiefern diese Bilder Praktiken der Überwachung zeigen und welche Einblicke sie in Machtverhältnisse, Zensur und soziale Kontrolle eröffnen. Zudem war die Intention ausschlaggebend, durch ein kreisförmiges, nach Farbton und Helligkeit sortiertes Metabild relevante Informationen und Bedeutungsstrukturen in Bezug auf die anfänglich fokussierten Forschungsfragen deutlich sichtbar machen zu können. Dabei wurde unter anderem vermutet, dass die bereits auf Netzkamerabildern aus anderen Kontinenten beobachtete visuelle Praxis der Verfremdung von Bildsegmenten durch Abdeckungen, Unschärfen oder Verpixelungen – etwa in Form von Balken, Weichzeichnungen oder groben Pixeln (vgl. Caviezel, 2015, S. 195 ff.)– besonders effizient identifiziert werden könnte, um einerseits deren konkrete Umsetzung in Asien zu erkunden und andererseits bestehende Forschungsfragen zu schärfen sowie neue zu generieren. Die interaktive Bedienbarkeit des exemplarischen Metabilds wiederum unterstützt sowohl den Explorationsprozess nach relevanten Informationen und Bedeutungsstrukturen als auch die Entwicklung und Präzisierung von Forschungsfragen, indem sie einen fließenden Übergang von einem Gesamtüberblick über das visuelle Datenkorpus, wie in Abbildung 1 dargestellt, bis hin zu detaillierten Einblicken in spezifische Ausschnitte einzelner Bilddaten ermöglicht, wie in Abbildung 2 exemplarisch anhand von Screenshots einer Zoom-Bewegung illustriert. Auf diese Weise bietet das Beispielbild nicht nur eine umfassende Übersicht über alle 40 000 Bilder, sondern zeigt auch tiefere Strukturen, wie etwa eine Anomalie im Regenbogen-Farbverlauf: eine unterbrochene, gelb leuchtende Linie, die sich vom Mittelpunkt des Kreises bis zum Rand erstreckt. Zugleich ermöglicht es die gezielte Fokussierung auf einzelne Netzkamerabilder, was beim Hineinzoomen auch die besagte Anomalie näher erklärt: Die gelbe Linie resultiert aus identischen Bildern, auf denen schwarze chinesische Schriftzeichen auf gelbem Grund den Hinweis „Maschine wird gewartet“ zeigen. Im Gegensatz zur Nutzung des Metabildes im Kontext des transdisziplinären Forschungsprojekts, wo es vor allem dazu dient, mit Blick auf das zentrale Forschungsinteresse an der Beobachtung der Welt via Netzkameras soziologisch relevante und/ oder künstlerisch vielversprechende Strukturen in den 40 000 Netzkamerabildern zu identifizieren – sei es durch die Fokussierung auf die (Un-)Sichtbarkeit sozialer Situationen, Beziehungen und Räume, sei es durch den Vergleich mit Metabildern, die Bilder anderer Kontinente oder Zeitpunkte zeigen, um regionale Unterschiede und zeitliche Entwicklungen zu analysieren, oder sei es durch die Bewertung der 6 Diese Vielzahl und Vielfalt wird im Rahmen des Forschungsprojekts zusätzlich dadurch erweitert, dass entsprechende Metabilder auch für Netzkameras von anderen Kontinenten oder geografischen Standorten erstellt werden können.
Metabilder als Forschungswerkzeuge … 343 SJS 51 (2), 2025, 337–359 Potenziale und Grenzen des Metabildes als analytisches Werkzeug einschließlich technischer Fehler und ethischer Herausforderungen–, möchten wir im vorliegenden Beitrag anhand des exemplarischen Metabildes verdeutlichen, dass das, was wir als Metabild wahrnehmen, auf komplexen statistischen Berechnungen, algorithmischen Verfahren, technischen Bedingungen und einer Vielzahl von Entscheidungen basiert, die grundsätzlich kontingent und veränderbar sind. Und speziell an dieser Stelle: Es geht uns darum, aufzuzeigen, dass Forschende im Herstellungsprozess eines Abbildung 2 Exemplarisches Hineinzoomen in das Metabild Quelle: Eigene Visualisierung [Grabner und Hoggenmüller].
344 Sebastian W. Hoggenmüller und Harald Klinke SJS 51 (2), 2025, 337–359 Metabildes an verschiedenen Stellen Wahlmöglichkeiten haben und damit Entscheidungszwängen unterliegen, die die Datenvisualisierung entscheidend beeinflussen. In unserem Beispiel betrifft diese Kontingenz etwa die Auswahl der Bilddaten (hier eine zufällige Stichprobe von 40 000 in Asien an einem Tag produzierten Bildern), die Wahl der Analysemodelle (eine Kombination aus drei Technologien), die Festlegung der Sortierkriterien (radiale Anordnung nach Farbton und Helligkeit), die Form der visuellen Darstellung (kreisförmig und zweidimensional), die Einbindung von Metadaten (Kamerastandort) sowie die Interaktivität und Benutzer*innenführung (Schwenkund Zoomfunktion). Es ist leicht vorstellbar, dass eine andere Auswahl der Bilddaten (z. B. eine Variation der Zeitspanne oder eine gezielte inhaltliche Fokussierung auf Naturbilder oder Porträts), eine abweichende Wahl der Analysemodelle (z. B.Objektdetektion oder Szenenklassifikation), eine alternative Festlegung der Sortierkriterien (z. B. allein nach Bildinhalt oder allein nach Metadaten), eine veränderte visuelle Darstellung (z. B. als Clusteroder Streudiagramm), der Einbezug anderer Metadaten (wie Zeitstempel oder Aufnahmewinkel) sowie variierende Interaktionsmöglichkeiten und Benutzer*innenführungen (z. B. Drehfunktion oder Ebenenverwaltung) zu grundlegend anderen Metabildern geführt hätten. Dadurch wäre maßgeblich beeinflusst worden, welche Muster sichtbar und welche Bedeutungsebenen hervorgehoben werden und wie flexibel die Bilddaten exploriert werden können. All diese vorausliegenden Wahlmöglichkeiten und Entscheidungen bleiben in der Regel in Black Boxes verborgen – zumindest werden sie bei der Beschreibung von Forschungsergebnissen nur selten thematisiert und noch seltener kritisch hinterfragt –, sodass Metabilder trotz ihres kontingenten und entscheidungsabhängigen Herstellungsprozesses meist als evidente und objektive Darstellungen präsentiert werden. Dies ist umso bemerkenswerter, wenn man bedenkt, dass Metabilder nicht einfach bestehende Zusammenhänge abbilden, sondern Bedeutungsstrukturen wie etwa wiederkehrende Muster, verwandte Cluster oder isolierte Ausreißer in großen digitalen visuellen Datenbeständen überhaupt erst hervorbringen. Um diese Hervorbringung in ihrer Eigenheit genauer zu verstehen und darauf aufbauend den epistemischen Gehalt von Metabildern angemessen bewerten zu können, ist ein tieferes Verständnis der zugrunde liegenden Prozesse, Algorithmen, Standards und Praktiken sowie ihres Zusammenspiels notwendig. Anders formuliert: Wenn man Metabilder als Forschungswerkzeuge zur Erkenntnisgewinnung nutzen möchte, ist es notwendig, die Black Box der Metabilder zu öffnen, um besser zu verstehen, was Metabilder zeigen, was sie verbergen und was sie glauben machen. Im folgenden Abschnitt werden wir dies speziell mit Blick auf einen Aspekt tun, der in der Forschungspraxis oft schwer fassbar ist, weil er sich der Wahrnehmung in ganz besonderer Weise entzieht: die algorithmische Bedingtheit von Metabildern.7 7 Ergänzend dazu wird derzeit eine ethnografische Untersuchung vorbereitet, mit der im Kontext des Projekts Watching the World das kommunikative Handeln und dessen soziotechnische Bedingtheit bei der Herstellung, Verwendung und Interpretation von Metabildern in den Fokus genommen werden soll.
Metabilder als Forschungswerkzeuge … 345 SJS 51 (2), 2025, 337–359 3 Einblicke in die Black Box – die Rolle von Algorithmen Bei der Beschreibung der algorithmischen Bedingtheit von Metabildern liegt unser Schwerpunkt weniger auf den mathematischen Grundlagen der Algorithmen. Vielmehr möchten wir einen Einblick in die Pipeline8 der Datenanalyse geben und die algorithmische Analyse visueller Ähnlichkeiten näher beleuchten: Nach welchen Logiken der Automation kommen Metabilder zustande? Welche epistemologischen Implikationen ergeben sich daraus? Oder konkret am Beispiel des im Abschnitt zuvor gezeigten Metabildes gefragt: Was bedeutet visuelle Ähnlichkeit in der computergestützten und algorithmusbasierten Datenvisualisierung der 40 000 Netzkamerabilder? Beim Erkunden des Metabildes erkennt man beispielsweise Cluster von Kamerabildern mit geringem Kontrast, während kontrastreiche Kamerabilder an anderer Stelle zu finden sind. Doch bei genauerer Betrachtung drängt sich die Frage auf, was die räumlich nahe beieinander liegenden Bilder tatsächlich gemeinsam haben, und umgekehrt, warum ähnlich erscheinende Bilder mitunter weit voneinander entfernt angeordnet sind. Diese Fragen leiten unsere Ausführungen zur Rolle von Algorithmen bei der Herstellung von Metabildern an, die wir anhand des mathematisch-statistischen Konzepts des Merkmalsraums (Abschnitt3.1) und des Verfahrens der Dimensionsreduktion (Abschnitt3.2) erläutern. 3.1 Merkmalsraum (Feature Space) Zur Herstellung von Metabildern müssen Algorithmen zunächst sogenannte Features aus dem zugrunde liegenden Datensatz extrahieren. Im Kontext der Bildverarbeitung sind Features messbare Merkmale, die wesentliche Eigenschaften der Bildobjekte erfassen. Ziel dieses Prozesses ist es, die wichtigsten Informationen abstrahiert und komprimiert darzustellen, um die Datenmenge zu reduzieren und die Komplexität zu verringern. Dies schafft eine handhabbare Grundlage für die weitere Analyse der Daten (vgl. Bishop, 2006). Für die Merkmalsextraktion (Feature Extraction) gibt es verschiedene Ansätze (vgl. als Überblick Balan P & Sunny, 2018). Diese können erstens auf handcodierte, das heißt vom Menschen definierte Features (klassische Bildfeatures) fokussieren, wobei grundsätzlich zwischen Features auf niedriger und auf hoher Abstraktionsebene unterschieden wird, auch bekannt als Low-Level-Features und High-Level-Features. Low-Level-Features umfassen grundlegende visuelle Eigenschaften, die direkte, pixelnahe Informationen wie Farben oder Kanten repräsentieren. Kanten bezeichnen dabei Stellen im Bild, an denen sich die Helligkeit oder Farbe stark ändert – sie markieren oft die Grenzen von Objekten (Bildelementen). Bei High-Level-Features 8 Der Begriff Pipeline bezeichnet strukturierte Arbeitsabläufe, die Daten durch eine festgelegte Abfolge von Schritten verarbeiten. Diese Schritte, die oft automatisiert ablaufen, zielen darauf ab, Daten zu sammeln, zu bereinigen, zu transformieren und schließlich zu analysieren. Das Ergebnis eines jeden Schritts dient dabei als Ausgangspunkt für den nächsten, wodurch ein durchgängiger und effizienter Prozess entsteht.
346 Sebastian W. Hoggenmüller und Harald Klinke SJS 51 (2), 2025, 337–359 handelt es sich hingegen um komplexere Strukturen, die durch fortgeschrittene Bildverarbeitungsmethoden wie Segmentierung oder die Kombination mehrerer Merkmale aus den Bildpixeln gewonnen werden, beispielsweise Objekterkennung oder semantische Features, die Bedeutungen aus dem Bildinhalt ableiten (etwa die Erkennung von Emotionen in Gesichtern). Zweitens können Features durch Convolutional Neural Networks (CNN), eine Klasse von Deep-Learning-Netzwerken, aus den Bilddaten extrahiert werden, wobei die Features in Form von Embeddings – komprimierten numerischen Repräsentationen der Bilddaten – ausgegeben werden. CNNs bestehen aus sogenannten Convolutional Layers, speziellen Schichten, die lokale Bildbereiche mithilfe von Filtern (auchKernelsgenannt) analysieren und charakteristische Muster wie Kanten oder Texturen in numerische Werte umwandeln. Diese Schichten sind hierarchisch angeordnet, sodass CNNs während des Trainingsprozesses zunehmend komplexere Features lernen, die oft über klassische, handcodierte Bildfeatures hinausgehen. Im Unterschied zu Letzteren erfolgt die Feature Extraction bei CNNs automatisiert und datengetrieben, also ohne, dass die Features explizit definiert werden müssen. Vielmehr lernen die Netzwerke während des Trainingsprozesses, welche Features für eine bestimmte Aufgabe am nützlichsten sind. Gleichzeitig bleibt die interne Struktur des Modells dabei weitgehend unzugänglich: Die Prozesse, durch die das Netzwerk Features aus den Bilddaten extrahiert und in Embeddings überführt, sind aufgrund ihrer datengetriebenen Natur nicht direkt steuerbar. Ferner hängen die resultierenden Embeddings sowohl von den Trainingsdaten als auch der Netzwerkarchitektur ab und ermöglichen daher lediglich die Reproduktion der während des Trainings gelernten Repräsentationen, ohne dass die internen Repräsentationen vollständig kontrolliert werden können. Drittens können Features auch aus vorhandenen Metadaten extrahiert werden, also aus Informationen, die nicht unmittelbar aus den Bildpixeln stammen, sondern die Bilddaten kontextuell ergänzen. Im Zusammenhang mit visuellen Daten aus Open-Data-Quellen wie dem bereits erwähnten Netzwerkkamera-Projekt oder Museumssammlungen können Metadaten verschiedene Informationen bereitstellen: im Fall der Netzwerkkameras Informationen wie den Kamerastandort, die Aufnahmezeit und die Wetterbedingungen, im Fall der Museumssammlungen den*die Künstler*in, das Entstehungsjahr und die Größe eines Kunstwerks. Diese und weitere Metadaten können in beiden Fällen als Features verwendet werden, um Ähnlichkeiten zwischen Netzwerkkamerabildern bzw. zwischen Kunstwerken zu erkennen, Muster und Trends im zeitlichen Verlauf der kamerabasierten Beobachtung der Welt bzw. in der Kunstgeschichte zu analysieren und Netzwerkkamerabilder bzw. Kunstwerke nach bestimmten Kriterien zu kategorisieren oder zu filtern. Die Wahl zwischen Low-Level-Features und High-Level-Features, datengetriebenen Embeddings aus CNNs und metadatenbasierten Features hängt in der kultur-, geistesund sozialwissenschaftlichen Forschungspraxis stark vom spezifischen
Metabilder als Forschungswerkzeuge … 347 SJS 51 (2), 2025, 337–359 Forschungskontext, von den verfügbaren Daten und von den Zielen der Analyse ab. Grundsätzlich ist allen drei Ansätzen jedoch gemein, dass jedes Bildobjekt in einen n-dimensionalen Merkmalsraum, den sogenannten Feature Space, projiziert wird, wobei n die Anzahl der extrahierten Features oder die Dimension der Embeddings zur Beschreibung eines Bildobjekts angibt. Was hier abstrakt klingt – die Projektion von Bildobjekten in einen Feature Space –, möchten wir an einem Beispiel veranschaulichen. Wir nutzen dafür die vom Museum of Modern Art (MoMA) als Open Data bereitgestellten Sammlungsdaten (vgl. MoMA o. J.). Dieser Datensatz ist für unsere Illustrationszwecke besonders geeignet, da er neben den Digitalisaten der Kunstwerke systematisch gepflegte und kuratierte Metadaten von hoher Qualität und Konsistenz bietet. Konkret enthält er eine Vielzahl textbasierter Features (wie Titel, Künstler*in, Abteilung) und numerischer Features (wie Entstehungsund Erwerbsdatum, Höhe, Breite). Dank dieser reichhaltigen und vielfältigen Features können die Kunstwerke präzise und differenziert in einen mehrdimensionalen Feature Space projiziert werden. Um das Verfahren der Projektion verständlich darzustellen, beschränken wir uns bewusst auf die in der Sammlung enthaltenen Fotografien als Bildobjekte und wählen nur zwei Features: das Entstehungsjahr, das aus den Metadaten entnommen wird, und die Durchschnittssättigung, die als visuelles Feature direkt aus den Bilddaten abgeleitet wird und den mittleren Sättigungswert über alle Pixel eines Bildes beschreibt. Diese beiden Features ermöglichen es, die Bildobjekte in einem zweidimensionalen XY-Diagramm abzubilden (Abb.3). Abbildung 3 Metabild nach Entstehungsjahr und Durchschnittssättigung, Fotografien des MoMA Quelle: Eigene Visualisierung [mit R erstelltes Metabild, Klinke].
348 Sebastian W. Hoggenmüller und Harald Klinke SJS 51 (2), 2025, 337–359 Das Metabild in Abbildung 3 zeigt die Verteilung der Fotografien des MoMA entlang der beiden Achsen Entstehungsjahr und Durchschnittssättigung. Auffällig ist dabei, dass bereits zwei Features genügen, um eine plausible, aussagekräftige Anordnung hinsichtlich der visuellen Ähnlichkeit der Bildobjekte zu erzielen. So lässt sich in dem Metabild beispielsweise das Aufkommen der Sepia-Fotografie in der Mitte des 19. Jahrhunderts erkennen. Erst später wird die Schwarz-Weiß-Fotografie zunehmend dominant, insbesondere das Gelatine-Silber-Verfahren, das am unteren Ende der Sättigungsskala angesiedelt ist. In der zweiten Hälfte des 20. Jahrhunderts kommen schließlich Farbabzüge hinzu, die sich über den gesamten Sättigungsbereich verteilen. Darüber hinaus offenbart die Visualisierung eine Sammlungslücke zwischen etwa 1870 und 1890, die möglicherweise überhaupt erst durch diese quantitative Analyse sichtbar wird und deren Ursachen weiterführend mit qualitativen Methoden eruiert werden könnten. Allgemein ist festzuhalten: Obwohl die erkennbaren Cluster ausschließlich auf den beiden von uns exemplarisch ausgewählten Features basieren, weisen sie an unterschiedlichen Stellen semantische Zusammenhänge auf, die zwar nicht explizit in den Features codiert sind, aber mit ihnen korrelieren. Eine ausführlichere Diskussion hierzu folgt im nächsten Abschnitt (3.2). Jenseits unseres Beispiels ist ein n-dimensionaler Feature Space ein abstraktes mathematisch-statistisches Konzept, das dazu dient, die Eigenschaften von Datenpunkten in einem definierten Raum zu beschreiben. Jede Dimension dieses Raums steht für ein spezifisches Feature der Daten. Analysiert man beispielsweise Bilder nicht (wie in unserem Beispiel) anhand von zwei, sondern anhand von drei Features (etwa Helligkeit, Kontrast und Anzahl bestimmter Kanten), so lässt sich jedes Bildobjekt als Datenpunkt in einem dreidimensionalen Raum darstellen, wobei die x-, yund z-Achse jeweils eines dieser Features repräsentiert. Werden für jedes Bildobjekt wiederum mehrere solcher Features ausgewählt, so kann es als Datenpunkt in einem n-dimensionalen Raum verstanden werden. Wenn zum Beispiel 4096 Features aus einem Bild extrahiert werden, wie es bei vortrainierten neuronalen Netzen wie VGG16 der Fall ist, wird das Bildobjekt in einem 4096-dimensionalen Raum positioniert. Dies ist zwar nichts, was man sich visuell vorstellen muss (oder kann), es bedeutet aber: Je mehr Features verwendet werden, desto höher ist die mathematische Präzision bei der Beschreibung eines Bildobjekts und der Differenzierung zu anderen Bildobjekten. Diese Präzision ist jedoch unmittelbar an den Korpus des vortrainierten Modells gebunden, das heißt, die Genauigkeit der Beschreibung bleibt auf jene Features und Kategorien beschränkt, die das Modell während des Trainings kennengelernt hat und als Embeddings im hochdimensionalen Raum darstellt.9 Entscheidend für unsere Beschreibung der algorithmischen Bedingtheit von Metabildern ist hierbei, dass die Positionierung der Datenpunkte im Feature 9 Wenn beispielsweise Schlagwörter vergeben werden, können nur die Schlagwörter aus dem Korpus des Netzes genutzt werden. Ebenso wird das Modell Schwierigkeiten haben, Kategorien zu erkennen, die nicht im Trainingskorpus enthalten sind.
Metabilder als Forschungswerkzeuge … 349 SJS 51 (2), 2025, 337–359 Space weiterführende Berechnungen ermöglicht. Beispielsweise können damit die Abstände zwischen Objekten berechnet werden, die eine zentrale Rolle bei der Bestimmung ihrer Ähnlichkeit spielen. Die Datenpunkte werden dabei als Vektoren im n-dimensionalen Raum dargestellt, deren Beziehungen durch verschiedene mathematische Operationen wie Abstandsmessungen oder Winkelberechnungen analysiert werden können. Der euklidische Abstand misst die direkte Distanz zwischen zwei Punkten und eignet sich besonders, wenn die physische Nähe der Datenpunkte von Bedeutung ist. Alternativ kann die Kosinus-Ähnlichkeit verwendet werden, die den Winkel zwischen den Vektoren berücksichtigt, was sinnvoll ist, wenn die Richtung der Vektoren wichtiger ist als ihre absolute Länge. Die so berechneten Abstandswerte geben an, wie stark sich zwei Objekte in Bezug auf ihre Features ähneln oder unterscheiden. Ein kleiner Abstand deutet darauf hin, dass die Objekte ähnliche Features haben, während ein großer Abstand anzeigt, dass die Objekte stark voneinander abweichen. Solche Abstandsmessungen sind zentral für viele algorithmusbasierte Anwendungen, darunter Klassifizierungen (vgl. etwa im Bereich Public Health Rösch, 2022), Clusterings (vgl. z. B. in der Literarkritik Tschuggnall et al., 2016) und Empfehlungssystemen (vgl. u. a. aus soziologischer Perspektive Unternährer, 2024). Im Zusammenhang mit Metabildern helfen sie wiederum, Beziehungen visuell darzustellen und dadurch Muster sowie Strukturen in großen visuellen Datenbeständen zu identifizieren. 3.2 Dimensionsreduktion (Dimension Reduction) Wie im vorherigen Abschnitt dargelegt, eröffnet eine größere Anzahl von Features, die zur Beschreibung von Datenobjekten verwendet werden, potenziell vielfältigere statistische Möglichkeiten auf der Ebene der Daten. Doch auf der Ebene der Visualisierung stoßen wir schnell an Grenzen: Menschen sind daran gewöhnt, nur eine bestimmte Anzahl von Dimensionen wahrzunehmen. Die physische Welt, die wir erleben, umfasst im Wesentlichen drei Dimensionen, die sich in einem dreidimensionalen Raum gut visualisieren lassen, während Zeitlichkeit zusätzlich durch Animationen dargestellt werden kann.10 Bei höherdimensionalen Datensätzen hingegen wird es zunehmend schwieriger, sie intuitiv zu erfassen oder visuell darzustellen. Eine Möglichkeit, mit dieser Herausforderung umzugehen, sind mathematischstatistische Verfahren der Dimensionsreduktion. Diese Verfahren zielen darauf ab, hochdimensionale Daten, also Daten, bei denen jedes einzelne Datenobjekt durch eine sehr große Anzahl an Features beschrieben wird, in einen niedrigdimensionalen Raum zu projizieren, wobei die wesentlichen Eigenschaften der Daten, insbesondere ihre Strukturen und Muster, erhalten bleiben. Dabei entstehen neue, repräsentative Dimensionen, die häufig als Kombinationen oder Transformationen der ursprüng10 Das Potenzial solcher Visualisierungen hat Hans Rosling eindrucksvoll gezeigt; vgl. etwa seine TED-Talks (https://www.ted.com/playlists/474/the_best_hans_rosling_talks_yo).
350 Sebastian W. Hoggenmüller und Harald Klinke SJS 51 (2), 2025, 337–359 lichen Features gebildet werden. Dies verbessert nicht nur die Interpretierbarkeit, sondern auch die Effizienz und Anwendbarkeit der Daten für maschinelle Lernmodelle. Eine Alternative wäre, nur die für die Datenanalyse relevantesten Features beizubehalten und irrelevante oder weniger informative Features zu eliminieren, mithin eine bestimmte Untermenge der ursprünglichen Dimensionen auszuwählen, ein Prozess, der auch als Merkmalsauswahl (Feature Selection) bezeichnet wird. Die Dimensionsreduktion ist prinzipiell vergleichbar mit der Projektion des Schattens eines dreidimensionalen Gegenstands auf eine flache Ebene: Wird ein dreidimensionaler Gegenstand beleuchtet, entsteht eine zweidimensionale Projektion, die eine vereinfachte Darstellung des Gegenstands ist, da in ihr die Tiefeninformation fehlt. Dabei können jedoch auch andere relevante Eigenschaften des Gegenstands verloren gehen. Um einen solchen Informationsverlust zu minimieren, wurden verschiedene Algorithmen entwickelt, die die Dimensionalität reduzieren, indem sie einen neuen, kompakteren Feature Space schaffen, der einerseits eine visuelle Darstellung ermöglicht, die von Menschen interpretiert werden kann, und andererseits die anderen Dimensionen so weit wie möglich bewahrt. Oder anders ausgedrückt: Die 4096 visuellen Features, die beispielsweise durch VGG16 extrahiert werden, werden in einen niedrigdimensionalen (üblicherweise zweidimensionalen) Raum projiziert, während bestimmte relevante Informationen erhalten werden. Das Ergebnis dieser Transformation ist eine XY-Position des Bildobjekts im neuen Feature Space. Es gibt eine Reihe von Algorithmen zur Dimensionsreduktion, die sich unter anderem in ihren mathematischen Grundlagen, den Annahmen über die Datenstruktur und in der Art und Weise, wie sie mit Daten umgehen, unterscheiden. Zu den bekanntesten gehören die Principal Component Analysis (PCA), das t-Distributed Stochastic Neighbor Embedding (t-SNE) und die Uniform Manifold Approximation and Projection (UMAP), wobei alle drei jeweils einen eigenen Ansatz bieten, um die hochdimensionalen Daten in eine Form zu überführen, die sowohl der menschlichen Wahrnehmung zugänglich ist als auch von maschinellen Lernprozessen effizienter verarbeitet werden kann.11 Der Algorithmus t-SNE hat sich in der kultur-, geistesund sozialwissenschaftlichen Forschung als besonders nützliches Werkzeug für die explorative Datenanalyse und die Visualisierung komplexer Datensätze erwiesen, weshalb wir uns im Folgenden auf diesen Algorithmus konzentrieren. Der Dimensionsreduktionsalgorithmus t-SNE basiert auf dem Stochastic Neighbor Embedding (SNE), das 2002 von Geoffrey Hinton und Sam Roweis 11 Dabei hat jeder dieser Ansätze seine Vorund Nachteile, insbesondere in Bezug auf die Verarbeitungsgeschwindigkeit und die Art des Informationsverlusts. PCA erfasst nur lineare Zusammenhänge, ist jedoch der schnellste Ansatz, da er effizient berechnet wird und sich gut für große Datensätze skalieren lässt. UMAP bietet eine gute Balance zwischen Geschwindigkeit und der Erhaltung lokaler Strukturen, wobei es globale Zusammenhänge besser bewahrt als t-SNE. t-SNE wiederum ist langsamer, liefert jedoch eine sehr gute Visualisierung der lokalen Datenstruktur, indem es Cluster und Nachbarschaftsbeziehungen bewahrt, jedoch oft auf Kosten der globalen Struktur (vgl. für eine kritische Perspektive auf PCA Shen et al., 2012; auf UMAP Damrich &Hamprecht, 2021; auf t-SNE Wattenberg et al., 2016).
Metabilder als Forschungswerkzeuge … 351 SJS 51 (2), 2025, 337–359 entwickelt wurde. Ziel von SNE war es, die lokalen Nachbarschaftsbeziehungen optimal zu erhalten, indem ähnliche Datenpunkte im hochdimensionalen Raum auch im niedrigdimensionalen Raum nahe beieinander platziert werden (vgl. Hinton &Roweis, 2003). 2008 wurde das SNE-Verfahren dann von Laurens van der Maaten und Geoffrey Hinton zu t-SNE weiterentwickelt, wobei sie im niedrigdimensionalen Raum eine t-Verteilung anstelle der Gaußschen Verteilung einführten. Da die t-Verteilung extreme Distanzen zwischen Datenpunkten besser erfasst, mindert diese Modifikation das Überfüllungsproblem (crowding problem) und verbessert die Visualisierung von Clustern in großen Datensätzen (vgl. van der Maaten & Hinton, 2008). Um die Nachbarschaftsbeziehungen zwischen den Datenpunkten weitgehend zu erhalten, wandelt t-SNE die Ähnlichkeiten im hochdimensionalen Raum (mithilfe einer Gaußschen Verteilung) in Wahrscheinlichkeiten um und versucht, diese Wahrscheinlichkeiten im niedrigdimensionalen Raum (mithilfe einer t-Verteilung) möglichst genau nachzubilden. Dabei werden die Datenpunkte zunächst zufällig im niedrigdimensionalen Raum platziert. Anschließend passt der Algorithmus die Position der Datenpunkte schrittweise an, um die Diskrepanz zwischen den Wahrscheinlichkeiten im hochund im niedrigdimensionalen Raum zu minimieren. Dieser Prozess wird so lange fortgesetzt, bis die bestmögliche Übereinstimmung zwischen den Wahrscheinlichkeiten erreicht ist. Generell liefert t-SNE dabei kein festes, deterministisches Ergebnis. Stattdessen variiert die genaue Anordnung der Datenpunkte bei jedem Durchlauf, da der Algorithmus mit zufällig gewählten Startpositionen beginnt und durch Zufallsfaktoren im Optimierungsprozess beeinflusst wird.12 Dies bedeutet, dass der Algorithmus verschiedene Versionen der hochdimensionalen Datenpunkte in der niedrigdimensionalen Darstellung erzeugt, während die übergeordneten Muster in der Regel erhalten bleiben. Kurz gesagt: Wiederholte Ausführungen des Algorithmus führen zu unterschiedlichen visuellen Darstellungen. Dies lässt sich erneut anhand der MoMA-Daten veranschaulichen: Zu Demonstrationszwecken berücksichtigen wir diesmal nur die ersten 1 000 Bildobjekte des Datensatzes und nutzen ausschließlich Features aus den Metadaten, verzichten also auf aus den Bildern abgeleitete Features. Mithilfe der Programmiersprache und statistischen Umgebung R werden daraufhin aus den 29 Spalten der MoMA-Daten, die die verschiedenen Bildobjekte beschreiben, elf numerische Features13 ausgewählt und in numerische Werte umgewandelt. Anschließend reduzieren wir diese Daten 12 Zumindest theoretisch könnte dieses Problem durch die Verwendung eines fixierten Startwerts (Random Seed) gelöst werden. In der Praxis nutzen jedoch die meisten Implementierungen Parallelisierungen auf der Central Processing Unit (CPU) oder Graphics Processing Unit (GPU), wodurch Berechnungen in variierender Reihenfolge und Geschwindigkeit ablaufen. Dies führt dazu, dass die exakte Reproduzierbarkeit aufgrund der parallel ablaufenden Prozesse oft nicht gewährleistet ist. 13 Die elf numerischen Features sind folgende: ObjectID, BeginDate, EndDate, DateAcquired, Depth, Diameter, Height, Weight, Width, Duration und AspectRatio.