scieee AI-readable full text Open interactive document viewer

ERC_Cog PROMENADE - WP1: GPT-generated ratings

IUSS Neurolinguistics and Experimental Pragmatics Laboratory (NEPLab)

Abstract

GPT ratings is a project that aims to test the validity and reliability of machine-generated norms for metaphorical stimuli. The project is part of the activities of the ERC Consolidator Grant project PROMENADE (PROcessing MEtaphors: Neurochronometry, Acquisition and DEcay, ID: 101045733), awarded to Prof. Valentina Bambini (University School for Advanced Studies IUSS Pavia, Italy) and funded by the European Research Council. The project is related to the Figurative Archive resource (10.5281/zenodo.14924803), which provides one of the human benchmark used in the study to assess the validity of GPT-generated ratings.

Full text

Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors Veronica Mangiaterra1, Hamad Al-Azary2, Chiara Barattieri di San Pietro1, Paolo Canal1, Valentina Bambini1 1) Laboratory of Neurolinguistics and Experimental Pragmatics (NEPLab), Department of Humanities and Life Sciences, University School for Advanced Studies IUSS, Pavia, Italy 2) Department of Humanities, Social Sciences and Communication, College of Arts and Sciences, Lawrence Technological University, Southfield, MI, USA Supplementary Materials These supplementary materials contain: i) Information about participants of each rating study ii) Examples of prompt provided to GPT models iii) Validity of GPT-generated metaphor ratings presented separately for each study (see Supplementary Table 1 and Supplementary Figure 1) iv) Validity of GPT-generated ratings for anomalous and literal statements (see Supplementary Table 2) v) Reliability of GPT-generated metaphor ratings presented separately for each study (see Supplementary Table 3) Supplementary table 1. The table shows demographic information on the sample of raters for each study. Study Participants Al-Azary & Buchanan (2017) 52 fluent speakers of English (age > 18) Bambini et al. (2013) 85 native speakers of Italian (42F; age: M = 26.85, SD = 3.80; education in years: M = 18.02, SD = 2.04) Bambini et al. (2014) 105 native speakers of Italian (83F; age: M = 23.00, SD = 4.31) Bambini et al. (2024) 122 native speakers of Italian (68F, age: M = 24.34, SD = 1.97) Campbell & Raney (2016) 90 fluent speakers of English Canal et al. (2022) 53 native speakers of Italian (40F; age: M = 23.91, range: 21–32; education in years: M = 15.83, range: 13–18) Cardillo et al. (2017) 40 native speakers of English (21F; age: M = 21.3, education in years: M = 15.25) IUSS NEPLab MetaBody study 49 native speakers of Italian (27F; age: M = 27.35, SD = 3.55; education in years: M = 15.82, SD = 2.76) Supplementary Table 2. Example of prompt provided to models compared to human instructions. Changes are highlighted in bold. Instructions to human participants Prompt to Large Language Models Your task will be to rate how suitable or natural a series of statements are. The statements will either be nonsensical, literal, or figurative. For example, “A Sheep is a Hill” is a nonsensical statement; “A Circle is a Shape” is a literal statement; “Love is a Journey” is a figurative statement. Use the number pad to indicate your rating from 1 to 6 (1 being very unsuitable/unnatural and 6 being very natural/suitable). Please read the statements carefully before making a response. Press the space bar to begin. There will be three practice trials followed by the rest of the experiment. When you are finished, a thank-you message will appear. Afterward, you may exit the room. Your task will be to rate how suitable or natural a series of statements are. The statements will either be nonsensical, literal, or figurative. For example, “A Sheep is a Hill” is a nonsensical statement; “A Circle is a Shape” is a literal statement; “Love is a Journey” is a figurative statement. Answer with your rating from 1 to 6 (1 being very unsuitable/unnatural and 6 being very natural/suitable). Answer with only the number from 1 to 6, do not add more. Supplementary table 3. Validity of GPT-generated metaphor ratings for each study. Correlations between human-generated ratings and ratings generated by GPT3.5-turbo, GPT4o-mini (prompted through API and ChatGPT interface), and GPT4o for the three dimensions (comprehensibility, imageability, and familiarity). Measure Language Study GPT3.5-turbo GPT4o-mini GPT4o ChatGPT Comprehensibility English Al-Azary & Buchanan (2017) 0.69*** 0.74*** 0.79*** 0.78*** Imageability English Campbell & Raney (2016) 0.39** 0.42** 0.36** 0.56*** Italian Bambini et al. (2024) 0.38*** 0.46*** 0.65*** 0.45*** Familiarity Italian Bambini et al. (2013) 0.02 0.37** 0.56*** 0.31* Italian Bambini et al. (2014) 0.50*** 0.55*** 0.61*** 0.29** Italian Bambini et al. (2024) 0.49*** 0.57*** 0.64*** 0.45*** English Campbell & Raney (2016) 0.65*** 0.67*** 0.62*** 0.73*** Italian Canal et al. (2022) 0.32*** 0.54*** 0.70*** -0.09 English Cardillo et al. (2017) 0.65*** 0.62*** 0.63*** 0.39*** Italian IUSS NEPLab MetaBody study -0.02 0.47*** 0.68*** -0.11 Supplementary Figure 1. Scatter plots showing the relationship between GPT ratings (y-axis) and human ratings (x-axis) across three dimensions: Imageability (blue panel), Comprehensibility (purple panel), and Familiarity (green panel), in Italian or English, with marginal density plots illustrating rating distributions. Regression lines are color-coded by model: GPT-3.5 (red), GPT-4o-mini (teal), and GPT-4o (orange) Supplementary table 4. Validity of GPT ratings for literal and anomalous statements. Correlations between human-generated ratings and ratings generated by GPT3.5-turbo, GPT4o-mini (prompted through API and ChatGPT interface), and GPT4o Measure Study Items Language GPT3.5-turbo GPT4o-mini (API) GPT4o-mini (Interface) GPT4o Familiarity Cardillo et al. (2017) Literal English 0.53*** 0.58*** 0.58*** 0.61*** Bambini et al. (2013) Anomalous and literal Italian 0.83*** 0.82*** 0.90*** 0.93*** Comprehe nsibility Al-Azary & Buchanan (2017) Anomalous and literal English 0.97*** 0.98*** 0.98*** 0.96*** Supplementary table 5. Reliability in each study. Correlations between ratings generated by GPT3.5-turbo, GPT4o-mini (prompted through API and ChatGPT interface), and GPT4o obtained in two independent sessions for the three dimensions (comprehensibility, imageability, and familiarity). Measure Language Study GPT3.5-turbo GPT4o-mini GPT4o ChatGPT Comprehensibility English Al-Azary & Buchanan (2017) 0.99*** 0.99*** 0.99*** 0.91*** Imageability English Campbell & Raney (2016) 0.99*** 0.98*** 0.97*** 0.66*** Italian Bambini et al. (2024) 0.99*** 0.99*** 0.98*** 0.67*** Familiarity Italian Bambini et al. (2013) 0.98*** 0.99*** 0.96*** 0.67*** Italian Bambini et al. (2014) 0.97*** 0.99*** 0.96*** 0.88*** Italian Bambini et al. (2024) 0.98*** 0.99*** 0.99*** 0.87*** English Campbell & Raney (2016) 0.98*** 0.99*** 0.99*** 0.86*** Italian Canal et al. (2022) 0.97*** 0.99*** 0.99*** 0.74*** English Cardillo et al. (2017) 0.98*** 0.99*** 0.98*** 0.63*** Italian IUSS NEPLab MetaBody study 0.93*** 0.99*** 0.98*** 0.47***