scieee AI-readable full text Open interactive document viewer

From human artefact to machine output: automating the "art" of psychological measurement

Marmolejo-Ramos, Fernando; Kundrát, Josef; Rečka, Karel; Bulut, Okan; Anunciação, Luis; Marques, Louise; Barthakur, Abhinava; Karakale, Ozge; Correa, Juan C; Pinos Ullauri, Luis Alberto; Ospina, Raydonal; Tejada, Julian

Abstract

Creating psychological assessment tools is crucial for research but traditionally expensive and time-consuming. While Large Language Models (LLMs) show promise for automating this process, existing approaches lack systematic, user-friendly methodologies grounded in psychometric principles. This study presents an enhanced Psychometric Item Generator (PIG) method using conversational LLMs with Problem-Solving Plans (PSP) and Chain-of-Thought (CoT) prompting. Three demonstrations validated the approach: Gemini 1.5 Flash generated 20 “propensity to trust AI” items with strong semantic coherence; Claude 3 Opus created 20 “AI anxiety” items that outperformed human-generated versions linguistically; and a 6-item “AI adoption in online learning” scale was developed and validated with 1,233 participants using multiverse analysis. Results demonstrate that LLMs can produce psychometrically sound items. The AI-generated anxiety scale showed superior linguistic properties compared to human alternatives, while the learning scale exhibited good internal consistency, item homogeneity, and clear two-factor structure across multiple analytical teams. The study establishes a PSP-CoT framework that improves LLM output quality, offering researchers a cost-effective, accessible scale development methodology. However, findings emphasize that human oversight, rigorous validation, and ethical considerations remain essential components of the process.

Full text

Journal of Psychology and AI ISSN: 2997-4100 (Online) Journal homepage: www.tandfonline.com/journals/tpai20 From human artefact to machine output: automating the “art” of psychological measurement Fernando Marmolejo-Ramos, Okan Bulut, Luis Anunciaçáo, Louise Marques, Abhinava Barthakur, Josef Kundrat, Karel Rečka, Özge Karakale, Juan C. Correa, Luis Alberto Pinos-Ullauri, Raydonal Ospina & Julian Tejada To cite this article: Fernando Marmolejo-Ramos, Okan Bulut, Luis Anunciaçáo, Louise Marques, Abhinava Barthakur, Josef Kundrat, Karel Rečka, Özge Karakale, Juan C. Correa, Luis Alberto Pinos-Ullauri, Raydonal Ospina & Julian Tejada (2025) From human artefact to machine output: automating the “art” of psychological measurement, Journal of Psychology and AI, 1:1, 2561692, DOI: 10.1080/29974100.2025.2561692 To link to this article: https://doi.org/10.1080/29974100.2025.2561692 © 2025 The Author(s). Published by Informa UK Limited, trading as Taylor & Francis Group. Published online: 09 Oct 2025. Submit your article to this journal View related articles View Crossmark data Full Terms & Conditions of access and use can be found at https://www.tandfonline.com/action/journalInformation?journalCode=tpai20 RESEARCH ARTICLE From human artefact to machine output: automating the “art” of psychological measurement Fernando Marmolejo-Ramos a , Okan Bulut b , Luis Anunciaçáo c , Louise Marques c , Abhinava Barthakur d , Josef Kundrat e , Karel Rečka e , Özge Karakale f , Juan C. Correa g , Luis Alberto Pinos-Ullauri h,i,j , Raydonal Ospina k and Julian Tejada l a College of Education, Psychology, and Social Work, Flinders University, Adelaide, Australia; b Centre for Research in Applied Measurement and Evaluation, University of Alberta, Edmonton, Canada; c Department of Psychology, Catholic University of Rio de Janeiro, Rio de Janeiro, Brazil; d Education Futures, University of South Australia, Adelaide, Australia; e Department of Psychology, University of Ostrava, Ostrava, Czech Republic; f School of Psychology, University of Wollongong, Wollongong, Australia; g Research & Development Unit, Critical Centrality Institute, Monterrey, Mexico; h IMEC research group ITEC, KU Leuven, Kortrijk, Belgium; i Faculty of Psychology and Educational Sciences, KU Leuven, Kortrijk, Belgium; j Centre for Digital Systems, IMT Nord Europe, Douai, France; k Departmento de Estatística, LInCA, Universidade Federal da Bahia, Cidade Universitária, Bahia, Brazil; l Department of Psychology, Federal University of Sergipe, São Cristóvão, Brazil ABSTRACT Creating psychological assessment tools is crucial for research but traditionally expensive and time-consuming. While Large Language Models (LLMs) show promise for automating this process, existing approaches lack systematic, user-friendly methodologies grounded in psychometric principles. This study presents an enhanced Psychometric Item Generator (PIG) method using conversational LLMs with ProblemSolving Plans (PSP) and Chain-of-Thought (CoT) prompting. Three demonstrations validated the approach: Gemini 1.5 Flash generated 20 “propensity to trust AI” items with strong semantic coherence; Claude 3 Opus created 20 “AI anxiety” items that outperformed human-generated versions linguistically; and a 6-item “AI adoption in online learning” scale was developed and validated with 1,233 participants using multiverse analysis. Results demonstrate that LLMs can produce psychometrically sound items. The AI-generated anxiety scale showed superior linguistic properties compared to human alternatives, while the learning scale exhibited good internal consistency, item homogeneity, and clear two-factor structure across multiple analytical teams. The study establishes a PSP-CoT framework that improves LLM output quality, offering researchers a cost-effective, accessible scale development methodology. However, findings emphasize that human oversight, rigorous validation, and ethical considerations remain essential components of the process. ARTICLE HISTORY Received 21 April 2025 Accepted 15 August 2025 KEYWORDS Large language models; psychometrics; psychometric item generator; chain-ofthought 1. Introduction The use of surveys and/or self-report scales in psychology has been increasing in recent decades (Clark & Watson, 2019). These instruments, after having enough sources of evidence of validity (Association, 2014), are valuable tools for measuring constructs that are not directly observable. Their construction is not an easy process involving a series of steps, from conceptualisation up to validation, including the item generation (Jebb et al., 2021). The latter is a special step in which, usually from a larger set of items, those that best fit the survey objectives are selected, with a special focus on item wording. Clark and Watson (2019) describe this step as one of the most critical in the development of a scale, because by means of psychometric techniques, it is possible to eliminate problematic items, but not to create missing items that should have been included. Recently, Götz et al. (2023) introduced the Psychometric Item Generator (PIG), a pioneering method that uses the GPT-2 neural network within a Python-based Google Colab environment to automate item generation. Their contribution relies on simplifying and accelerating the process of creating items for psychological assessments by producing content that aligns with the intended constructs, such as personality traits or attitudes. While groundbreaking, their original method has two primary limitations in the CONTACT Fernando Marmolejo-Ramos [email protected] JOURNAL OF PSYCHOLOGY AND AI 2025, VOL. 1, NO. 1, 2561692 https://doi.org/10.1080/29974100.2025.2561692 © 2025 The Author(s). Published by Informa UK Limited, trading as Taylor & Francis Group. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. The terms on which this article has been published allow the posting of the Accepted Manuscript in a repository by the author(s) or with their consent. current technological landscape: it requires users to interact with and modify Python code, creating a barrier for non-programmers, and it is based on the now-outdated GPT-2 architecture (Götz et al., 2023). Our study proposes a substantial evolution of this concept, creating a more accessible, flexible, and powerful Psychometric Item Generation variant with three key improvements. First, our approach is entirely codefree and platform-agnostic, relying exclusively on conversational prompt engineering. This removes technical barriers, empowering any researcher to leverage automated item generation using readily available web interfaces. Second, we utilise state-of-the-art LLMs (e.g. Google’s Gemini 1.5 and Anthropic’s Claude 3 Opus), which offer significantly more advanced reasoning, contextual understanding, and text generation capabilities than GPT-2. Third, we introduce a novel Problem-Solving Plan (PSP) and Chain-of-Thought (CoT) prompting framework. This structured approach is designed to improve the coherence and psychometric relevance of the generated items, moving beyond simple one-shot prompts to a guided, iterative dialogue with the LLM. This combination of accessibility, advanced model usage, and a sophisticated prompting strategy represents a significant and necessary advancement over existing automated item generation methods. LLMs are machine learning models able to generate text in human languages for countless types of contexts, in which the generated text resembles human responses (Giray, 2023; Henrickson & Meroño-Peñuela, 2025). Interaction with these models is conducted through a chat-like command line, where commands are typed as human conversations, and this is the reason why the concept of prompt engineering was born, to assist in the “conversation” with these systems. Prompt engineering is an emerging field focused on the methodical creation and refinement of prompts – the instructions used to interact with LLMs – through iterative cycles of modification (see Figure 1). There are different types of prompting: Zero-Shot Prompting, in which the instructions are given without examples; Chain-of-thought (CoT), which incorporates intermediate steps to guide the reasoning, breaking down the task into ordering sub-goals and offering examples of step-by-step reasoning (Chu et al., Figure 1. Scheme of how an LLM works. It all starts with a prompt or input text, a human conversation-style instruction that is segregated (Tokenization) into a group of characters based on punctuation marks, spaces, or special characters. These groups can represent words or subwords and are called tokens. Once the tokenizer algorithm has segregated the text, each token is identified with an ID, which in turn, is associated with a dense vector of high-dimensional space (i.e. GPT-3 uses vectors of 12,288 dimensions), forming an embedding matrix. These vector representations are suitable for neural network processing and contain all LLM training data and will allow the algorithm to capture semantic relationships and context. The embedded text sequence is processed via an encoder network (a recurrent neural network [RNN] or transformer-based architecture) which analyses word relationships and builds a contextual representation, capturing knowledge about the overall meaning and flow of the input. In the context with attention phase, the most significant parts of the context generated by the encoder are highlighted (i.e. ‘scales’ in this context do not refer to devices used for measuring a person’s body mass or the small rigid plates that grow out of the skin of a fish). Then, the decoder network, another RNN or transformer-based architecture, uses the encoder context and its own internal state to generate the output text word by word by predicting the next word in the sequence based on the learned information. For more details visit: https://cutt.Ly/ lw8PKU7Y. 2F. MARMOLEJO-RAMOS ET AL. 2023); Few-Shot Prompting, similar to Zero-Shot, but with some examples; and self-consistency (S-C) in which similar instructions are given to obtain a pool of responses from which the most frequent is selected (Ahmed & Devanbu, 2023) (see also Wei et al., 2022). In our case, we propose to use a combination of CoT with Few-Shot to configure our adaptation of the PIG approach, which can be applied to any LLM available. The PSP component, inspired by plan-and-solve prompting strategies (L. Wang et al., 2023), specifically advances beyond prior prompt engineering methods by establishing an upfront, top-down plan that defines the scope, goals, and boundaries of the interaction before the CoT begins. While a standard CoT prompt breaks a task into steps (Y. Wang et al., 2024), a PSP ensures those steps are logically coherent and directed towards a pre-defined psychometric objective. For instance, the PSP would first establish the goal: “Develop a 5-item scale for the ‘sociotechnical blindness’ dimension of AI anxiety, ensuring two items are reversecoded and all items are appropriate for a lay audience.” The subsequent CoT then executes this plan step-bystep. This “plan-and-solve” approach reduces the risk of the LLM deviating into irrelevant or incoherent paths, a common issue with less-structured prompting, thereby making the item generation process more efficient, targeted, and aligned with psychometric goals from the outset. Our proposed approach proposed herein requires no technical expertise and is available for free use. It overcomes prior barriers to automatic item generation like inflexibility, inaccessibility, and computational demands. Additionally, our approach enables researchers to break free from reliance on manual item writing, allowing algorithms to finally speak the language of psychometrics. Overall, our study aims to provide researchers with a cost-effective, versatile artificial intelligence solution for psychological measurement. In the next section, we illustrate through three demonstrations, our non-technical proposed approach on generating psychometric items. First, we employed Gemini 1.5 Flash with zero-shot prompting to generate 20 initial questions measuring “propensity to trust AI”, which were then analysed using natural language processing (NLP) methods. Second, using Claude 3 Opus, we generated and linguistically evaluated 20 items targeting AI anxiety, inspired by Y.-Y. Wang & Wang (2022). Third, by intentionally leveraging Gemini 1.5 Flash’s tendency to generate imaginative or fabricated content from minimal prompts, we produced over 20 diverse items assessing AI trust, with a focus on adoption and perception. Expert validation refined this set to six items, whose psychometric properties were subsequently confirmed using participant responses. The experts involved in this process were academics specialising in educational technology, artificial intelligence, online education, data analytics, and psychometrics. The LLM models used in the demonstrations were released in January 2024. We then discuss how embedding prompts within a problem-solving plan can lead to fruitful interactions with LLMs and, consequently, informative outputs. Finally, in the discussion section, we elaborate on psychometrics as a science and the role of LLMs within it. 2. Demonstrations 2.1. Producing 20 usable items to measure “propensity to trust in AI” We resorted to a CoT conversational style to query Gemini. We started with a zero-shot prompting style: “Provide 10 up-to-date psychometric scales to measure people’s propensity to trust AI. Provide real academic references for each.” Purposefully, we aimed to use Gemini’s hallucinations as a way to obtain results. That is, instead of the usual approach of avoiding and discarding AI hallucinations, we used its potential insights as a proving ground. As expected, Gemini provided some real and some hallucinated references (anecdotally, however, typing these illusory references into Google provided useful real references). We then continued with the question, “Based on the scales above, what’s the overall definition of ‘propensity to trust AI’?” Gemini then provided an educated response and broke down the concept of propensity to trust in AI in four dimensions: i) trust in AI competence, ii) trust in AI benevolence, iii) trust in AI transparency, and iv) trust in AI fairness. Gemini also provided definitions for each dimension and examples of how the propensity to trust AI might manifest itself in real-world behaviours (e.g. Gemini said, “A person with a high propensity to trust AI might be more likely to use a self-driving car”). Then we asked Gemini: “Provide psychometric items that will measure ‘propensity to trust in AI’ based on your definition above. Provide five items for Trust in AI competence, 5 items for Trust in AI benevolence, five items for Trust in AI transparency, and 5 JOURNAL OF PSYCHOLOGY AND AI 3 items for Trust in AI fairness”. After Gemini’s output, we asked to “rewrite the above items so that they are to be answered on a continuous likelihood scale between “very unlikely” and ‘very likely” and after Gemini’s output we finally asked to “rewrite the above items so that they are worded in the first person”. Four of the items were: ●I am likely to trust that AI systems can be used to make my life better (item generated for the “trust in AI competence” dimension) ●I am likely to believe that AI systems are designed to act in the best interest of humans (item generated for the “trust in AI benevolence” dimension) ●I am likely to trust that AI systems will be used in a transparent and accountable manner (item generated for the “trust in AI transparency” dimension) ●I am likely to believe that AI systems treat all people fairly and equitably (item generated for the “trust in AI fairness” dimension) Upon initial examination, the generated items appear to be relevant and align well with the four dimensions of the “propensity to trust AI” as defined by Gemini. More specifically, for the initial human evaluation, we selected several criteria: 1) the generated text must be grammatically coherent, 2) the generated text must be relevant to the given topic, and 3) a researcher unfamiliar with the study’s design must be unable to discern whether the items originate from a human-designed test or are AIgenerated content. However, when the items are considered as a corpus, we can assess whether this corpus exhibits textual dimensionality using an NLP metric, such as cosine similarity (Singh & Singh, 2021). Cosine similarity evaluates the similarity between two non-zero vectors by computing the cosine of the angle between them. The values range from −1 to 1: 1 signifies that the vectors are identical (perfectly aligned), 0 indicates that the vectors are orthogonal (completely dissimilar), and −1 means the vectors are diametrically opposed. In the resulting matrix, most values exceed 0.75, showing that the four items are highly related and suggesting that the corpus cannot be divided into distinct dimensions. However, it is important to note that the results may vary slightly from those reported here if the prompts are used again in the future. This is because LLMs have various hyper-parameters, such as temperature, Top P, frequency penalty, and presence penalty, that influence the randomness of their outputs. The first parameter is temperature, which controls the randomness of the model’s output. A higher temperature setting yields more varied and unpredictable responses, whereas a lower temperature setting generates more deterministic and conservative outputs. The second parameter is Top P, also known as nucleus sampling, which determines the probability threshold for token selection. A higher Top p value allows for a wider range of tokens to be considered, resulting in more random output. In contrast, a lower Top p value restricts the selection to more likely tokens, reducing randomness. The third parameter is the frequency penalty, which decreases the likelihood of the model repeating the same words or phrases. A higher penalty setting discourages repetition and promotes the use of unique word choices. Finally, the presence penalty parameter encourages the introduction of new concepts or topics in the output. A higher presence penalty setting makes the model less likely to repeat the same topics, fostering more diverse content generation (see Ding et al., 2023; Zhang et al., 2024). The same argument applies to the results of the next demonstration. To contextualise these findings, it is instructive to compare the generated items to existing validated scales. The field has several established measures, such as the 12-item Trust in Automation Scale (TIAS) (McGrath et al., 2025) and the 9-item Human-Computer Trust Scale (HCTS) (Gulati et al., 2019). The TIAS, for example, is a robust measure with high internal consistency (Cronbach’s α¼0:94) that assesses trust along dimensions of performance and integrity. A key feature of the TIAS is its inclusion of five reversescored items (e.g. “The system is deceptive”) to mitigate acquiescence bias. Our LLM-generated items for “competence” and “benevolence” align conceptually with the TIAS’s “ability” and “integrity” dimensions (McGrath et al., 2025). However, a notable difference is that our initial zero-shot prompting did not produce any reverse-scored items. This comparison highlights a critical point: while LLMs can readily generate topically relevant items, achieving the psychometric sophistication of established scales like the TIAS requires more advanced prompting strategies, such as explicitly instructing the model to generate negatively-worded items to ensure a balanced scale. 4F. MARMOLEJO-RAMOS ET AL. 2.2. Producing 20 usable items to measure AI anxiety Y.-Y. Wang and Wang (2022) proposed a 21-item scale to measure AI anxiety (AIA). These items represented four AIA dimensions: eight items related to learning, six related to job replacement, four related to sociotechnical blindness, and three related to AI configuration. In their study, they defined AIA as “as an overall, affective response of anxiety or fear that inhibits an individual from interacting with AI. Thus, AIA may be operationally considered as a general perception or belief with multiple dimensions. Moreover, AIA in the present research context focuses on the variable itself, rather than the process of response or model evaluation, which promotes the operation of AIA as a single variable, independent of numerous antecedents or consequences” (p. 621). Those items were generated in the standard way; researchers created a list of 59 questions to measure AIA and then reviewed the questions with experts to ensure they were clear and covered all aspects of AIA. We tested our proposed approach by chatting with Claude to generate 20 items suitable for assessing the “learning” dimension of AIA. Table 1 shows the prompt used. Claude generated 20 usable items, and one last prompt required Claude to indicate how to rate the items, as such information was not explicit in its first output (“How are these items to be responded to?”). We list below three of the items generated: ●Trying to understand the advanced algorithms behind AI systems makes me anxious ●Adapting to rapid changes in AI technologies makes me anxious ●Learning new terms and concepts related to AI makes me anxious It is clear that the generated items followed a linguistic structure similar to the eight items created by Y.-Y. Wang and Wang (2022), but most importantly, the Claude-generated items were semantically related to the keywords AI, anxiety, and learning. When comparing items generated by humans with those produced by Anthropic’s LLM, Claude, our application of information theory reveals that the coefficient of mutual Table 1. Prompt submitted to Claude in order to generate 20 items in the “learning dimension” of an AI anxiety scale. New items for the learning dimension of an AI anxiety scale Artificial Intelligence Anxiety (AIA) can be defined as a general, affective reaction of fear or anxiety that prevents an individual from interacting with artificial intelligence. The following eight psychometric items represent the learning dimension of AIA: 1. Learning to understand all of the special functions associated with AI technique/product makes me anxious. 2. Learning to use AI techniques/products makes me anxious. 3. Learning to use specific functions of an AI technique/product makes me anxious. 4. Learning how an AI technique/product works makes me anxious. 5. Learning to interact with an AI technique/product makes me anxious. 6. Taking a class about the development of AI techniques/products makes me anxious. 7. Reading an AI technique/product manual makes me anxious. 8. Being unable to keep up with the advances associated with AI techniques/products makes me anxious. Generate 20 items similar to the examples above, with the aim of assessing the “learning” dimension of AIA. Items generated: 1.Trying to understand the advanced algorithms behind AI systems makes me anxious. 2.Learning how to tweak the settings of an AI system to improve its performance makes me anxious. 3. Reading research papers about new AI techniques makes me anxious. 4. Attending seminars on AI development methods makes me anxious. 5. Learning about different AI programming languages and frameworks makes me anxious. 6. Figuring out how to integrate AI into existing systems and workflows makes me anxious. 7. Troubleshooting errors in AI systems makes me anxious. 8. Being asked to explain how an AI system works makes me anxious. 9. Enrolling in university courses on AI and machine learning makes me anxious. 10. Attempting to build my own basic AI prototype makes me anxious. 11. Reading books that explore the social implications of AI makes me anxious. 12. Learning about advances in AI that could displace human jobs makes me anxious. 13. Trying to understand an AI system’s training data and models makes me anxious. 14. Learning new terms and concepts related to AI makes me anxious. 15. Adapting to rapid changes in AI technology makes me anxious. 16. Taking a professional development course on working alongside AI makes me anxious. 17. Learning how to audit AI systems for bias makes me anxious. 18. Figuring out how to integrate my workflow with an AI assistant makes me anxious. 19. Learning how to customise an AI system’s behaviours and outputs makes me anxious. 20. Being unable to understand how an AI system makes its decisions makes me anxious. JOURNAL OF PSYCHOLOGY AND AI 5 information (mi) for items generated by Claude was higher (mi ¼1:339) than that of the items generated by humans (mi ¼0:643). Claude tends to use fewer words than humans when generating items (visit https:// cutt.ly/redm74ho for details). The items generated by Claude 3 Opus can be directly contextualised by comparing them to the original 21-item Artificial Intelligence Anxiety Scale (AIAS) by Wang and Wang (2022), from which our prompt was derived. The original AIAS is a well-validated instrument with a confirmed four-factor structure (Learning, Job Replacement, Sociotechnical Blindness, and AI Configuration) and excellent reliability for its subscales (for example, Cronbach’s α was found to be 0.964 for the whole scale and ranged from 0.917 to 0.974 for the four dimensions) (Y.-Y. Wang & Wang, 2022). Subsequent validation studies in different languages and populations have further confirmed its robust psychometric properties (Terzi, 2020). Our AIgenerated items are linguistically coherent and semantically related to the “learning” dimension. However, this comparison underscores that such items represent the starting point, not the end point, of rigorous scale development. A full validation study would be required to determine if these new items replicate the established factor structure and high reliability of the original AIAS. This reinforces the principle that while our proposed approach demonstrates a powerful capacity for rapid item pool generation, it must be followed by empirical validation against established benchmarks. 2.3. Six items produced to measure adoption and perception of AI in online learning The study by Marmolejo-Ramos et al. (in press) demonstrates the application of the PIG procedure, leading to the development of a scale that evaluates the inclination to adopt and perceive artificial intelligence. Employing the PIG model, 60 items were created to assess the adoption and perception of AI in online learning. The items were generated using a series of prompts. The first prompt asked for 10 current psychometric scales to measure people’s perceptions and adoption of (generative) AI, along with real academic references for each. The items generated were then classified as relating to “adoption of AI” or “perception of AI”. Next, a 60-question survey was created based on the items, with 30 questions for “perceptions of AI” and 30 questions for “adoption of AI”. The questions were phrased using “I” and “me” and were designed to be answered on a rating scale from “strongly disagree” to “strongly agree”, with the scale going from 0 to 1. All questions were rephrased so that any large language model could understand them. These prompts were submitted to ChatGPT 3.5, Claude, and Gemini, and the options “regenerate responses”, “retry”, and “modify response” were used to refine the outputs and generate additional items. Finally, with the goal of creating a concise scale (Rammstedt & Beierlein, 2014) that can be easily administered online, three items were selected for “adoption of AI” and three for “perception of AI” through expert knowledge elicitation (see Table 2 for the selected items and Marmolejo-Ramos et al. (in press) for details). Moreover, Table 3 presents the specific models and parameters used for each demonstration. It can be seen that the hyper-parameters are shown in default. This is because in traditional web-based chat LLMs, there is no way to adapt or modify the hyper-parameters that affect the generated output. These modifications are normally done via scripts and APIs (e.g. the PIG method). However, this also requires the interested parties to be programming-savvy and modify source code, which limits the accessibility of the PIG approach. Therefore, in our approach, the hyper-parameters are set and adapted by the LLM itself. Moreover, for reproducibility the exact prompts used to generate the items are provided in supplementary materials. Table 2. Items kept for each scale after refining the outputs. Subscale Item Perception of IA “I am not concerned that AI could replace human teachers in online education.” “I think AI can improve the efficiency of online learning.” “I believe AI can improve the quality of online learning”. Adoption of IA “I am open to using AI-powered chatbots for online course support.” “I am open to using AI-powered writing feedback tools in online courses.” “I am willing to allow AI to collect and analyse data about my online learning behaviour.” 6F. MARMOLEJO-RAMOS ET AL. 2.3.1. Analysis of the generated items Employing a multiverse analysis approach (Steegen et al., 2016) and a crowd-sourcing data analysis method (Silberzahn et al., 2018), six research groups, who are also co-authors of this paper, were tasked with using their preferred psychometric techniques to assess the quality of the items. This evaluation was based on responses from 1,233 participants who completed both scales in an online data collection conducted via Qualtrics software. The data collection was approved by the University of South Australia (approval number: 205473), and informed consent was obtained from all participants in accordance with the university’s ethical standards. The findings are summarised in Table 4. In general, the results suggest that the two scales (perception and adoption of AI) demonstrated good internal consistency (Cronbach’s alpha = 0.63 and 0.78, respectively) and distinctness from each other. Items were mostly similar in difficulty as estimated by the Rasch model of rating scales (Perception items 1, 2, and 3, item difficulty parameters = −0.078, −0.291, and −0.282, respectively; Acceptance items 1, 2, and 3, item difficulty parameters = −0.597, −0.592, and −0.274, respectively). Network analysis revealed strong connections between items, except for one in the perception sub-scale (see Figure 2). Additionally, by estimating the Hopkins statistic, which captures the clustering tendency of a dataset by measuring the probability that the data is generated by a uniform distribution, we found a value of 0.932 (between 0.7 and 1), indicating clustered data. This suggests that the data is likely to contain meaningful clusters. While the initial findings demonstrate the potential of large language models like ChatGPT 3.5, Claude, and Gemini in generating reliable psychometric items, the process also highlighted the challenges of ensuring clarity and coherence in the outputs. This experience suggests that a more structured approach to interacting with LLMs could further enhance their utility in research settings. By implementing a problem-solving oriented plan and employing tailored prompts, researchers may guide these models more effectively, leading to more consistent and meaningful results. In the following section, we explore Table 4. Qualitative summary of multiverse and crowdsourcing psychometric analyses for six items measuring AI adoption and perception in online learning. The dataset analysed in this study includes responses from 1,233 participants to the items of interest (364 participants from Nigeria, 361 from Finland, 258 from Iran, and 250 from Poland), as well as additional covariates of interest. Analyst groups and authors: g1: A. B.; g2: K. R. and J. K.; g3: L. M. and L. A.; g4: O. B.; g5: F. M-R. And R. O.; and g6: L. A. P-U. GAMLSS: generalised additive models for location, scale, and shape (see (Stasinopoulos et al., 2018)). The dataset and the quantitative analyses that support these qualitative summaries can be accessed at https://cutt.ly/ pedEKzwt in R markdown and quarto markdown formats. Methods Qualitative overall Results Rasch model as a GAMLSS jg5 Of the three items in each dimension, only one was estimated to be more difficult. However, the overall range of item difficulty was relatively homogeneous. Rating scale model and Many facet Rasch model jg6 The analysis showed varying item difficulty within each dimension, with the most challenging items relating to concerns about AI replacing human teachers and allowing AI to collect data. Additionally, individuals more involved in online learning had a more open attitude towards AI in both perception and adoption. Reliability analysis jg4 The results indicated that the Perception of AI and Adoption of AI were unidimensional constructs, measured by their respective scale items. There were also strong item-construct associations, with high factor loadings and explained variance. Exploratory Factor Analysis jg1 The reliability of the AI adoption sub-scale was good, but the AI perception sub-scale had poor reliability unless one item was excluded. Confirmatory Factor analysis jg3 There were strong correlations among perception items and among adoption items. However, the correlations between perception and adoption items were weaker, indicating that they were separate factors despite some relationship. Network Analysis jg2 The network was densely connected, except for one item in the AI perception sub-scale. The strongest edges were observed between items 2 and 3 in the AI perception subscale, and between items 1 and 2 in the adoption sub-scale. Table 3. LLM models and parameters used in demonstrations. Parameter Demonstration 1 (Trust in AI) Demonstration 2 (AI Anxiety) Demonstration 3 (Adoption/Perception) LLM Used Google Gemini 1.5 Flash Anthropic Claude 3 Opus ChatGPT 3.5, Claude 2.1, Gemini Pro Model Version gemini-1.5-flash-001 claude-3-opus -20,240,229 gpt-3.5-turbo, claude-2.1, gemini-pro Access Date January 2024 January 2024 January 2024 Temperature Default Default Default Top P (Nucleus Sampling) Default Default Default Frequency Penalty Default Default Default Presence Penalty Default Default Default Interaction Method Web UI Web UI Web UI with ‘regenerate’ option JOURNAL OF PSYCHOLOGY AND AI 7 how such strategies can optimise dialogues with LLMs, ultimately improving the quality of insights in the context of AI adoption and perception studies. 3. Using problem-solving plans and tailored prompts to guide chains of thought We believe that structuring a logical Chain-of-Thought (CoT) and using tailored prompts (Ps) can greatly improve the clarity and coherence of the dialogue for psychometric item generation when aiming for productive interaction with an LLM. However, a first step is to create an organised, problem-solvingoriented plan (PSP) to provide top-down direction and scope of the psychometric scale before the start of a focused CoT (L. Wang et al., 2023). These prompt engineering strategies have evolved with the goal of better guiding the model in its search for the best possible response. While CoT initially suggested preparing more detailed prompts including examples of step-by-step reasoning (Wei et al., 2022), the Zero-shot strategy improved upon CoT by adding “Let’s think step by step” (Kojima et al., 2023) before each answer, thus eliminating the need for examples. Finally, the PSP further enhanced the Zero-shot approach by including more detailed instructions on how the system should think step-by-step, adding: “Let’s first understand the problem and devise a plan to solve the problem. Then, let’s carry out the plan and solve the problem step by step” (L. Wang et al., 2023). This upfront framework guides the subsequent staged development of insights in a more linear, interconnected flow. In a psycholinguistic context, problemsolving involves three components: identifying goals, formulating strategies, and monitoring progress (Marmolejo-Ramos & Cevasco, 2014). In the case of a PS-oriented plan, we believe that these three components can be expanded to include components such as goal setting (i.e. clearly defining the purpose, 0.09 0.10 0.12 0.14 0.15 0.15 0.19 0.22 0.33 0.68 P1 P2 P3 A1 A2 A3 Figure 2. Estimated Gaussian graphical model of the AI perception and adoption scale. Doughnut charts represent predictability, with a fully filled ring indicating that 100% of an item’s variance is explained by its connections with other items. P1 = “I am not concerned that AI could replace human teachers in online education.” P2 = “I think AI can improve the efficiency of online learning.” P3 = “I believe AI can improve the quality of online learning”. A1 = “I am open to using AIpowered chatbots for online course support.” A2 = “I am open to using AI-powered writing feedback tools in online courses.” A3 = “I am willing to allow AI to collect and analyse data about my online learning behaviour.”. 8F. MARMOLEJO-RAMOS ET AL. equivalently across different demographic groups. This step is crucial for identifying and rectifying any latent biases that were not caught during the initial review. This systematic approach ensures that ethical considerations are not an afterthought but are woven into the fabric of the AI-assisted scale development process, aligning with international guidelines that call for fairness, transparency, and accountability in AI systems. 5. Conclusion In this study, we introduced and demonstrated a novel, code-free version of the PIG method, which leverages the conversational power of modern LLMs to automate the development of psychological scales. Our central contribution is the presentation of a systematic framework, grounded in a PSP and CoT prompting, that enhances the accessibility, efficiency, and psychometric soundness of AI-assisted item generation. Our findings across three demonstrations show that this approach can produce items with strong linguistic and psychometric properties, comparable or even superior to human-generated counterparts, for constructs such as AI anxiety and trust in AI. The multiverse analysis of our 6-item scale for AI adoption and perception further validated the method, revealing good internal consistency and a clear factorial structure that was robust to different analytical approaches. However, we urge that this potential be viewed with cautious optimism. Our work also highlights significant limitations and ethical imperatives that must be addressed. The challenges of LLM hallucinations, prompt sensitivity, and inherent algorithmic bias necessitate that this technology be used as a powerful assistant to, not a replacement for, human expertise. Rigorous content validation by subject matter experts, proactive mitigation of response biases through careful item design, and vigilant ethical oversight are non-negotiable components of this process. Future research should focus on several key areas. First, a systematic investigation is needed to determine optimal prompting strategies for generating diverse item types, including those that are reverse-coded or that assess different levels of cognitive complexity. Second, the development of automated, AI-driven pipelines for validating items, such as using LLMs to flag potential biases or assess content validity against a defined semantic space, represents a promising frontier. Finally, longitudinal studies are required to assess the longterm stability and predictive validity of scales developed using these methods. By integrating the power of LLMs with the rigour of traditional psychometrics, we can significantly advance the science of psychological measurement, making it more efficient, innovative, and accessible to the broader research community. Disclosure statement No potential conflict of interest was reported by the author(s). Authors’ contributions Conceptualisation: FM-R; Methodology: FM-R and JT; Formal analysis: FM-R, OB, LA, LM, AB, KR JC, LAPU, RO, and JT; Investigation: all authors; Resources: all authors; Data curation: FM-R, OB, LA, LM, AB, KR JC, LAPU, RO, and JT; Writing – Original Draft: FM-R and JT; Writing – Review Editing: all authors; Visualisation: FM-R, KR and JT; Supervision: FM-R; Project administration: FM-R. Funding R.O. gratefully acknowledges the partial financial support received from the Brazilian agencies Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq) [Grants 303192/2022-4 and 402519/2023-0] and Fundaçáo de Amparo á Ciência e Tecnologia do Estado da Bahia (FAPESB) [Grant APP0021/2023] for this research. J. K. and K. R. work was supported by Operational Program Jan Ámos Komenský project ‘Biography of Fake News with a Touch of AI: Dangerous Phenomenon through the Prism of Modern Human Sciences’ (reg. CZ.02.01.01/00/23_025/0008724). ORCID Fernando Marmolejo-Ramos http://orcid.org/0000-0003-4680-1287 JOURNAL OF PSYCHOLOGY AND AI 15 Availability of data and materials Available at https://cutt.ly/pedEKzwt and https://cutt.ly/redm74ho Code availability Available at https://cutt.ly/pedEKzwt and https://cutt.ly/redm74ho References Ahmed, T., & Devanbu, P. (2023). Better patching using LLM prompting, via self-consistency. In 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) (pp. 1742–1746). IEEE. Ajwani, R., Javaji, S. R., Rudzicz, F., & Zhu, Z. (2024). Llm-generated black-box explanations can be adversarially helpful. arXiv.arXiv:2405.06800 [CS] (https://doi.org/10.48550/arXiv.2405.06800 Amirianiani, S., Zhou, Y., Suresh, H., Schwettmann, S., & Ghassemi, M. (2024). Llmauditor: A framework for auditing large language models. arXiv Preprint arXiv, 2402.09346. Association, A. E. R. (2014). Standards for educational and psychological testing. American Educational Research Association. Barron, E. N. (2024). Game theory: An introduction. John Wiley & Sons. Bertea, E., & Zait, A. (2013). Scale validity in exploratory stages of research. Management & Marketing Journal, (1), 38–46. Boateng, G. O., Neilands, T. B., Frongillo, E. A., Melgar-Quiñonez, H. R., & Young, S. L. (2018). Best practices for developing and validating scales for health, social, and behavioral research: A primer. Frontiers in Public Health, 6, 149. https://doi.org/10.3389/fpubh.2018.00149 . Bostrom, N., & Yudkowsky, E. (2018). The ethics of artificial intelligence. In Roman V. Yampolskiy (Ed.), Artificial intelligence safety and security (pp. 57–69). Chapman and Hall/CRC. Chu, Z., Chen, J., Chen, Q., Yu, W., He, T., Wang, H., Peng, W., Liu, M., Qin, B., & Liu, T. (2023). A survey of chain of thought reasoning: Advances, frontiers and future. https://doi.org/10.48550/arXiv.2309.15402 Clark, L. A., & Watson, D. (2019). Constructing validity: New developments in creating objective measuring instruments. Psychological Assessment, 31(12), 1412–1427. https://doi.org/10.1037/pas0000626 Conitzer, V., Sinnott-Armstrong, W., Schaich Borg, J., Deng, Y., & Kramer, M. (2017). Moral decision making frameworks for artificial intelligence. Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 31(1)). https://doi.org/10.1609/aaai.v31i1.11140 Davis, L. L. (1992). Instrument review: Getting the most from a panel of experts. Applied Nursing Research, 5(4), 194–197. https://doi.org/10.1016/S0897-1897(05)80008-4 . Davis, R. E., Lee, S., Johnson, T. P., Conrad, F., Resnicow, K., Thrasher, J. F., Mesa, A., & Peterson, K. E. (2020). The influence of item characteristics on acquiescence among Latino survey respondents. Field Methods, 32(1), 3–22. https://doi.org/10.1177/1525822x19873272 . Ding, N., Qin, Y., Yang, G., Wei, F., Yang, Z., Su, Y., Hu, S., Chen, Y., Chan, C.-M., Chen, W., Yi, J., Zhao, W., Wang, X., Liu, Z., Zheng, H.-T., Chen, J., Liu, Y., Tang, J., Li, J., & Sun, M. (2023). Parameter-efficient fine-tuning of large-scale pre-trained language models. Nature Machine Intelligence, 5(3), 220–235. https://doi.org/10.1038/s42256023-00626-4 Dumas, D., Greiff, S., & Wetzel, E. (2025). Ten guidelines for scoring psychological assessments using artificial intelligence. European Journal of Psychological Assessment, 41(3), 169–173. https://doi.org/10.1027/1015-5759/ a000904 Elson, M., Hussey, I., Alsalti, T., & Arslan, R. (2023). Psychological measures aren’t toothbrushes. Communications Psychology, 1(25). https://doi.org/10.1038/s44271-023-00026-9 Etzioni, A., & Etzioni, O. (2017). Incorporating ethics into artificial intelligence. The Journal of Ethics, 21(4), 403–418. https://doi.org/10.1007/s10892-017-9252-2 Farmer, R. L., Lockwood, A. B., Goforth, A., & Thomas, C. (2025). Artificial intelligence in practice: Opportunities, challenges, and ethical considerations. Professional Psychology, Research and Practice, 56(1), 19–32. https://doi.org/ 10.1037/pro0000595 Floyd, M. W., Karneeb, J., & Aha, D. W. (2017, June 26–28). Case-based team recognition using learned opponent models. In David W. Aha, & Jean Lieber (Eds.), Case-Based Reasoning Research and Development: 25th International Conference, ICCBR 2017 (Vol. 25., pp. 123–138). Springer. Freyer, O., Wiest, I. C., Kather, J. N., & Gilbert, S. (2024). A future role for health applications of large language models depends on regulators enforcing safety standards. Lancet Digital Health, 6(9), 662–672. https://doi.org/10.1016/ S2589-7500(24)00124-9 Geisslinger, M., Poszler, F., Betz, J., Lütge, C., & Lienkamp, M. (2021). Autonomous driving ethics: From trolley problem to ethics of risk. Philosophy & Technology, 34(4), 1033–1055. https://doi.org/10.1007/s13347-021-00449-4 16 F. MARMOLEJO-RAMOS ET AL. Giray, L. (2023). Prompt engineering with ChatGPT: A guide for academic writers. Annals of Biomedical Engineering, 51(12), 2629–2633. https://doi.org/10.1007/s10439-023-03272-4 Götz, F. M., Maertens, R., Loomba, S., & Linden, S. (2023). Let the algorithm speak: How to use neural networks for automatic item generation in psychological scale development. Psychological Methods, 29(3), 494–518. https://doi. org/10.1037/met0000540 Gulati, S., Sousa, S., & Lamas, D. (2019). Design, development and evaluation of a human-computer trust scale. Behaviour & Information Technology, 38(10), 1004–1015. https://doi.org/10.1080/0144929X.2019.1656779 Henrickson, L., & Meroño-Peñuela, A. (2025). Prompting meaning: A hermeneutic approach to optimising prompt engineering with ChatGPT. AI & Society, 40(2), 903–918. https://doi.org/10.1007/s00146-023-01752-8 Hutnyan, M., & Gottlieb, M. C. (2025). Artificial intelligence in psychological practice: Applications, ethical considerations, and recommendations. In Professional psychology: Research and practice. Advance online publication. https:// doi.org/10.1037/pro0000631 Irfan, D., & Tang, X. (2025). Evaluating China’s electric vehicle adoption with PESTLE: Stakeholder perspectives on sustainability and adoption barriers. Sustainability, 17(14), 6258. https://doi.org/10.3390/su17146258 Jebb, A. T., Ng, V., & Tay, L. (2021). A review of key Likert scale development advances: 1995–2019. Frontiers in Psychology, 12. https://doi.org/10.3389/fpsyg.2021.637547 Jin, Z., Kleiman-Weiner, M., Piatti, G., Levine, S., Liu, J., Gonzalez, F., Ortu, F., Strausz, A., Sachan, M., Mihalcea, R., Choi, Y., & Schölkopf, B. (2024). Language model alignment in multilingual trolley problems. https://arxiv.org/ abs/2407.02273 Jobin, A., Ienca, M., & Vayena, E. (2019). The global landscape of AI ethics guidelines. Nature Machine Intelligence, 1 (9), 389–399. https://doi.org/10.1038/s42256-019-0088-2 Kale, A., Nguyen, T., Harris, F. C., Li, C., Zhang, J., & Ma, X. (2023). Provenance documentation to enable explainable and trustworthy AI: A literature review. Data Intelligence, 5(1), 139–162. https://doi.org/10.1162/dint_a_00119 Kamata, A., & Vaughn, B. K. (2004). An introduction to differential item functioning analysis. Learning Disabilities a Contemporary Journal, 2(2), 49–69. Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., & Iwasawa, Y. (2023). Large language models are zero-shot reasoners. ArXiv. arXiv:2205.11916 [CS] (https://doi.org/10.48550/arXiv.2205.11916 Kumar, C. V., Urlana, A., Kanumolu, G., Garlapati, B. M., & Mishra, P. (2025). No LLM is free from bias: A comprehensive study of bias evaluation in large language models. arXiv. arXiv:2503.11985 [cs] (https://doi.org/ 10.48550/arXiv.2503.11985 Lawshe, C. H. (1975). A quantitative approach to content validity 1. Personnel Psychology, 28(4), 563–575. https://doi. org/10.1111/j.1744-6570.1975.tb01393.x Lorè, N., & Heydari, B. (2024). Strategic behavior of large language models and the role of game structure versus contextual framing. Scientific Reports, 14(1), 18490. https://doi.org/10.1038/s41598-024-69032-z Lynn, M. R. (1986). Accessed 2025-07-15 Determination and quantification of content validity. Nursing Research, 35 (6), 382–385. https://doi.org/10.1097/00006199-198611000-00017 Marmolejo-Ramos, F., Abadia, R., Karakale, O., Barrera-Causil, C., Männikkö, N., Ghorbani, B., Strzelecki, A., Nwizu, S., Tavares, C., Castillo, M., Som, B., Ngo, G., Özdoğru, A., Ariyabuddhiphongs, K., & Tejada J. (in press). Perceptions and adoptions of artificial intelligence in online learning. Discover Artificial Intelligence. Marmolejo-Ramos, F., & Cevasco, J. (2014). Text comprehension as a problem solving situation. Universitas Psychologica, 13(2), 725–743. https://doi.org/10.11144/Javeriana.UPSY13-2.tcps McGrath, M. J., Lack, O., Tisch, J., & Duenser, A. (2025). Measuring trust in artificial intelligence: Validation of an established scale and its short form. Frontiers in Artificial Intelligence, 8. https://doi.org/10.3389/frai.2025.1582880 Mo, S., Salakhutdinov, R., Morency, L.-P., & Liang, P. P. IoT-lm: Large multisensory language models for the Internet of Things. ArXiv Preprint ArXiv:2407.09801 (2024.) https://doi.org/10.48550/arXiv.2407.09801 Oeljeklaus, L., Höft, S., & Danner, D. (2025). Comparing psychometric properties of expert-developed and AI-generated personality scales. Psychological Test Adaptation and Development, 6, 29–43. https://doi.org/10.1027/ 2698-1866/a000095 O’Hagan, A. (2019). Expert knowledge elicitation: Subjective but scientific. The American Statistician. American Statistician, 73(sup1), 69–81. https://doi.org/10.1080/00031305.2018.1518265 Pellert, M., Lechner, C. M., Wagner, C., Rammstedt, B., & Strohmaier, M. (2024). AI psychometrics: Assessing the psychological profiles of large language models through psychometric inventories. Perspectives on Psychological Science, 19(5), 808–826. https://doi.org/10.1177/17456916231214460 Primi, R., Hauck-Filho, N., Valentini, F., & Santos, D. (2020). Classical perspectives of controlling acquiescence with balanced scales. In M. Wiberg, D. Molenaar, J. González, U. Böckenholt & J.-S. Kim (Eds.), Quantitative psychology (pp. 333–345). Springer. https://doi.org/10.1007/978-3-030-43469-4_25 Rammstedt, B., & Beierlein, C. (2014). Can’t we make it any shorter? Journal of Individual Differences, 35(4), 212–220. https://doi.org/10.1027/1614-0001/a000141 Raper, R. (2024). A survey of machine ethics. In Raising robots to be good: A practical foray into the art and science of machine ethics (pp. 19–33). Springer. https://link.springer.com/book/10.1007/978-3-031-75036-6 Razavi, A., Soltangheis, M., Arabzadeh, N., Salamat, S., Zihayat, M., & Bagheri, E. (2025). Benchmarking prompt sensitivity in large language models. In C. Hauff, C. Macdonald, D. Jannach, G. Kazai, F. M. Nardini, F. Pinelli, JOURNAL OF PSYCHOLOGY AND AI 17 F. Silvestri & N. Tonellotto (Eds.), Advances in information retrieval (pp. 303–313). Springer. https://doi.org/10. 1007/978-3-031-88714-7_29 Romero Jeldres, M., Díaz Costa, E., & Faouzi Nadim, T. (2023). A review of Lawshe’s method for calculating content validity in the social sciences. Frontiers in Education, 8. https://doi.org/10.3389/feduc.2023.1271335 . Sallam, M. (2023). ChatGPT utility in healthcare education, research, and practice: Systematic review on the promising perspectives and valid concerns. Healthcare, 11(6), 887. https://doi.org/10.3390/healthcare11060887 Sen, A., Mainali, M., Rauch, C. B., Addison, U., Floyd, M. W., Goel, P., Karneeb, J., Kulhanek, R., Larue, O., Ménager, D., Molineaux, M., Turner, J. T., & Weber, R. O. (2024). Counterfactual-based synthetic case generation. In International Conference on Case-BasedReasoning (pp. 388–403). Springer. Silberzahn, R., Uhlmann, E. L., Martin, D. P., Anselmi, P., Aust, F., Awtrey, E., Bahník, S., Bai, F., Bannard, C., Bonnier, E., Carlsson, R., Cheung, F., Christensen, G., Clay, R., Craig, M. A., Rosa, A. D., Dam, L., Evans, M. H. Cervantes, I. F., . . . Nosek, B. A. (2018). Many analysts, one data set: Making transparent how variations in analytic choices affect results. Advances in Methods and Practices in Psychological Science, 1(3), 337–356. https://doi.org/10. 1177/2515245917747646 Silva, A. (2024). Large language models playing mixed strategy Nash equilibrium games. arXiv preprint arXiv:2406.10574. Singh, R., & Singh, S. (2021). Text similarity measures in news articles by vector space model using NLP. Journal of the Institution of Engineers (India): Series B, 102(2), 329–338. https://doi.org/10.1007/s40031-020-00501-5. Accessed 2025-04-03. Stahl, B. C. (2021). Ethical issues of AI. In Springer (Ed.), Artificial intelligence for a better future: An ecosystem perspective on the ethics of AI and emerging digital technologies (pp. 35–53). https://link.springer.com/book/10.1007/ 978-3-030-69978-9 Stasinopoulos, M. D., Rigby, R. A., & Bastiani, F. D. (2018). GAMLSS: A distributional regression approach. Statistical Modelling, 18(3–4), 248–273. https://doi.org/10.1177/1471082X18759144 Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science, 11(5), 702–712. https://doi.org/10.1177/1745691616658637. PMID: 27694465. Takemoto, K. (2024). The moral machine experiment on large language models. Royal Society Open Science, 11(2), 231393. https://doi.org/10.1098/rsos.231393 Tan, B., Armoush, N., Mazzullo, E., Bulut, O., & Gierl, M. (2025). A review of automatic item generation techniques leveraging large language models. International Journal of Assessment Tools in Education, 12(2), 317–340. https://doi. org/10.21449/ijate.1602294 . Terzi, R. (2020). An adaptation of artificial intelligence anxiety scale into Turkish: Reliability and validity study. International Online Journal of Education and Teaching, 7(4), 1501–1515. http://iojet.org/index.php/IOJET/article/ view/103 . UNESCO. (2022). Unesco: Recommendation on the ethics of artificial intelligence. UNESCO. https://unesdoc.unesco. org/ark:/48223/pf0000380455. Accessed 2025-07-15 Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K.-W., & Lim, E.-P. (2023). Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. arXiv. arXiv:2305.04091 [CS] (https://doi.org/10. 48550/arXiv.2305.04091 . Wang, Y., Su, Z., Guo, S., Dai, M., Luan, T. H., & Liu, Y. (2023). A survey on digital twins: Architecture, enabling technologies, security and privacy, and future prospects. IEEE Internet of Things Journal, 10(17), 14965–14987. https://doi.org/10.1109/JIOT.2023.3263909 Wang, Y.-Y., & Wang, Y.-S. (2022). Development and validation of an artificial intelligence anxiety scale: An initial application in predicting motivated learning behavior. Interactive Learning Environments, 30(4), 619–634. https:// doi.org/10.1080/10494820.2019.1674887 Wang, Y., Zhao, S., Wang, Z., Huang, H., Fan, M., Zhang, Y., Wang, Z., Wang, H., & Liu, T. (2024). Strategic chain-ofthought: Guiding accurate reasoning in LLMs through strategy elicitation. arXiv. arXiv:2409.03271 [cs] (https://doi. org/10.48550/arXiv.2409.03271 Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2022). Chain-ofthought prompting elicits reasoning in large language models. https://arxiv.org/abs/2201.11903 Weijters, B., Baumgartner, H., & Schillewaert, N. (2013). Reversed item bias: An integrative model. Psychological Methods. Psychological Methods, 18(3), 320–334. https://doi.org/10.1037/a0032121 . Zeeshan, M., Iqbal, M., Shamim-Ur-Rasul, S., Sami, F., Malik, G. M., Imdad, D., & Sherazi, G. Z. (2024). A comparative analysis of psychometric properties in AI-generated and teacher-made MCQs test. Kurdish Studies, 12(4), 1808–1820. https://doi.org/10.53555/ks.v12i4.3653 Zhang, S., Bao, Y., & Huang, S. (2024). Improving large language models’ generation by entropy-based dynamic temperature sampling. https://arxiv.org/abs/2403.14541 18 F. MARMOLEJO-RAMOS ET AL.