Psychiatrists' Experiences and Opinions of Generative AI: An Exploratory Online Mixed Methods Survey in Germany
Full text
Psychiatrists’ Experiences and Opinions of Generative AI: An Exploratory Online Mixed Methods Survey in Germany Authors Full name Title Mail Affiliation(s) ORCID Julian Schwarz MD [email protected] 1,2 0000-0001-7306-7909 Vera Goer MSc [email protected] 1,2 0009-0009-1904-7777 Olga Guzhova MD [email protected] 1,2 0009-0003-1593-4199 Lena Holtz BSc [email protected] 1,2 0009-0000-5610-9572 Felix Mühlensiepen PhD [email protected] 2,3 0000-0001-8571-7286 Charlotte Blease PhD [email protected] 4,5 0000-0002-0205-1165 1. Department of Psychiatry and Psychotherapy, Center for Mental Health, Immanuel Hospital Rüdersdorf, Brandenburg Medical School Theodor Fontane, Rüdersdorf, Germany 2. Faculty of Health Sciences Brandenburg, Brandenburg Medical School Theodor Fontane, Neuruppin, Germany 3. Center for Health Services Research, Brandenburg Medical School Theodor Fontane, Rüdersdorf, Germany 4. Department of Women’s and Children’s Health, Uppsala University, Uppsala, Sweden 5. Division of General Medicine Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, US Corresponding author(s): Julian Schwarz Department of Psychiatry and Psychotherapy, Center for Mental Health, Immanuel Hospital Rüdersdorf, Brandenburg Medical School Theodor Fontane Seebad 82/83, Rüdersdorf, DE, Germany Email: [email protected] 1
Abstract Background: Generative artificial intelligence is increasingly discussed as a tool to support psychiatric practice, particularly by improving clinical workflows. Despite growing interest, little is known about how psychiatrists use genAI and how they assess its usefulness, especially within the German healthcare system. Objective: This study examines the experiences and attitudes of psychiatrists in Berlin and Brandenburg toward genAI based chatbots in clinical practice, focusing on applications, perceived benefits, and concerns. Methods: An online mixed methods survey was conducted between September 19, 2024 and March 14, 2025. Psychiatrists working in public psychiatric hospitals, outpatient practices, and members of the German Association for Psychiatry, Psychotherapy, Psychosomatics and Neurology were invited. Of 11,754 contacted psychiatrists, 126 completed the survey, resulting in a response rate of 1.07 percent. The survey assessed sociodemographic characteristics, prior genAI experience, perceived effects on clinical work, and use cases such as documentation, diagnostics, and communication. Quantitative data were analyzed descriptively and qualitative responses using summarizing content analysis. Results: Slightly more than half of respondents, 52.0 percent, reported using genAI chatbots, most commonly ChatGPT. GenAI was primarily used for administrative tasks, especially writing medical letters and accessing medication information. Many participants reported potential benefits in reducing bureaucratic workload and improving documentation efficiency. At the same time, skepticism regarding clinical usefulness was common. Respondents highlighted clear limitations in diagnostic and therapeutic decision making and emphasized the need for targeted training. Language support was identified as a key advantage, particularly for non native German speakers. The limited sample size restricts generalizability. Conclusions: Psychiatrists expressed cautious optimism toward genAI, mainly for administrative support, while raising concerns about diagnostic use and insufficient training. These findings underline the need for evidence based integration of genAI in psychiatry, particularly in light of the EU AI Act. Trial Registration: This study was reviewed by the Ethics Committee of the Brandenburg Medical School, which issued a waiver of ethical approval (protocol no. 228072024-ANF). Keywords: Large language models; psychiatry; genAI chatbots; clinical practice; documentation; diagnostic support 2
1. Introduction Generative artificial intelligence (genAI) - defined as large language model (LLM) - based systems such as OpenAI’s GPT-4/5, Google Gemini, and Microsoft Copilot - are rapidly reshaping how clinicians across specialties think about day-to-day practice. Emerging studies show that these tools can summarise complex information, support clinical reasoning, translate or adjust communication for different audiences, generate drafts of correspondence or educational material, and interact conversationally with both clinicians and patients [1,2]. Such broad capability has positioned genAI as a potentially transformative aid in mental healthcare, where clinical work is deeply text-based, communication-intensive, and characterised by diagnostic uncertainty and high emotional stakes. Early research across medicine suggests that genAI may enhance the completeness of clinical histories, support hypothesis generation for differential diagnosis, assist with empathic phrasing, and streamline workflow in areas such as administrative correspondence and, importantly, clinical note-drafting [1–7]. Although these tools remain imperfect - and often error-prone - there is growing evidence that clinicians are experimenting with them across a wide range of tasks, from background research to patient education and decision support [1]. Mental healthcare appears to be an especially early domain of adoption. The conversational nature of LLMs aligns closely with the communicative core of psychiatric work, and uptake by patients has grown at an unprecedented rate. In October 2025, OpenAI reported that more than 1.2 million people per week were using its tools to disclose suicidality - an indicator of both public demand and the shifting digital landscape in which psychiatrists now practise. Against this backdrop, understanding clinicians’ own experiences and attitudes is increasingly urgent [3]. Yet despite intensifying interest, empirical research on mental health professionals’ use of genAI remains sparse. Existing studies involving general practitioners, psychologists, and other healthcare providers highlight both enthusiasm for potential benefits and concern regarding accuracy, bias, liability, and privacy risks. However, psychiatrists - who frequently manage complex histories, high-risk presentations, and sensitive patient data - are markedly under-studied. To date, the study by Blease et al. remains the only published research focused specifically on psychiatrists’ attitudes toward genAI [4]. That work provided valuable early insights but did not fully explore the breadth of clinical tasks for which psychiatrists may be adopting these tools, nor how their perspectives compare with those of other mental health or medical professionals. This gap is significant. GenAI systems can produce fluent yet inaccurate or misleading content, and they can encode or amplify gender, racial, disability, and other biases already documented in psychiatric care [8–10]. Privacy concerns are especially salient in mental healthcare, where entering identifiable or highly sensitive data into non-secure platforms risks breaches of confidentiality [11,12]. Understanding how psychiatrists are engaging with genAI - across documentation, diagnostic reasoning, patient interaction, communication, and administrative work - is therefore essential for developing safe, ethical, and evidence-based guidance for practice. Regulatory oversight remains in flux. In Europe, the Artificial Intelligence Act entered into force in 2024, with a staged roll-out extending until 2027 [5,6]. The Act introduces a risk-based framework that classifies AI systems used in healthcare as “high-risk,” imposing stringent requirements for transparency, data governance, and human oversight. However, it remains unclear how these regulations will be operationalized in psychiatric practice or how they will shape day-to-day clinical workflows. Complementary international work is also advancing rapidly, with proposals for structured implementation pathways, including education strategies and digital navigator support [7], as well as technical infrastructures for evaluating genAI tools in clinical mental health contexts [8]. For example, in Germany, 3
early applied research has begun to examine real-world use cases, including studies on genAI-assisted clinical documentation and ambient AI scribes in psychiatric and psychotherapeutic settings [9,10]. These emerging studies highlight both usability potential and persistent risks related to accuracy and privacy. Despite these policy developments, empirical evidence on psychiatrists’ engagement with genAI remains scarce. Existing studies are set in English-speaking countries. Only one recent APA-affiliated survey of U.S. psychiatrists found that while awareness was high, clinical use was limited and accompanied by concerns about accuracy, ethics and patient impact [4]. However, little is known about perspectives in other healthcare systems, where regulatory environments, technological adoption rates, and cultural attitudes toward digital innovation may differ. To our knowledge, there are currently no published mixed-methods surveys examining psychiatrists’ engagement with genAI in EU countries. Germany ,we argue, presents a highly relevant case, in the EU context, and one ripe for investigation. While efforts to digitalize healthcare are ongoing, adoption of genAI in psychiatry has been cautious, and the profession operates within strict data protection laws (e.g., GDPR) and a mixed public–private healthcare model. Understanding German psychiatrists’ experiences and attitudes about genAI can inform both national policy and broader international discussions about responsible genAI integration in mental healthcare. Therefore this exploratory mixed-methods study aimed to assess psychiatrists’ experiences with and attitudes toward genAI-based chatbots in psychiatric practice in Germany. 2. Methods 2.1. Subjects Participants in this online survey were recruited using a stepwise sampling approach. In the first step, we contacted public psychiatric hospitals (n = 8), psychiatric departments within general hospitals (n = 30), and outpatient psychiatric practices (n = 57) located in the federal states of Berlin and Brandenburg between September 19 and 25, 2024. These regions were selected to represent both urban (Berlin) and rural (Brandenburg) areas of Germany. Institutions were contacted via email and asked to disseminate the survey among their employed psychiatrists. Through this outreach, we aimed to reach approximately 840 psychiatrists. The study used an open convenience sampling strategy based on voluntary participation. In the second step, we collaborated with the German Association for Psychiatry, Psychotherapy, Psychosomatics and Neurology (DGPPN), the largest professional organization for psychiatry in Germany, with 11,754 members as of May 27, 2024. The DGPPN included a link to the survey in its monthly newsletter, distributed on December 18, 2024 and again on March 14, 2025. All invited participants were informed that their responses would remain anonymous and that no identifying information would be shared with the research team. Informed consent was obtained from all participants prior to their participation. The Ethics Committee of the Brandenburg Medical School reviewed the study and issued a waiver of ethical approval (protocol no. 228072024-ANF), as no personally identifiable or sensitive data were collected. Based on a pretest, survey completion took approximately four to five minutes. No compensation was provided for participation. 2.2. Procedures The online survey was based on an adapted version of a questionnaire originally developed by one of the authors [4]. It was translated into German and adapted to the specific conditions of the German psychiatric care system. To ensure content validity, three board-certified psychiatrists completed and reviewed the 4
survey, and their feedback was incorporated into a revised version. The core questions from the original version remained unchanged. The questionnaire consisted of two parts (see Appendix 1). In the first section, participants provided sociodemographic information. The second section included seven substantive questions. First, participants were asked whether they had prior clinical experience with genAI, such as ChatGPT, Google Gemini, Microsoft Bing AI, or other systems, using a multiple-choice format. Participants were also asked which specific tasks these AI tools had been used for. Then, participants rated the extent to which AI-powered chatbots might influence six core areas of psychiatric clinical work using a Likert scale. These areas included gathering patient information, diagnostic and prognostic accuracy, treatment planning, conveying empathy, and clinical documentation. Another question used a Likert scale to assess the degree of agreement about the potential impact of genAI use on caring for individuals with mental illness. Five subdomains were addressed: risk of patient harm, potential inequalities in care, genAI as substitute for medical treatment, training needs for clinicians, and potential efficiency gains in the healthcare system. All Likert scales offered four response options: "Strongly disagree," "Somewhat disagree," "Somewhat agree," and "Strongly agree." Each item also included the options "I don't know" and "No response.". An additional closed-ended question asked participants to estimate whether the clinical use of genAI might influence the likelihood of legal action against clinicians. Finally, two open-ended questions invited participants to share additional comments regarding the use of genAI in psychiatric practice and chatbots in general. 2.3. Analysis We used descriptive statistics to analyze the closed-ended questions regarding physicians' experiences with and attitudes toward genAI in psychiatry. These analyses were conducted using the survey software R (version 4.3.2; R Core Team, 2023) and Microsoft Excel (Microsoft 365, version 2408). Given the exploratory objective and limited sample size, only descriptive statistics were calculated. No inferential statistical models were used, as the study was not powered to detect subgroup differences or predictors. For Likert-scale items, response options 'I don’t know' and 'No response' were treated as missing values and excluded from percentage calculations. Denominators for each item are therefore based on valid responses only. All free-text responses (47 comments; 2,152 words) were analyzed using summarizing content analysis [11,12]. Due to the limited nature of the dataset – often consisting of brief phrases or sentence fragments – a full thematic analysis was deemed inappropriate [13]. Two researchers (JS, OG) independently coded all comments and developed an inductive codebook (see Appendix 2) through interactive comparison and consensus meeting. Because response length varied, our coding combined two procedures: we assigned a single primary code to each response; for longer responses that contained clearly separable thematic units, the response was segmented and each segment was coded according to a shared response code (for example, comment #150, was split into two code segments). Theme prevalence counts and all quantitative summaries therefore refer to coded segments rather than raw comments counts. Inter-rater agreement for the initial double-coding was substantial (percent agreement = 87%; Cohen´s k = 0.78); remaining disagreements were resolved through discussion. The final codebook comprised three main themes and eight subthemes (see Appendix 2). 3. Results 3.1. Participant characteristics A total of 126 psychiatrists completed the survey. Based on institutional outreach alone, this corresponds to an estimated minimum response rate of 1.07%. However, due to additional open recruitment via DGPPN newsletter, the true response rate cannot be precisely determined. 51 partially completed responses were 5
excluded from the analysis. Among the 126 respondents, 68 (54.0%) identified as male, 55 (43.7%) as female, 1 (0.8%) as non-binary, and 2 (1.6%) chose not to disclose their gender. See Table 1 for the complete sociodemographic details of both the subset of the sample that provided qualitative responses and the entire sample. Tab. 1. Sociodemographic characteristics of the qualitative subsample (n = 47) and the total study sample (N = 126) Parameter All Participants, n (%) Participants with Qualitative Responses, n (%) Specialist title* Psychiatry and psychotherapy Neurology Child and adolescent psychiatry and psychotherapy Psychosomatics and psychotherapy n = 111 (88.1%) n = 12 (9.5%) n = 7 (5.6%) n = 4 (3.2%) n = 39 (83.0%) n = 4 (8.5%) n = 3 (6.4%) n = 1 (2.1%) Psychotherapeutic specialty Cognitive behavioral therapy Psychodynamic psychotherapy Systemic psychotherapy Psychoanalysis n = 65 (51.6%) n = 43 (34.1%) n = 8 (6.4%) n = 4 (3.2%) n = 23 (48.9%) n = 18 (38.3%) n = 3 (6,.4%) n = 0 (0%) Training status/role Resident Specialist Consultant Chief physician n = 37 (29.4%) n = 27 (21.4%) n = 37 (29.4%) n = 16 (12.7%) n = 14 (29.8%) n = 8 (17.0%) n = 16 (34.0%) n = 5 (10.6%) Age <= 35 36 - 45 46 - 55 => 56 n = 34 (26.9%) n = 39 (30.9%) n = 26 (20.6%) n = 24 (19.1%) n = 12 (25.5%) n = 10 (21.3%) n = 11 (23.4%) n = 13 (27.7%) Year of medical license 1970s 1980s 1990s 2000s 2010s 2020s n = 3 (2.4%) n = 7 (5.6%) n = 26 (20.6%) n = 22 17.5%) n = 45 (35.7%) n = 20 (15.9%) n = 0 (0%) n = 5 (10.6%) n = 13 (27.7%) n = 5 (10.6%) n = 16 (34.0%) n = 7 (14.9%) Institution Hospital Office based practice n = 98 (77.8%) n = 15 (11.9%) n = 36 (76.6%) n = 6 (12.8%) Population area Metropolitan region Large city Medium-sized city Small town Village/town n = 66 (52.4%) n = 15 (11.9%) n = 19 (15.1%) n = 11 (8.7%) n = 4 (3.2%) n = 24 (51.2 %) n = 9 (19.2%) n = 5 (10.6%) n = 3 (6.4%) n = 0 (0%) Number of Patients Treated Annually per 6
Institution Don’t know < 1000 patients 1.001 - 5.000 patients 5.001 - 10.000 patients > 10.000 patients n = 30 (24.6%) n = 23 (18.9%) n = 33 (27.0%) n = 13 (10.6%) n = 23 (18.9%) n = 7 (15.9%) n = 9 (20.5%) n = 7 (15.9%) n = 11 (25.0%) n = 10 (22.7%) Treatment setting* Inpatient Day-care Outpatient Home Treatment n = 57 (45.2%) n = 17 (13.5%) n = 49 (38.9%) n = 16 (12.7%) n = 21 (44.7%) n = 6 (12.8%) n = 20 (42.6%) n = 3 (6.4%) * Multiple responses were allowed. Percentages for these categories may therefore exceed 100%. 3.2. Quantitative responses 3.2.1. Use of genAI Approximately half of the respondents (52%, n = 66) reported having used genAI tools at least once to answer clinical questions, with ChatGPT being by far the most commonly used tool (50%, n = 63). Other genAI chatbots mentioned by participants included Perplexity (n = 4), ClaudeAI (n = 2), You (n = 1), Mistral (n = 1), and Aya (n = 1). In total, 48% (n = 60) of respondents stated that they had not used genAI chatbots in psychiatric practice: See Fig. 2. Participants reported using genAI tools for a range of clinical and administrative purposes. The most common uses included obtaining information on medications and potential interactions (n = 33, 22%) and creating medical reports (n = 28, 19%). Other frequently reported applications were suggesting treatment options (n = 15, 10%), supporting differential diagnostic considerations (n = 21, 14%), creating progress entries (n = 15, 10%), and generating clinical summaries from prior documentation (n =18, 12%) (see Fig. 3). Fig. 2. Use of generative AI to answer clinical questions. Percentages reflect valid responses (missing data excluded). N may vary by item. 7
Fig. 3. Clinical tasks supported by genAI by participants. Percentages reflect valid responses (missing data excluded). Multiple responses were allowed. Percentages may therefore exceed 100%. 3.2.2. Opinions about effect on practice of genAI See Fig.4. Participants were generally positive about the potential impact of genAI on administrative and informational aspects of psychiatric work. The majority either agreed or somewhat agreed that genAI could improve medical reports (77%) and progress documentation (76%). However, only a small minority (17%) believed that genAI could improve the implementation of empathy in clinical care. Disagree Somewhat disagree Somewhat agree Agree Don’t know Will improve findings and medical reports 7 (5.8%) 9 (7.4%) 38 (31.4%) 55 (45.5%) 12 (9.9%) Will improve progress documentation 7 (5.8%) 11 (9%) 42 (34.4%) 51 (41.8%) 11 (9%) Will improve the implementation of empathy 68 (55.3%) 23 (18.7%) 19 (15.4%) 2 (1.6%) 11 (8.9%) Will improve the prognostic accuracy 24 (19.7%) 29 (23.8%) 39 (32%) 12 (9.8%) 18 (14.8%) Will improve the creation of personalized treatment plans 28 (22.6%) 15 (12.1%) 50 (40.3%) 17 (13.1%) 14 (11.3%) Will improve the diagnostic accuracy 20 (16.1%) 22 (17.7%) 51 (41.1%) 15 (12.1%) 16 (12.9%) Will improve the collection of patient information 12 (9.7%) 13 (10.5%) 47 (37.9%) 31 (25%) 21 (16.9%) Fig. 4. Opinions about effect on practice of generative AI. Percentages reflect valid responses (missing data excluded). N may vary by item. 8
3.2.3. Opinions about patient’s use of genAI See Fig.5. Regarding patients’ use of genAI tools, responses reflected mixed attitudes. A majority (70%) agreed that clinicians would require additional training to handle genAI use in clinical contexts, and 63% believed that patients may increasingly rely on genAI instead of consulting a physician. Concerns about potential risks were common, with 59% agreeing that genAI use could increase the risk of patient harm. Disagree Somewhat disagree Somewhat agree Agree Don’t know Will lead to the increase of efficiency in the healthcare system 8 (6.5%) 13 (10.6%) 37 (30.1%) 49 (39.8%) 16 (13%) Will result in clinicians needing more training to understand AI tools 3 (2.4%) 16 (12.8%) 26 (20.8%) 78 (62.4%) 2 (1.6%) Will lead to more patients relying on AI tools instead of seeing a doctor 11 (8.9%) 23 (18.5%) 52 (41.9%) 29 (23.4%) 9 (7.3%) Will increase the risk of inequality in care 26 (21.1%) 31 (25.2%) 30 (24.4%) 16 (13%) 20 (16.3%) Will increase the risk of patient harm 13 (10.6%) 24 (19.5%) 50 (40.7%) 23 (18.7%) 13 (10.6%) Fig. 5. Opinions about patient’s use of genAI. Percentages reflect valid responses (missing data excluded). N may vary by item. 3.2.4. Connection Between GenAI Use and Legal Action See Fig.6. Participants were uncertain about the legal implications of genAI use in psychiatric care. Approximately one quarter of respondents believed that genAI use would increase the likelihood of legal action against clinicians (n = 32; 25%), while a third expected no change (n = 41; 32%). Notably, more than one third of participants indicated uncertainty (n =48; 38%), suggesting that medico-legal risks of genAI remain unclear among psychiatrists. Fig. 6. Connection Between genAI Use and Legal Action. Percentages reflect valid responses (missing data excluded). Multiple responses were allowed. Percentages may therefore exceed 100%. 9
in psychiatric settings. The strong call for training underscores the urgent need for structured, evidence-based guidance. This transitional moment offers an opportunity for policymakers, professional bodies, and developers to work with clinicians to ensure that any future genAI adoption in psychiatry is both ethically sound and aligned with patient care priorities. Authors’ Contributions Charlotte Blease: Conceptualization, Supervision, Writing – original draft, Writing – review & editing. Vera Goer: Writing – review & editing. Olga Guzhova: Qualitative data analysis, Writing – original draft, Writing – review & editing. Lena Holtz: Data collection, quantitative data analysis, software, visualization, writing – original draft. Felix Mühlensiepen: Writing – review & editing. Julian Schwarz: Conceptualization, supervision, data collection, qualitative data analysis, resources, Writing - original draft, Writing – review & editing. Funding Statement The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was funded by the Brandenburg Medical School publication fund supported by the German Research Foundation and the Ministry of Science, Research and Cultural Affairs of the State of Brandenburg. Conflicts of Interest None declared. Data Availability Due to GDPR and ethics constraints, the original non-anonymized dataset and analysis scripts cannot be shared. The individual-level survey dataset has been fully anonymized by removing all free-text responses, geographic information, and any potentially identifying details to ensure participant privacy in accordance with GDPR and institutional ethics requirements. The anonymized dataset, along with the full survey instrument (German original version and English translation), is openly available in a persistent repository on the Open Science Framework (OSF) at: DOI: https://doi.org/10.17605/OSF.IO/WCEGD. Guarantor JS Acknowledgement The authors thank all participating psychiatrists for taking part in the survey and sharing their experiences and perspectives. We also acknowledge the support of the institutions and professional networks that assisted with the distribution of the survey. We would also like to thank Lena Holtz for editorial assistance in the submission process. 5. References 1. Yim D, Khuntia J, Parameswaran V, Meyers A. Preliminary evidence of the use of generative AI in health care clinical services: Systematic narrative review. JMIR Med Inform JMIR Publications Inc.; 2024 Mar 20;12(1):e52073. PMID:38506918 2. Brodeur PG, Buckley TA, Kanjee Z, Goh E, Ling EB, Jain P, Cabral S, Abdulnour R-E, Haimovich 16
AD, Freed JA, Olson A, Morgan DJ, Hom J, Gallo R, McCoy LG, Mombini H, Lucas C, Fotoohi M, Gwiazdon M, Restifo D, Restrepo D, Horvitz E, Chen J, Manrai AK, Rodman A. Superhuman performance of a large language model on the reasoning tasks of a physician. arXiv [csAI]. 2024. Available from: http://arxiv.org/abs/2412.10849 3. Jamali L. OpenAI shares data on ChatGPT users with suicidal thoughts, psychosis. BBC BBC News; 2025 Oct 27; Available from: https://www.bbc.com/news/articles/c5yd90g0q43o [accessed Nov 17, 2025] 4. Blease C, Worthen A, Torous J. Psychiatrists’ experiences and opinions of generative artificial intelligence in mental healthcare: An online mixed methods survey. Psychiatry Res Elsevier BV; 2024 Mar;333(115724):115724. PMID:38244285 5. Smuha NA. Regulation 2024/1689 Eur. Parl. & Council. International Legal Materials Published online 2024;2025:1–148. doi: 10.1017/ilm.2024.46 6. EU AI Act: first regulation on artificial intelligence. Topics | European Parliament. Available from: https://www.europarl.europa.eu/topics/en/article/20230601STO93804/eu-ai-act-first-regulation-on-arti ficial-intelligence [accessed Aug 19, 2025] 7. Torous J, Ledley KT, Gorban C, Strudwick G, Schwarz J, Choudhary S, Emerson M, Patriquin M, Dempsey A, Bantjes J, Ospina-Pinillos L, Hornick J, Kochhar S. Accelerating digital mental health: The society of digital psychiatry’s three-pronged roadmap for education, digital navigators, and AI. (preprint). JMIR Ment Health JMIR Publications Inc.; 2025 Sep 20; doi: 10.2196/84501 8. Dwyer B, Flathers M, Sano A, Dempsey A, Cipriani A, Gazi AH, Gorban C, Rodriguez CI, Stromeyer C IV, King D, Rozenblit E, Strudwick G, Linardon J, Cheong J, Firth J, Herpertz J, Schwarz J, Emerson M, Paulus MP, Patriquin M, Hua Y, Choudhary S, Siddals S, Pinillos LO, Bantjes J, Scheuller S, Xu X, Duckworth K, Gillison DH, Wood M, Torous J. MindBenchAI: An actionable platform to evaluate the profile and performance of large language models in a mental healthcare context. arXiv [csHC]. 2025. doi: 10.48550/arXiv.2510.13812 9. Ozkara Menekseoglu P, Weibezahl M, Ellingsen M, Sterkenburg J, Kharko A, Hochwarter S, Schwarz J. Errors in generative AI-based patient-centered mental health documentation: Psychiatrists’ qualitative pre-post comparison (preprint). JMIR Preprints. 2025. doi: 10.2196/preprints.78351 10. Goer V, Schwarz J. Ambient AI Scribes in Psychotherapy Documentation: A Qualitative Study on Psychotherapists' User Experiences (preprint). SpringerNature. Die Psychotherapie 2025. 11. Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, Faix DJ, Goodman AM, Longhurst CA, Hogarth M, Smith DM. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med 2023 Jun 1;183(6):589–596. PMID:37115527 12. Kharko A, McMillan B, Hagström J, Muli I, Davidge G, Hägglund M, Blease C. Generative artificial intelligence writing open notes: A mixed methods assessment of the functionality of GPT 3.5 and GPT 4.0. Digit Health 2024 Jan;10:20552076241291384. PMID:39493632 13. Patton MQ. Qualitative Research and Evaluation Methods. 3rd ed. Thousand Oaks, Calif: Sage Publications Ltd; 2002. Available from: https://www.amazon.de/Qualitative-Research-Evaluation-Methods-Michael/dp/0761919716 14. Haupt CE, Marks M. AI-generated medical advice-GPT and beyond. JAMA American Medical Association (AMA); 2023 Apr 25;329(16):1349–1350. PMID:36972070 15. Cross S, Bell I, Nicholas J, Valentine L, Mangelsdorf S, Baker S, Titov N, Alvarez-Jimenez M. Use of AI in mental health care: Community and mental health professionals survey. JMIR Ment Health 2024 Oct 11;11:e60589. PMID:39392869 16. Blease C, Hagström J, Sanchez CG, Kharko A, McMillan B, Gaab J, Brulin E, Locher C, Hägglund M, 17
Riggare S, Mandl KD. General practitioners’ experiences with generative artificial intelligence in the UK: An online survey. Research Square. 2025. doi: 10.21203/rs.3.rs-6196250/v1 17. Kharko A, Garcia Sanchez C, Hagström J, Gaab J, Locher C, McMillan B, Sundemo D, Blease C. General practitioners’ opinions of generative artificial intelligence in the UK: An online survey. Digit Health 2025 Jul 17;11:20552076251360863. PMID:40688576 18. Blease C, Rodman A. Generative artificial intelligence in mental healthcare: An ethical evaluation. Curr Treat Options Psychiatry Springer Science and Business Media LLC; 2024 Dec 9;12(1). doi: 10.1007/s40501-024-00340-x Appendix 1. Survey instrument General Information S1 - What is your area of expertise - Psychiatry and psychotherapy - Child and adolescent psychiatry and psychotherapy - Psychosomatics and psychotherapy - Neurology S2 - What is your psychotherapeutic specialty? - Behavioral therapy - Depth psychology-based psychotherapy - Psychoanalysis - Systemic psychotherapy S3 - Which of the following best describes your medical role/function? - Assistant doctor 18
- Medical specialist - Senior physician - Chief physician S4 - In which institution do you work? - Clinic - Office based practice S5 - In which setting do you mainly work? - Outpatient - Day-care - Inpatient - Outreach (home treatment) D2 - Gender - are you… - female - male - diverse/other D3 - Age - are you… - 35 or younger - 36 - 45 - 46 - 55 - 55 or older D4 - When did you qualify as a doctor (state examination)? - 1960 (1960) - …. - 2024 (2024) D5 - Which of the following best describes the size of your supply region? - Metropolitan region (< 100,000 inhabitants, e.g. Berlin, Munich) - Large city (at least 100,000 inhabitants, e.g. Bonn, Wuppertal) - Medium-sized city (< 20,000 inhabitants, e.g. Solingen, Cottbus) - Small town (> 20,000 inhabitants, e.g. Tegernsee, Heiligenhafen) - Village/town D6 - How many patients are treated by your healthcare facility each year (please estimate!)? - up to 500 patients - 501 - 1.000 patients - 1.001 - 2.500 patients - 2.501 - 5.000 patients - 5.001 - 7.500 patients - 7.501 - 10.000 patients - 10.001 - 12.500 patients - 12.501 patients or more 19
- i don’t know __________________________________________________________________________________________ The following questions relate to your expectations and any experience with ChatGPT (or with Google Gemini, Microsoft Bing AI or similar) Q1 - Have you ever used any of the following AI tools in clinical practice? - ChatGPT - Google Gemini - Microsoft Bing AI - None Q1a - What do you use AI tools for in your clinical practice? - Suggesting differential diagnoses - Suggesting treatment options - Information on medication (interactions) - Creation of progress entries according to patient contacts - Creation of medical reports - Creation of summaries from previous documentation, findings and/or medical reports Q2 - Please think about how the following tasks can be influenced by ChatGPT/Google Gemini/Bing AI. Indicate how strongly you agree or disagree with the following statements. "The use of these AI tools improves... Disagree Somewhat disagree Somewhat agree Agree Don’t know findings and medical reports progress documentation the implementation of empathy the prognostic accuracy the creation of personalized treatment plans the diagnostic accuracy the collection of patient information Q3 -Indicate how much you agree or disagree with the following statements about the use of these AI tools by doctors and patients. "The use of these AI tools... Disagree Somewhat disagree Somewhat agree Agree Don’t know will increase the risk (e.g. due to AI-related misinformation) on the part of patients Will increase the risk of inequality in care Will lead to more patients 20
relying on AI tools instead of seeing a doctor WIll result in clinicians needing more training/support to understand AI tools Will lead to the increase of efficiency in the healthcare system Q4 - Please think about how your practice might be affected by ChatGPT/Google Gemini/Bing AI. "In your opinion, these AI tools will... - reduce the risk of legal action being taken against me - increase the risk of legal action being taken against me - do not influence any risks - I don't know Q5 - Please post further comments on the use of ChatGPT/Google Gemini/Bing AI in clinical practice. _________________________________ Q6 - Please add further comments on the use of ChatGPT/Google Gemini/Bing AI in general. _________________________________ Appendix 2. Codebook Main Theme Subtheme Description Illustrative Quotation Frequency (n) % of Total Comments (N = 47) 1. Applicatio ns Improving Documentati on and Reducing Bureaucracy Mentions of administrative relief, efficiency gains, or AI use for medical reports and correspondence “The administrative tasks alone would be made much easier.” (#68, F, No) 18 38 % Diagnostic and Therapeutic Support AI as tool for differential diagnosis, treatment planning, or information retrieval “Helpful for differential diagnostic considerations.” (#147, M, Yes) 9 19 % 21
Enhancing Communicati on and Collaboratio n Linguistic or collaborative support, esp. for non-native speakers “Linguistic support in preparing and correction of reports.” (#128, M, Yes) 7 15 % 2. Challenge s and Limitation s Implementati on Barriers Structural, organizational, or technical constraints (e.g. IT restrictions) “Our hospital IT even blocks ChatGPT.” (#54, M, Yes) 10 21 % AI Constraints in Diagnostic and Therapeutic Practice Concerns about misinformation, hallucinations, loss of critical reflection “ChatGPT invents sources that do not exist.” (#64, M, Yes) 12 26 % 3. Data Privacy and Ethical Risks Data Security and Patient Confidentiali ty Data protection, anonymization, non-EU data transfers “A topic of utmost relevance would be data protection when transferring patient cases, especially when transferring them to AI outside the EU. ” (#163, M, No) 11 23 % Dehumanizat ion and Doctor–Patie nt Relationship Fears of reduced empathy, loss of human connection or training quality “I fear a deterioration in the doctor–patient relationship.” (#81, F, No) 7 15 % Potential Environment al Impacts A perspective on the potential risks and benefits associated with the use of genAI “The use of AI consumes considerable amounts of energy and water.” (#122, F, No), 2 4% 22