scieee AI-readable full text Open interactive document viewer

Informatics Research Evaluation, 2025 Revised Report

Informatics Europe; Baquero, Carlos; CARRO LIÑARES, MANUEL; Hitz, Martin; Orponen, Pekka; Peroni, Silvio; Rossmanith, Peter; Goedicke, Michael; Nardelli, Enrico; Paradinas, Pierre; Teniente, Ernest; Trautmann, Heike

Abstract

Evaluation enhances research quality and impact but must follow established principles, discipline-specific criteria, and responsible methodologies to be effective. This report expands on the 2008 and 2018 Informatics Europe reports, aligning with the CoARA Agreement on Reforming Research Assessment (CoARA, 2022). It provides updated recommendations on key topics such as responsible bibliometrics, the evaluation of artefacts, Open Science, interdisciplinary research, and the role of AI in assessment. The report underscores the unique nature of Informatics, the importance of selective conferences, the growing relevance of open archives and overlay journals, and the necessity of transparent, criteria-driven evaluation. It also highlights that quantitative metrics should never replace expert judgment and that AI should enhance, not replace, human evaluators. This update guides researcher evaluations while ensuring fairness, accuracy, and alignment with the evolving Informatics research landscape.

Full text

INFORMATICS RESEARCH EVALUATION, 2025 REVISED REPORT An Informatics Europe report endorsed by National Informatics Association members of Informatics Europe Informatics Research Evaluation, 2025 Revised Report An Informatics Europe Report Prepared by the Research Evaluation Recommendations Panel of Informatics Europe. Endorsed by National Informatics Associations of Austria, France, Germany, Italy, Netherlands, Spain, Switzerland and United Kingdom. Editors: ● Carlos Baquero, Universidade do Porto and INESC TEC, Portugal ● Manuel Carro, Universidad Politécnica de Madrid and IMDEA Software Institute, Spain, and Informatics Europe ● Martin Hitz, Alpen-Adria-Universität Klagenfurt and Informatik Austria ● Pekka Orponen, Aalto University, Finland, and Informatics Europe ● Silvio Peroni, University of Bologna and GRIN – Gruppo di Informatica, Italy ● Peter Rossmanith, RWTH Aachen University and Fakultätentag Informatik, Germany Advisers: ● Michael Goedicke, Universität Duisburg-Essen and Gesellschaft für Informatik, Germany ● Enrico Nardelli, Università di Roma “Tor Vergata”, Italy, and Informatics Europe ● Pierre Paradinas, Conservatoire National des Arts et Métiers and Société Informatique de France ● Ernest Teniente, Universitat Politècnica de Catalunya and Sociedad Cientifica Informática de España, Spain ● Heike Trautmann, Paderborn University and Fakultätentag Informatik, Germany © Informatics Europe, 2025, CC BY-SA 4.0|Page 1 Informatics Research Evaluation, Revised Report March 2025 Published by: Informatics Europe Binzmhlestrasse 14/54 8050 Zurich, Switzerland www.informatics-europe.org [email protected] © Informatics Europe, 2025, CC BY-SA 4.0 Other Informatics Europe Reports ● Survey about Diversity and Inclusion Initiatives (2024, Elisabetta Di Nitto, Antinisca Di Marco and the Informatics Europe Diversity & Inclusion Working Group) ● Proceedings of the 1st Early Career Researchers Workshop at ECSS 2021 (2021, Elisabetta Di Nitto and Standa Živný) ● Bridging the Digital Talent Gap: Towards Successful Industry-University Partnerships (2020, Enrico Nardelli, Cristina Pereira and rapporteurs in Artificial Intelligence, Cyber Security and Software Engineering, with the support of the European Commission, DG CONNECT) ● Informatics Education in Europe: Institutions, Degrees, Students, Positions, Salaries — Key Data 2013-2018 (2019, Svetlana Tikhonenko, Cristina Pereira) ● Ethical/Social Impact of Informatics as a Study Subject in Informatics University Degree Programs. (2019, Paola Mello, Enrico Nardelli, with contribution from the Working Group Members) ● The Wide Role of Informatics at Universities. (2019, Elisabetta Di Nitto, Susan Eisenbach, Inmaculada García Fernández, Eduard Gröller) ● Industry Funding for Academic Research in Informatics in Europe. Pilot Study. (2018, Data Collection and Reporting Working Group of Informatics Europe) ● Informatics Education in Europe: Institutions, Degrees, Students, Positions, Salaries — Key Data 2012-2017 (2018, Svetlana Tikhonenko, Cristina Pereira) ● Informatics Research Evaluation (2018, Research Evaluation Working Group of Informatics Europe) ● Informatics for All: The Strategy (2018, Michael E. Caspersen, Judith Gal-Ezer, Andrew McGettrick, Enrico Nardelli. Joint report with ACM Europe) ● When Computers Decide: Recommendations on Machine-Learned Automated Decision Making (2018, James Larus, Chris Hankin, Siri Granum Carson, Markus Christen, Silvia Crafa, Oliver Grau, Claude Kirchner, Bran Knowles, Andrew McGettrick, Damian Andrew Tamburri, Hannes Werthner Joint Report with ACM Europe) ● Informatics Education in Europe: Are We All In The Same Boat? (2017, The Committee on European Computing Education. Joint report with ACM Europe) All reports can be obtained from Informatics Europe at: www.informatics-europe.org © Informatics Europe, 2025, CC BY-SA 4.0|Page 2 Executive Summary Evaluation is an indispensable instrument for improving research quality and impact. To achieve the intended effects, research evaluation should follow established and widely accepted principles, be benchmarked against appropriate criteria, and be sensitive to disciplinary differences. This report addresses the principles and criteria that are to be followed when individual researchers are evaluated for their research in the field of Informatics. This report builds on and updates the outcomes of the 2008 and 2018 Informatics Europe reports on Research Evaluation for Computer Science/Informatics, while aligning its recommendations with other recent documents on research evaluation, most notably the CoARA Agreement on Reforming Research Assessment (CoARA 2022). The report also contains updated analyses and recommendations in four topical areas of concern in Informatics: the responsible use of bibliometrics and credit assignment in contributions, assessing artefacts, Open Science, and interdisciplinary research, together with a discussion on the role of AI in research evaluation. Our key messages are the following: 1. Informatics is an original discipline that combines aspects of mathematics, science, and engineering. Researcher evaluation must recognise and respect its specificity. 2. A distinctive feature of publication in Informatics is the importance of highly selective conferences. Journals have complementary advantages but do not necessarily carry more prestige. Publication models that couple conferences and journals, where the papers of a conference are published directly in a journal, are a growing trend that may bridge the current gap between these two forms of publishing. 3. Open archives and overlay journals are recent innovations in the Informatics publication culture that offer improved tracking in evaluation. 4. The impact of artefacts such as software, open datasets, and other research products such as trained machine learning models can be as great as publications. The evaluation of such objects, which is now conducted by many conferences, should be encouraged and accepted as an established component of research assessment. Another important indicator of impact is advances that lead to commercial exploitation or adoption by industry or standardisation bodies. 5. Open Science and its research evaluation practices are highly relevant to Informatics. Informatics has played a key enabling role in the Open Science revolution and should remain at its forefront. 6. Numerical measurements such as citation and publication counts must never be used as the sole evaluation instrument. They must be filtered through human interpretation, specifically to avoid errors, and complemented by peer review and assessment of outputs other than publications. In particular, numerical measurements must not be used to compare researchers across scientific disciplines, including across subfields of Informatics. 7. In Informatics, the order of authors often holds little significance and varies across subfields. Without clear guidelines, it should not be a factor in researcher evaluation. Instead, authors should be encouraged to clearly state the scope and role of their individual contributions to multi-author works. 8. In assessing institutions, researchers, publications, and citations, the use of open research information provided by Open Science infrastructures should be favoured and supported. When using ranking and benchmarking services provided by for-profit companies, respect for open access criteria is mandatory. Journal-based or journal-biased ranking services are inadequate for most of Informatics and must not be used. 9. Any evaluation, especially quantitative, must be based on clear, published criteria. Furthermore, assessment criteria must themselves undergo assessment and revision. 10. Any use of generative AI in research evaluation should increase the quality of the assessments and reduce the effort of the human evaluators. AI must not be used to reduce the number of human experts in assessment panels and their collective responsibility for the panels’ recommendations. © Informatics Europe, 2025, CC BY-SA 4.0|Page 3 Table of Contents 1. Research Evaluation ................................................................................................................................. 4 2. Informatics and Its Specificity .................................................................................................................. 5 2.1 Characteristics of Informatics ............................................................................................................ 5 2.2 The Informatics publication culture and its evolution....................................................................... 5 3. Research Evaluation for Increased Quality and Impact ........................................................................... 6 3.1 Assessing the quality of research....................................................................................................... 6 Other indicators for quality ...................................................................................................................... 7 3.2 Assessing the impact of research....................................................................................................... 7 4. Responsible Use of Indicators and Credit Assignment in Contributions ................................................. 8 5. Assessing Artefacts .................................................................................................................................. 9 5.1 Software artefacts ............................................................................................................................ 10 5.2 Why evaluate software systems ...................................................................................................... 10 5.3 Caveats ............................................................................................................................................. 11 5.4 Recommendations ........................................................................................................................... 12 6. Open Science .......................................................................................................................................... 13 7. Interdisciplinary Research ...................................................................................................................... 15 8. The Role of AI in Research Evaluations .................................................................................................. 16 9. Conclusions ............................................................................................................................................ 17 Acknowledgements.................................................................................................................................... 17 Endorsements ............................................................................................................................................ 17 References ................................................................................................................................................. 18 © Informatics Europe, 2025, CC BY-SA 4.0|Page 4 1. Research Evaluation Evaluation is innately linked with research. The work of researchers is subject to evaluation from the very beginning of their careers. Research results submitted for assessment are subject to a peer review process that scrutinises their scientific quality and potential for impact. The researchers themselves are constantly being evaluated: for recruitment, for promotion, for funding of research proposals, when being considered for specific roles in committees, for recognition by awards, etc. Researchers also often act as evaluators of their peers. By and large, these evaluation processes are internal to the research community and aim to guarantee its internal fairness and integrity through self-regulation. Increasingly, evaluation is also mandated and regulated by exogenous entities and scaled from individuals to entire institutions. In many cases, governments have developed national evaluation standards and processes. Often, their end results are institutional rankings. The main motivation for these efforts is to guarantee that taxpayers’ money is spent on research efficiently and leads to societal benefits. Research evaluation can thus be performed for different goals. It can target a specific piece of research (documented by a single or several artefacts), an individual researcher, a research group, or an organisational unit (e.g., a department). It may even generalise to entire institutions (e.g., universities) or even countries. In any case, the goal of evaluation is to assess some explicit or implicit notion of value, or quality, of research. To achieve the positive objectives of research evaluation, the specific goals of any such evaluation effort must be clearly stated, and the way it is conducted must be aligned with these goals. The evaluation must follow established principles and practical criteria, known and shared by evaluators and researchers, and take into account any specificities of the scientific field and area involved. Evaluation can have a tremendously positive effect in improving research quality and productivity. It is vital to recognise and support research that can lead to advances in knowledge and impact on society. At the same time, the effect of following ill-conceived criteria or practices in research evaluation, or misusing metrics and indicators, can have seriously negative long-term effects. In particular, it may greatly damage the potential of future generations of researchers. Furthermore, very frequent evaluation can have a detrimental effect on fundamental research, as the assessment may then accentuate low-risk, short-term developments at the expense of potentially high-gain, long-term work. This report focuses on the main principles and criteria that should be followed when individual researchers 1 are evaluated for their research activity in the field of Informatics, 2 addressing the specificities of this area. This subsumes the evaluation of a specific piece of research and can, to some extent, be generalised to departments since their research performance is largely determined by their individuals. This report reasserts the recommendations in previous Informatics Europe reports on the subject (Informatics Europe 2008, 2018), to which it incorporates a number of observations concerning the topical areas of bibliometrics and credit assignment in contributions, assessing artefacts, Open Science, interdisciplinary research, and the role of AI in research evaluation. An important development since the publication of the 2018 report has been the convergence of several reports and initiatives on revising research evaluation practices into the CoARA Agreement on Reforming Research Assessment (CoARA 2022). The recommendations in this report have been reviewed and updated to align with the CoARA recommendations. 1 Some aspects of department evaluation are addressed in the publication (Informatics Europe 2013). 2 Interchangeably “Computer Science” or “Computing”. © Informatics Europe, 2025, CC BY-SA 4.0|Page 5 2. Informatics and Its Specificity 2.1 Characteristics of Informatics Informatics is a relatively young science that is rapidly evolving in close connection with technology. Beyond the two basic and universal research pillars of theory and experimentation, which in Informatics research are often both present in varying proportions, a third important component in Informatics is the creation of artefacts, which provide new designs, tools or otherwise improved support for information processing tasks. Informatics research covers an extremely wide and methodologically diverse family of areas: from the development of new computing devices to the mathematical theory of algorithms and complexity, from human factors to big data, machine learning, and artificial intelligence, from studies of programmers’ productivity to secure encryption methods. Interdisciplinarity, as in the case of bioinformatics, medical informatics, geo-informatics, or cognitive science, brings in even more diversity. With the continuing digitalisation of society and the emergence of potentially disruptive computational technologies such as artificial intelligence and quantum computing, Informatics has an extremely high societal and economic impact. Measuring up to these responsibilities not only requires continuous progress in Informatics research, but also calls for more emphasis on Informatics education and the influence of Informatics on society, including ethical concerns. Research in Informatics, as in any other science, must be evaluated according to criteria that take into account the field’s specificity. Universal criteria do not exist for evaluating research quality. This is also true for the different subfields of Informatics. Differences between fields must be taken into account, and the temptation to adopt simplistic, one-size-fits-all criteria should be resisted. 2.2 The Informatics publication culture and its evolution The publication culture within Informatics differs from most other fields in the prominent role played by conference publications. Many subfields of Informatics have leading conferences with status, visibility, and impact comparable to or higher than their respective leading journals. Conference papers at these meetings undergo a highly selective peer-review process that makes these venues very competitive and often leads to lower acceptance rates than in the best journals. Conferences in Informatics provide a faster turnaround time than journals to publish research results, get feedback from peers and build upon it, and also often have higher standards of novelty than journals. These factors are crucial in a rapidly evolving field like Informatics, and as a result, in Informatics, journals do not necessarily carry more prestige than conferences. Journal publications are, of course, also important, especially for gathering and reworking antecedent research into an archival presentation that is carefully prepared, has guaranteed long-term availability, and is free of the space limitations imposed by conferences. Towards bridging the dichotomy between conferences and journals, new alternatives are now changing the publication culture: ● Coupled conferences and journals, implemented in different ways: VLDB-style with continuous submission to the journal and presentation of the accepted papers at the conference; ICLP-style with the proceedings of the conference being published as a special issue of a standard journal (in this case TPLP); and its variant PACMPL-style, 3 where a dedicated journal is used to publish 3 VLDB: "International Journal on Very Large Databases"; ICLP: "International Conference on Logic Programming"; TPLP: "Journal on Theory and Practice of Logic Programming"; PACMPL: "Proceedings of the ACM on Programming Languages". © Informatics Europe, 2025, CC BY-SA 4.0|Page 6 proceedings of several different, related conferences. These hybrid combinations of conferences and journals are a promising and growing trend that combines the advantages of timely publication of conferences with the impact tracking of journals (Dagstuhl 2012). ● Open Archives (arXiv, HAL, Zenodo etc.) provide opportunities to publish first versions of papers and protect intellectual property of new results, at the same time giving online access to all proofs and materials, including data and software, that sustain these results. After possible feedback from peers, improved versions can be submitted to overlay journals according to their publication constraints. In this model, reviewers have access to the complete history of the results and can better evaluate their quality and impact. Books remain a specific and important format for providing a comprehensive view of a given topic and for contributing to education. Here again, prepublication in an open archive can help in sharing drafts, getting feedback from peers before official publication, managing versions, etc. Another specificity of the Informatics publishing culture is that unlike in other sciences, such as Physics or Medicine, Informatics does not have a generally adopted convention for interpreting the order in which the authors of a publication are listed, and does not especially distinguish the “corresponding author(s)”. Thus, in the absence of specific indications, author order should not serve as a factor in individual researchers’ evaluations. This issue is discussed more extensively in Section 4 of this report. 3. Research Evaluation for Increased Quality and Impact The fundamental goal of research evaluation is to assess the quality and impact of research for the eventual improvement of both. Quality is an elusive intrinsic characteristic for which a commonly accepted assessment method, even if imperfect, is peer review by a panel of informed experts. Impact is an observable external characteristic that takes many forms and can to some extent be measured by numerical indicators, but even then, only with human expert interpretation. Quality is mostly a good predictor of impact, and impact is mostly a good indicator of quality, but the two are not coextensive. The CoARA agreement recommends to "focus research assessment criteria on quality [and] recognise the contributions that advance knowledge and the (potential) impact of research results" (CoARA 2022, p. 3). Assessing impact can be quite difficult, and is often even infeasible in a short timeframe, because the impact of a body of research can be indirect and it may take years before the eventual impact can be recognised. 3.1 Assessing the quality of research As recognised by the CoARA agreement, “research assessment should rely primarily on qualitative assessment for which peer review is central, supported by responsibly used quantitative indicators where appropriate” and “it is important that peer review processes are designed to meet the fundamental principles of rigour and transparency” (CoARA 2022, p. 5). With the increasing availability of publicly available bibliometric data, research assessments have been resorting more and more to bibliometric indicators provided by a variety of sources, often even without questioning the sources’ soundness and trustability. While we acknowledge the usefulness of quantitative data and bibliometric indicators when used responsibly (see Section 4), we emphasise the following: © Informatics Europe, 2025, CC BY-SA 4.0|Page 7 Research assessment should not take quantity as a proxy for quality and impact (cf. Friedman & Schneider 2015). Any policy that tends to favour quantity over quality has potentially disruptive effects and can mislead young researchers with very negative long-term effects. Such policies can lead researchers to focus on publishing the least publishable increments, and in this respect, some European countries have established national evaluation practices with indicators that raise serious concerns. To stress the importance of quality and impact, researcher evaluations should preferably focus on a relatively small number of high-quality publications and artefacts, trying to identify also their impact carriers such as novelty, supporting artefacts, and influence on other researchers’ work. Quantitative data and bibliometric indicators must be interpreted in the specific context of the research being evaluated. They should never constitute the sole evaluation criterion. Different research areas and even subfields of Informatics have very different characteristics. Even within a homogeneous set, bibliometric indicators only provide a very coarse assessment. For example, although a very high number of citations may indicate a potentially impactful piece of work, it may also indicate a widely referenced survey rather than a novel original contribution (which, of course, also has its own value). Human appraisal with respect to established criteria is needed to interpret data and discern quality and impact. Numbers can only help; they are not a substitute. In addition to being established, known, and shared by evaluators and researchers, assessment criteria must themselves undergo assessment and revision in order to follow the evolution of science. Other indicators for quality Major conferences and scholarly societies in Informatics often grant “best paper awards” to contributions perceived to have the highest value of those accepted at a conference. A further development is the granting of “most influential paper” or “test-of-time” awards for contributions perceived to have had the most influence in the area over some past number of years. These distinguishing, peer-reviewed awards should be taken into account in evaluations. 3.2 Assessing the impact of research The impact of research can be assessed along many different dimensions. First, one can distinguish between external and community impact. The external impact of research concerns its effect on society at large. The new GenAI techniques, for example, will potentially have an enormous impact on society. Likewise, the invention of a new secure protocol may lead to the development of more secure networks on which society relies, and a new automated development environment can improve the industry’s productivity. Advances that lead to commercial exploitation, spin-off activities, or adoption by industry (e.g., software licenses) or standardisation bodies, are some highly valued indicators of external impact. Community impact, on the other hand, refers to the impact of one’s research on other researchers. This means that other researchers can build on top of one’s research. Evaluation often tries to capture this kind of impact by, e.g., citation indices or research software download counts. A different kind of indicator could be provided by software competitions that are regularly organised to assess progress in tools (e.g., SAT solvers or learning tools). Such competitions are typically highly appreciated and contribute to increasing the quality of the tools, and success in them could be taken as an indicator, or at least a predictor, of community impact. An orthogonal dimension of impact concerns the outcomes of research. An outcome can have an impact because it advances knowledge in a given area or because it advances practice. New knowledge can be created by theoretical research, similar to mathematics. A typical example is research on algorithms and computational complexity. New knowledge can also be generated by research that pursues empirical © Informatics Europe, 2025, CC BY-SA 4.0|Page 14 research assessment practices to align them with the principles of Open Science. Indeed, the overall direction suggested is to build assessment processes on existing and well-known practices and guidelines – such as the San Francisco Declaration on Research Assessment (DORA 2012), the Leiden Manifesto (Hick et al. 2015), and the Coalition for Advancing Research Assessment (CoARA 2022). The main action points of these guidelines are: ● Avoid using metrics developed for measuring one kind of entity (e.g. Journal Impact Factor for journals) as a proxy measure for assessing other kinds of entities (e.g. researchers). ● Make explicit the criteria used in the assessment, and prefer peer-review evaluation, supported by quantitative indicators when appropriate – as discussed in detail earlier in Sections 4 and 5. ● Recognise the diversity of research contributions in assessment exercises (datasets, databases, software, and other artefacts), in addition to classic print-like publications (journal articles, books, book chapters, etc.) – as largely discussed in Section 5. ● Apply openness and transparency when providing data and methods used in the assessment exercise. ● Enable anyone to verify both the data and the analysis. ● Fight against manipulations and incorrect uses of metrics – as anticipated in Section 4. ● Consider the variability of different types of research output and subject areas when comparing an entity against another entity (e.g. researchers). ● Focus on qualitative judgement of research outputs when assessing researchers – as discussed in Section 4. Several of these actions are not implementable without the availability of appropriate Open Science infrastructures such as open bibliographic and citation databases and metadata repositories, institutional research information systems, and open bibliometrics and scientometrics systems for assessing and analysing scientific domains. Indeed, UNESCO insists that investing in Open Science infrastructures and services is essential, and that the scholarly community should retain control and ownership over these infrastructures. They are a crucial means for providing the data, i.e., research information used for devising metrics and indicators that may be used to support peer-review assessment. Research information refers to all the metadata related to the conduct and communication of research, such as bibliographic metadata of research outcomes (articles, software, datasets, methodologies, etc.), information on funding and grants, and information on organisations and research contributors. Several initiatives, such as the Barcelona Declaration (DORI 2024), and several Open Science infrastructures push for making research information openly available and freely reusable for the scholarly community because various activities, including research assessment, are characterised by using such information. Among the relevant infrastructures, OpenCitations (https://opencitations.net), OpenAIRE (https://openaire.eu), Software Heritage (https://www.softwareheritage.org), and DBLP (https://dblp.org) also provide information about software and other Informatics-related artefacts in addition to classic publications in journals and conference proceedings. Other infrastructures, such as PREreview (https://prereview.org), help in addressing related assessment tasks in a very transparent way, e.g. by enabling open peer review practices. These open reviewing processes permit the disclosure of the identity of the reviewers, the publication of reviews, and thus, the recognition of the effort that scholars put into reviewing, and the possibility for a broader community to provide comments and participate in the assessment process (UNESCO 2021). Supporting these kinds of Open Science and community-guided infrastructures is key to maintaining an environment that enables transparent assessment workflows. Indeed, there is an urgent need to make research information to advance responsible research assessment and operationalise Open Science and promote unbiased, transparent, and high-quality decision-making. The Open Science Career Assessment Matrix (OS-CAM) (European Commission 2017) represents a possible, practical move toward a more comprehensive approach to evaluating researchers through the lens of Open Science. In addition, other © Informatics Europe, 2025, CC BY-SA 4.0|Page 15 templates for Open Science assessment that involve Informatics as a domain of application have been studied in the context of specific research projects, such as GraspOS (https://graspos.eu/). Informatics should acknowledge Open Science practices in its research evaluation. Informatics thus also has a prominent role to play in the adoption and development of the Open Science approaches and infrastructures, and its support is key to keep them sustainable in the long term. R6.1 Acknowledge, support, and adopt Open Science practices in Informatics. R6.2 Apply principles derived from Open Science practices in the evaluation of research. R6.3 Strive to keep the control of Open Science infrastructures and processes within the research, academic, and educational Informatics communities. 7. Interdisciplinary Research Informatics has for a long time been an important support provider for research in other fields, in the form of software tools for modelling, optimisation and visualisation, data management and analysis, etc. In an interdisciplinary collaboration this kind of work is often considered mundane by the companion area partners, and not really computing research by the Informatics community. And indeed, for instance simply applying a known analysis method to a new dataset without a clear scientific advance on the Informatics side is just a technical support task. There are however also more ambitious collaborations, where the Informatics contribution is an integral element of the research agenda, and pursuing it requires both competence in the companion area and an ability to adapt existing and develop new methods in Informatics for the needs of the specific collaboration. A high-profile example of the latter kind of collaboration is e.g. the work on computational protein design that was awarded one of the 2024 Nobel Prizes in Chemistry. Even in such cases, however, credit is commonly assigned along disciplinary lines, so that recognition goes primarily to the companion area, and also on the Informatics side the computational methods contribution is considered as “an application”. An added challenge is that in assessments that are based mechanically on publication venues and/or indicators, Informatics contributions to other areas than core Informatics are easily overlooked, because the relevant venues are either not indexed at all in Informatics databases, or are considered as “application area outlets”, no matter how prestigious in the companion area. The challenge of overcoming disciplinary barriers in the assessment of integrative interdisciplinary research is a widely recognised problem (e.g. McLeish & Strang 2016). Some key recommendations for the responsible evaluation of Informatics research in this context are: R7.1 Recognise the value of integrated interdisciplinary research in its own terms, not as an “application” of Informatics. R7.2 Assess (i) the depth of the integration and (ii) the novelty and significance of the Informatics contribution to the totality of the work. R7.3 Be wary of numerical indicators and “top venue” lists oriented towards assessing Informatics disciplinary work (such as e.g. the CSRankings conference list). © Informatics Europe, 2025, CC BY-SA 4.0|Page 16 8. The Role of AI in Research Evaluations The emergence of powerful AI methods and tools influences research evaluation in two ways: first, how to regard and assess AI being used in research, and second, how AI can be used to perform or assist in carrying out research evaluations. While the first point is being intensely discussed, much less discussion exists regarding the second point. We shall address the latter issue here. Whenever a task is tedious, the question arises of how it can be automated or at least made easier by delegating parts to an automated process. A good example in the context of research evaluations is the h-index, which is widely used but also heavily criticised. Its advantage lies in the ease with which it can be obtained. It is much harder to judge the content of publications than simply count them and their citation numbers. Until now, there was no easy way to delegate the refereeing of publications to an automated process. This has changed with the advent of Large Language Models (LLMs). It is certainly tempting to use such models to evaluate the output of researchers or entire departments (or even universities). At the time of writing this position publication, the generally available LLMs are hesitant to write entire evaluations. For example, Gemini primarily describes the process of writing an evaluation but does not provide one. ChatGPT can write an evaluation but is very cautious in comparing different Informatics departments. Moreover, ChatGPT bases its judgments on citation numbers and very vague identification of key strengths. Nevertheless, we can safely assume that LLMs can, in principle, be used for research evaluations and that they will be used for this purpose in the future, if not already today. The question for us is to understand the potential dangers and benefits of this development. As with any other use of AI, some might call for restrictions or an outright ban on using AI methods in evaluations that form the basis of important decisions, such as funding large projects or granting tenure. It is short-sighted to ban AI altogether: it might actually help in the decision-making process when used with discretion, and a ban cannot be enforced anyway. Rather than a ban, we recommend the following policies for the responsible use of such technologies, if they are used at all: R8.1 The use of generative AI in research evaluations should be communicated openly, including detailed logs of the communications and the way the information was used in the decision-making. R8.2 Decisions must still be made by human experts, and the use of AI should be restricted to the lower levels of the decision process. R8.3 Efforts should be made to verify critical data obtained from a computer. R8.4 The use of AI must not be used to reduce the number of human experts in decision panels and their responsibility for the outcome of the panel. The use of AI is a moving target, and today we are just beginning to harness its power. While AI has been instrumental in many areas for some time, its applications in decision-making are still in their early stages. Research evaluation is among these areas, and further discussions in the near future will be important as the field continues to evolve. Therefore, our conclusions at this time can only be considered preliminary. © Informatics Europe, 2025, CC BY-SA 4.0|Page 17 9. Conclusions Research evaluation is essential for maintaining the vigour of any scientific field, and the system of expectations and incentives created by evaluation practices can have a significant effect on a research field’s development. Informatics has many special characteristics, including methodological diversity, close connection between topical research and socio-economic impact, prominence of top conferences as primary publication venues, artefacts as first-class research outcomes, etc. These may not always be apparent from outside of the field, yet need special consideration when assessing research in Informatics. In a judicious evaluation process, these specificities must be taken into account, together with the general principles of responsible research evaluation as advanced in, e.g., the CoARA recommendations. In this document, we are reconfirming and updating the general guidelines for Informatics research evaluation presented in the 2008 and 2018 Informatics Europe reports on the topic. We are also providing analyses and recommendations in some areas that were previously treated in less detail or have increased in prominence since the preparation of the earlier reports: bibliometrics and credit assignment in contributions, assessing artefacts, Open Science, interdisciplinary research, and the role of AI in research evaluation. As with any other aspects of science and research, the dynamic and changing landscape will certainly warrant revisiting these recommendations again in some years, especially as pertains to topics that are relatively recent. Acknowledgements We would like to acknowledge the input from a number of researchers who have generously devoted their time to answer our questions: Erwan Bousse, Carlos Esteban Budde, Karine Even-Mendoza, Hadar Frenkel, Lynda Hardman, Arnd Hartmanns, Tobias Kappé, Raphaël Monat, Benoit Montagu, Marius Muench, Roberto Natella, Michael Rawson, Robert Ricci, Markku-Juhani Saarinen, Jesús Sánchez Cuadrado, Alexander Serebrenik, Caleb Stanford, and Quentin Stiévenart. Endorsements This report is endorsed by the following National Informatics Association members of Informatics Europe: ● Informatik Austria (Austria) ● SIF – Societé Informatique de France (France) ● FTI – Fakultätentag Informatik (Germany) ● GI – German Informatics Society (Germany) ● GRIN – GRuppo di INformatica (Italy) ● IPN - ICT-Research Platform Netherlands (Netherlands) ● CODDII – Conferencia de Directores y Decanos de Ingeniería Informática (Spain) ● SCIE – Sociedad Científica Informática de España (Spain) ● SIRA – Swiss Informatics Research Association (Switzerland) ● CPHC – Council of Professors and Heads of Computing (United Kingdom) © Informatics Europe, 2025, CC BY-SA 4.0|Page 18 References (ACM 2020) ACM, Artifact Review and Badging Version 1.1., 2020. https://www.acm.org/publications/policies/artifact-review-and-badging-current (Baker 2016) M. Baker, 1,500 scientists lift the lid on reproducibility. Nature 533 (2016), 452-454. https://doi.org/10.1038/533452a (Ball 2023) P. Ball, Is AI leading to a reproducibility crisis in science? Nature 624 (2023), 22-25. https://doi.org/10.1038/d41586-023-03817-6 (Canteaut 2021) A. Canteaut et al., Software Evaluation (Research Report). INRIA, 2021. https://inria.hal.science/hal-03110728/document (CoARA 2022) Coalition for Advancing Research Assessment: Agreement on Reforming Research Assessment, 2022. https://coara.eu/agreement/the-agreement-full-text/ (Dagstuhl 2012) Publication Culture in Computing Research -- Position Papers. Dagstuhl Perspective Workshop 12452, 2012. Eds. K. Mehlhorn, M. Y. Vardi, M. Herbstritt. https://doi.org/0.4230/DagRep.2.11.20 (Di Cosmo 2019) R. Di Cosmo. How to use Software Heritage for archiving and referencing your source code: guidelines and walkthrough. Hal-02263344, 2019. https://doi.org/10.48550/arXiv.1909.10760 (Di Cosmo et al. 2020) R. Di Cosmo, M. Gruenpeter, S. Zacchiroli. Referencing source code artifacts: a separate concern in software citation. Computing in Science and Engineering 22:2 (2020), 33-43. https://doi.org/10.48550/arXiv.2001.08647 (DORA 2013) San Francisco Declaration on Research Assessment, 2013. https://sfdora.org/read/ (DORI 2024) Barcelona Declaration on Open Research Information, 2024. https://barcelona-declaration.org/ (European Commission 2017) European Commission. Directorate General for Research and Innovation, Evaluation of research careers fully acknowledging Open Science practices: Rewards, incentives and/or recognition for researchers practising Open Science, 2017. https://doi.org/10.2777/75255 (Friedman & Schneider 2015) B. Friedman, F. B. Schneider, Incentivizing Quality and Impact: Evaluating Scholarship in Hiring, Tenure, and Promotion. CRA Best Practice Memo, 2015. https://www.cra.org/cra/wp-content/uploads/sites/10/2016/02/BP_Memo.pdf (Hicks et al. 2015) D. Hicks, P. Wouters, L. Waltman, S. de Rijcke, I. Rafols, Bibliometrics: The Leiden Manifesto for research metrics. Nature 520 (2015), 429–431. https://doi.org/10.1038/520429a (Informatics Europe 2008) Research Evaluation for Computer Science, Informatics Europe Report, 2008. Eds. B. Meyer, C. Choppy, J. van Leeuwen and J. Staunstrup. https://www.informatics-europe.org/services/publications/reports.html (Informatics Europe 2013) Protocol for Research Assessment in Informatics, Computer Science and IT Departments and Research Institutes. Informatics Europe Report, 2013. Ed. Manfred Nagl. https://www.informatics-europe.org/services/publications/reports.html (Informatics Europe 2018) Informatics Research Evaluation, Informatics Europe Report, 2018. Eds. F. Esposito, C. Ghezzi, M. Hermenegildo, H. Kirchner, L. Ong. https://www.informatics-europe.org/services/publications/reports.html © Informatics Europe, 2025, CC BY-SA 4.0|Page 19 (McLeish & Strang 2016) T. McLeish, V. Strang, Evaluating interdisciplinary research: the elephant in the peer-reviewers’ room. Nature Palgrave Communications 2 (2016), 16055. https://doi.org/10.1057/palcomms.2016.55 (Rous 2017) B. Rous, The ACM Task Force on Data, Software, and Reproducibility in Publication, 2017. https://www.acm.org/publications/task-force-on-data-software-and-reproducibility (UNESCO 2021a), UNESCO Recommendation on Open Science, 2021. https://www.unesco.org/en/open-science (UNESCO 2021) UNESCO Recommendation on Open Science (Programme and Meeting Document SCPCB-SPP/2021/OS/UROS), 2021. https://unesdoc.unesco.org/ark:/48223/pf0000379949 (USENIX 2025) USENIX Security '25 Call for Papers. https://www.usenix.org/conference/usenixsecurity25/call-for-papers (Winter et al. 2022) S. Winter et al., A retrospective study of one decade of artifact evaluations. Proc. ESEC/FSE 2022. https://doi.org/10.1145/3540250.3549172 For enquiries and feedback about this report, please contact [email protected] www.informatics-europe.org © Informatics Europe, 2025 CC BY-SA 4.0