scieee AI-readable full text Open interactive document viewer

[ECSS 2024] Posters Session (28 October 2024)

Informatics Europe

Abstract

ECSS 2024 poster session welcomed poster submissions from all levels of Informatics researchers, expanding beyond PhD students in previous years, and offering enriched knowledge exchange and networking opportunities for everyone.

Full text

20th European Informatics Leaders Summit (ECSS 2024) Poster Session 28 October 2024 A Framework for Integrating Patient Generated Health Data (PGHD) with Electronic Health Records (EHR) in a Fast Health Care Interoperability Resource (FHIR) supported EHR System. Supervisory team: Dympna O’Sullivan, TUD Lucy Hederman, TCD PGHD refer to health-related information gathered by patients or caregivers outside of clinical settings [1]. •Collection Methods: Diaries, apps, wearables etc. •FHIR is used for exchanging data with EHRs [2] •Integrating PGHD includes the use of extensions – albeit unsupported in some EHRs. •This framework enhance this through; -the FHIR Resource Provenance, and a category metadata in the Observation FHIR Resource. Key Enables of the Framework Adoption of the Provenance Resource Capability & Conformance Statement Adopt the metadata category: patientreported. Extension Methodology: The Framework Background: Abdullahi Abubakar Kawu, Dympna O’Sullivan, Lucy Hederman Fake Papers in Science: Paper Mills, Characteristics and Detection Strategies Ahmar K. Hussain, Marcus Thiel, Bernhard A. Sabel, Andreas Nürnberger Fake scientific articles, or fake papers, are a significant issue and a threat to research integrity. They spread misinformation, which might have a severe impact on research and ultimately also undermine the public's trust in science. This research aims to raise awareness among the research community about fake papers and highlight strategies to detect them. Fake papers contain various issues, including fabricated data, plagiarized content, manipulated images, and AI-generated content. Often, such papers are produced by professionals, i.e., so-called "paper mills" that sell authorship or manuscripts for a profit. Those fabricated papers pollute the scientific record and increase the burden on reviewers and other scientists. The rise of Large Language Models (LLMs) like ChatGPT and Gemini, as well as the increase in publication volume over the years, additionally contributes to the issue. Therefore, the need for detection tools to ensure scientific integrity is becoming more critical. Specific tools exist for publishers to detect fake papers, including Integrity Hub from STM, Snapshot, and Gepetto. However, these tools are not publicly accessible, thus hindering researchers to evaluate their impact. We present preliminary results from our ongoing research using text mining techniques and metadata features. Our initial findings suggest that there are indicators to flag papers for closer scrutiny and that further research is promising. What are Papermills? Characteristics of Fake Papers Papermills are organizations that produce academic manuscripts for researchers for money [1]. They have professional writers for different domains in science and charge hefty fees from researchers to appear as an author on their paper, as shown in Fig. 1. For researchers whose promotion or PhD requires them to publish frequently, this is an easy way out for them. Fig 1: An illustration on how papermills work Papermill-produced research has been increasing steadily since the beginning of the year 2000, as shown in an article by Nature [2] in Fig. 2. The abundance of papermill products in the scientific literature undermines the hard work of authentic researchers and poses a threat to scientific integrity. Preliminary Results Fig 2: The increase in papermill activity over the years[2] Fake papers could contain one or multiple different reasons that make them inauthentic. The following list states some reasons as to why a paper could be flagged as being fake. Fabricated data Manipulated images AI-generated text Non-institutional email address Plagiarized content Figure 3 shows an example of duplication in different parts of an image in an actual research paper [3], where the similar colors represent duplications. Fig 3: Image duplications in a research article [3] How to Detect Faked Papers The detection of fake papers is a challenging task because fraud in publications can be in different forms; therefore, multiple detection methods are required to verify the authenticity of the research carried out. Secondly, due to the emergence of LLMs like ChatGPT and Gemini, detection of AI-generated texts is becoming more difficult. Commercial tools exist for publishers such as Integrity Hub, Problematic Paper Screener, Snapshot and Gepetto to weed out papers with integrity issues, however, these tools are not available for the public to use. There are different research papers that propose detection techniques using machine learning (Decision Trees) [4], BERT models [5], image manipulation detection [6], and LSTM models [7]. However, each method has its limitations and shortcomings related to a high false positive rate and manual detection methods etc. In order to develop a detection method, we collected a dataset of publications from the biomedicine domain containing fake and authentic papers. Different types of features and machine learning algorithms to detect fraudulent papers. The features used for the machine learning model were a combination of metadata and TF-IDF based features from the abstracts of the papers. The recall score was used as an evaluation metric, and the Gradient boosting classifier achieved a score of 83%. The word cloud in Fig. 4 presents the most common terms found in fake papers, including 'cell', 'protein' and 'mir' (referring to DNA), especially in the ones from papermills. The papermill products have a common writing style as well as image format. Another observation was that most of the fake papers in biomedicine are in the field of molecular biology. Fig 4: A word cloud of common terms in fake papers Fig 5: : Stacked bar plots of proportion of important binary features across classes (a) Hospital affiliation (b) ORCID availability There are also significant differences in metadata features of fake papers, including ORCID availability and hospital affiliation, as can be seen in Fig. 5. Impact on Scientific Community The mass production of fake papers could make the public question the validity of scientific papers and ultimately loose trust in science. Especially in the field of medicine and biology, where authentic knowledge and research is directly related to human health and well-being, fake papers would have a huge adverse effect. Research by Nature [2] shows the distribution of likely papermill products across different domains, with medicine and biology having the highest number (around 3% of all papers), followed by chemistry and computer science. Similarly, Sabel et al. [1] recently reported the rate of red-flagged potential fake papers at 11%. Fig 6: Distribution of papermill activity across domains [2] Fig. 7 shows the distribution of fake papers by country based on a pool of fake papers selected and analyzed by [8]. The bottom figure is a distribution of 12 proven fakes, whereas the top one is a reference group consisting of 733 papers. Another huge impact of this phenomenon includes the wastage of time and resources spent on funding to produce fradulent research. Producing fake research by a researcher can also be harmful to the reputation of the research group and the university producing it. Fig 7: Distribution of fake papers across countries [8] Conclusion October 2024Otto von Guericke University [email protected] The aim of this research is to spread awareness in the scientific community about the presence of fradulent research and present preliminary results of our contribution to detect them. Our findings suggest that there exists certain metadata and domain-related features that separate fakes from nonfakes. The future direction of the research will include refinement of the detection methods by adding further relevant features from the full text of papers and organization of workshops to raise awareness among researchers about the presence of fake papers in science. References [1] Sabel, B.A., Knaack, E., Gigerenzer, G., Bilc, M.: Fake publications in biomedical science: Red-flagging method indicates mass production. medRxiv pp. 2023–05 (2023) [2] Noorden, R.V.: How big is science's fake-paper problem? Nature 623, 466–467 (2023). https://doi.org/10.1038/d41586-023-03464-x [3] Bik, Elisabeth: comments on pubpeer: Inhibition of the receptor tyrosine kinase ROR1 by anti-ROR1 monoclonal antibodies and siRNA induced apoptosis of melanoma cellsduplicated images (https://pubpeer.com/publications/3823967D4947674E9BDF3B0C219BF5) [4] Dadkhah, M., Oermann, M.H., Hegedüs, M., Raman, R., Dávid, L.D.: Detection of fake papers in the era of artificial intelligence. Diagnosis 10(4), 390–397 (2023). https://doi.org/doi:10.1515/dx-2023-0090, https://doi.org/10.1515/dx-2023-0090 [5] Razis, G., Anagnostopoulos, K., Metaxas, O., Stefanidis, S.D., Zhou, H., Anag-nostopoulos, I.: Papermill detection in scientific content. pp. 1–6 (09 2023). https://doi.org/10.1109/SMAP59435.2023.10255173 [6] Bucci, E.M. Automatic detection of image manipulations in the biomedical literature. Cell Death Dis 9, 400 (2018). https://doi.org/10.1038/s41419-018-0430-3 [7] heocharopoulos, P.C., Anagnostou, P., Tsoukala, A., Georgakopoulos, S.V., Tasoulis, S.K., Plagianakos, V.P.: Detection of fake generated scientific abstracts. In: 2023 IEEE Ninth International Conference on Big Data Computing Service and Applications (BigDataService). IEEE (Jul 2023). https://doi.org/10.1109/bigdataservice58306.2023.00011, http://dx.doi.org/10.1109/BigDataService58306.2023.00011 [8] Wittau, J., Seifert, R. Metadata analysis of retracted fake papers in Naunyn-Schmiedeberg’s Archives of Pharmacology. NaunynSchmiedeberg's Arch Pharmacol 397, 3995–4011 (2024). https://doi.org/10.1007/s00210-023-02850-6 Process Execution and Monitoring Funded by the EU Chips Joint Undertaking project AIMS5.0 and the Dutch National Funding Agency RVO under grant agreement number 101112089. BPMN4ES: A BPMN extension for specifying key environmental indicators (KEIs) from different categories in process models Next Steps • Complete BPMN extension: first version includes indicators for energy and one indicator for other categories • Continue working towards attaining full life cycle coverage, starting with the execution and monitoring phase • Further develop calculator services (prototype for carbon emissions realised) • Optimisation of multiple KEIs or KPIs • Account for impact of execution infrastructure • Adaptation for process performance improvement References 1. Brocke, J. vom, Seidel, S., & Recker, J. (2012). Green Business Process Management: Towards the Sustainable Enterprise. 2. Fritsch, A., von Hammerstein, J., Schreiber, C., Betz, S., & Oberweis, A. (2022). Pathways to Greener Pastures: Research Opportunities to Integrate Life Cycle Assessment and Sustainable Business Process Management Based on a Systematic Tertiary Literature Review. Sustainability, 14(18), Article 18. 3. Bogdan Popescu. (2024). Environmental Sustainability Calculator Service for Business Processes. 4. Idil Oksuz. (2024). Modeling Sustainability in Business Processes. Process Analysis Process Mining Process Simulation Conformance Checking Adaptation Learning Process Modelling Process Adaptation energy water emissions waste KEI Calculator Services Worklist ManagerWeb Services Sensors Sustainability Monitoring Component BPM Enactment System Organisations are increasingly concerned with environmental sustainability for various reasons Societal Legislative Economic Green BPM Problem Introduction Quantifying sustainability performance across different dimensions is necessary for fulfilling legislative requirements and evaluating improvement efforts Sustainability as additional performance dimension alongside traditional economic dimensions Existing Green BPM research and initiatives focus on specific life cycle phases or particular environmental performance indicators such as carbon emissions or energy consumption Ecological Time Cost Quality Flexibility Sustainability Task scheduling, resource selection Process Life Cycle Events Green BPM Life Cycle Environmental Sustainability in Business Process Management Michel Medema1, Vasilios Andrikopoulos2 and Dimka Karastoyanova1 Information Systems Group1 and SEARCH Group2Bernoulli Institute University of Groningen [email protected] 10 kWh Manufacture product Inspect quality Ship product Order processed Quality ok? Yes No X Reject product Order received X Reporting BACKGROUND COMPUTER SCIENCE THEMES DESIGN STRATEGIES FOR EXPLAINABLE AI: Two expert evaluations with participants from HCI, AI, and data science Concurrent think-aloud method and “I like, I wish, What if?” discussion framework Evaluated XAI interfaces from IBM’s Explainability 360 toolkit in FinTech context [3] Data collected via audio recordings, transcribed, and analyzed using conventional content analysis METHODOLOGY We recruited seven experts in AI, HCI, and data science, spanning various age groups. Participants self-assessed their expertise, with one novice and others intermediate or above. PARTICIPANTS Interactive elements should be clear and enhance understanding. Descriptions must use natural language and complement visuals. Charts/graphs should clarify explanations and align with text. Users need transparency on data used in AI decisions. Example-based XAI should clarify reasoning behind examples. Features should use natural language, highlight importance, and provide counterfactuals. With the growing demand for transparent AI systems, the EU's AI Act emphasizes the need for accessible explanations to foster user trust and ethical AI use [1]. Our research explores the "gulf of explanation" in XAI, evaluating how well current systems align with users' mental models and engaging experts to improve human-centred XAI design [2]. This work particularly considers the application of a new and novel HCI framework, a set of 12 key design principles for human-centred XAI. CONCLUSION 12 guiding design principles for human-centred XAI 01 OBVIOUS BUT UNOBTRUSIVE INTERACTIVITY 12 EXPLAIN IF THE DATA CHANGES OVER TIME 02 USE NATURAL LANGUAGE & REALWORLD SCENARIOS 03 BE CONSISTENT 04 VISUALS AND TEXT SHOULD REINFORCE EACH OTHER 05 USE NATURAL LANGUAGE FOR DATA AND FEATURES 06 SHOW FEATURE IMPORTANCE 07 SHOW COUNTERFACTUAL EXPLANATIONS 08 VISUALS SHOULD AID UNDERSTANDING 09 USE MIXED MODALITIES FOR ACCESSIBILITY 10 CLEAR RELATIONSHIPS SHOULD BE SHOWN WITH EXAMPLE BASED XAI 11 DATA USAGE AND ITS CONNECTION TO FEATURES SHOULD BE SHOWN Interactive elements should be obvious but not intrusive and should enhance user understanding. Text / Language descriptions should use natural language and/or real-world scenarios. Information should be consistent across all XAI representations. Written descriptions should enhance understanding of visuals and visuals should enhance understanding of written descriptions. Data used, and descriptions of features should be explained in natural language. Features should give option to show importance related to the system decision Features should show counterfactuals, so users know what they need to improve or what they are able to improve to gain a different result. The format of visuals should enhance explanation. visuals and text should use mixed modalities and should be accessible. Example based XAI are useful when well executed showing clearly how the examples relate to each other and the reason for the matched example. The user should know what data was used in AI decision and how the data is linked to features. The user should know what data was used in the AI decision and if the data changed over time. Study evaluates and informs the design of human-centred XAI. Introduces a novel HCI framework with 12 key design principles for XAI. Highlights the importance of XAI for non-technical users, especially in high-risk domains. Identifies gaps between current XAI design and user understanding. Advances HCI practice with actionable guidelines for enhancing AI transparency. Acknowledges limitations and plans further validation of design principles through iterative design with non-technical users. [1] Madiega, T., 2021. Artificial intelligence act. European Parliament: European Parliamentary Research Service. [2] Sheridan, H., Murphy, E. and O’Sullivan, D., 2023, July. Exploring Mental Models for Explainable Artificial Intelligence: Engaging Cross-disciplinary Teams Using a Design Thinking Approach. In International Conference on Human-Computer Interaction (pp. 337-354). Cham: Springer Nature Switzerland. [3] IBM, AI Explainability 360 – Demo. 2024 Retrieved October 1, 2024 from https://aix360.res.ibm.com/data Helen Sheridan, Dympna O’Sullivan & Emma Murphy School of computer science, TU Dublin, Ireland 200 7 Integrating Large Language Models with Digital Twins: Supporting Modelling and Design of Sustainable Systems 1. Introduction 2. Problem Statement Dr. Kunal Suri Université Paris-Saclay, CEA, List, F-91120, Palaiseau, France [email protected] | https://orcid.org/0000-0002-2341-5343 3. Research Questions 5. Conclusion & Future Works The integration of Large Language Models (LLMs) and Digital Twins technologies (virtual replicas of physical systems) opens up new possibilities for process analysts and engineers by (semi-)automating process discovery and (re)design within the Business Process Management (BPM) lifecycle. These critical steps form the foundation for developing any sustainable Process-Aware Information System (PAIS). Traditionally, these steps are tedious and prone to errors, but the advancements in Generative AI, particularly LLMs like those from OpenAI (ChatGPT), Llama, and Mistral, show potential. This is because LLMs have significantly become better in understanding the semantics and context present in large volumes of unstructured data such as documents, emails, maintenance logs, operational videos, and chat exchanges. •The integration of LLMs with systems used to develop Digital Twins (including Process and Product Digital Twins) will significantly enhance traditional modelling techniques, addressing key challenges in fast-evolving process-aware sectors such as manufacturing (Industry 5.0), healthcare, and smart cities. •These systems offer (semi-)automated, domain-specific, data-driven optimizations by utilizing both structured and unstructured data. This aids decision-making, facilitates compliance with new regulations (e.g., GDPR, AI Act), and helps bridge Europe’s widening skills gap, which is intensified by the retirement of experienced workers and the increasing demand for digital expertise. •As this research advances, it will further explore the benefits and limitations of these AI agents, laying the technological foundation and offering practical insights for integrating LLMs with systems used to develop Digital Twins to support upskilling, resource optimization, and sustainable operations. Q1. How can LLM-based agents be optimized by combining model fine-tuning with Retrieval-Augmented Generation (RAG) to deliver outputs that align with user requirements and real-world constraints in process-aware environments, such as manufacturing and healthcare? Q2. How to address challenges related to data scarcity and privacy (e.g., GDPR) during the domain-specific data processing to enhance the quality and accuracy of the system, while supporting the stakeholders in making informed decisions? This research is partially funded by the RAASCEMAN project, supported by Research and Innovation Action (RIA) under the Horizon Europe funding program of the European Commission, through Grant Agreement No. 101138782. 4. Experimentation This study explorers the integration of LLMs to systems used to model and design Digital Twins (both Process and Product Digital Twins), which are integral to a Digital Product Passport (DPP). We developed custom Generative Pre-trained Transformers (GPTs) based on OpenAI’s ChatGPT, focusing on refining prompts (prompt engineering) and incorporating internal knowledge from standards like BPMN, UML, SysML, and related documentation (see Fig. 1). Our approach underlines rapid prototyping, leveraging OpenAI’s intuitive tools to integrate domain knowledge and enable Retrieval-Augmented Generation (RAG) for quick assimilation of updated reports (see Fig. 2). However, a key challenge lies is the black-box nature of internal knowledge management in these custom GPTs. This makes them impractical for handling proprietary data due to confidentiality constraints, thus restricting testing to open data and standards. Despite these limitations, the custom GPTs have demonstrated promising results. Furthermore, this approach of leveraging OpenAI for initial prototyping has significantly reduced both GPU resource costs and development efforts. The insights gained from this approach will ultimately support the integration of Open LLMs with an RAG system in the FactoryIA initiative at CEA-List. LinkedIn Europe faces a major challenge in strengthening the resilience of its systems while managing the socio-economic impacts of digital transformation, especially in sectors such as manufacturing and healthcare: •The growing skills gap between traditional expertise and digital proficiency highlights the need for leveraging emerging technologies •The skill gap is widening as experienced workers retire or transition, making workforce upskilling and reskilling crucial to meet new demands Instead of replacing workers, systems leveraging AI present a promising solution to bridge the skills gap by assisting in the comprehension and design of complex processes and systems: •Supporting analysts in navigating emerging regulations such as AI Act, Chip Act •Enhancing operational efficiency and sustainability by identifying opportunities to optimize energy and resource use during the (re)design or (re)analysis phases •Enabling less experienced workers (or new hires) to grasp and operate complex processes without needing in-depth knowledge from the start Fig. 1: Architecture illustrating components of a LLM-integrated System Fig. 2: Example of a Custom GPT Enhancing Industry 5.0 Processes: Prioritizing Wellbeing with Smart Task-Worker Assignment 1. Research Motivation 3. Approach Internet of Things Worker well-being Business process management Economic benefits More flexibility and resiliency in business processes Worker well-being improvement • More motivated and productive workers • Less employee turnover • Improved business image • Short term: run processes under more diverse circumstances • Long term: avoiding long-term absence of workers • Less heavy, boring, stressful, … work • More versatile work and on-thejob training 4. Project Impact Worker assignment in BPMS lack flexibility to take real-time well-being into account Worker to task assignment Human aspect Manual Static Negative impact on well-being No flexibility in assignment Difficult to retain/ attract workers Heavy, stressful work Dynamic worker to task assignment mechanism for BPM systems that uses real-time well-being data from wearable sensors •IoT: identify appropriate wearables (e.g.: headband, t-shirt, smartwatch) for industry settings to measure sensor data (e.g.: heart rate, sweat, temperature) in real-time •Well-being: select relevant well-being factors (e.g.: stress, fatigue) that can be derived from real-time sensor data •BPM: improve business process management systems to dynamically adjust worker assignments considering well-being factors, business goals and worker characteristics Research Center for Information Systems Engineering, KU Leuven, Campus Brussels, Belgium [email protected] 2. Contribution Mathis Wyffels, Irene Vanderfeesten, Estefanía Serral Asensio Freepik.com