Describing Data for Discovery: RSEs, Croissant, and the Future of Research Metadata
Abstract
Many research projects collect or create datasets as part of their outcomes, but these are often hard to find, interpret, or reuse. A key reason is the lack of structured, machine-readable metadata that can support both human understanding and automated tooling. This creates friction with data discovery, integration, and reuse—especially in machine-learning workflows.This talk makes a case for adopting Croissant, a metadata format developed by MLCommons to address these challenges. It provides a standardised vocabulary and structure to describe datasets in ways that can be validated, shared, and used programmatically. Although Croissant originated from the ML community, its design is broadly applicable to any research data ecosystem. RSEs are in a unique position to help the widespread adoption of such standards. Since we often work with tools and platforms that manage datasets, we can embed good metadata practices by design, not as an afterthought.The talk elaborates on what it looks like in practice to integrate Croissant into a real-world project workflow. To this end, we draw on EnergySHR, a project supporting AI for energy and sustainability. It aims to help researchers share energy-related datasets and references. During submission, users provide metadata to generate structured, machine-readable descriptions of each dataset. These are validated and embedded into the dataset's page. This helps support structured web search and enables automated loading of datasets in downstream workflows.By embedding standards like Croissant into the tools we build and support, RSEs can make research outputs more visible and reusable. These are not just technical improvements but key contributions to research excellence.A recording of this session is available on YouTube: https://youtu.be/h0M66R7Oa80