scieee AI-readable full text Open interactive document viewer

microbetag: simplifying microbial network interpretation through annotation, enrichment tests, and metabolic complementarity analysis

Zafeiropoulos, Haris; Michail Delopoulos, Ermis Ioannis; Erega, Andi; Schneider, Aline; Geirnaert, Annelies; Morris, John; Faust, Karoline

Full text

Open Access © The Author(s) 2025. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http:// creativecommons.org/licenses/by/4.0/. SOFTWARE Zafeiropoulosetal. Genome Biology (2025) 26:292 https://doi.org/10.1186/s13059-025-03769-2 Genome Biology microbetag: simplifying microbial network interpretation throughannotation, enrichment tests, andmetabolic complementarity analysis Haris Zafeiropoulos1*, Ermis Ioannis Michail Delopoulos1, Andi Erega2, Aline Schneider2, Annelies Geirnaert2, John Morris3 and Karoline Faust1* Abstract Microbial co-occurrence network inference is often hindered by low accuracy and tool dependency. We introduce microbetag, a comprehensive software ecosystem designed to annotate microbial networks. Nodes, representing taxa, are enriched with phenotypic traits, while edges are enhanced with metabolic complementarities, highlighting potential cross-feeding relationships. microbetag’s online version relies on microbetagDB, a database of 34,608 annotated representative genomes. microbetag can be applied to custom (metagenome-assembled) genomes via its stand-alone version. MGG, a Cytoscape app designed to support microbetag, offers a streamlined, userfriendly interface for network retrieval and visualization. microbetag effectively identified known metabolic interactions and serves as a robust hypothesis-generating tool. Keywords: Microbial associations, Enrichment analysis, Data integration, Pathway complementarity, Seed set, Phenotypic traits Background Most microbial species live in communities and most natural microbial communities consist of hundreds or even thousands of species. Each species exhibits a unique repertoire of biochemicalreactions and adapts to various niches, each with specific nutrient and environmental requirements. Depending on the net fitness effects that result for the taxa involved, interactions range from cooperation, competition, parasitism, commensalism, and amensalism [1]. In addition to other mechanisms, microorganisms can interact by competing for or exchanging metabolites. The latter interaction mechanism can involve either one-way (unidirectional) or two-way (bidirectional) exchanges of metabolites. High-throughput sequencing has provided insight into the diversity and composition of microbial communities. Uncultivated species can now be detected, and their traits can be predicted based on their genomic information [2]. Moreover, the composition of *Correspondence: haris.zafeir[email protected]; karoline[email protected] 1 Department of Microbiology, Immunology and Transplantation, Rega Institute for Medical Research, Laboratory of Molecular Bacteriology, KU Leuven, Herestraat 49, Leuven 3000, Belgium 2 Institute of Food, Nutrition and Health, Laboratory of Food Biotechnology, ETH Zurich, Schmelzbergstrasse 7, Zurich 8092, Switzerland 3 Resource for Biocomputing, Visualization, and Informatics, University of California San Francisco, 16th Street, San Francisco, CA 94143, USA Page 2 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 thousands of microbiome samples is now accessible, allowing for the inference of associations across large sets of samples. A widely used approach to extract such patterns is the creation of microbial co-occurrence (i.e., association) networks based on microbial sequencing data (amplicon and/or shotgun) [3, 4]. Several approaches are available for co-occurrence network inference based on the assessment of similarity, dissimilarity, and/or correlation (e.g., CoNet [5], SparCC [6]) and conditional dependency identification, which allows reducing the number of indirect edges (e.g., SpiecEasi [7], FlashWeave [8]). Nevertheless, microbial network inference encounters various challenges [9, 10]. Each inference approach comes with its own assumptions and parameter settings, leading to variations in network structure. The choice of the association measure and data preprocessing techniques as well as the handling of sparsity and zero inflation all influence the resulting network. As a consequence, the result of network construction is tool-dependent [11, 12]. Moreover, microbial network inference inherits the challenges of sequencing data and analysis (e.g., sampling scale, compositionality, tuning of parameters linked to sequencing data processing) and the returned network is often a “hairball” of densely interconnected taxa. Thus, additional analysis is necessary to generate testable hypotheses [10]. The comparison of interactions predicted by microbial networks with a collection of known interactions has underscored their low accuracy for this task [11, 13]. Data integration has been suggested to help interpret edges in microbial networks [10]. In addition, clusters in microbial networks have been demonstrated to detect key drivers of community composition [14] and several algorithms and implementations are available to identify them (e.g.[15]). However, data integration approaches available for microbial networks are so far limited. Metabolic networks are comprehensive representations, in a mathematical form, of biochemical reactions occurring within an organism [16–18]. Ametabolic network can be considered as a knowledge-base of its corresponding strain, capable of integrating genomic, regulatory, and phenotypic information. Over the last decade, (semi-) automated approaches support the fast generation of such reconstructions (e.g., [19, 20]). A key concept for the topological analysis of such metabolic networks is the seed set. Based on the original definition, a seed set is “the minimal subset of the occurring compounds that cannot be synthesized from other compounds in the network (and hence are exogenously acquired) and whose existence permits the production of all other compounds in the network” [21]. Seeds are a useful proxy for the essential nutrients of an organism [21, 22], and based on the seed concept, several graph theory-based metrics (indices) have been described to predict species interactions directly from their metabolic networks’ topologies [23–26]. As Lam etal. [27] highlight in their study, seed sets “represent a baseline of metabolites that in theory enable a given bacterium to produce any metabolite in their predicted metabolic network.” Seed and non-seed compound sets, as defined above, can be used to compute complementarity and overlap indices. Metabolic complementarity between two species reflects their potential for cooperation through cross-feeding. In contrast, metabolic competition refers to the metabolic overlap between two species leading to exploitative competition. Page 3 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 The examination of such indices can indicate metabolic interactions that may drive the patterns observed in co-occurrence networks. To explore whether a species may benefit from a partner, it is helpful to move from the network to the pathway level and check whether their pathways complement each other. For this, we rely on a naive approach that enumerates all possible complements between a pair of species based on their KEGG ORTHOLOGY (KOs) annotations and the KEGG MODULES database [28]. Here, we present microbetag, a software ecosystem for microbial network annotation that exploits several sources of information to enhance the confidence in the associations suggested by the network, thereby generating hypotheses for further investigation both at the taxon pair and the community level. microbetag serves as a comprehensive platform that provides information about taxa along with their potential metabolic interactions from multiple channels (see “Methods”). The key concept here is the reverse ecology approach [29]. Reverse ecology leverages genomics to explore community ecology with no a priori assumptions about the taxa involved. The reverse ecology framework enables the prediction of ecological traits for less-understood microorganisms and their interactions with others [30]. microbetag annotates a user’s co-occurrence network by integrating phenotypic traits of the taxa present in the network (nodes) and by mapping potential metabolic interactions onto their associations (edges). microbetag is accompanied by a graphical user interface (GUI) implemented as a Cytoscape [31] app [32] providing a user-friendly environment to investigate annotations in a straightforward way. Its online version depends on precalculated annotations for ~ 35,000 highquality reference genomes stored in the microbetagDB. All annotations present in microbetagDB are also available through an application programming interface (API). microbetag’s source code is distributed under a GNU GPL v3 license and available on GitHub (https:// github. com/ msysb io/ micro betag). Documentation and further support on how to use microbetag is available through a ReadTheDocs documentation page (https:// micro betag. readt hedocs. io/ en/ latest). To the best of our knowledge, there is no software with which microbetag could be compared to directly. To validate our annotations, we used a recently published network with partially known interactions [33]. We then present the main features of microbetag’s interface and discuss two extra application examples. Results Overview ofthemicrobetag workflow andsoftware ecosystem microbetag integrates four sources of annotations—two at the node level and two at the edge level. It can operate in two modes: (a) on-the-fly, where user-provided taxa are mapped to their closest Genome Taxonomy Database (GTDB) representative genomes for downstream annotation; or (b) locally, using custom genome inputs. At the node level, phenotypic traits are assigned based on either the genome associated with each taxon (see “Genome-based node annotation” in Methods) or their taxonomic identity, using the curated literature-based FAPROTAX database (see “Literature-based node annotation” in Methods). At the edge level, microbetag infers potential metabolic interactions via two complementary approaches: (a) pathway Page 4 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 complementarity and (b) seed complementarity (see “Pathway complementarity” and “Seed scores and complementarity” in Methods). The annotated network can then be visualized on Cytoscape thanks to a dedicated Cytoscape app (MGG). For a complete overview of the workflow, refer to “The microbetag workflow” section in Methods. The microbetag software ecosystem consists of five main modules: (a) microbetagDB consisting of the precalculations for 34,608 reference genomes and their pairwise combinations (see Methods), (b) a web server hosting microbetagDB and the microbetag app to annotate a co-occurrence network on-the-fly or to retrieve annotations through an API, (c) a Cytoscape app called MGG that enables users to easily invoke the workflow and investigate the annotated network, (d) a preprocessing workflow for data sets with more than 1000 sequence identifiers (OTUs/ASVs/bins, etc.), and (e) a stand-alone version of the microbetag precalculation steps combined with the microbetag workflow that supports annotating a network based on the user’s reference genomes. Figure1 shows a simplified overview of microbetag’s on-the-fly annotation approach. Through the MGG Cytoscape app, the user may load an abundance table, and optionally its corresponding network, and send a job to the microbetag server. If not provided, microbetag will then infer a network using FlashWeave, and it will annotate its nodes (i.e., taxa) with phenotypic traits and its edges (i.e., significant co-occurrences or mutual exclusions) with potential metabolic complementarities. The annotated network is then returned to the user as a response to their job-query and is automatically loaded into Fig. 1 Simplified overview of the microbetagDB-dependent version of microbetag using an illustrative abundance table and its bins. From abundance data, microbial co-occurrence networks can be built (node colors represent different bins). Such networks consist of associations (edges) that are usually not directed and do not provide any interaction mechanisms, among other limitations. To tackle the challenge of interpreting microbial co-occurrence networks, microbetag adds a series of annotations at the node and edge levels. When used on-the-fly, microbetag makes use of microbetagDB, which stores precalculated phenotypic traits and metabolic complementarities among all possible pairwise combinations of high-quality, representative GTDB genomes. We have also developed a Cytoscape app called MGG to enable a user-friendly visualization of the microbetag derived annotations. Combined with a network clustering algorithm, microbetag also supports enrichment analysis of these traits within clusters. Node colors in the toy networks stand for different genera assigned to the sequences. Page 5 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 Cytoscape’s main window. The user can then investigate its annotations further thanks to the MGG features (see “Annotating networks with microbetag” section). For thorough instructions on how to use MGG and microbetag, answers to frequently asked questions and hints to address the idiosyncrasies of various data sets, the reader may visit the documentation web site (https:// micro betag. readt hedocs. io). For more information on how microbetag-annotated networks can be viewed through the MGG Cytoscape app, see also microbetag’s ReadTheDocs. Finally, microbetag’s Matrix community (https:// matrix. to/ micro betag commu nity: matrix. org) offers a dedicated communication platform where users can directly report issues, ask questions, exchange feedback, and suggest new features. The microbetagDB resource Currently, microbetagDB includes more than 34,000 genomes (Table1) along with their corresponding annotations. Most of these genomes represent bacterial taxa (364 taxa are archaea). The presence/absence of more than 30 phenotypic traits have been predicted for those genomes; in Additional file2: TableS1, we list those traits along with their corresponding abbreviations on the MGG Cytoscape app. About 1.4 billion potential metabolic interactions leading to pathway or seed complementarities have been precomputed as well. In the microbetag framework, pathway complementarity occurs when, based on their genomes, a species (donor) provides missing enzymes or pathways to another species (beneficiary), enabling the beneficiary to complete a KEGG module (see “Pathway complementarity” in “Methods”). A total of 341,568 unique pathway complementarities, i.e., sets of KEGG terms supporting the completion of a module, gave rise to the 184.2 million pairwise pathway complementarities observed. Similarly, seed complementarity, following the definition in Lam etal. [27], emerges when a beneficiary’s seed metabolite(s) are produced by the donor’s metabolic network, enabling potential crossfeeding. Using the sets of the metabolites a species can produce on its own (non-seed set) and of its seeds, cooperation/competition scores were calculated (see “Seed scores and complementarity” in “Methods”). Currently, microbetagDB contains seed and nonseed sets for 33,755 GTDB representative genomes, which are being used to compute seed complementarities for any pairwise combination of these genomes on-the-fly. All annotations can be accessed directly from microbetagDB through an API. Table 1 Summary of the data in microbetagDB Description Entries GTDB representative genomes 34,608 Phen-model-oriented metabolic functions 32 FAPROTAX functions 92 Metabolic networks 33,755 Unique pathway complements 341,568 Pairwise pathway complementarities 184,184,548 (non-)seed sets 33,755 Page 6 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 Annotating networks withmicrobetag The microbetag annotation pipeline can be performed either locally using custom genomes or on-the-fly, after mapping taxa to their closest GTDB genomes. In both cases, the annotated network will be viewed in Cytoscape through the MGG Cytoscape app. The MGG app is also the one supporting running microbetag on-the-fly, by importing data in the app and submitting a job to microbetag’s web server. More specifically, MGG was developed to simplify the use of microbetag; thus, running microbetag on-the-fly is straightforward and requires no bioinformatics skills. In the simplest case, microbetag only requires an abundance table with taxonomic assignments as input. When the user-provided taxonomy scheme is not one among the GTDB, Silva, or the GTDB-oriented taxonomy for 16S rRNA amplicon data (see “microbetag preprocessing”), microbetag maps the user’s taxonomy to an NCBI Taxonomy Id and from that to GTDB representative genomes; this mapping is the most time-consuming step of the workflow. Network inference can be a computationally intensive step too, particularly as the number of taxa in the abundance table increases. To enable annotation of large data sets, a stand-alone preprocessing workflow is provided with microbetag. The user can either assign their amplicon data to the GTDB-oriented taxonomy and/or reconstruct a network locally. Once a network is available and the taxonomy provided is among the standard ones for microbetag, the computational time required for annotation ranges from several seconds to just a few minutes based on the user’s settings. An annotated network in.cx2 format is then returned, which can be viewed in Cytoscape. To enable the microbetag approach with custom/local genomes, a microbetagDBfree version is also available. In this case, the user needs to run the microbetag pipeline locally, either installing microbetag and its dependencies from source or as a container. microbetag will first perform all the genome annotation steps and the metabolic model reconstruction if needed, depending on the user’s settings. For example, in case a user has already annotated their genomes with KEGG, or they have already reconstructed their corresponding metabolic networks, these precalculation steps may be omitted. Then, the microbetag workflow (see “Methods”) will be performed to return an annotated network. The computing resources for such a task can be quite extensive (see “Run times”). Once completed, the annotated network will be saved as a.cx2 file and can be visualized in Cytoscape using the MGG app. MGG allows the user to import data, retrieve an annotated network, and investigate the annotations through a series of CyPanels both for node and edge annotations. Figure2 shows an example of the Nodes CyPanel. The node name, taxonomy, NCBI Taxonomy Id, and GTDB genome to which the sequence was mapped (in case of an on-the-fly analysis) can be viewed. Depending on the user’s settings and the available annotations for a node, genomeor literature-based predictions may be presented. Further, the trait groups mentioned in “Groups of phenotypic traits” are displayed on top of this panel allowing for the selection of the nodes carrying either one among several attributes (OR logical relationship) or all of them (AND) (Fig.2). Similarly, in the Edges panel (Fig.3), the set of beneficiary taxa is specified, along with their corresponding GTDB representative sequence identifiers, while pathway and seed complementarities are listed in collapsible tables. Potential metabolic interactions are shown in a sub-table, bearing as title the genome pair under consideration Page 7 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 as several GTDB genomes may have been assigned to a node. In case of pathway complementarities, these tables consist of six columns: (a) the KEGG MODULE id of the module to be completed, (b) its description, (c) a more general metabolic category to which the module belongs, (d) the complement itself as a list of KEGG terms, (e) the alternative that now represents a complete module in the beneficiary, and (f) a URL that points to a colored KEGG map highlighting the complement (see “Pathway complementarity” in “Methods”). If clicked, the user’s default browser pops up displaying a colored KEGG map as shown in an example in Fig.4A. Similarly, in case of seed complementarities, a nested sub-table is again provided, but this time having as a header of each sub-table their corresponding PATRIC ids, which were used to reconstruct metabolic networks (see “Seed scores and complementarity” in “Methods”). In this case, the KEGG metabolism category is mentioned for each complement and the potential compounds being cross-fed are shown in red, while metabolites being produced by the beneficiary species are shown in blue (Fig.4B). Fig. 2 Nodes CyPanel of the MGG Cytoscape app. Once a node is selected, the Nodes panel displays its taxonomy, the NCBI Taxonomy id to which it was mapped, and its corresponding taxonomic level, and if available, the GTDB genome(s) to which it was mapped. In this example, an ASV that was assigned as Mycobacterium fragae was mapped to the GCA_002102185.1 GTDB representative genome, based on which a set of phenotypic traits was predicted to be present along with a prediction score. Also, based on FAPROTAX literature-based annotations, the ASV was annotated as chemoheterotroph. When multiple nodes are selected, a collapsible panel—similar to the one shown for ASV0128 in this example—will appear. On top of the panel, there is a “control area” with a set of features that allow the user to filter the nodes to be displayed. The “Show species” button filters taxa so that only those that have been matched to a genome are shown in the main panel, while the “Highlight first neighbors” highlights all nodes that are directly connected with a selected node. After clicking on the “PhenDB/Faprotax Filters” collapsible panel, the complete list of phenotypic traits across all taxa in the network is shown, organized in groups (see “Groups of phenotypic traits” in “Methods” for more). From there, the user may select to show only nodes (taxa) carrying either at least one (OR) or all (AND) of a set of traits. Page 8 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 Last, MGG allows testing for enrichment or depletion of the phenotypic traits assigned to the nodes in each of the network clusters. Clusters are optionally assigned bymanta [15] while performing the microbetag workflow or assigned by users either manually or with another network clustering algorithm (see use case on subgingival plaque microbiome data). In the following two sections, we present a test case and two use cases, highlighting our approach’s potential. For a thorough description of MGG’s panels, features, and visual style, readers are referred to the corresponding page in microbetag’s online documentation. Test case: complementarities predicted bymicrobetag foranetwork withknown interactions We tested microbetag using the correlation network of Hessler etal. [33], which describes mine tailing-derived (materials remaining after extracting the valuable Fig. 3 Edges CyPanel of the MGG Cytoscape app. Once an edge is selected, the edge panel displays the names of the two taxa involved and their corresponding sequence ids, as well as the interaction type. In case of a complementarity edge, a sub-panel for each complementarity type is provided. For pathway complementarities, the KEGG module being complemented is shown in the first column, followed by its description and the metabolic category it belongs to. The corresponding KOs related to the suggested cross-feeding compounds are also listed, along with the completed version of the module based on these additions. A URL is provided linking to a color-coded KEGG map that highlights the module under study (Fig. 4A). Likewise, for seed complementarities, a collapsible table is provided showing the category type and the associated KEGG map to which the seed compound contributes. It includes the seeds that can potentially be cross-fed from the donor, listed in both ModelSEED and KEGG namespaces. A URL is provided linking to a color-coded KEGG map highlighting both the beneficiary’s related compounds and the potential seeds it could acquire (Fig. 4B). Seeds linked to multiple KEGG maps may appear repeatedly in the table (see C00544). As in the Nodes panel, edges are filtered via the top control area. They can be either of “co-occurrence” type (undirected, unless using a custom network inference method) or “complementarity” (directed, with the source as the donor and the target as the beneficiary). The “Edges with Seed Complementarities” and “Edges with Pathway Complementarities” hide any edges that do not carry such annotations. The sliders below filter edges based on a score, allowing the user to set a threshold that can include values either above or below it. Slidebars can be combined. Page 9 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 fraction from an ore) consortia cultivated in bioreactors under varying conditions. In this study, Variovorax, a thiamine producer, and its co-occurrence with a series of thiamine auxotrophs are discussed. The authors tested network predictions by performing co-culture experiments measuring thiamine production. Sequence bins corresponding to network nodes and the original network were obtained from the authors (personal communication). Using GTDB-tk [34], GTDB taxonomies were assigned to the bins and added to the original network, which was then annotated with the on-the-fly version of microbetag. In addition, the local version of microbetag Fig. 4 Colored KEGG maps illustrating the suggested completed mechanism and highlighting the impact of the corresponding potential metabolic interaction with the overall species metabolism. microbetag constructs URLs using the KEGG API to visualize the complementarities found between two taxa. A In case of “pathway complementarity,” a set of missing KOs is provided to complete a specific KEGG module. In this example, module M00565 (trehalose biosynthesis) would be complete after acquiring a starch synthase, represented by K00703 (EC:2.4.1.21). This approach is unaware of alternative ways for the species to achieve the same task. For example, in this case, we do not know whether the species could use another pathway to synthesize trehalose. B In contrast, “seed complementarity” identifies as potential cross-fed metabolites those compounds that the beneficiary cannot produce on its own by any means. The returned complements may contribute to multiple parts of the species’ metabolism. In this example, the strain donor could potentially provide the beneficiary with two seed compounds: 4-phosphonooxy-L-threonine (C06055) and pyridoxol 5′-phosphate (C00627) Page 16 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 of compounds are seeds in small sets of genomes with only 81 compounds being a seed in more than 50% of the genomes (Fig.7C), while a total of 270 compounds could be potentially produced by more than 75% of the genomes (Fig.7D). Theoretically, if a genome’s non-seed compounds could be cross-fed to complete all seeds across the rest of the genomes, this would result in approximately 15.3 M overlaps. As shown in Fig.7E, the percentage of potential cross-feedings follows a normal distribution, with an average of 3.2% of a genome’s non-seeds appearing as seeds in other genomes. This suggests that most genomes have a similar potential to provide seed compounds, with only 16 genomes exceeding this average by more than three standard deviations and just one genome falling below it by the same margin. Notably, this one genome represents Ureaplasma diversum NCTC 246, a facultative intracellular strain of the Mycoplasmataceae family, which includes several parasitic and saprotrophic taxa. Table3 highlights the most and least common metabolites found as part of the seed set of the microbetagDB genomes. The presence of essential compounds—such as ATP and phenylpyruvate—that no organism can survive without, as the least frequent seed Fig. 7 Distributions of seed and non-seed metabolites, including only those that are part of at least one KEGG module, across the ModelSEED reconstructions in microbetagDB v1.0.1, as well as of their overlap. A The number of seeds per genome is a symmetric distribution ranging from 13 to 51 seeds. B The number of non-seed compounds per genome has a left-skewed distribution. C and D Distributions of the relative frequency of a metabolite being a seed or non-seed, with the seed distribution decaying quickly from 0 and the non-seed distribution decaying towards 1. E Normal distribution of the percentage of the non-seeds of a genome found to overlap with seeds of all the other microbetagDB genomes. Page 17 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 compounds serves as a positive validation of our pipeline. L-aspartate may be among the least common seeds because cells typically make aspartate and glutamate from TCA intermediates using transamination (an amino group is transferred from an amino acid to a keto acid) and since transaminase annotation is often challenging, most metabolic network reconstruction tools assume their presence by default [44]. It is also not surprising that cThz-P and L-2-amino-6-oxo pimelate are among the most common seeds since a large number of taxa are auxotrophic for thiamine or thiazole intermediates, and/or lack parts of the diaminopimelate pathway. The other three compounds are part of pathways that are relatively rare among bacterial species; for instance, most bacteria do not degrade purines completely to allantoin, often utilizing alternative degradation routes instead. Run times Microbetag offers various modules tailored to the unique characteristics of each dataset. In the following sections, we discuss run times for different scenarios to guide users in selecting the most suitable approach for their dataset. Run times range from seconds (s) to minutes (min) and hours (h). For an overview of microbetag’s workflow, see “The microbetag workflow” in the “Methods” section. On‑the‑fly Using the abundance table from the test case [33] (Table4, vitAbund.tsv, 120 bins) annotated with the GTDB taxonomy scheme and its corresponding network file (edgelist. Table 3 Most and least frequent seeds across the metabolic networks in microbetagDB. For the most frequent (TOP) seeds, we skipped the first four fatty acid metabolism-related compounds, which are artifacts of the ModelSEED reconstructions and are not being used except in rare cases Compound ModelSEED ID KEGG metabolism categories TOP 2-(2-Carboxy-4-methylthiazol-5-yl)ethyl phosphate (cThz-P)) cpd21480 Thiamine biosynthesis Geranylgeranyl diphosphate cpd00289 Ubiquinone and other terpenoid-quinone biosynthesis 2-Oxo-4-hydroxy-4-carboxy-5-ureidoimidazoline cpd09027 Purine degradation, xanthine = > urea L-2-amino-6-oxopimelate cpd02414 Lysine biosynthesis, DAP dehydrogenase pathway, aspartate = > lysine 5-Hydroxyisourate cpd08625 Purine degradation, xanthine = > urea Least ATP cpd00002 Amino acids biosynthesis Choline cpd00098 Nucleotide metabolism Phenylpyruvate cpd00143 Phenylalanine biosynthesis ADP-L-glycero-D-mannoheptose cpd03831 Biosynthesis of nucleotide sugars L-aspartate cpd00041 Carbon metabolism, biosynthesis of nucleotide sugars Page 18 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 Table 4 Run times of the on-the-fly and the preprocessing microbetag modules Abundance table Number of sequences Number of samples Taxonomy scheme Network available Run time Number of nodes Number of annotated nodes Number of edges Number of annotated edges Notes On-the-fly vitAbund.tsv 120 Invariant GTDB edgelist.tsv 40 s 102 55 251 93 testAbundMeta.tsv 78 16 Silva NA 67 s 23 7 14 1 metadata.tsv testAbund.csv 78 84 Silva NA 74 s 78 19 118 5 Same ASVs as in the row above (testAbundMeta.tsv) 78 84 Other edgelist.tsv 52 s 78 16 118 3 Network from the row above Preprocessing seq_ab_tab.tsv 16 997 GTDB NA 232 s + 94 s 240 105 219 49 As returned from previous row (+ 40 s) Page 19 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 tsv, 102 nodes, 251 edges), it took 40 s for microbetag to return an annotated network; microbetag was able to annotate 55 of its nodes and 93 of its edges. Given a network, the computing time is independent of the number of samples. Next, using a different abundance table containing 16 samples and 78 Silva-annotated ASVs (testAbundMeta. tsv), microbetag took 67 s to infer a co-occurrence network with 23 nodes and 14 edges using the FlashWeave sensitive approach and annotated 7 of its nodes and 1 of its edges. We then ran the same analysis with an abundance table that had the same ASVs as in the previous example but with 84 samples (testAbund.tsv). This time, it took 74 s for microbetag to return a co-occurrence network of 78 nodes of which 19 were annotated and 118 edges of which 5 were annotated. Thus, as expected, the number of samples affects the time required to infer the co-occurrence network. To highlight the impact of the taxonomy scheme on the run time, we provided the returned co-occurrence network from this last case as input, setting the taxonomy scheme to “Other” (instead of “Silva” as selected initially). microbetag returned an annotated network with only 16 annotated nodes and 3 annotated edges, while it took an extra time of 52 s to map the taxonomies against NCBI. Overall, since the on-the-fly module is capable of running jobs that require no more than a few minutes, it is recommended for cases involving small datasets or for scenarios where a co-occurrence network already exists and the taxonomy scheme is either Silva or GTDB. Preprocessing tool The microbetag preprocessing tool provides a GTDB annotated abundance table and its corresponding co-occurrence network, both tailored for seamless integration with the on-the-fly version, facilitating the analysis of larger datasets. We used an abundance file (seq_ab_tab.tsv) of 16 samples and 997 ASVs and annotated them with a GTDBbased taxonomy. We also performed the FlashWeave step using the sensitive approach. In total, using a personal computer of 15 CPUs, it took 3 min 52 s real time, while the user time was 24 min 19 s implying a degree of parallelization; on average 6 CPUs were used. When providing an abundance table with taxonomies only, it took 94 s to return a network of 240 nodes, 105 of which annotated, and 219 edges, 49 of which annotated. When we used both the GTDB annotated table and the network, it took 40 s; the 54 s difference was the time FlashWeave required to build the co-occurrence network on the fly (Table4). In addition, locally, one can further exploit the parallelization potential of FlashWeave thanks to its Julia implementation; in our tests, we did not make use of it. Stand‑alone version Real-world data may lead to large co-occurrence networks consisting of thousands of nodes and edges. The on-the-fly version of microbetag is not suitable for such large networks. For this reason, we developed a stand-alone version, which not only handles large networks better but also supports genome-scale metabolic reconstruction for custom genomes. Custom genomes require an important amount of computing resources, especially in two steps: gene prediction and KEGG annotation of the genomes. In case genome-scale reconstructions are carried out with ModelSEEDpy, then RAST annotation would also be required. For instance, in our demo case, with only 7 bins to annotate, Page 20 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 it took 1 h for Open Reading Frame (ORF) detection (running prodigal through DiTing) and about 2 h for the KOs assignment using a personal computer with 2 CPUs. Reconstructing the metabolic models with carveme is faster and performs in a more robust way since there is no need for establishing a connection with external servers as in the ModelSEED case, where a connection with the RAST server is required. Overall, the stand-alone version supports large datasets, with computing times varying significantly depending on the user’s parameter settings. Discussion Potential andlimitations The previous paragraphs illustrate the potential of microbetag in the interpretation of co-occurrence networks and how it can be used to generate new hypotheses derived from those. However, microbetag benefits the microbiome community in several other ways. microbetagDB provides a vast number of annotations: 31 predicted traits for more than 30,000 genomes, their metabolic networks along with their corresponding seed sets, potential metabolic complementarities, and cooperation/competition scores. Such a resource may support a range of studies; from a more theoretical perspective regarding the distribution of the complements among taxonomic groups or how often a complement potentially appears, to applications such as eco-evolution studies and the investigation of interactions. For example, the “Pentose phosphate pathway” module (md:M00004) was the one with the most alternatives [40] that could be completed by another species. At the same time, “De novo purine biosynthesis, PRPP + glutamine = > IMP” (md:M00048) is the module with the lowest percentage of alternatives completed by any donor (0.06%). A closer look to the latter shows that as part of the definition of the module, there are KOs that can be found only in non-bacterial taxa; e.g., among all the bacterial KEGG genomes, K11787 (phosphoribosylamine–glycine ligase) is present only in Defluviicoccus sp. SSA4 while K01587 (phosphoribosylaminoimidazole carboxylase) is only present in Rhodoplanes sp. Z2-YC6860 and Berkelbacteria bacterium GW2011_GWE1_39_12). In addition, K11787 is responsible for three steps of the module in case of Homo sapiens [M00048], suggesting a significantly different way of implementing the same task/module compared to prokaryotic taxa. For studies with a small number of sequence identifiers, the on-the-fly version of microbetag returned annotated networks in a couple of minutes while in cases where a network was also provided it only took a few seconds. The two most time-consuming steps were network inference and matching non-Silva, non-GTDB taxonomies to GTDB genomes. When microbetag’s preprocessing was used, and its results were provided as input, the computational time for the on-the-fly part was reduced to a couple of seconds even for larger networks. The computing requirements for running microbetag with custom genomes are strongly dependent on the number of bins/MAGs/genomes and whether annotation steps and GEM reconstruction have been already performed or not. Yet, there are several challenges involved in our approach. First, microbetag inherits all the biases and drawbacks of both the data and the software it is based on. Functional annotation comes with its own limitations. Some functional domains boast richer annotations and more comprehensive descriptions compared to others, which may be partly Page 21 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 due to processes being studied at different depths (e.g., glycolysis versus biosynthesis ofsecondary metabolites). In the test case (data from [33]), the bin representing the Variovorax strain was mapped to a genome that is supposed to contain the pantothenate KEGG modules. Thus, the fact that it requires pantothenate to grow, as the authors mention, would not have been predicted in the microbetag framework. This highlights the importance of strain-level variety and also the difference between genomic potential versus expression and synthesis of enzymes. Beyond the sequencing and annotation challenges, we also need to consider the fact that a pathway may not be fully represented in a KEGG module or that an organism may have alternative pathways for its product(s). Pathway complementarity can only be as accurate as the KEGG MODULE database and as precise as the software annotating genomes with KO terms. In addition, pathway complementarity per se does not guarantee that intracellular metabolites are indeed exchanged; microbetag does not check whether metabolites can be excreted or consumed. If the donor lacks a transporter and the metabolite cannot cross the cell wall, sharing could still occur through lysis, but that requires a sufficiently high death rate. With respect to seed complementarity, it is well-known that automated metabolic reconstruction comes with a number of challenges, and different tools for this task have their own limitations [45]. As a result, different reconstruction approaches can lead to different metabolic networks and thus, to different seed and non-seed sets of each species. In microbetagDB, seed complementarities have been precalculated using metabolic networks built with ModelSEED and a complete medium, which may limit potential metabolic interactions but the retrieved ones will be more reliable because using a complete medium to gapfill a metabolic network reduces the number of reactions that need to be added. In general, interaction prediction based on the seed approach does not take into account the environment. To some extent, this can be addressed when using the stand-alone version of microbetag, where the user can work with their own metabolic models or reconstruct them using specific in silico media. The fact that the environment is currently not considered also means that seeds are consistent across networks, since different seed sets would arise in different environments. Furthermore, we do not check whether any of the seed-derived metabolites is required for growth. Seeds not required in rich environments or only needed for non-essential compounds lead to false positives and thus an overestimation of potential cross-feeding interactions. These are all reasons why pathway and network complementarities are predictions of potential cross-feeding relationships and do not necessarily reflect actual interactions. In addition, pairwise relationships do not capture higher-order ecological interactions, in which species depend on (or are influenced by) multiple other species [1]. However, since microbetag is decoupled from network inference, it could annotate a network with hyperedges (i.e., edges connecting more than two taxa) produced by a future tool capable of inferring higher-order interactions. Last, the limited number of archaea in microbetagDB is also the result of a software limitation. As shown in ([46]; Fig.6b), the original version of CheckM [47] that is still being used by GTDB returns lower completeness scores for genomes that correspond to phyla known for having smaller genomes in general, e.g., Patescibacteria representative genomes in GTDB have an average completeness of ∼65%. Thus, only a few Page 22 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 representatives from these taxonomic groups passed our filters leading to an important under-representation of archaea. Context Microbetag is among a number of tools that are based on the reverse ecology approach, whose goal is to derive ecological insight, in particular on interactions, from the genomes of community members [25, 48–50]. microbetag goes beyond these previous tools and methods by combining interaction prediction based on metabolic networks not only with microbial network inference but also with the systematic annotation of taxa with phenotypic properties. In addition, it makes all these approaches accessible to researchers without bioinformatics skills. As such, it is a unique resource of potential use to everyone working with microbiome data. Future work In the near future, we plan to develop two main features: (a) the integration of transcriptomics data provided by the user, which would enhance or lower the probability for a potential metabolic interaction to occur based on whether the KO terms involved are present or not, and (b) the integration of spatial data; it is well-known that the distance between cells determines whether an interaction occurs [51]. For this, we intend to support data with spatial dimensions. We may also consider the integration of additional phenotype prediction tools, such as bacLIFE [52]. Conclusions Co-occurrence networks are widely used in microbiome studies to explore associations. However, their inference and their interpretation come with several challenges. Metabolic exchanges among microbial taxa are considered ubiquitous [53] in a large number of environments. In our study, we applied the reverse ecology paradigm [23, 24, 29, 54, 55] and publicly available genomic data and software to predict phenotypic traits and construct metabolic networks to annotate co-occurrence networks derived from amplicon or shotgun data. Our annotation was in agreement with the study of Hessler etal. [33] predicting thiamine-related metabolic interactions among Variovorax and its closest neighbors, suggesting several ways to achieve them. Using the Cabrera etal. dataset [39], we showcased the potential of microbetag as a hypothesis-generating tool suggesting mechanisms for the increase of butyrate producers in infant microbiota supplemented with iron and GOSFOS. microbetag is the first one-stop-shop platform for the annotation of microbial co-occurrence networks, highlighting the potential of data integration combined with follow-up analyses for network interpretation. Methods Genomes included Using the GTDB v207 metadata files for bacteria and archaea, we retrieved the NCBI genome accessions of the high-quality representative genomes, i.e., completeness ≥ 95% and contamination ≤ 5%. A set of 34,608 genomes was obtained, representing 25,294 unique NCBI Taxonomy Ids, since there are cases where more than one GTDB species Page 23 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 maps to the same NCBI Taxonomy id. All genomes were annotated with COGs (see “Genome-based node annotation”). For 16,900 genomes with available amino acid sequence files (.faa), these files were utilized to identify potential pathway complementarities between genome pairs (see “Pathway complementarity”). Additionally, when available, corresponding annotations from the PATRIC database [56] were retrieved to reconstruct metabolic networks (see “Seed scores and complementarity”). The microbetag workflow As shown in Fig.8, the microbetag workflow expects an abundance table representing either amplicon or shotgun data. If a microbial network is already available, the user may provide it too as input. The microbetag workflow will first map the taxa present in the abundance table to their corresponding GTDB representative genomes if Fig. 8 Diagram of microbetag’s on-the-fly workflow. microbetag expects either an abundance table as input and infers a microbial network using FlashWeave or an abundance table along with an already inferred network. After mapping taxa to GTDB reference genomes, for those with sufficient taxonomic resolution, phenotypic attributes are assigned to the nodes. Literature-based annotations of the nodes are also added using FAPROTAX. On the edge level, microbetag assigns the precalculated potential complements based on the pathway and the seed complementarity approaches. microbetag supports optional network clustering with manta. The annotated network can then be parsed into Cytoscape using the MGG app. Page 24 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 that is possible, i.e., in case the taxonomy provided does reach the species or the strain level (see “Taxonomy schemes and genome assignment”). If a network is not provided, microbetag will build one using FlashWeave. Then, the abundance table will be used for a literature-based annotation using FAPROTAX. This is the only annotation step that is microbetagDB-independent within the web-service workflow. The nodes of the network will be further annotated with phenotypic traits based on phenotrex predictions [57]. Edges linking taxa assigned to the species or strain level will be annotated with pathway and seed complementarities and seed scores. Last, network clustering will be performed with manta, assigning each node to a cluster. The annotated network is then returned in a.cx2 format. The user may skip any of these annotation steps if not needed for their analysis. Taxonomy schemes andgenome assignment Microbetag’s on-the-fly version needs to map the taxonomy of an entry in the abundance table to its corresponding NCBI Taxonomy Id and, if available, its closest GTDB representative genome(s). Several GTDB representative genomes may map to the same NCBI Taxonomy Id. Two well established taxonomy schemes are supported: the GTDB [58] that is being broadly used for bins and/or MAGs (Metagenome Assembled Genomes) taxonomic classification and the Silva database [59] that is widely used in amplicon studies. Both taxonomy schemes link their taxonomies to NCBI Taxonomy IDs [60]. In case Silva, GTDB, or the GTDB-specific 16S rRNA database applied in the microbetag preprocessing tool (see “microbetag preprocessing” section) is used, microbetag requires an exact taxonomy match to assign a genome to the sequence/taxonomy under study. In case neither of those taxonomies is used, and the abundance table contains less than 1000 taxa, microbetag maps the user-provided taxonomies to the NCBI Taxonomy. To this end, microbetag makes use of the fuzzywuzzy library (v. 0.18.0) that implements the Levenshtein distance metric to get the closest NCBI taxon name and thus its corresponding NCBI Taxonomy Id; a high similarity score is used (90) to avoid false positives. Also, using the nodes dump file of NCBI Taxonomy, microbetag may retrieve the child taxa of a taxon in user data, along with their corresponding NCBI Taxonomy Ids, if requested by the user. If the user provides their abundance table with taxonomies already mapped to the GTDB taxonomy, microbetag will efficiently match them to corresponding entries in microbetagDB and return their associated annotations. Network inference When a microbial network is not provided by the user, microbetag relies on FlashWeave (v. 0.19.2) [8] to build one on the fly. microbetag supports the annotation of networks built from any algorithm/software, in any format Cytoscape can load. microbetag preprocessing To aid the user to map their sequences to the GTDB taxonomy, DADA2-formatted 16S rRNA gene sequences for both bacteria and archaea [61] were used to train the IDTAXA classifier of the DECIPHER package (v. 2.14.0) [62] and are available through the microbetag preprocessing Docker image. Likewise, when the abundance Page 25 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 table consists of more than 1000 taxa and/or the taxonomy scheme is not among those that are automatically mapped, providing a network as an input is mandatory. The microbetag preprocessing Docker image also supports the inference of a network using FlashWeave. Literature‑based node annotation Using a set of Tara Ocean samples [63], FAPROTAX [64] estimates the functional potential of the bacterial and archaeal communities, by classifying each taxonomic unit into functional group(s) based on the current literature, descriptions of cultured representatives, and/or manuals of systematic microbiology. In this manually curated approach, a taxon is associated with a function only if all the cultured species within the taxon have been shown to exhibit that function. In the version microbetag makes use of, FAPROTAX (v. 1.2.4) includes more than 80 functions based on 7600 functional annotations and covers more than 4600 taxa. Contrary to gene-contentbased approaches, e.g., PICRUSt2 [65], FAPROTAX estimates metabolic phenotypes based on experimental evidence. microbetag invokes the accompanying script of FAPROTAX and converts the taxonomic microbial community profile of the samples included in the user’s abundance table or of the taxa present in the provided network, into putative functional profiles. Then, it parses FAPROTAX’s sub-tables to annotate each taxonomic unit present in the user’s data with all the functions for which they had a hit. FAPROTAX annotations are not part of the microbetagDB but are computed on the fly. Genome‑based node annotation phenDB [57] is a publicly available resource that supports the analysis of bacterial (meta)genomes to identify distinct functional traits, e.g., whether a species is producing butanol or has a halophilic lifestyle. It relies on support vector machines (SVM) trained with manually curated datasets based on gene presence/absence patterns for trait prediction. More specifically, the model for a particular trait is trained using a collection of EggNOG annotated genomes [66] where the knowledge of whether that trait is present or absent among its members is available. These models (classifiers) are used to predict presence/absence of their corresponding traits in non-studied species. For microbetagDB, classifiers were re-trained using the genomes provided by phenDB for each trait to sync with the latest version of EggNOG (v. 5.0.0) [66] and the phenotrex (v. 0.6.0) [57] software tool. Genomes were downloaded from NCBI using the Batch Entrez program. Then, genotype files were produced for all the high-quality GTDB representative genomes. Each model was then used against all the GTDB genotype files to annotate each with the presence or the absence of the trait. A list of all the phenotypic traits tested for the genomes present in microbetagDB is available on microbetag’s documentation site. The updated models are also available. Figure9 summarizes these precalculation steps. An exhaustive list of the traits whose presence or absence is predicted for each genome is available in Additional file2: TableS1. Page 32 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 5. Faust K, Sathirapongsasuti JF, Izard J, Segata N, Gevers D, Raes J, et al. Microbial co-occurrence relationships in the human microbiome. PLoS Comput Biol. 2012;8(7):e1002606. 6. Friedman J, Alm EJ. Inferring correlation networks from genomic survey data. PLoS Comput Biol. 2012;8(9):e1002687. 7. Kurtz ZD, Müller CL, Miraldi ER, Littman DR, Blaser MJ, Bonneau RA. Sparse and compositionally robust inference of microbial ecological networks. PLoS Comput Biol. 2015;11(5):e1004226. 8. Tackmann J, Rodrigues JFM, von Mering C. Rapid inference of direct interactions in large-scale ecological networks from heterogeneous microbial sequencing data. Cell Syst. 2019;9(3):286-296.e8. 9. Cao HT, Gibson TE, Bashan A, Liu YY. Inferring human microbial dynamics from temporal metagenomics data: pitfalls and lessons. Bioessays. 2017;39(2):1600188. 10. Faust K. Open challenges for microbial network construction and analysis. ISME J. 2021;15(11):3111–8. 11. Weiss S, Van Treuren W, Lozupone C, Faust K, Friedman J, Deng Y, et al. Correlation detection strategies in microbial data sets vary widely in sensitivity and precision. ISME J. 2016;10(7):1669–81. 12. Kishore D, Birzu G, Hu Z, DeLisi C, Korolev KS, Segrè D. Inferring microbial co-occurrence networks from amplicon data: a systematic evaluation. MSystems. 2023;8(4):e00961-22. 13. Deutschmann IM, Lima-Mendez G, Krabberød AK, Raes J, Vallina SM, Faust K, et al. Disentangling environmental effects in microbial association networks. Microbiome. 2021;9(1):232. 14. Guidi L, Chaffron S, Bittner L, Eveillard D, Larhlimi A, Roux S, et al. Plankton networks driving carbon export in the oligotrophic ocean. Nature. 2016;532(7600):465–70. 15. Röttjers L, Faust K. Manta: a clustering algorithm for weighted ecological networks. mSystems. 2020;5(1):e00903-19. 16. Durot M, Bourguignon PY, Schachter V. Genome-scale models of bacterial metabolism: reconstruction and applications. FEMS Microbiol Rev. 2009;33(1):164–90. 17. Thiele I, Palsson BØ. A protocol for generating a high-quality genome-scale metabolic reconstruction. Nat Protoc. 2010;5(1):93–121. 18. Cerk K, Ugalde-Salas P, Nedjad CG, Lecomte M, Muller C, Sherman DJ, et al. Community-scale models of microbiomes: articulating metabolic modelling and metagenome sequencing. Microb Biotechnol. 2024;n/a(n/a):e14396. 19. Machado D, Andrejev S, Tramontano M, Patil KR. Fast automated reconstruction of genome-scale metabolic models for microbial species and communities. Nucleic Acids Res. 2018;46(15):7542–53. 20. Seaver SMD, Liu F, Zhang Q, Jeffryes J, Faria JP, Edirisinghe JN, et al. The ModelSEED Biochemistry Database for the integration of metabolic annotations and the reconstruction, comparison and analysis of metabolic models for plants, fungi and microbes. Nucleic Acids Res. 2021. Available from: https:// acade mic. oup. com/ nar/ advan ceartic le/ doi/ 10. 1093/ nar/ gkaa1 143/ 59768 30. Cited 2020 Nov 23. 21. Borenstein E, Kupiec M, Feldman MW, Ruppin E. Large-scale reconstruction and phylogenetic analysis of metabolic environments. Proc Natl Acad Sci U S A. 2008;105(38):14482–7. 22. Parter M, Kashtan N, Alon U. Environmental variability and modularity of bacterial metabolic networks. BMC Evol Biol. 2007;7(1):169. 23. Kreimer A, Doron-Faigenboim A, Borenstein E, Freilich S. NetCmpt: a network-based tool for calculating the metabolic competition between bacterial species. Bioinformatics. 2012;28(16):2195–7. 24. Levy R, Carr R, Kreimer A, Freilich S, Borenstein E. Netcooperate: a network-based tool for inferring host-microbe and microbe-microbe cooperation. BMC Bioinformatics. 2015;16(1):164. 25. Zelezniak A, Andrejev S, Ponomarova O, Mende DR, Bork P, Patil KR. Metabolic dependencies drive species cooccurrence in diverse microbial communities. Proc Natl Acad Sci U S A. 2015;112(20):6449–54. 26. Belcour A, Frioux C, Aite M, Bretaudeau A, Hildebrand F, Siegel A. Metage2metabo, microbiota-scale metabolic complementarity for the identification of key species. eLife. 2020;9:e61968. 27. Lam TJ, Stamboulian M, Han W, Ye Y. Model-based and phylogenetically adjusted quantification of metabolic interaction between microbial species. PLoS Comput Biol. 2020;16(10):e1007951. 28. Muto A, Kotera M, Tokimatsu T, Nakagawa Z, Goto S, Kanehisa M. Modular architecture of metabolic pathways revealed by conserved sequences of reactions. J Chem Inf Model. 2013;53(3):613–22. 29. Levy R, Borenstein E. Reverse ecology: from systems to environments and back. In: Soyer OS, editor. Evolutionary systems biology [Internet]. New York, NY: Springer; 2012. p. 329–45. (Advances in Experimental Medicine and Biology). Available from: https:// doi. org/ 10. 1007/ 978-146143567-9_ 15. Cited 2024 Jan 21. 30. Levy R, Borenstein E. Metagenomic systems biology and metabolic modeling of the human microbiome. Gut Microbes. 2014;5(2):265–70. 31. Shannon P, Markiel A, Ozier O, Baliga NS, Wang JT, Ramage D, et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 2003;13(11):2498–504. 32. Saito R, Smoot ME, Ono K, Ruscheinski J, Wang PL, Lotia S, et al. A travel guide to Cytoscape plugins. Nat Methods. 2012;9(11):1069–76. 33. Hessler T, Huddy RJ, Sachdeva R, Lei S, Harrison STL, Diamond S, et al. Vitamin interdependencies predicted by metagenomics-informed network analyses and validated in microbial community microcosms. Nat Commun. 2023;14(1):4768. 34. Chaumeil PA, Mussig AJ, Hugenholtz P, Parks DH. GTDB-tk: a toolkit to classify genomes with the genome taxonomy database. Bioinformatics. 2020;36(6):1925–7. 35. Carbajal-Rodríguez I, Stöveken N, Satola B, Wübbeler JH, Steinbüchel A. Aerobic degradation of mercaptosuccinate by the gram-negative bacterium Variovorax paradoxus strain B4. J Bacteriol. 2011;193(2):527–39. 36. Han JI, Choi HK, Lee SW, Orwin PM, Kim J, LaRoe SL, et al. Complete genome sequence of the metabolically versatile plant growth-promoting endophyte Variovorax paradoxus S110. J Bacteriol. 2011;193(5):1183–90. 37. Sun J, Matsumoto K, Nduko JM, Ooi T, Taguchi S. Enzymatic characterization of a depolymerase from the isolated bacterium Variovorax sp. C34 that degrades poly(enriched lactate-co-3-hydroxybutyrate). Polym Degrad Stab. 2014;110:44–9. 38. Astafyeva Y, Gurschke M, Qi M, Bergmann L, Indenbirken D, de Grahl I, et al. Microalgae and bacteria interaction— evidence for division of diligence in the alga microbiota. Microbiol Spectr. 2022;10(4):e00633–22. Page 33 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 39. Cabrera Momo Rachmühl C, Derrien M, Bourdet-Sicard R, Lacroix C, Geirnaert A. Comparative prebiotic potential of galactoand fructo-oligosaccharides, native inulin, and acacia gum in Kenyan infant gut microbiota during iron supplementation. ISME Commun. 2024;4(1):ycae033. 40. Louis P, Flint HJ. Formation of propionate and butyrate by the human colonic microbiota. Environ Microbiol. 2017;19(1):29–41. 41. Huttenhower C, Gevers D, Knight R, Abubucker S, Badger JH, Chinwalla AT, et al. Structure, function and diversity of the healthy human microbiome. Nature. 2012;486(7402):207–14. 42. Gonzalez A, Navas-Molina JA, Kosciolek T, McDonald D, Vázquez-Baeza Y, Ackermann G, et al. Qiita: rapid, webenabled microbiome meta-analysis. Nat Methods. 2018;15(10):796–8. 43. Ng HM, Kin LX, Dashper SG, Slakeski N, Butler CA, Reynolds EC. Bacterial interactions in pathogenic subgingival plaque. Microb Pathog. 2016;1(94):60–9. 44. Ramoneda J, Jensen TBN, Price MN, Casamayor EO, Fierer N. Taxonomic and environmental distribution of bacterial amino acid auxotrophies. Nat Commun. 2023;14(1):7608. 45. Mendoza SN, Olivier BG, Molenaar D, Teusink B. A systematic assessment of current genome-scale metabolic reconstruction tools. Genome Biol. 2019;20(1):158. 46. Chklovski A, Parks DH, Woodcroft BJ, Tyson GW. CheckM2: a rapid, scalable and accurate tool for assessing microbial genome quality using machine learning. Nat Methods. 2023;20(8):1203–12. 47. Parks DH, Imelfort M, Skennerton CT, Hugenholtz P, Tyson GW. CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 2015;25(7):1043–55 Jul 1; 48. Levy R, Borenstein E. Metabolic modeling of species interaction in the human microbiome elucidates communitylevel assembly rules. Proc Natl Acad Sci U S A. 2013;110(31):12804–9. 49. Mendes-Soares H, Mundy M, Soares LM, Chia N. Mminte: an application for predicting metabolic interactions among the microbial species in a community. BMC Bioinformatics. 2016;17(1):343. 50. Brunner JD, Gallegos-Graves LA, Kroeger ME. Inferring microbial interactions with their environment from genomic and metagenomic data. PLoS Comput Biol. 2023;19(11):e1011661. 51. Dal Co A, van Vliet S, Kiviet DJ, Schlegel S, Ackermann M. Short-range interactions govern the dynamics and functions of microbial communities. Nat Ecol Evol. 2020;4(3):366–75. 52. Guerrero-Egido G, Pintado A, Bretscher KM, Arias-Giraldo LM, Paulson JN, Spaink HP, et al. bacLIFE: a user-friendly computational workflow for genome analysis and prediction of lifestyle-associated genes in bacteria. Nat Commun. 2024;15(1):2072. 53. Kost C, Patil KR, Friedman J, Garcia SL, Ralser M. Metabolic exchanges are ubiquitous in natural microbial communities. Nat Microbiol. 2023;8(12):2244–52. 54. Tal O, Selvaraj G, Medina S, Ofaim S, Freilich S. Netmet: a network-based tool for predicting metabolic capacities of microbial species and their interactions. Microorganisms. 2020;8(6):840. 55. Tal O, Bartuv R, Vetcos M, Medina S, Jiang J, Freilich S. NetCom: a network-based tool for predicting metabolic activities of microbial communities based on interpretation of metagenomics data. Microorganisms. 2021;9(9):1838. 56. Wattam AR, Davis JJ, Assaf R, Boisvert S, Brettin T, Bun C, et al. Improvements to PATRIC, the all-bacterial bioinformatics database and analysis resource center. Nucleic Acids Res. 2017;45(D1):D535–42. 57. Feldbauer R, Schulz F, Horn M, Rattei T. Prediction of microbial phenotypes based on comparative genomics. BMC Bioinformatics. 2015;16(14):S1. 58. Parks DH, Chuvochina M, Rinke C, Mussig AJ, Chaumeil PA, Hugenholtz P. GTDB: an ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Res. 2022;50(D1):D785-94. 59. Quast C, Pruesse E, Yilmaz P, Gerken J, Schweer T, Yarza P, et al. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res. 2013;41(D1):D590–6. 60. Schoch CL, Ciufo S, Domrachev M, Hotton CL, Kannan S, Khovanskaya R, et al. NCBI Taxonomy: a comprehensive update on curation, resources and tools. Database J Biol Databases Curation. 2020;2020:baaa062. 61. Alishum A. DADA2 formatted 16S rRNA gene sequences for both bacteria & archaea. Zenodo; 2022. Available from: https:// zenodo. org/ recor ds/ 66556 92. Cited 2024 Aug 8. 62. Murali A, Bhargava A, Wright ES. IDTAXA: a novel approach for accurate taxonomic classification of microbiome sequences. Microbiome. 2018;6(1):140. 63. Sunagawa S, Coelho LP, Chaffron S, Kultima JR, Labadie K, Salazar G, et al. Structure and function of the global ocean microbiome. Science. 2015;348(6237):1261359–1261359. 64. Louca S, Parfrey LW, Doebeli M. Decoupling function and taxonomy in the global ocean microbiome. Science. 2016;353(6305):1272–7. 65. Douglas GM, Maffei VJ, Zaneveld JR, Yurgel SN, Brown JR, Taylor CM, et al. PICRUSt2 for prediction of metagenome functions. Nat Biotechnol. 2020;38(6):685–8. 66. Huerta-Cepas J, Szklarczyk D, Heller D, Hernández-Plaza A, Forslund SK, Cook H, et al. eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Res. 2019;47(D1):D309–14. 67. Aramaki T, Blanc-Mathieu R, Endo H, Ohkubo K, Kanehisa M, Goto S, et al. Kofamkoala: KEGG ortholog assignment based on profile HMM and adaptive score threshold. Bioinformatics. 2020;36(7):2251–2. 68. Kanehisa M, Goto S, Sato Y, Furumichi M, Tanabe M. KEGG for integration and interpretation of large-scale molecular data sets. Nucleic Acids Res. 2012;40(D1):D109-14. 69. Kanehisa M, Sato Y, Kawashima M. Kegg mapping tools for uncovering hidden features in biological data. Protein Sci. 2022;31(1):47–53. 70. Kanehisa M, Sato Y. Kegg mapper for inferring cellular functions from protein sequences. Protein Sci. 2020;29(1):28–35. 71. Henry CS, DeJongh M, Best AA, Frybarger PM, Linsay B, Stevens RL. High-throughput generation, optimization and analysis of genome-scale metabolic models. Nat Biotechnol. 2010;28(9):977–82. Page 34 of 34 Zafeiropoulosetal. Genome Biology (2025) 26:292 72. Overbeek R, Olson R, Pusch GD, Olsen GJ, Davis JJ, Disz T, et al. The SEED and the Rapid Annotation of microbial genomes using Subsystems Technology (RAST). Nucleic Acids Res. 2014;42(Database issue):D206–214. 73. Brettin T, Davis JJ, Disz T, Edwards RA, Gerdes S, Olsen GJ, et al. RASTtk: a modular and extensible implementation of the RAST algorithm for building custom annotation pipelines and annotating batches of genomes. Sci Rep. 2015;5(1):8365. 74. Choudhary K, Meng EC, Diaz-Mejia JJ, Bader GD, Pico AR, Morris JH. scNetViz: from single cells to networks using Cytoscape [Internet]. F1000Research; 2021. Available from: https:// f1000 resea rch. com/ artic les/ 10448. Cited 2024 Aug 8. 75. Merkel D. Docker: lightweight Linux containers for consistent development and deployment. Linux J. 2014;2014(239):2:2. 76. Reese W. Nginx: the high-performance web server and reverse proxy. Linux J. 2008;2008(173):2:2. 77. Zafeiropoulos H. msysbio/microbetag. GitHub. 2025. https:// github. com/ msysb io/ micro betag. 78. Zafeiropoulos H. msysbio/microbetag: v1.0.4. Zenodo. 2025. https:// doi. org/ 10. 5281/ zenodo. 15619 540. 79. Zafeiropoulos H. msysbio/microbetagApp-public. GitHub. 2025. https:// github. com/ msysb io/ micro betag Apppublic. 80. Zafeiropoulos H. msysbio/microbetagApp-public: v1.0.1 - sync to zenodo. Zenodo. 2025. https:// doi. org/ 10. 5281/ zenodo. 16905 743. 81. Morris S, Michail Delopoulos EI, Zafeiropoulos H, Pico A, Mo. msysbio/MGG: v1.0.2 - sync to zenodo. GitHub. 2025. https:// github. com/ msysb io/ MGG. 82. Morris S, Michail Delopoulos EI, Zafeiropoulos H, Pico A, Mo. msysbio/MGG: v1.0.2 - sync to zenodo. Zenodo. 2025. https:// doi. org/ 10. 5281/ zenodo. 16905 768. 83. Mitchell AL, Scheremetjew M, Denise H, Potter S, Tarkowska A, Qureshi M, et al. EMG produced TPA metagenomics assembly of the Metagenomic analysis of thiocyanate and cyanide biodegrading microbial consortia Metagenome (bioreactor metagenome) data set. 2019. https:// www. ncbi. nlm. nih. gov/ biopr oject/ 513540. 84. Kantor RS, van Zyl AW, van Hille RP, Thomas BC, Harrison ST, Banfield JF. Metagenomic analysis of thiocyanate and cyanide biodegrading microbial consortia Metagenome. 2015. https:// www. ncbi. nlm. nih. gov/ biopr oject/ PRJNA 279279. 85. Huddy RJ, Sachdeva R, Kadzinga F, Kantor RS, Harrison STL, Banfield JF. Thiocyanate and organic carbon inputs drive convergent selection for specific autotrophic Afipia and Thiobacillus strains within complex microbiomes. Front Microbiol. 2021. https:// doi. org/ 10. 3389/ fmicb. 2021. 643368. 86. Momo Cabrera P, Rachmühl C, Derrien M, Bourdet-Sicard R, Lacroix C, Geirnaert A. Distinct fiber effect in ex vivo Kenyan infant gut microbiota. 2023. https:// www. ebi. ac. uk/ ena/ brows er/ view/ PRJEB 67393. Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.