Network-based integration of metabolomics data from large-scale repositories.

Publication date: Jul 15, 2026

Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github. com/EloisaRL/Metabolomic-data-analysis-app/tree/main . As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Open Access PDF

Concepts Keywords
Biocuration
Data integration
Databases, Factual
Harmonised annotation
Humans
Metabolomics
Metadata
Networks
Public data reuse
Repositories
Software

Semantics

Type Source Name
drug DRUGBANK Coenzyme M
disease MESH COVID-19
drug DRUGBANK Pentaerythritol tetranitrate
drug DRUGBANK Pidolic Acid
pathway REACTOME Metabolism
pathway REACTOME Digestion
pathway REACTOME Reproduction
disease MESH CVs
disease MESH PCA
drug DRUGBANK L-Valine
pathway REACTOME Release
drug DRUGBANK Methionine
drug DRUGBANK Alpha-1-proteinase inhibitor
disease MESH image
drug DRUGBANK Hyaluronic acid
disease MESH dis
drug DRUGBANK Proline
drug DRUGBANK Indoleacetic acid
disease MESH hepatitis
disease MESH tuberculosis
pathway KEGG Tuberculosis
disease MESH obesity
disease MESH colorectal cancer
pathway KEGG Colorectal cancer
drug DRUGBANK Ilex paraguariensis leaf
drug DRUGBANK L-Phenylalanine
drug DRUGBANK Amino acids
drug DRUGBANK Water
drug DRUGBANK Arachidonic Acid
drug DRUGBANK Taurine
drug DRUGBANK Citric Acid
drug DRUGBANK L-Tryptophan
drug DRUGBANK L-Alanine
pathway KEGG Tryptophan metabolism
disease MESH long COVID
disease MESH MS2
drug DRUGBANK Phenformin
disease MESH included
drug DRUGBANK Sparfosic acid
disease MESH inflammation
disease MESH insulin resistance
pathway KEGG Insulin resistance
disease MESH chronic fatigue syndrome
disease MESH Cap
disease MESH Ger
disease MESH dysbiosis
drug DRUGBANK L-Aspartic Acid
drug DRUGBANK Kale
disease MESH severe acute respiratory syndrome
drug DRUGBANK Tocilizumab
drug DRUGBANK Dosulepin
disease MESH Park
drug DRUGBANK Guanosine
drug DRUGBANK Troleandomycin
disease MESH Chai

Original Article

(Visited 8 times, 1 visits today)

Leave a Comment

Your email address will not be published. Required fields are marked *

Network-based integration of metabolomics data from large-scale repositories.

Publication date: Jul 15, 2026

Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github. com/EloisaRL/Metabolomic-data-analysis-app/tree/main . As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Open Access PDF

Concepts Keywords
Biocuration
Data integration
Databases, Factual
Harmonised annotation
Humans
Metabolomics
Metadata
Networks
Public data reuse
Repositories
Software

Semantics

Type Source Name
drug DRUGBANK Coenzyme M
disease MESH COVID-19
drug DRUGBANK Pentaerythritol tetranitrate
drug DRUGBANK Pidolic Acid
pathway REACTOME Metabolism
pathway REACTOME Digestion
pathway REACTOME Reproduction
disease MESH CVs
disease MESH PCA
drug DRUGBANK L-Valine
pathway REACTOME Release
drug DRUGBANK Methionine
drug DRUGBANK Alpha-1-proteinase inhibitor
disease MESH image
drug DRUGBANK Hyaluronic acid
disease MESH dis
drug DRUGBANK Proline
drug DRUGBANK Indoleacetic acid
disease MESH hepatitis
disease MESH tuberculosis
pathway KEGG Tuberculosis
disease MESH obesity
disease MESH colorectal cancer
pathway KEGG Colorectal cancer
drug DRUGBANK Ilex paraguariensis leaf
drug DRUGBANK L-Phenylalanine
drug DRUGBANK Amino acids
drug DRUGBANK Water
drug DRUGBANK Arachidonic Acid
drug DRUGBANK Taurine
drug DRUGBANK Citric Acid
drug DRUGBANK L-Tryptophan
drug DRUGBANK L-Alanine
pathway KEGG Tryptophan metabolism
disease MESH long COVID
disease MESH MS2
drug DRUGBANK Phenformin
disease MESH included
drug DRUGBANK Sparfosic acid
disease MESH inflammation
disease MESH insulin resistance
pathway KEGG Insulin resistance
disease MESH chronic fatigue syndrome
disease MESH Cap
disease MESH Ger
disease MESH dysbiosis
drug DRUGBANK L-Aspartic Acid
drug DRUGBANK Kale
disease MESH severe acute respiratory syndrome
drug DRUGBANK Tocilizumab
drug DRUGBANK Dosulepin
disease MESH Park
drug DRUGBANK Guanosine
drug DRUGBANK Troleandomycin
disease MESH Chai

Original Article

Leave a Comment

Your email address will not be published. Required fields are marked *