Abstract
Motivation: Over the last two decades, transcriptomics has become a standard technique in biomedical research. We now have large databases of RNA-seq data, accompanied by valuable metadata detailing scientific objectives and the experimental procedures used. The metadata is crucial in understanding and replicating published studies, but so far has been underutilized in helping researchers to discover existing datasets. Results: We present SampleExplorer, a tool allowing researchers to search for relevant data using both text and gene set queries. SampleExplorer embeds sample metadata and uses a transformer-based language model to retrieve similar datasets. Extensive benchmarking (see /3 (Supplementary Materials and Methods) provides detailed methodological information, including an algorithmic description of the retrieval process and data preparation steps.
Cite
CITATION STYLE
Chin, W. L., & Lassmann, T. (2025). SampleExplorer: using language models to discover relevant transcriptome data. Bioinformatics, 41(1). https://doi.org/10.1093/bioinformatics/btae759
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.