Public bioinformatics databases are online resources that store biological data and make it available to researchers for study and analysis. These databases contain different types of data, such as DNA sequences, RNA sequencing reads, gene expression data, and cancer-related information. Researchers use these databases to study biological processes, compare samples, and analyze existing experimental data without collecting new samples for every research project.
In this article, we will learn about four important public bioinformatics databases: SRA (Sequence Read Archive), GEO (Gene Expression Omnibus), TCGA (The Cancer Genome Atlas), and ENA (European Nucleotide Archive). We will also explore how to find, download, and organize datasets using command-line tools and R packages.
Understanding Public Bioinformatics Databases
Public bioinformatics databases provide access to biological data collected and submitted by researchers. It depends on the database, you can find raw sequencing reads, processed gene expression tables, cancer-related genomic data, and information about the experiments. Each database serves a different purpose, so choosing the right one depends on the type of data you need.
Sequence Read Archive (SRA)
The Sequence Read Archive (SRA) is maintained by the National Center for Biotechnology Information (NCBI). It stores raw sequencing reads submitted by researchers. SRA is useful when you want to perform your own analysis, such as quality control, read alignment, or RNA-seq analysis.
For example, you can download RNA-seq reads from SRA and process them using tools such as FastQC, HISAT2, STAR, or other suitable bioinformatics software.
Gene Expression Omnibus (GEO)
The Gene Expression Omnibus (GEO) is an NCBI database that stores functional genomics data from different experiments. It provides information about studies and samples, along with processed data and supplementary files when available. Depending on the study, you may find gene expression matrices, normalized values, raw data files, or links to sequencing data in SRA.
GEO is especially useful when you want to explore an existing study or begin analyzing processed gene expression data.
The Cancer Genome Atlas (TCGA)
The Cancer Genome Atlas (TCGA) is a major cancer research program that collected molecular and clinical data from thousands of cancer cases. Its datasets include gene expression measurements, somatic mutations, copy number alterations, DNA methylation data, and clinical information. You can access TCGA datasets through the National Cancer Institute's Genomic Data Commons (GDC) Data Portal. Some data are available to everyone, while other files require approved access.
TCGA is useful for studying cancer-related gene expression, comparing tumor samples, and exploring molecular changes associated with different cancer types.
European Nucleotide Archive (ENA)
The European Nucleotide Archive (ENA), maintained by EMBL-EBI, stores nucleotide sequencing data, including raw reads and related information. ENA is part of the international network of nucleotide sequence archives that exchange data with resources such as NCBI SRA and DDBJ. One useful feature is that ENA provides downloadable FASTQ files for many sequencing runs. This can save you the step of converting SRA files into FASTQ format, provided the required files are available.
Difference Between SRA, GEO, TCGA, and ENA
Several differences between the SRA, GEO, TCGA and ENA, which are as follows:
| Database | Main purpose | Common data |
|---|---|---|
| SRA | Access raw sequencing reads. | Sequencing runs |
| GEO | Explore functional genomics studies. | Expression matrices, sample information, supplementary files |
| TCGA | Study cancer genomics. | Gene expression, mutations, clinical data |
| ENA | Access nucleotide and sequencing data. | Raw reads, FASTQ files, sequence metadata |
Downloading Raw Sequencing Data Using the SRA Toolkit
The SRA Toolkit is a collection of command-line tools used to access and process data from SRA. It includes tools such as prefetch and fasterq-dump. The first downloads sequencing data in SRA format, while the second converts the downloaded data into FASTQ files.
Understanding SRA Accession Numbers
SRA records use accession numbers to identify studies, experiments, and sequencing runs.
| Accession | Meaning |
|---|---|
| SRP | An SRA study or project |
| PRJNA | An NCBI BioProject |
| SRX | An SRA experiment |
| SRR | An SRA sequencing run |
For example, SRR1039508 identifies a sequencing run. You can use this accession to locate and download the corresponding sequencing data.
A project can contain multiple experiments, and an experiment can contain multiple sequencing runs. Always check the database record to understand how the samples and runs are related.
Checking the SRA Toolkit Installation
After installing the SRA Toolkit, open a terminal and run:
prefetch --version
If the command displays the installed version, the tool is available in your terminal. If the command is not recognized, check the installation and make sure the toolkit's executable directory is included in your system's PATH.
You can find installation instructions in the official NCBI SRA Toolkit documentation.
Downloading an SRA run with prefetch
The prefetch command downloads sequencing data associated with an SRA accession.
For example:
prefetch SRR1039508
By default, the command creates a directory named after the accession and downloads the required data into it. Using prefetch is a convenient approach because an interrupted download can often be resumed by running the same command again. Before downloading large datasets, check the available storage space on your computer.
Converting SRA Files into FASTQ Format
Many bioinformatics tools use FASTQ files for sequencing analysis. If you have downloaded an SRA file, you can convert it into FASTQ format using the fasterq-dump command.
For paired-end sequencing data, use the following command:
mkdir -p fastq fasterq-dump SRR1039508 --split-files -e 8 -O fastq/
Command options:
- SRR1039508: Specifies the sequencing run.
- --split-files: Saves paired-end reads in separate files.
- -e 8: Uses eight worker threads.
- -O fastq/: Sets the output directory.
For paired-end data, the command may create SRR1039508_1.fastq and SRR1039508_2.fastq. Single-end data usually produces one FASTQ file.
Run the command from the appropriate directory or provide the correct path to the downloaded SRA data. Make sure enough disk space is available before starting the conversion.
Compressing FASTQ Files
FASTQ files can take up a large amount of disk space. You can compress them using gzip:
gzip fastq/*.fastq
This produces compressed files with the .fastq.gz extension, which many bioinformatics tools can read directly.
Important: fasterq-dump can require substantial temporary storage during conversion. Check the expected output size and available disk space before processing large sequencing runs.
Finding and Downloading Datasets from GEO
GEO is useful when you want to explore an existing gene expression study or obtain processed data without downloading and processing every raw sequencing read.
Understanding GEO Accession Numbers
GEO uses three common accession types.
| Accession | Meaning |
|---|---|
| GSE | A study or series |
| GSM | An individual sample |
| GPL | A platform used in the experiment |
For example, a GSE record may contain several GSM sample records. Each sample record provides information about the sample and its role in the experiment. The GPL record describes the platform associated with the experiment, such as a microarray platform.
Searching for a GEO Dataset
Follow these steps to find a suitable dataset:
- Open the NCBI GEO website.
- Search for a topic such as breast cancer gene expression or RNA-seq.
- Open a study that matches your research question.
- Read the study summary and check the organism, sample types, and experimental design.
- Review the sample records and supplementary files.
- Check whether the study provides processed expression data or links to raw sequencing reads.
For example, if you want to compare gene expression between tumor and normal samples, look for a study containing both groups and enough information to identify each sample correctly.
Use an actual accession from the study you select rather than assuming every study has the same files.
Downloading GEO Data Manually
On a GEO study page, look for the supplementary files and series matrix files. Depending on the experiment, you may find:
- Processed gene expression tables.
- Normalized expression data.
- Sample annotation files.
- Raw or supplementary data files.
- Links to sequencing runs in SRA.
Not every study provides all these files. Check the file descriptions before downloading them.
Downloading GEO Data Using R
You can also use the GEOquery package in R to retrieve data from GEO.
First, install the package:
if (!requireNamespace("BiocManager", quietly = TRUE)) { install.packages("BiocManager") } BiocManager::install("GEOquery")
Load the package and retrieve a GEO series:
library(GEOquery) gse <- getGEO("GSE12345", GSEMatrix = TRUE)
GSE12345 is an example accession placeholder. Replace it with a real GEO series accession.
If the study provides a compatible series matrix, you can inspect the expression data and sample information:
expr <- exprs(gse[[1]]) pheno <- pData(gse[[1]]) head(expr) head(pheno)
Here, exprs() extracts the expression matrix, while pData() retrieves the sample information associated with the expression dataset.
A GEO series may contain multiple expression datasets, so check the returned objects before selecting one for analysis. Also, do not assume that all expression values are raw counts; they may be normalized or transformed.
To download supplementary files, use:
getGEOSuppFiles("GSE12345")
Replace the example accession with the real series ID. Review the downloaded files to determine which one is suitable for your analysis.
For more details, see the official GEOquery package documentation.
Accessing TCGA Data Through the GDC Portal
The Genomic Data Commons (GDC) provides access to TCGA datasets through a web portal and command-line tools. It is useful when you need cancer-related molecular data along with information about samples and clinical cases.
Searching for a TCGA Dataset
To find TCGA data, follow these steps:
- Open the GDC Data Portal.
- Navigate to the available data repository.
- Select a project, such as TCGA-BRCA, which represents the breast invasive carcinoma project.
- Choose the data category and data type relevant to your research.
- Review the available files and their associated sample information.
- Add the required files to your cart.
For example, you can look for gene expression quantification files under Transcriptome Profiling. The available workflows and file types may differ between datasets, so check the metadata before downloading.
Downloading TCGA Files Using the GDC Client
For multiple files, the GDC Data Transfer Tool provides a convenient way to download data using a manifest. First, download the manifest from the GDC Data Portal and save it as gdc_manifest.txt.
Then run:
gdc-client download -m gdc_manifest.txt -d tcga_data/
Install the GDC Data Transfer Tool before running this command. For controlled-access files, you also need the appropriate authorization and authentication.
Understanding Open and Controlled-Access Data
TCGA data is divided into two main categories: open-access data and controlled-access data.
Open-access data is data that researchers can download and use without special permission. It includes many processed genomic datasets and data that do not require restricted access.
Controlled-access data is data that requires approval before researchers can download or use it. It may contain sensitive genomic or individual-level information that needs additional privacy protection.
The access requirements depend on the type of data and its usage conditions. Before downloading a dataset, check its access status on the GDC portal.
For more information, read the official GDC Data Access Guide.
Accessing TCGA Data Using R
R users can use the TCGAbiolinks package to search for, download, and prepare TCGA datasets.
Install the package using Bioconductor:
if (!requireNamespace("BiocManager", quietly = TRUE)) { install.packages("BiocManager") } BiocManager::install("TCGAbiolinks")
Next, create a query for gene expression data from the TCGA-BRCA project:
library(TCGAbiolinks) query <- GDCquery( project = "TCGA-BRCA", data.category = "Transcriptome Profiling", data.type = "Gene Expression Quantification", workflow.type = "STAR - Counts" )
Download the files and prepare the data:
GDCdownload(query) data <- GDCprepare(query)
The query selects files matching the specified project and data options. GDCdownload() downloads the matching data, while GDCprepare() prepares the downloaded files for analysis. The exact files available can change, so verify that the selected workflow and data type are supported by the current dataset. The package may also require additional setup or authentication for particular data.
Downloading FASTQ Files Directly from ENA
ENA is another useful resource for accessing sequencing data. It often provides direct links to compressed FASTQ files, which can simplify the download process.
This is especially helpful when you want raw sequencing reads for an RNA-seq or another sequencing workflow.
Finding FASTQ Download Links
ENA provides a file-report API that returns information about sequencing runs in a tabular format.
For example, the following URL requests run accessions and FASTQ file links for an NCBI BioProject:
https://www.ebi.ac.uk/ena/portal/api/filereport?accession=PRJNA123456&result=read_run&fields=run_accession,fastq_ftp,sample_alias&format=tsv
PRJNA123456 is a placeholder. Replace it with a real BioProject accession.
The returned table can include:
- run_accession: the sequencing run ID.
- fastq_ftp: links to available FASTQ files.
- sample_alias: a sample name or alias when provided.
Review the returned table and confirm that the project contains the data you need.
Downloading FASTQ Files
After obtaining the FASTQ links, you can download a file using wget.
For example, copy a real FASTQ URL from the ENA response and run:
wget "PASTE_THE_FASTQ_URL_HERE"
Replace the placeholder with the complete URL returned by ENA.
For paired-end sequencing, the response may contain two FASTQ links for each run. Download both files and keep their relationship clear.
ENA files are often already compressed as .fastq.gz, so you generally do not need to convert an SRA file or compress the downloaded FASTQ files again.
Tip: Use the ENA API to retrieve the exact links for your selected project rather than guessing the directory structure of a FASTQ URL.
Managing Bioinformatics Metadata
Metadata is information that describes a dataset, including sample details, experimental conditions, and sequencing methods. It helps researchers understand what each sample represents and how the data was collected.
For example, when comparing tumor and normal samples, metadata helps identify which files belong to each group. Proper metadata management reduces errors and helps ensure accurate data analysis.
Where to Find Metadata
You can collect metadata from several sources:
- SRA: Use the SRA Run Selector to review sample details and download available metadata and accession lists.
- GEO: Check the sample records, series matrix files, or use pData() in R.
- GDC: Review the sample, case, clinical, and file information available through the portal.
- ENA: Use the browser or file-report API to retrieve run and sample information.
You can also use tools such as pysradb to retrieve SRA metadata from the command line.
For example:
pysradb metadata SRP012345 --detailed
Replace SRP012345 with the study accession you want to investigate. Install pysradb before running the command.
6.2 Creating a Sample Sheet
A sample sheet is a table that connects sample names, accession IDs, experimental groups, and file paths.
For example:
| sample_id | run_id | condition | layout | file_path |
|---|---|---|---|---|
| S1 | SRR1039508 | control | paired | fastq/SRR1039508_1.fastq.gz |
| S1 | SRR1039508 | control | paired | fastq/SRR1039508_2.fastq.gz |
| S2 | SRR1039509 | treated | paired | fastq/SRR1039509_1.fastq.gz |
| S2 | SRR1039509 | treated | paired | fastq/SRR1039509_2.fastq.gz |
This is an illustrative example. Use the actual sample relationships and file paths from your downloaded dataset.
For paired-end data, both FASTQ files belong to the same sequencing run and biological sample. Some analysis pipelines expect one row per sample with separate read-file columns, while others accept a long-format table like the example above. Follow the format required by your chosen tool.
Best Practices for Metadata Management
Before starting the analysis, check the following details.
1. Match sample IDs and run IDs.
A single sample can have multiple sequencing runs. Keep the relationships between sample accessions and run accessions clear.
2. Check the sequencing layout.
Determine whether the data are single-end or paired-end. The layout affects how you supply the reads to downstream tools.
3. Confirm the organism and sequencing method.
Make sure the dataset matches your research question. For example, human RNA-seq data should not be confused with mouse RNA-seq or ChIP-seq data.
4. Keep sample labels consistent.
Labels such as Tumour, tumor, and TUMOR may refer to the same group. Standardize them carefully while preserving the original metadata for reference.
5. Keep the original metadata.
Save the original downloaded tables and create a separate cleaned sample sheet. This makes it easier to trace your decisions and repeat the analysis.
6. Check the files before analysis.
Confirm that all expected files were downloaded successfully and that the filenames match the sample sheet.
A Step-by-Step Workflow for Accessing Public Bioinformatics Data
You can use the following workflow when starting a new bioinformatics project.
Step 1: Define Your Research Question
Decide what you want to investigate. For example, you may want to compare gene expression between tumor and normal samples.
Step 2: Find a Suitable Study
Search GEO, SRA, ENA, or TCGA based on the type of data you need. Check the organism, experimental design, sample groups, and available files.
Step 3: Review the Metadata
Download the sample information and identify the relationship between study IDs, sample IDs, run IDs, and experimental conditions.
Step 4: Choose the Right Data Format
Use processed expression data if it is suitable for your research question. If you need to perform your own raw-read analysis, download the relevant FASTQ files.
Step 5: Download the Dataset
Choose the appropriate method:
- Use prefetch and fasterq-dump for SRA data.
- Use GEO's supplementary files or GEOquery for GEO studies.
- Use the GDC Data Portal, GDC client, or TCGAbiolinks for TCGA data.
- Use ENA's file-report API for direct sequencing file links.
Step 6: Organize and Verify the Files
Store the data in clearly named directories, check the downloaded files, and create a sample sheet that links the files to their samples and conditions.
Step 7: Record the Dataset Information
Save the accession IDs, download details, software versions, and relevant study information. Cite the original research study and follow any applicable data-use requirements.
Conclusion
Public bioinformatics databases such as SRA, GEO, TCGA, and ENA provide access to biological data for research and analysis. Tools like the SRA Toolkit, GEOquery, and TCGAbiolinks help you download and work with these datasets.
Before starting your analysis, check the metadata, organize your sample files, and verify the dataset information. This helps you avoid errors and makes your research easier to reproduce.