Home › Bioinformatics Tutorial › Bioinformatics File Format

Bioinformatics File Format

⏱ 11 min read Updated: 08 Oct 2026

A bioinformatics file format is a standard way of storing biological data so that it can be read and processed by different software and tools. Different formats are used for different types of data, such as DNA sequences, sequencing reads, gene annotations, genetic variants, and protein structures.

In this article, we will learn about the most commonly used bioinformatics file formats, their structure, common uses, and important differences between them.

FASTA File Format

FASTA is a simple text-based file format used to store DNA, RNA, or protein sequences. It is commonly used because its structure is easy to read and process.

A FASTA record has two main parts:

  • Header: Starts with the > symbol and contains the sequence identifier and optional description.
  • Sequence: Contains the DNA, RNA, or protein sequence.

For example:

>sample_gene1 Homo sapiens example gene ATGGCGTACGTTAGCTAGCTAGGCTAACGT TTAGCCGATCGATCGTAGCTAGCTAGCTAA

Here, sample_gene1 is the sequence identifier. The text after the identifier provides additional information about the sequence.

A FASTA file can contain one or many sequences. Long sequences can be written across multiple lines, and programs read them as a single continuous sequence.

Common extensions: .fa, .fasta, .fna for nucleotide sequences and .faa for protein sequences.

Common uses: FASTA files are used for reference genomes, gene sequences, protein sequences, and sequence searches such as BLAST.

Important: FASTA does not contain sequencing quality scores. If quality information is required, FASTQ is normally used.

FASTQ File Format

FASTQ is a text-based file format that stores sequencing reads along with a quality score for each base. It is commonly produced by DNA and RNA sequencing machines.

Each FASTQ record contains four lines:

@read1 GATTACAGGT + IIIIHHGG5#

  • Line 1: Starts with @ and contains the read identifier.
  • Line 2: Contains the sequence.
  • Line 3: Starts with + and may optionally contain the read identifier.
  • Line 4: Contains the quality scores for the sequence.

The sequence and quality string must have the same number of characters.

What is a Phred Quality Score?

A Phred quality score shows the probability that a sequencing base has been called incorrectly.

The formula is:

Q = −10 × log₁₀(P)

Here, P is the probability that the base call is wrong.

Phred ScoreError ProbabilityAccuracy
Q101 in 1090%
Q201 in 10099%
Q301 in 1,00099.9%
Q401 in 10,00099.99%

A higher Phred score means a lower probability of an incorrect base call.

GenBank File Format

GenBank is a sequence record format used by NCBI to store biological sequences together with detailed information and annotations.

Unlike FASTA, which mainly stores sequence information, a GenBank record can contain information about the organism, sequence, genes, coding regions, references, and other features.

A simplified GenBank record looks like this:

LOCUS       DEMO0001      60 bp    DNA     linear
DEFINITION  Example bacterial gene for teaching.
ACCESSION   DEMO0001
VERSION     DEMO0001.1
SOURCE      Escherichia coli
 ORGANISM  Escherichia coli

FEATURES             Location/Qualifiers
    source          1..60
                    /organism="Escherichia coli"
    gene            1..51
                    /gene="demA"
    CDS             1..51
                    /gene="demA"
                    /product="demonstration protein"

ORIGIN
       1 atggcgtacg ttagctagct aggctaacgt
      31 ttagccgatc gatcgtagct agctagctaa
//

Important parts include:

  • LOCUS: Provides basic information about the record.
  • ACCESSION: Provides the record's accession identifier.
  • FEATURES: Contains annotations such as genes and coding sequences.
  • ORIGIN: Contains the biological sequence.
  • // marks the end of the record.

A location such as 1..51 means that the feature starts at position 1 and ends at position 51.

Common use: GenBank is useful when you need both the biological sequence and its associated annotations.

GFF and GTF File Formats

GFF (General Feature Format) and GTF (Gene Transfer Format) are text-based formats used to store genomic annotations such as genes, transcripts, and exons.

They usually store annotation separately from the actual genome sequence.

A GFF3 file uses nine columns:

  1. Sequence name
  2. Source
  3. Feature type
  4. Start
  5. End
  6. Score
  7. Strand
  8. Phase
  9. Attributes

For example:

chr1    demo    gene    1000    2200    .    +    .    ID=gene1;Name=demA
chr1    demo    mRNA    1000    2200    .    +    .    ID=tx1;Parent=gene1
chr1    demo    exon    1000    1300    .    +    .    ID=ex1;Parent=tx1

The Parent attribute shows the relationship between features. For example, an exon can belong to a transcript, and the transcript can belong to a gene.

GTF uses a similar structure but has different rules for its attribute field. A common GTF record looks like this:

chr1  demo  exon  1000  1300  .  +  .  gene_id "gene1"; transcript_id "tx1";

Important: GFF and GTF coordinates are 1-based and inclusive. This means both the start and end positions are included.

Common uses: GFF and GTF files are widely used for genome annotation, gene prediction, RNA-seq analysis, and transcript analysis.

BED File Format

BED (Browser Extensible Data) is a simple text format used to describe regions of a genome.

The first three BED columns are:

  1. Chromosome
  2. Start position
  3. End position

For example:

chr1  100  200 chr1  450  520 chr2  900  1000

BED can also contain additional columns such as a name, score, and strand.

Common uses: BED files are used for genomic regions such as genes, exons, sequencing peaks, and other regions of interest. They are also commonly used with genome browsers and tools such as bedtools.

Important: BED uses 0-based, half-open coordinates. The start position is included, but the end position is not.

Difference Between 0-Based and 1-Based Coordinates

Coordinate systems are important when working with genomic data because different formats use different systems.

  • GFF/GTF: 1-based and inclusive
  • BED: 0-based and half-open
  • SAM: Alignment position (POS) is 1-based
  • VCF: Variant position (POS) is 1-based

For example, consider the sequence:

ACGTACGTAC

If we want the bases from position 3 to position 6:

1  2  3  4  5  6  7  8  9  10 A  C  G  T  A  C  G  T  A  C

In GFF/GTF, the region is written as:

3  6

In BED, the same region is written as:

2  6

The BED length is calculated as:

end − start = 6 − 2 = 4

This coordinate difference is important when converting or comparing genomic files.

SAM and BAM File Formats

SAM (Sequence Alignment/Map) is a text format used to store sequencing reads aligned to a reference sequence.

BAM stores the same type of alignment information in a compressed binary format. BAM files are smaller and faster to process than SAM files.

A SAM file contains header lines followed by alignment records.

A simplified alignment record looks like this:

read1    99    chr1    1000    60    10M    =    1050    60    GATTACAGGT    IIIIHHGG5#    NM:i:0

The main SAM columns include:

ColumnMeaning
QNAMERead name
FLAGInformation about the read and its alignment
RNAMEReference sequence name
POSAlignment start position
MAPQMapping quality
CIGARDescribes the alignment
RNEXTReference name of the mate
PNEXTPosition of the mate
TLENTemplate length
SEQRead sequence
QUALBase quality scores

CIGAR String

The CIGAR string describes how a read is aligned to the reference.

For example:

10M

means that 10 read bases are aligned using the M operation.

Other common CIGAR operations include:

  • I - Insertion
  • D - Deletion
  • S - Soft clipping
  • N - Skipped region

For example:

5M2I3M

means 5 aligned bases, followed by a 2-base insertion, followed by 3 aligned bases.

SAM FLAG

The FLAG field stores information about the alignment using a set of numeric values.

Some common values are:

ValueMeaning
1Read is paired
2Read is part of a proper pair
4Read is unmapped
16Read is mapped to the reverse strand
32Mate is mapped to the reverse strand
64First read in a pair
128Second read in a pair
1024Read is a duplicate

For example, FLAG 99 is made from:

1 + 2 + 32 + 64 = 99

Tools such as samtools can be used to read and interpret SAM/BAM files.

Basic samtools Commands

samtools view -b sample.sam > sample.bam samtools sort sample.bam -o sorted.bam samtools index sorted.bam samtools flagstat sorted.bam

These commands can be used to convert SAM to BAM, sort a BAM file, create an index, and generate a basic alignment summary.

VCF File Format

VCF (Variant Call Format) is a text-based format used to store genetic variants identified in a sample.

A variant can be a single-base change, insertion, deletion, or another type of genetic difference.

A VCF file contains header lines followed by variant records.

For example:

##fileformat=VCFv4.2
##INFO=<ID=DP,Number=1,Type=Integer,Description="Total depth">
#CHROM  POS    ID      REF  ALT  QUAL  FILTER  INFO   FORMAT  SAMPLE1
chr1    12345  rs123   A    G    50    PASS    DP=30  GT:DP   0/1:28
chr1    20400  .       AT   A    42    PASS    DP=25  GT:DP   1/1:25

Important VCF columns include:

  • CHROM: Chromosome or reference sequence
  • POS: Position of the variant
  • REF: Reference allele
  • ALT: Alternate allele
  • QUAL: Variant quality score
  • FILTER: Filtering status
  • INFO: Additional information
  • FORMAT: Describes sample-specific fields

Genotype Values

The GT field describes the genotype of a sample.

GenotypeMeaning
0/0Both alleles match the reference
0/1One reference allele and one alternate allele
1/1Both alleles are alternate
0|1Phased genotype

For example:

chr1  12345  rs123  A  G

means that the reference allele is A and the alternate allele is G.

VCF files are commonly compressed and indexed for faster processing. Tools such as bcftools, bgzip, and tabix are widely used with VCF files.

PDB and mmCIF File Formats

PDB and PDBx/mmCIF are file formats used to store information about the three-dimensional structures of proteins, nucleic acids, and other biological molecules.

Structure files contain information such as atom names, residue names, chains, and the three-dimensional coordinates of atoms.

PDB Format

The PDB format is a fixed-width text format. Important records include ATOM and HETATM.

For example:

ATOM      1  N   MET A   1      27.340  24.430   2.614  1.00 11.99           N
ATOM      2  CA  MET A   1      26.266  25.413   2.842  1.00 11.85           C
HETATM  900  O   HOH A 301      15.100  10.220   7.400  1.00 20.10           O

Important information includes:

  • Atom number
  • Atom name
  • Residue name
  • Chain identifier
  • Residue number
  • X, Y, and Z coordinates
  • Occupancy
  • B-factor
  • Element

The X, Y, and Z coordinates are usually measured in ångströms (Å).

PDBx/mmCIF

PDBx/mmCIF is the modern format used for representing structural data in the Protein Data Bank. Unlike the fixed-width PDB format, mmCIF stores information using named fields and tables.

A simple example is:

loop_
_atom_site.group_PDB
_atom_site.id
_atom_site.label_atom_id
_atom_site.label_comp_id
_atom_site.label_asym_id
_atom_site.Cartn_x
_atom_site.Cartn_y
_atom_site.Cartn_z
ATOM 1 N  MET A 27.340 24.430 2.614
ATOM 2 CA MET A 26.266 25.413 2.842

The loop_ keyword starts a table, followed by the names of the fields and their values.

mmCIF is better suited for large and complex structures because it does not have the same field and identifier limitations as the older PDB format.

Common Conversion Tools

Bioinformatics workflows often require converting data from one format to another. Several command-line tools and libraries are available for this purpose.

TaskToolExample
FASTQ to FASTAseqkitseqkit fq2fa reads.fq -o reads.fa
FASTQ to FASTAseqtkseqtk seq -a reads.fq > reads.fa
GenBank to FASTABiopythonSeqIO.convert()
GFF3 to GTFgffreadgffread in.gff3 -T -o out.gtf
GTF to GFF3gffreadgffread in.gtf -o out.gff3
BAM to BEDbedtoolsbedtools bamtobed -i sorted.bam
BED regions to FASTAbedtoolsbedtools getfasta -fi genome.fa -bed regions.bed
SAM to BAMsamtoolssamtools view -b in.sam > out.bam
BAM to SAMsamtoolssamtools view -h in.bam > out.sam
BAM to FASTQsamtoolssamtools fastq sorted.bam > reads.fq
VCF to BCFbcftoolsbcftools view -O b in.vcf -o out.bcf
Compress VCFbgzipbgzip in.vcf
Index VCFtabixtabix -p vcf in.vcf.gz

Typical Bioinformatics Workflow

Different file formats are often used at different stages of a bioinformatics workflow.

A basic variant analysis workflow can look like this:

  1. FASTQ files contain sequencing reads from the sequencing machine.
  2. Quality control is performed on the reads.
  3. The reads are aligned to a reference FASTA file.
  4. The alignment is stored in SAM/BAM format.
  5. A variant caller identifies variants and produces a VCF file.
  6. GFF/GTF or BED files can be used to identify genes and genomic regions related to the variants.
  7. If a variant affects a protein-coding gene, its protein structure can be studied using PDB/mmCIF data.

Quick Comparison of Bioinformatics File Formats

FormatMain purposeCoordinate system
FASTADNA, RNA, or protein sequencesNot applicable
FASTQSequences with quality scoresNot applicable
GenBankSequences with detailed annotations1-based
GFF/GTFGene and feature annotations1-based, inclusive
BEDGenomic regions0-based, half-open
SAM/BAMAligned sequencing readsSAM POS is 1-based
VCFGenetic variants1-based
PDB/mmCIF3D molecular structuresX, Y, Z coordinates

Conclusion

Bioinformatics file formats are used to store different types of biological data. FASTA stores sequences, FASTQ stores sequences with quality scores, GenBank stores sequences with annotations, GFF/GTF and BED store genomic features and regions, SAM/BAM store alignments, VCF stores variants, and PDB/mmCIF store molecular structures.

Understanding the basic structure and purpose of these formats makes it easier to work with bioinformatics tools and analyze biological data.