Pairwise Sequence Alignment is a method of comparing two biological sequences, such as DNA, RNA, or protein sequences, to identify matching regions and measure their similarity. It helps researchers study evolutionary relationships and identify possible similarities in biological functions.
Global vs Local Sequence Alignment
- Global alignment stretches across the full length of both sequences, end to end. Use it for sequences of similar length that you expect to be related throughout, such as the same gene from two species.
- Local alignment finds only the best-matching region. Use it when sequences differ in length or share only a part, such as a protein domain inside a much larger protein.
Example: When comparing a short motif with a long protein, a global alignment would force the motif against the whole protein and give a poor score. A local alignment finds the region where the motif actually fits.
Dot Plots: A Quick Visual Way to Compare Sequences
A dot plot is a simple visual way to inspect sequence similarity. Write one sequence along the top and the other down the side. Place a dot wherever the compared symbols match (or meet the chosen similarity criterion).
- A diagonal line of dots indicates a similar region between the sequences.
- Parallel diagonal lines can indicate repeated regions within or between sequences.
- An opposite-slope diagonal in a DNA dot plot can indicate an inverted or reverse-complement relationship, depending on how the plot is constructed.
- Random isolated dots may represent noise. Applying a window threshold (for example, 5 matches out of 7 positions) can reduce random matches, although it may also miss weaker similarities.
Dot plots give a quick overview, but they do not give a score. For that, we need alignment algorithms.
Scoring Matrices: BLOSUM and PAM Explained
For DNA, a simple example scoring scheme is +1 for a match and −1 for a mismatch, although other schemes are used. Protein scoring is more nuanced because some amino-acid substitutions are chemically conservative (for example, leucine and isoleucine), whereas others are less conservative. A scoring matrix assigns a score to each pair of amino acids.
BLOSUM Matrices
BLOSUM (BLOcks SUbstitution Matrix) matrices are derived from conserved blocks in protein families.
- The BLOSUM number refers to the sequence-identity threshold used to cluster sequences when the matrix was built. It is not the identity percentage of the sequences you are aligning.
- BLOSUM62 is a common default. BLOSUM80 is generally useful for closer evolutionary relationships, whereas BLOSUM45 can be useful for more distant relationships.
- Examples from BLOSUM62: W–W = 11, A–A = 4, and L–I = 2. These scores show that identical or chemically similar amino acids can receive positive scores; the meaning of a score depends on the matrix.
PAM Matrices
PAM (Point Accepted Mutation) matrices are based on an evolutionary model of accepted amino-acid substitutions over time.
- PAM1 represents one accepted point mutation per 100 amino-acid positions on average in the model used to construct the matrix.
- Higher PAM numbers represent greater evolutionary distances, so PAM250 is commonly used for more distant relationships.
Remember: a higher BLOSUM number generally corresponds to a matrix designed for closer relationships, while a higher PAM number represents a greater evolutionary distance. The numbering conventions therefore run in opposite directions.
Gap Penalties: Linear vs Affine
Insertions and deletions are represented by gaps. Without a gap penalty, an algorithm could introduce excessive gaps to improve the alignment score. Gap penalties discourage unnecessary gaps.
- Linear gap penalty: every gap position has the same cost, so a gap of length 3 costs 3 times the per-position penalty.
- Affine gap penalty: opening a gap incurs a larger cost, while extending an existing gap incurs a smaller cost. This models the idea that a single insertion or deletion event may affect several residues.
A common setting is gap open = 10 and gap extend = 0.5. Under this convention, a gap of length 4 costs 10 + (3 × 0.5) = 11.5. Software may differ in the precise gap-cost convention, so check the documentation of the program you use.
Needleman–Wunsch Algorithm (Global Alignment)
This algorithm uses dynamic programming. It fills a table in which each cell stores the best score for an alignment ending at that position, then traces back through the table to construct an optimal alignment.
For each cell, take the maximum of these three candidate scores:
- The diagonal cell + the match/mismatch score
- The cell above + the gap penalty
- The cell to the left + the gap penalty
Initialize the first row and first column with 0, −1, −2, … because each additional position represents another gap under the linear gap penalty used here.
Worked Example
Align GATC (rows) with GTC (columns), using match = +1, mismatch = −1 and gap = −1.
| – | G | T | C | |
|---|---|---|---|---|
| – | 0 | -1 | -2 | -3 |
| G | -1 | 1 | 0 | -1 |
| A | -2 | 0 | 0 | -1 |
| T | -3 | -1 | 1 | 0 |
| C | -4 | -2 | 0 | 2 |
Check one cell: the T/T cell is the maximum of the diagonal (0 + 1 = 1), the cell above (0 − 1 = −1) and the cell to the left (−1 − 1 = −2), which gives 1.
Traceback starts at the bottom-right cell (score 2) and follows optimal predecessor choices back to the top-left. If there are ties, more than one optimal traceback may be possible.
G A T C | . | | G - T C
Final score = +1 − 1 + 1 + 1 = 2. The A in GATC is aligned with a gap in GTC.
Smith–Waterman Algorithm (Local Alignment)
The Smith–Waterman algorithm is similar to Needleman–Wunsch, with two key changes:
- No cell can have a score below 0. Any negative candidate score is replaced by 0, which allows a local alignment to start at different positions.
- Traceback starts at a cell with the highest score in the table and stops when it reaches a cell with score 0. If several cells share the maximum, there may be multiple optimal local alignments.
Worked Example
Align GGCTGA (rows) with TCTGT (columns), using match = +2, mismatch = −1 and gap = −1.
| – | T | C | T | G | T | |
|---|---|---|---|---|---|---|
| – | 0 | 0 | 0 | 0 | 0 | 0 |
| G | 0 | 0 | 0 | 0 | 2 | 1 |
| G | 0 | 0 | 0 | 0 | 2 | 1 |
| C | 0 | 0 | 2 | 1 | 1 | 1 |
| T | 0 | 2 | 1 | 4 | 3 | 3 |
| G | 0 | 1 | 1 | 3 | 6 | 5 |
| A | 0 | 0 | 0 | 2 | 5 | 5 |
The highest score is 6. Trace back from this cell, stopping at 0.
C T G | | | C T G
The local alignment is the shared CTG block, with a score of 6. The nonmatching ends are excluded. A global alignment would instead account for the full lengths of both sequences.
Time Complexity and BLAST
The standard dynamic-programming versions of both algorithms take O(mn) time, where m and n are the lengths of the two sequences. If the sequences have similar length n, this is O(n²). This cost can be high for large database searches, so tools such as BLAST use heuristic methods to identify promising matches efficiently.
EMBOSS Needle and Water: Free Alignment Tools
You do not need to fill in tables by hand. EMBOSS offers ready-made tools:
- Needle runs Needleman–Wunsch (global).
- Water runs Smith–Waterman (local).
Both are available through web interfaces and command-line installations; availability and interface details may vary by service.
needle -asequence seq1.fasta -bsequence seq2.fasta -gapopen 10 -gapextend 0.5 -outfile result.needle water -asequence seq1.fasta -bsequence seq2.fasta -gapopen 10 -gapextend 0.5 -outfile result.water
With standard EMBOSS defaults, protein alignments use EBLOSUM62 and nucleotide alignments use EDNAFULL, with gap open = 10 and gap extension = 0.5. Defaults can vary with program versions or options, so verify them for your installation.
How to Read EMBOSS Alignment Output
- Identity: the percentage of exactly matching positions.
- Similarity: the proportion of matching and conservatively substituted residues in protein alignments; the exact reporting depends on the scoring matrix and output format.
- Gaps: the percentage of gap positions.
- Score: the total alignment score.
- Match symbols: in the usual EMBOSS display, | marks an identical pair, : a strong similarity, . a weaker similarity, and a blank generally means no reported similarity. Interpret the symbols in the context of the selected output format.
Conclusion
Pairwise alignment is a foundation of sequence comparison. A dot plot gives a quick visual overview. Global alignment (Needleman–Wunsch) compares sequences end to end, while local alignment (Smith–Waterman) finds the best-matching region. Results depend on choices such as the scoring matrix and gap penalties; BLOSUM62 is a common starting point for protein alignments, and affine gap penalties are widely used.