sidebar-main
title
griff
banner-button_0 banner-button_Layer-7 banner_button_03 banner-button_Layer-4 banner-button_05 banner-button_Layer-5 banner-button_07
banner-button_08 banner-button_09 banner-button_10
Bioinformatics Unit banner
   homeoffresearchoffconfonprogoffmemoffpuboffvacofflinkoff
   tabfoot tabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bgtabfoot-bg

Practical 4: Multiple Sequence Alignment

Imagine you have been given part of a DNA sequence and with your outstanding experience you have managed to find a matching protein and some homologous sequences to it. However, you have been unlucky and there is very little information on these proteins and you are stuck with a bunch of sequences with no known motifs and no structure.

One of the most important and basic steps to begin an analysis on a set of sequences is multiple sequence alignment. It is used for almost everything (sequence searching, prediction, modelling etc). So in this part you will learn how to interpret and produce accurate multiple sequence alignments.

What to do:
  1. Open the PRALINE Web Server page in a new window at http://ibi.vu.nl/programs/pralinewww/
    Make sure you understand the layout and the different options available.
  2. To begin, right click on the first example (Actin) from the list below and open it in a new window with any text editor. This is a file of five actin sequences in FASTA format.

    Alignment 1 - set of equidistant sequences of low divergence:
    Alignment 2 - set with two distant groups of divergent sequences:
    The following steps need to be carried out for both alignments.
    We suggest you carry out the steps each time for both alignments as this gives you a better appreciation of differences in behaviour arising when when aligning a distant or non-distant sequence set.

  3. Select and copy all the sequences (NB. one family at the time!).
  4. Now switch to the PRALINE Web Server page and paste them into the text area provided on the PRALINE page.
  5. Select the 'Standard progressive strategy' and give your job a name in the space provided (e.g. Actin alignment)
  6. Run the alignment with BLOSUM62 and gap penalties 12 and 1 using the Standard global alignment strategy.

    Question 1: Study the different colour schemes available. What can you say in a few words about what each colour scheme represents?
    How could this help your analysis of the alignment? (hint: conservation, hydrophobicity, residue type).


  7. Now vary the gap penalties (use extreme values for a larger effect -- for instance use the minimum penalty=0 and the maximum=99 for both gap-open and extension penalties). How does the alignment change when you increase and decrease each value?
  8. Go to the raw output link at the top of the PRALINE alignment page and look under the alignment at the line resembling:

    Iteration 0 SP= 136168.00 AvSP= 10.708 SId= 3975 AvSId= 0.313 Nseq= 14 Len= 191 Naa= 2171 Ngaps= 503

    where SP is the sum-of-pairs (BLOSUM62) scores, AvSP is the average SP per pair, SId is the total number of identical pairs, AvSId is the average identity per pair, Nseq is number of sequences, Len is alignment length, Naa and Ngaps is the total number of amino acids and gaps in the alignment, respectively.

    Question 2: Note the SId for comparison with the results in the next step.

  9. Now align the same set of sequences using the PAM250 matrix (historically used matrix) with its recommended penalties open=10 and extension=1. Find the SId for this alignment and compare it to that for the BLOSUM62 alignment.

    Question 3: Are the two SId`s the same? What does this tell you about the matrices?

  10. Compare the two matrix/gap-penalty combinations.

    Question 4: Can you think of a matrix that would maximise the number of identical aa pairs in any alignment?

    Question 5: Is there a relation between matrix type and gaps? (e.g. look at the Ngaps)

  11. Have a look at the MSF file from the link at the top of the alignment page. This is a standard format for multiple sequence alignment files.

    Question 6: What are the main characteristics of the file (e.g. compared to FASTA file format)?

  12. Now we are going to use information from all sequences right from the start during progressive alignment using global profile pre-processing.
    In addition, each time, you should include the tree representation of the alignment by selecting 'yes' at 'Tree representation of the final alignment'
    • Use pre-processing with a cutt-off value of 0 to start with.

      Question 7a: What does a cut-off value of 0 mean? Go to the 'raw output file'. What extra output do you get?

      Study the tree to get an idea about the relationships between the sequences.

      Question 7b: What does this tree tell you? (NB. This is not an evolutionary tree, i.e. it does not tell you the phylogeny of the sequences)

    • Now use different cut-off values based on the new information you have been given from the values in the preprocessing scores table.

      Question 7c: What changes can you see in the alignment? Is the tree different? What has changed?
      Can you improve the alignment using pre-processing with a sensible cut-off value?

  13. Question 8: Summarize the observed differences between both alignments 1 (actin) and 2 (kinase).

  14. Align the two sequence sets using another program out of the following three: ClustalW, T-COFFEE and MUSCLE.

    Question 9: Compare the output alignments with those you got from PRALINE.

  15. The structural alignment links given next to the sequence sets (step 2) give you the 'true' alignment (standard of truth). These are taken from the BAliBASE benchmark set where sequences have been aligned with the help of structural knowledge. Within the structural alignments, so-called core blocks are indicated in 'uppercase'. Core blocks are the regions that are the most trusted by the BAliBASE authors.

    Question 10: What is a structural alignment? What can you say about the quality of the output alignments from each method?
    Is the 'distant' sequence set indeed harder to align? How can you tell?

Send your answers to: pirovano [at] few.vu.nl

(c) IBIVU 2026. If you are experiencing problems with the site, please contact the webmaster.