CAPÍTULO 2: El modelo clásico de regresión lineal: especificación y estimación
2.7. Introducción al uso de EViews (I)
A virion contains the genome of a virus in the form of one or more molecules of nucleic acid. For any one virus the genome is composed of either RNA or DNA.
If a new virus is isolated, one way to determine whether it is an RNA virus or a DNA virus is to test its suscep-tibility to a ribonuclease and a deoxyribonuclease. The virus nucleic acid will be susceptible to degradation by only one of these enzymes.
Each nucleic acid molecule is either single-stranded (ss) or double-stranded (ds), giving four categories of virus genome: dsDNA, ssDNA, dsRNA and ssRNA.
The dsDNA viruses encode their genes in the same kind of molecule as animals, plants, bacteria and other cellular organisms, while the other three types of genome are unique to viruses. It interesting to note that most fungal viruses have dsRNA genomes, most plant viruses have ssRNA genomes and most prokaryotic viruses have dsDNA genomes. The reasons for these distributions presumably concern diverse origins of the viruses in these very different host types.
A further categorization of a virus nucleic acid can be made on the basis of whether the molecule is linear, with free 5 and 3 ends, or circular, as a result of the strand(s) being covalently closed. Examples of each category are given in Figure 3.1. In this figure, and indeed throughout the book, molecules of DNA and
RNA are colour coded. Dark blue and light blue depict (+) RNA and (−) RNA respectively; these terms are explained in Section 6.2.
It should be noted that some linear molecules may be in a circular conformation as a result of base pairing between complementary sequences at their ends (see Figure 3.7 below). This applies, for example, to the DNA in hepadnavirus virions and to the RNA in influenza virions.
3.2.1 Genome size
Virus genomes span a large range of sizes. Porcine circovirus (ssDNA) and hepatitis delta virus (ssRNA) each have a genome of about 1.7 kilobases (kb), while at the other end of the scale there are viruses with dsDNA genomes comprised of over 1000 kilobase pairs (kbp). The maximum size of a virus genome is subject to constraints, which vary with the genome category.
As the constraints are less severe for dsDNA all of the large virus genomes are composed of dsDNA.
The largest RNA genomes known are those of some coronaviruses, which are 33 kb of ssRNA.
The largest virus genomes, such as that of the mimivirus, are larger than the smallest genomes of cellular organisms, such as some mycoplasmas (Figure 3.2).
3.2.2 Secondary and tertiary structure As well as encoding the virus proteins (and in some cases untranslated RNAs) to be synthesized in the infected cell, the virus genome carries addi-tional information, such as signals for the con-trol of gene expression. Some of this information is contained within the nucleotide sequences, while for the single-stranded genomes some of it is con-tained within structures formed by intramolecular base pairing.
In ssDNA complementary sequences may base pair through G–C and A–T hydrogen bonding; in ssRNA weaker G–U bonds may form in addition to G–C and A–U base pairing. Intramolecular base pairing results in regions of secondary structure with stem-loops and bulges (Figure 3.3(a)). In some ssRNAs intramolecular base pairing results in structures known
VIRUS GENOMES 33
Examples DNA genomes
RNA genomes
ss, linear
ds, linear
ss, circular ds, circular ss, circular
ds, linear ss, linear
Parvoviruses
Poxviruses
PhagejX174
Baculoviruses
Tobacco mosaic virus
Reoviruses
Hepatitis delta virus
Figure 3.1 Linear and circular viral genomes.
ss: single-stranded ds: double-stranded
There are no viruses known with circular dsRNA genomes
as pseudoknots, the simplest form of which is depicted in Figure 3.3(b).
Regions of secondary structure in single-stranded nucleic acids are folded into tertiary structures with specific shapes, many of which are important in molecular interactions during virus replication. For an example see Figure 14.6, which depicts the 5 end of poliovirus RNA, where there is an internal ribosome entry site to which cell proteins bind to initiate translation. Some pseudoknots have enzyme activity, while others play a role in ribosomal frameshifting (Section 6.4.2).
3.2.3 Modifications at the ends of virus genomes
It is interesting to note that the genomes of some DNA viruses and many RNA viruses are modified at one or both ends (Figures 3.4 and 3.5). Some genomes have a covalently linked protein at the 5end. In at least some viruses this is a vestige of a primer that was used for initiation of genome synthesis (Section 7.3.1).
Some genome RNAs have one or both of the mod-ifications that occur in eukaryotic messenger RNAs (mRNAs): a methylated nucleotide cap at the 5 end
34 VIRUS STRUCTURE
Cells Viruses
Hepatitis B virus 3.2
48.5
1200
580
4 700
3 200 000 Approximate
genome size (kbp DNA)
Mimivirus
Mycoplasma genitalium
Escherichia coli
Human Phageλ
Figure 3.2 Genome sizes of some DNA viruses and cells.
VIRUS GENOMES 35
(a) (b)
L2
L1 L2 L1 L1
Figure 3.3 Secondary structures resulting from intramolecular base-pairing in single-stranded nucleic acids. (a) Stem-loops and bulges in ssRNA and ssDNA. (b) Formation of a pseudoknot in ssRNA. A pseudoknot is formed when a sequence in a loop (L1) base-pairs with a complementary sequence outside the loop. This forms a second loop (L2).
Examples
Adenoviruses Phage PRD1 (E. coli)
Hepatitis B Virus
Parvoviruses Virus Genome
dsDNA
ssDNA protein
protein
(+) DNA (−) DNA
capped RNA protein
protein
5′ 5′
5′ 3′
3′ 3′ 5′
3′ 5′
3′
Figure 3.4 DNA virus genomes with one or both ends modified. The 5 end of some DNAs is covalently linked to a protein. One of the hepatitis B virus DNA strands (the (+) strand) is linked to a short sequence of RNA with a methylated nucleotide cap.
36 VIRUS STRUCTURE
Barley yellow dwarf virus
SARS coronavirus
Figure 3.5 RNA virus genomes with one or both ends modified. The 5end may be linked to a protein or a methylated nucleotide cap. The 3end may be polyadenylated or it may be folded like a transfer RNA.
(Section 6.3.4) and a sequence of adenosine residues (a polyadenylate tail; poly(A) tail) at the 3 end (Section 6.3.5).
The genomes of many RNA viruses function as mRNAs after they have infected host cells. A cap and a poly(A) tail on a genome RNA may indicate that the molecule is ready to function as mRNA, but neither structure is essential for translation. All the ssRNAs in Figure 3.5 function as mRNAs, but not all have a cap and a poly(A) tail.
The genomes of some ssRNA plant viruses are base paired and folded near their 3 ends to form structures similar to transfer RNA. These structures contain sequences that promote the initiation of RNA synthesis.
3.2.4 Proteins non-covalently associated with virus genomes
Many nucleic acids packaged in virions have proteins bound to them non-covalently. These proteins have
VIRUS PROTEINS 37
Cys
Cys His
His Zn++
Figure 3.6 A zinc finger in a protein molecule. A zinc finger has recurring cysteine and/or histidine residues at regular intervals. In this example there are two cysteines and two histidines.
regions that are rich in the basic amino acids lysine, arginine and histidine, which are negatively charged and able to bind strongly to the positively charged nucleic acids.
Papillomaviruses and polyomaviruses, which are DNA viruses, have cell histones bound to the virus genome. Most proteins associated with virus genomes, however, are virus coded, such as the HIV-1 nucleo-capsid protein that coats the virus RNA; 29 per cent of its amino acid residues are basic. As well as their basic nature, nucleic-acid-binding proteins may have other characteristics, such as zinc fingers (Figure 3.6);
the HIV-1 nucleocapsid protein has two zinc fingers.
In some viruses, such as tobacco mosaic virus (Section 3.4.1), the protein coating the genome con-stitutes the capsid of the virion.
3.2.5 Segmented genomes
Most virus genomes consist of a single molecule of nucleic acid, but the genes of some viruses are
encoded in two or more nucleic acid molecules. These segmented genomes are much more common amongst RNA viruses than DNA viruses. Examples of ssRNA viruses with segmented genomes are the influenza viruses (see Figure 3.20 below), which package the segments in one virion, and brome mosaic virus, which packages the segments in separate virions.
Most dsRNA viruses, such as members of the family Reoviridae (Chapter 13), have segmented genomes.
The possession of a segmented genome provides a virus with the possibility of new gene combinations, and hence a potential for more rapid evolution (Section 20.3.3.c). For those viruses with the segments packaged in separate virions, however, there may be a price to pay for this advantage. A new cell becomes infected only if all genome segments enter the cell, which means that at least one of each of the virion categories must infect.
3.2.6 Repeat sequences
The genomes of many viruses contain sequences that are repeated. These sequences include promot-ers, enhancpromot-ers, origins of replication and other ele-ments that are involved in the control of events in virus replication. Many linear virus genomes have repeat sequences at the ends (termini), in which case the sequences are known as terminal repeats (Figure 3.7). If the repeats are in the same ori-entation they are known as direct terminal repeats (DTRs), whereas if they are in the opposite orien-tation they are known as inverted terminal repeats (ITRs). Strictly speaking, the sequences referred to as ‘ITRs’ in single-stranded nucleic acids are not repeats until the second strand is synthesized dur-ing replication. In the sdur-ingle-stranded molecules the
‘ITRs’ are, in fact, repeats of the complemen-tary sequences (see ssDNA and ssRNA (−) in Figure 3.7).