The purpose of genome sequencing is to identify genes and functional ribonucleic acids (RNAs) and predict their function through annotation. In its most basic sense, a gene or an open reading frame (ORF) constitutes a length of DNA commencing with a start codon (typically ATG although occasionally this can be substituted for TTG or GTG; Clark & Marcker, 1966) followed by a variable number of nucleotide triplets (known as codons) each encoding a single amino acid, finishing with a stop codon (TAA, TAG or TGA; Kohli & Grosjean, 1981). Prokaryotic genomes are far less complex than those of eukaryotes and therefore gene identification and annotation is comparatively easier. Intergenic regions account for very little of the total genome (typically 10 - 20%; Salzberger et al. 1998). Prokaryotic open reading frames (ORFs) are intronless and short, typically averaging around 1 Kb (Ochman & Davalos, 2006). This reduced complexity gives some predictability allowing bioinformatics tools to identify key features of the genome (such as ORFs) with more reliable accuracy.
However not every stretch of DNA preceded by a start codon and finishing with a stop codon is actually transcribed. Therefore an ORF without any evidence of transcription, function or significant similarity to any gene previously sequenced is described as encoding a hypothetical protein. The proportion of hypothetical proteins varies considerably between bacterial genomes ranging from 0.2% (Shigenobu et al., 2000) to 53.4% (Casjens et al., 2000; average 15.3%; Fukuchi & Nishikawa, 2004). Organisms with the fewest hypothetical proteins tend to be those with reduced genomes such as endosymbionts or the parasitic Mycoplasma genitalium. While the number of hypothetical proteins tends to increase with the phylogenetic distance between the sequenced organism and its closest sequenced relative (Fukuchi & Nishikawa, 2004).
Currently the best known and most commonly used software packages for predicting ORFs are GeneMark (Borodovsky & McIninch, 1993, Besemer & Borodovsky, 2005) and GLIMMER (Gene locator and interpolated Markov modeller; Delcher et al., 1999). These programs detect more than 91% and 97% of ORFs from prokaryotic genomes, respectively, without manual curation (Salzberg et al., 1998). GeneMark
occurring by chance (Kellis et al., 2003). This probability increases with %G+C content (Skovgaard et al., 2001). These programs then determine the probability that each nucleotide of an ORF, and collectively the entire ORF, is in a coding (or non- coding) region using Markov models built by a training set of known genes from the organism (Salzberg et al., 1998). This ability to differentiate coding from non-coding regions is based on the finding that genomes contain genome-specific signatures (Karlin et al., 1997). These signatures vary between coding and non-coding regions (Sandberg et al., 2003). Such signatures include di-nucleotide, tri-nucleotide and tetra-nucleotide frequencies (Karlin et al., 1997), %G+C content (Chargaff, 1951), codon usage (Grantham et al., 1980a, Grantham et al., 1980b) and amino acid usage bias (Sueoka, 1961). GLIMMER and GeneMark differ in the length of the Markov- chains. GLIMMER does not have a fixed length (Delcher et al., 1999), whereas GeneMark uses a 5th order model (Borodovsky & McIninch, 1993). GeneMark additionally refines a retained ORF‟s start codon by searching for ribosome-binding sequences upstream of potential start codons (Lukashin & Borodovsky, 1998).
Ribosome binding sequences (RBS) are typically centered approximately 8 - 13 nucleotides upstream of a true start codon (Shine & Dalgarno, 1975). However, in some genes, typically those transcribed from alternative start codons (GTG or TTG), there appears to be a larger gap between the RBS and the start codon (Kozak, 1983). Ribosome binding sites typically have a sequence such as AGGA or GAGG, and is complimentary to the 3' end of the 16S rRNA chain (Shine & Dalgarno, 1975).
Several comparative analyses of genome sequences to date have identified a theoretical minimal set of between 169 and 206 genes essential for cellular life (Gil et al., 2004, Koonin, 2000). While over-annotation, particularly of small ORFs, makes it difficult to accurately determine an average gene size (Skovgaard et al., 2001, Fukuchi & Nishikawa, 2004, Huynen & Snel, 2000) it is generally considered to be approximately 1 Kb (Ochman & Davalos, 2006). However, a single gene can exceed 20 Kb (Reva & Tummler, 2008). The largest prokaryotic genes described to date were found in Chlorobium chlorochromatii CaD3 and encode proteins of 36, 806 and 20,
Genome analysis has shown the presence of a number of inactivated genes also, known as pseudogenes (Jacq et al., 1977, Vanin, 1985). Pseudogenes arise through the accumulation of mutations that disrupt and ultimately degrade their functional predecessors. These are commonly described in genome analyses. It is believed that all traces of pseudogenes are likely to completely erode over time. This erosion is due to prokaryotes having a mutational bias favouring deletions over insertions (Ochman & Davalos, 2006).
Following the identification and refinement of ORFs, they can be analysed and ascribed a putative function. A number of bioinformatics tools exist to facilitate this analysis. The most popular include Basic Local Alignment Search Tool (BLAST;
Altschul et al., 1997), Clusters of Orthologous Groups (COG; Tatusov et al., 1997),
SignalP (Bendtsen et al., 2004), LipoP (Juncker et al., 2003), Transmembrane HMM
(TmHMM; Krogh et al., 2001) and Interpro (Apweiler et al., 2001). BLAST
compares the similarity of a nucleotide or amino acid sequence to a sequence database
(typically GenBank; Benson et al. 2008) using a heuristic alignment approach
(Altschul et al., 1997). Targeting of the encoded proteins to secretory systems can be
predicted by SignalP and LipoP. SignalP uses neural network and hidden Markov model (HMM) algorithms to predict the classical signal peptidase I cleavage sites, and LipoP uses HMMs to predict signal peptidase II cleavage sites, characteristic of lipoproteins. The transmembrane topology can also be predicted by HMM, using
TmHMM (Krogh et al., 2001). Specific targeting signatures, such as carboxy-terminal
cell wall anchoring in Gram-positive bacteria, via LPXTG-like motifs (Navarre &
Schneewind, 1994), recognised by the protein Sortase (Ton-That et al., 1999), can be
predicted by specific HMMs available through PFam or TIGRFam libraries
(Boekhorst et al., 2005). Interpro integrates the major hidden Markov model searches:
PFam (Sonnhammer et al., 1997) and TIGRFam (Haft et al., 2001). It also includes
other protein signature databases, including Uniprot (Apweiler et al., 2004), Prosite
(Bairoch, 1991), ProDom (Corpet et al., 1998), Smart (Schultz et al., 2000, Ponting et
al., 1999, Schultz et al., 1998), Panther (Thomas et al., 2003), PIRSF (Wu et al.,
2004), Superfamily (Gough et al., 2001), SCOP (Murzin et al., 1995), CATH (Pearl et
functional domains using primary, secondary and tertiary structural prediction and HMM analyses.
Following such analysis each ORF is ascribed a putative function. This should be descriptive but conservative and, where possible, consistent with the type species (this being Bacillus subtilis for Gram-positive organisms). Several publications have pointed out the confusion that is often caused by inconsistencies in the assignment of gene names (Wang et al., 2005, Wyman et al., 2004, Mitchell et al., 2003). Konforti (2007) recently suggested that a gene name often reflects the laboratory where the genome was annotated rather than its function or even the organism it is derived from. Various methods have arisen to attempt to circumvent such inconsistencies, such as gene ontologies (Smith et al., 2005). Gene Ontologies are a finite list of descriptors that define a gene by its molecular function(s), thebiological process(es) where these
functions take place and what cellular componentry it resides in, or belongs to.
Collectively all above bioinformatics analyses can be managed on a genomic scale using an interface such as MANATEE (TIGR, 2001) or GAMOLA (Altermann & Klaenhammer, 2003).