• No se han encontrado resultados

7.5.7 The DGCR on Chromosome 22qlL

During the three years over which this study was conducted, the transcription map of the DGCR on chromosome 22ql 1 has increased in complexity and three minimally deleted regions have been identified (see chapter 4.2). This section, however, will concentrate on the structure of the map and the genes which had been isolated at the beginning of this study, as shown in Figure 1.5 (page 48).

L 5 .L 1 TUPLEl/HIRA. 1.5.1.1.1 TUPLEl

Conserved sequences from the cosmid scF5 (Lindsay et al 1993) were used to screen a human foetal brain cDNA library. Several independent positives were obtained after tertiary screening of the human library, of which the largest clone, C5, contained an insert of 3.1 kb. When hybridised to a northern blot containing human foetal mRNAs, C5 detected a 3.4kb transcript in all tissues, and a smaller 3.2kb transcript in liver. The missing 3’ untranslated sequences were obtained by 3’ RACE (Rapid

Amplification of cDNA ends). Hybridisation of C5 subfragments to cosmids and a PI clone from the region, orientated the direction of transcription of C5 towards the centromere. As shown in figure 1.5, C5 maps within the DGCR, but lies

approximately 300kb distal to the ADU breakpoint.

Sequence analysis o f C5 predicted an ORF of 766 amino acids. Database comparisons with known protein sequences indicated that CS shares significant homology to a yeast regulatory protein, TUPl (Williams & Trumbly 1990). Mutations in TUPl result in decreased repression of many genes (Williams et al

1991). TUPl comes from a large family of WD40-repeat proteins. These eukaryote proteins contain between four to eight repetitive units, the WD40 domain. This well studied protein domain consists of a region of variable length followed by a

characteristic GH (Glycine-Histidine), adjacent to another core o f between 23-41 residues which ends with the characteristic WD (Tryptophan-Aspartic Acid) (Neer et al 1994). The sequence similarity with TUPl is confined to the amino-terminus of C5, due to the presence of six WD40 repeats.

The amino-terminal WD40 repeat unit of C5 is actually most strongly related to that seen in the Drosophila gene E(sp) product (Enhancer of split)(Halford et al 1993). Human homologues of this protein have been isolated and named TLE proteins for Transducin-like/Enhancer of split (Stifani et al 1992). The gene encoding C5, due to the strongest homology to the TUPl gene, was named TUPLEl for Tup-

likeÆnhancer of split gene 1 (Halford et al 1993a).

Many WD40-repeat proteins function due to the formation of multiprotein complexes (Neer et al 1994). TUPl forms a complex with Ssn6, a TPR (tetratricopeptide repeat) snap helix-repeat protein (Keleher et al 1992). This novel motif contains a 34

amino acid repeat that is defined by a degenerate nucleotide consensus sequence. The repeat is thought to form jointed helical structures, each with a “knob” and “hole”, which allows association between repeats and other complexing proteins (Goebl and Yanagida 1991). Transcriptional regulation in yeast can occur due to binding of the TUPUSsnô complex with other specific proteins, such as the homeodomain protein complexes a2-Mcml and a l-a 2 (Keleher et al 1992).

Complex formation between WD40 proteins and TPR proteins also play roles in RNA processing and cell cycling (Neer et al 1994). TUPl is also known to bind directly to histone genes, regulating transcription (Edmundson et al 1996).

Features of the TUPLEl gene include an overrepresentation of the amino acid

sequence SP, (Serine-Proline), a potential nuclear localisation sequence at amino acids 235-242, and a carboxy-terminus rich in polar amino acids, particularly glutamine, serine and threonine. These three features are all common to transcription regulatory proteins (reviewed in Halford et al 1993). TUPLEl may therefore act as a

transcription regulator by forming a multimeric complex with other proteins, such as TRP snap helix proteins. A TUPLElIproitm complex might then interact with other DNA binding proteins, a situation analogous to regulation in yeast. Additionally, direct binding of TUPLEl to histone genes or other chromatin components, may alter the chromatin structure of neighbouring genes and maintain transcriptional regulation.

L5 .L L 2 H IR A .

A CpG rich region of the cosmid DAC30 assigned to chromosome 22ql 1, was used to screen a human foetal brain cDNA library (Lamour et al 1995). Three overlapping cDNAs were identified and sequenced. The full length cDNA detected a transcript of approximately 4.2-4.4kb in all adult and foetal tissues examined. The cDNA, named

HIRA (see below), encodes a 1017 amino acid protein which encompasses the entire length o f TUPLEL However, there are two differences between HIRA and TUPLEl. HIRA contains an additional, internal, 207 amino acid residues which contain a second potential nuclear localisation signal. The presence of these additional residues is due to a deletion of 621 bp in TUPLEL An additional 44 N-terminal residues are also encoded by the HIRA gene, due to the use of a more 5’ initiation codon. Two

the clone name MF2). However, Halford et al favoured the use of the downstream codon, due to the presence of a termination codon between the two methionine codons in the murine homologue (Halford et al 1993a). HIRA therefore contains seven WD40-repeats, as these 44 residues are a repeat motif.

Following analysis of the genomic organisation of TUPLEl and HIRA, it was observed that the 207 internal residues oïHIRA fell between two exons of TUPLEl

(Lorain et al 1996, Llevadot et al 1996). RT-PCR experiments, using primers

designed from TUPLEl and HIRA, and hybridisation o f the cDNAs to these products, showed that TUPLEl is actually a splice variant of HIRA. Furthermore, these

experiments showed that TUPLEl is expressed at very low levels in most foetal tissues analysed, in comparison to the high levels of expression of HIRA, deduced to be the major transcript from this locus (Llevadot et al 1996).

HIRA shares significant homology to two WD40-repeat proteins, H IR l and HIR2,

(Lamour et al 1995). The homology to these two proteins is higher than to TUPl.

The homology with H IRl is at the amino-terminus, over the WD40-repeats. With

HIR2 the homology is localised at the carboxy-terminus, over the two nuclear localisation motifs. There is a 63 amino acid stretch separating these two regions of similarity in HIRA. At both the nucleotide and protein levels, HIRA also shares similarity with the yeast chromatin assembly factor (CAFl) p60 subunit (Gutjahr et al

1995). Both H IR l and HIR2 have been identified as repressors o f core histone gene transcription in yeast (Sherwood et al 1993) and are thought to function by forming multimeric complexes with DNA-binding proteins, which will bind to repressors found within the histone genes, and thus regulate transcription. CAFl is also thought to function as a transcriptional regulator, by assembling histone H3 and acetylated H4 onto newly synthesised DNA and thus altering chromatin structure (Kaufman et al

1995).

1.5.1.1.3 HIRA as a Candidate Gene.

The murine and chick homologues have been isolated and named Hira and Chira

respectively (Halford et al 1993a, Roberts et al 1997). Preliminary expression data in the mouse showed that Hira is expressed in early development in the neural tube and branchial arches (C.Roberts, personal communication). More extensive expression

studies in the chick, using whole mount and sectioned whole mount embryos, showed that Chira is expressed very early in the embryo, predominantly in the

neuroepithelium, at all axial levels. High levels of transcripts were detected in regions containing neural crest cells, especially in populations of these cells migrating from rhombomeres 4 and 6. At later stages of development, expression could be observed in neural crest derived regions of the head, branchial arches and the phaiyngeal pouches. Transcripts were detected in the heart and limb buds (Roberts et al 1997).

Chira is also expressed strongly at rhombomere boundaries, at a stage after the boundary formation occurs. The expression pattern here is very similar to that seen with Pax6 (Heyman et al 1995), and may suggest a role in the maintenance o f the boundaries.

The overall expression data implies that, at least in the developing chick, HIRA may play an important role in the development of structures affected in DGS.

HIRA maps within the DGCR and is therefore hemizygously deleted in the majority of CATCH22 patients. Although no mutations of HIRA have been found in non-deleted patients (J.Goodship, U.Atif, unpublished data), HIRA remains a candidate gene for contributing to the phenotype of DGS, due to the putative function and expression pattern of the gene. As already mentioned, HIRA lies approximately 300kb distal to the ADU breakpoint. It has been suggested that the breakpoint may exert a positional effect on HIRA and thus result in DGS in patient ADU.

1.5.1.2 COMT,

Catechol-O-methyl-transferase {COMT) metabolises catecholamines, such as noradrenaline, adrenaline and dopamine, to inactive 0-methyl esters. The COMT

gene had previously been mapped to 22ql 1 (Winqvist et al 1992, Grossman et al

1992). During experiments to more finely map the region, it was found that COMT

was located within a 450kb YAC, TMYACHP500 (see figure 1.4a) that also contained the conserved marker HP500 (Scambler et al 1992). This placed COMT

within the 2Mb region deleted in most CATCH22 patients. However, COMT does not map within the DGCR.

been observed that the variation in activity has an important influence on the response of Parkinsonian patients to L-dopa (Dunham et al 1992). Low COMT

activity has also been reported in women with primary affective disorder (Cohn et al

1970). The presence of this variable activity led to the hypothesis that individuals hemizygously deleted for COMT, with the low-activity variant allele on the normal chromosome 22, would be predisposed to development of the psychotic features reported in VCFS patients (Dunham et al 1992). Molecular analysis of COMTm

VCFS patients with psychiatric disease is currently underway.

L 5 .i.J TIO.

The conserved marker HP500, which lies approximately midway between the scl 1.1 repeats, had failed to isolate cDNA clones when used extensively as a probe to screen cDNA libraries (Wadey et al 1993). Further attempts to isolate potential coding sequences involved the use of a subclone from the cosmid sc4.1, a cosmid that contains HP500 (Carey at al 1992b). This subclone, scl0.6 (figure 1.4a), detected several fragments in human genomic DNA indicating that it might be a member of a low-copy repeat family. However only one fragment was detected when sc l0.6 was used as a probe and hybridised to mouse DNA. Therefore sc 10.6 was used to screen mouse 8.5dpc. and 1 l.Sdpc. whole embryo cDNA libraries and a human foetal brain cDNA library. One positive clone was isolated from the murine 8.5dpc. library, but no human cDNA could be found after these screens. The murine recombinant, named

TIO, had an insert of l.Skb and was shown to hybridise to a subset of human genomic cosmids containing TMYAC500 sequences (Halford et al 1993b). Sequence analysis showed that TIO encodes a serine/threonine rich protein o f 276 amino acids.

However, no strong homologies to any other known proteins were identified after extensive database searching using the BLAST (Altschul et al 1990), BLOCKS and Maspar programs (HGMP). Individual fragments of TIO were used to orientate the direction of transcription, from telomere to centromere. HP500 was mapped distal to the 5’ end of TIO. It has been suggested that the conservation of HP500 is due to regulatory elements that map within this region of 22ql 1 in the human genome (Halford et al 1993b).

Expression studies detected high levels of a 2kb TIO transcript in foetal liver tissue. Lower levels were also seen in foetal lung, heart and kidney. Preliminary tissue section in situ hybridisation experiments of TIO, on murine embryos from 8.5dpc. to 15.5dpc. showed that TIO is expressed during early embryogenesis. High levels of expression were detected in several tissues, including the mouth, trachea, lung, velo­ pharyngeal region, liver and lower limbs. Expression was also detected in the arterial pole of the heart, the neural ganglia and the central nervous system, including the posterior pituitary gland, and the dorsal spinal cord (Halford et al 1993b).

These data suggest that the murine TIO gene has a human homologue that maps to the commonly deleted region of 22ql 1. However, TIO does not map within the DGCR (figure 1.5), although it has been suggested that haploinsufriciency for TIO

could bring about variability in the 22ql 1 deletion phenotype (Halford et al 1993b).

L5.1.4 ZNF74.

Using a candidate gene approach, chromosome 22 human libraries were screened for genes containing the zinc-finger motif of the Cys2-His2 Kruppel/TFIIIA type (Aubry et

al 1992). Evidence that the zinc finger protein Krox20 (a member of this Kruppel family) acts as an upstream transcriptional regulator of the homeobox gene HoxB2

(Sham et al 1993), has lead to the hypothesis that a similar regulatory pathway is involved in the development of structures affected in DiGeorge syndrome (Aubry et al

1993). The involvement of a zinc finger gene in the control of this complex

developmental pathway, would explain the DGS-like phenotype observed in newborn rats whose mothers had been exposed to the zinc-chealating agent bis-

dichloroacetylamine (WIN 18,446) (see section 1.3.2, 1.3.3).

Four genomic sequences that contained open reading frames potentially encoding zinc finger proteins were isolated. Each was mapped to human chromosome 22ql 1.2 using somatic cell hybrids. The longest transcribed sequence of 2.4kb, ZNF74,

encodes a predicted protein of 505 amino acids. ZNF74 contains twelve contiguous zinc finger motifs exhibiting sequence similarity with the consensus sequence reported for the Kruppel zinc finger family (Aubry et al 1993). Using a probe derived from a non zinc finger region of ZNF74, northern analysis indicated the presence of two

transcripts of 4.4kb and 3.1 kb (migrating as a doublet) in human foetal tissues, of which the highest levels of expression was detected in the brain.

ZNF74 was found to be hemizygously deleted by FISH analysis in 34/35 patients, who were also deleted for other markers within the DiGeorge commonly deleted region. However ZNF74 was mapped outside the distal limit of the DGCR, as shown in figure

1.5.

1.5.1.5 LZTR-1.

As part of a positional cloning strategy a 22ql 1-specific plasmid (microclone) library was constructed, using the techniques of microcloning and microdissection of cosmids isolated fi"om the commonly deleted DGS region (Kurahashi et al 1995). Microclones were analysed by hybridisation with total human genomic DNA to locate unique clones, which were then mapped back to chromosome 22. Southern blot

hybridisation, using three clones as probes (that mapped to the region between

21(\iQV-BCR, thus encompassing the DGS locus), detected a reduced intensity of signal in two CATCH22 patients and this result was confirmed using FISH. To expand these microclones into larger genomic fragments, a genomic cosmid library was screened with the clones, followed by one round of cosmid walking. This resulted in the formation of two distinct cosmid contigs within the DGS region (Kurahashi et al 1995).

To isolate putative coding sequences from these contigs, direct cDNA selection was performed separately on eight cosmids (six from contig 1, two from contig 2). PCR amplified cDNA clones derived from one cosmid, cos86, were then used to screen a human foetal brain cDNA library, and fifteen overlapping, recombinant clones were isolated. The longest clone had an insert of 4.3kb and was designated C17. In northern analysis, C17 detected two different transcripts of 6.0kb and 4.3kb in foetal brain, heart, kidney, lung and liver tissues. Sequence analysis of Cl 7 showed that the encoded protein had several characteristics seen in DNA binding proteins and more specifically, transcription factors. These included a basic leucine zipper-type domain, found in the family of b-ZIP proteins and which is thought to form a parallel, two stranded a-helical coiled coil with adjacent basic amino acids resulting in the formation of the DNA binding site. However, in Cl 7 highly conserved asparagine

and alanine residues are replaced with proline and glycine residues that may disrupt the a-helix and affect the putative DNA binding. The presence of another basic domain homologous to the b-ZIP proteins, at the N-terminus of C l7, and two S(T)PXX (serine or threonine-proline-X-X) sequences often associated with DNA binding proteins, imply that C l7 will function as a DNA binding protein and possibly play a role in the developmental pathways involved in DGS (Kurahashi et al 1995). C17 was named LZTR-1 due to the putative function.

Quantitative hybridisation of LZTR-I was performed on genomic DNA fi*om six DGS patients and two CTAF patients. LZTR-1 was found to be hemizygously deleted in all but one of the DGS patients. However there was no detectable deletion in patient GM00980 whose deletion, due to an unbalanced translocation, provides the distal limit of the DGCR As shown in figure 1.5, LZTR-1 maps distal to the DGCR.

\ . 5 . \ . 6 G p I b p

Glycoprotein Ib (Gplb) complexes with two other glycoproteins to form the major platelet receptor for von Willebrand factor. Defects in this receptor result in the autosomal recessive Bemard-Soulier syndrome (BSS) which is characterised by prolonged bleeding time, thrombocytopenia and very large platelets (Bernard J. & Soulier JP. 1948). Gplb is a dimer composed of two plasma membrane proteins, G plba and Gplbp. The Gplbp gene was recently assigned to 22ql 1, using somatic cell hybrids and FISH studies, to a region immediately distal to the DGCR (reviewed in Budarf et al 1995b).

A patient has been reported with the features of both BSS and DiGeorge syndrome. Subsequent FISH analysis using markers known to map within the DiGeorge critical region, including the region containing Gplbp, showed that the proband was in fact hemizygously deleted for this region. The authors suggest that haploinsufficiency in this region will contribute to the phenotype by unmasking any autosomal recessive mutations, allowing for the normal Gplbfi allele in this individual to be screened for possible mutations.

1(11:22) GM05401 GM05878 GM00980 ADU 00 LZTR-1 ZNF74 n o COMT Gplbp HIRA tel

f t -

5 0 6 H P500 KI-145 237 contig 1 sc F 5 SC4.1 S C 1 1 .1 B S C 1 1 .1 A 10.6 TM YAC500 D 0832 cen

DGCR

Figure 1.5 A transcription map of the DiGeorge commonly deleted region in 1993 (not to scale).

Chromosome 22ql 1 is represented by the long, horizontal line. The genes isolated from the region are indicated by the lines and arrowheads above the chromosome, the direction of the arrow representing the direction of transcription of each gene. The DNA clones used in the isolation of each gene (as described in the text) are indicated below the chromosomal line. The position of each gene relative to each translocation breakpoint, and the DGCR, is shown.