Download 3D protein structure - Diamantina Institute

Survey
yes no Was this document useful for you?
   Thank you for your participation!

* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project

Document related concepts
no text concepts found
Transcript
SPARQ-ed PROJECT
Mutations in the tumor suppressor gene p53
Pulari Thangavelu (PhD student)
April 5 2016
Chromosome Instability Group
UQ Diamantina Institute
Name of presentation Month 2009
Assignment: To find the mutation being introduced into the p53 protein, given
a particular primer sequence
DNA is made up of nucleotides.
A nucleotide is made up of deoxyribose sugar, a nitrogenous base (A/T/G/C) and a phosphate group.
The four bases found in DNA
EXAMPLE PRIMER SEQUENCE
tgatgctgtccccgcacgatattgaacaa
Name of presentation Month 2009
STEP 1: Obtain the p53 nucleotide sequence from PUBMED database.
Go to the PUBMED database and select ‘Nucleotide’ in the dropdown menu
Name of presentation Month 2009
Type in p53 in the search box and click search. Select Homo Sapiens in the ‘Results by taxon’
section.
Name of presentation Month 2009
Select the first p53 mRNA sequence shown (Note: this is a complete coding sequence for the
p53 protein), 2,451 bp linear mRNA, Accession: AB082923.1, GI: 23491728
Name of presentation Month 2009
At the very end of the page, you will see the nucleotide sequence of the p53 protein.
Immediately above the sequence, there is the amino acid sequence under the subheading
‘CDS’. Click CDS to highlight the coding region in the nucleotide sequence.
Name of presentation Month 2009
Now copy the nucleotide sequence into the space below and highlight the CDS region in red:
ORIGIN
1
61
121
181
241
301
361
421
481
541
601
661
721
781
841
901
961
1021
1081
1141
1201
1261
1321
1381
1441
1501
1561
1621
1681
1741
1801
1861
1921
1981
2041
2101
2161
2221
2281
2341
2401
cgtgctttcc
gccatggagg
tcagacctat
atggatgatt
ccagatgaag
cctacaccgg
cagaaaacct
aagtctgtga
acctgccctg
atggccatct
gagcgctgct
aatttgcgtg
tatgagccgc
agttcctgca
tccagtggta
agagaccggc
cccccaggga
aagaaaccac
atgttccgag
ggggggagca
cataaaaaac
gttccccact
tctttgaacc
tgtcccgggg
tggggagtag
atgtaagaaa
gcccacttca
aaggcccata
ttaatgaaat
tgaccccctt
agttgggcag
gtctgacaac
ctggaggatt
agggtcaatt
ctttgttgcc
ccggctcgag
atggccagcc
gtctcaaact
caattgtgag
tctgcatttt
atctcttatt
acgacggtga
agccgcagtc
ggaaactact
tgatgctgtc
ctcccagaat
cggcccctgc
accagggcag
cttgcacgta
tgcagctgtg
acaagcagtc
cagatagcga
tggagtattt
ctgaggttgg
tgggcggcat
atctactggg
gcacagagga
gcactaagcg
tggatggaga
agctgaatga
gggctcactc
tcatgttcaa
gacagcctcc
cttgcttgca
ctccactgaa
gacataccag
tgttcttgca
ccgtactaac
tctgtgaaat
aatgtacatc
gagggtgctt
ctggttaggt
ctcttggtga
tcatctcttg
tcttttttct
caggctggag
cagtcctgcc
aacttttgca
cctgggctca
ccaccacgtc
caccccaccc
ttacaataaa
cacgcttccc
agatcctagc
tcctgaaaac
cccggacgat
gccagaggct
accagccccc
ctacggtttc
ctcccctgcc
ggttgattcc
acagcacatg
tggtctggcc
ggatgacaga
ctctgactgt
gaaccggagg
acggaacagc
agagaatctc
agcactgtcc
atatttcacc
ggccttggaa
cagccacctg
gacagaaggg
cacccccatc
ataggtgtgc
caagttggcc
cttagatttt
gttaagggtt
cagggaagct
gctggcattt
tggccttgaa
gttccctctc
agagggagtt
accttagtac
tatatgatga
tttttttttt
tggagtggcg
tcagcctccg
tgttttgtag
ggcgatccac
cagctggaag
ttcccctcct
actttgctgc
tggattggcc
gtcgagcccc
aacgttctgt
attgaacaat
gctccccgcg
tcctggcccc
cgtctgggct
ctcaacaaga
acacccccgc
acggaggttg
cctcctcagc
aacacttttc
accaccatcc
cccatcctca
tttgaggtgc
cgcaagaaag
aacaacacca
cttcagatcc
ctcaaggatg
aagtccaaaa
cctgactcag
tctccctccc
gtcagaagca
tgcactggtg
aaggttttta
agtttacaat
gtccctcact
gcacctacct
accacctttt
cctgttggtc
gtcaagtctc
ctaaaaggaa
tctggatcca
ttttttcttt
tgatcttggc
gagtagctgg
agatggggtc
ctgtctcagc
ggtcaacatc
tctccctttt
caaaaaaaaa
agactgcctt
ctctgagtca
cccccttgcc
ggttcactga
tggcccctgc
tgtcatcttc
tcttgcattc
tgttttgcca
ccggcacccg
tgaggcgctg
atcttatccg
gacatagtgt
actacaacta
ccatcatcac
atgtttgtgc
gggagcctca
gctcctctcc
gtgggcgtga
cccaggctgg
agggtcagtc
actgacattc
ctgccatttt
cccaggactt
ttttgttgtg
ctgtgaggga
cagccacatt
gttgaatttt
cacagagtgc
attacatggg
ggtgggttgg
tgctggccca
atctcacccc
ccaagacttg
ttctttgaga
ttactgcagc
gaccacaggt
tcacagtgtt
ctcccagagt
ttttacattc
tatatcccat
aaaaaaaaaa
ccgggtcact
ggaaacattt
gtcccaagca
agacccaggt
accagcagct
tgtcccttcc
tgggacagcc
actggccaag
cgtccgcgcc
cccccaccat
agtggaagga
ggtggtgccc
catgtgtaac
actggaagac
ctgtcctggg
ccacgagctg
ccagccaaag
gcgcttcgag
gaaggagcca
tacctcccgc
tccacttctt
gggttttggg
ccatttgctt
gggaggagga
tgtttgggag
ctaggtaggg
ctctaacttc
attgtgaggg
gtctagaact
tagtttctac
gccaaaccct
atcccacacc
ttttatgctc
ctgggtctcg
ctttgcctcc
tcatgccacc
gcccaggctg
gctgggatta
tgcaagcaca
ttttatatcg
a
//
Name of presentation Month 2009
STEP 2: Compare and find out which region your particular primer binds to in
the p53 coding sequence.
Using google, search for ‘nucleotide blast’ – an NCBI sequence alignment program and open
the following page.
Name of presentation Month 2009
Copy and paste your primer sequence in the QUERY box and your p53 nucleotide sequence in
the SUBJECT box. Click on BLAST.
The colourful box shows a visual representation of your results while the exact nucleotide
matches are shown below.
Name of presentation Month 2009
The query sequence (1-29 nucleotides in length) matched exactly to the subject sequence from its
nucleotide sequence starting from 194-222. Now go back to your nucleotide sequence and highlight in
yellow, the corresponding region (194-222). NOTE: There will be a mismatch of one nucleotide (which is
your mutation). Also, the reverse primers will map in the 3’ to 5’ direction.
Name of presentation Month 2009
STEP 3: Find out which amino acid is being mutated and what the mutation is.
Copy your CDS (coding sequence highlighted in red) from the nucleotide sequence below.
121
181
241
301
361
421
481
541
601
661
721
781
841
901
961
1021
1081
1141
1201
atggagg
tcagacctat
atggatgatt
ccagatgaag
cctacaccgg
cagaaaacct
aagtctgtga
acctgccctg
atggccatct
gagcgctgct
aatttgcgtg
tatgagccgc
agttcctgca
tccagtggta
agagaccggc
cccccaggga
aagaaaccac
atgttccgag
ggggggagca
cataaaaaac
agccgcagtc
ggaaactact
tgatgctgtc
ctcccagaat
cggcccctgc
accagggcag
cttgcacgta
tgcagctgtg
acaagcagtc
cagatagcga
tggagtattt
ctgaggttgg
tgggcggcat
atctactggg
gcacagagga
gcactaagcg
tggatggaga
agctgaatga
gggctcactc
tcatgttcaa
agatcctagc
tcctgaaaac
cccggacgat
gccagaggct
accagccccc
ctacggtttc
ctcccctgcc
ggttgattcc
acagcacatg
tggtctggcc
ggatgacaga
ctctgactgt
gaaccggagg
acggaacagc
agagaatctc
agcactgtcc
atatttcacc
ggccttggaa
cagccacctg
gacagaaggg
gtcgagcccc
aacgttctgt
attgaacaat
gctccccgcg
tcctggcccc
cgtctgggct
ctcaacaaga
acacccccgc
acggaggttg
cctcctcagc
aacacttttc
accaccatcc
cccatcctca
tttgaggtgc
cgcaagaaag
aacaacacca
cttcagatcc
ctcaaggatg
aagtccaaaa
cctgactcag
ctctgagtca
cccccttgcc
ggttcactga
tggcccctgc
tgtcatcttc
tcttgcattc
tgttttgcca
ccggcacccg
tgaggcgctg
atcttatccg
gacatagtgt
actacaacta
ccatcatcac
atgtttgtgc
gggagcctca
gctcctctcc
gtgggcgtga
cccaggctgg
agggtcagtc
actga
ggaaacattt
gtcccaagca
agacccaggt
accagcagct
tgtcccttcc
tgggacagcc
actggccaag
cgtccgcgcc
cccccaccat
agtggaagga
ggtggtgccc
catgtgtaac
actggaagac
ctgtcctggg
ccacgagctg
ccagccaaag
gcgcttcgag
gaaggagcca
tacctcccgc
Name of presentation Month 2009
A codon is a group of three nucleotides coding for a single amino acid. Please refer to the
codon table provided below to know the composition of each amino acid. As you will notice,
ATG is a start codon (first codon of your sequence in red) and TGA is a stop codon (last codon
of your sequence in red).
Amino
acid
Ala/A
Arg/R
Codons
GCT, GCC,
GCA, GCG
CODON TABLE
Compresse
Amino acid
d
GCN
Leu/L
CGN,
MGR
Lys/K
Asn/N
Asp/D
Cys/C
CGT, CGC,
CGA,
CGG,
AGA, AGG
AAT, AAC
GAT, GAC
TGT, TGC
AAY
GAY
TGY
Met/M
Phe/F
Pro/P
Gln/Q
CAA, CAG
CAR
Ser/S
Glu/E
GAA, GAG GAR
Thr/T
Gly/G
GGT,
GGN
GGC,
GGA, GGG
CAT, CAC CAY
ATT, ATC, ATH
ATA
ATG
Trp/W
His/H
Ile/I
START
Tyr/Y
Val/V
STOP
Codons
Compresse
d
YTR, CTN
TTA, TTG,
CTT, CTC,
CTA, CTG
AAA, AAG AAR
ATG
TTT, TTC
CCT, CCC,
CCA, CCG
TCT, TCC,
TCA, TCG,
AGT, AGC
ACT, ACC,
ACA, ACG
TGG
TAT, TAC
GTT, GTC,
GTA, GTG
TAA,
TGA, TAG
TTY
CCN
TCN, AGY
ACN
TAY
GTN
TAR, TRA
Name of presentation Month 2009
Next, split your above sequence into groups of three (i.e. into each codon) as shown below.
1
2
3
4
5
6
7
8
9
10
11
atg
tca
gca
cca
cca
tct
cat
ttt
ccc
gtt
cct
aac
tgt
cgg
cgg
gaa
aag
gat
gag
agc
aaa
gag
gac
atg
ggt
gca
gtc
tct
tgc
ggc
gtg
cag
act
acc
agg
aac
gag
cga
gga
ctg
agg
aaa
gag
cta
gat
cca
gct
cct
ggg
caa
acc
agg
cat
ttt
acc
ccc
agc
aat
gca
gaa
aat
gct
ctc
ccg
tgg
gat
gat
cct
tcc
aca
ctg
cgc
cgc
ctt
cga
atc
atc
ttt
ctc
ctg
tat
gag
cac
atg
cag
aaa
ttg
gaa
aca
cag
gcc
gcc
gtc
tgc
atc
cat
cac
ctc
gag
cgc
tcc
ttc
gcc
tcc
ttc
tca
cta
atg
gct
ccg
aaa
aag
aag
cgc
ccc
cga
agt
tac
acc
gtg
aag
aac
acc
ttg
agc
aag
gat
ctt
ctg
ccc
gcg
acc
tct
acc
gcc
cac
gtg
gtg
aac
atc
cat
aaa
aac
ctt
gaa
cac
aca
cct
cct
tcc
aga
gcc
tac
gtg
tgc
atg
cat
gaa
gtg
tac
atc
gtt
ggg
acc
cag
ctc
ctg
gaa
agc
gaa
ccg
atg
cct
cag
act
cct
gcc
gag
gga
gtg
atg
aca
tgt
gag
agc
atc
aag
aag
ggg
gtc
aac
gac
cca
gca
ggc
tgc
gtg
atc
cgc
aat
ccc
tgt
ctg
gcc
cct
tcc
cgt
gat
tcc
cct
gag
aac
gat
gag
cca
agc
acg
cag
tac
tgc
ttg
tat
aac
gaa
tgt
cac
tct
ggg
gcc
aaa
gac
12 13
ccc
gtt
att
gct
gcc
tac
tac
ctg
aag
tca
cgt
gag
agt
gac
cct
cac
ccc
cgt
cag
aag
tca
cct
ctg
gaa
gct
ccc
ggt
tcc
tgg
cag
gat
gtg
ccg
tcc
tcc
ggg
gag
cag
gag
gct
ggt
gac
14 15
ctg
tcc
caa
ccc
tcc
ttc
cct
gtt
tca
agc
gag
cct
tgc
agt
aga
ctg
cca
cgc
ggg
cag
tga
agt
ccc
tgg
cgc
tgg
cgt
gcc
gat
cag
gat
tat
gag
atg
ggt
gac
ccc
aag
ttc
aag
tct
16
17
18
19
cag
ttg
ttc
gtg
ccc
ctg
ctc
tcc
cac
ggt
ttg
gtt
ggc
aat
cgg
cca
aag
gag
gag
acc
gaa
ccg
act
gcc
ctg
ggc
aac
aca
atg
ctg
gat
ggc
ggc
cta
cgc
ggg
aaa
atg
cca
tcc
aca
tcc
gaa
cct
tca
ttc
aag
ccc
acg
gcc
gac
tct
atg
ctg
aca
agc
cca
ttc
ggg
cgc
ttt
caa
gac
gca
tct
ttg
atg
ccg
gag
cct
aga
gac
aac
gga
gag
act
ctg
cga
ggg
cat
Name of presentation Month 2009
Now, copy below the highlighted primer region and pick out the mutation site. In this particular
sequence, the codon being mutated is as follows:
GAC With ‘G’ being mutated
Using your codon table, identify which amino acid ‘GAC’ codes for and write it below.
GAC codes for the amino acid Aspartic acid.
From your primer sequence, you can notice that ‘G’ has been replaced by ‘C’. This
mutation will now change the codon to encode for: CAC which codes for the amino
acid Histidine.
EXAMPLE PRIMER SEQUENCE
tgatgctgtccccgcacgatattgaacaa
Name of presentation Month 2009
Results
Thus we know that at the 48 amino acid position, the codon ‘gac’ is being mutated to
‘cac’ which means that an aspartic acid is being converted into a histidine amino acid.
This can be written in notations as below:
D48H
Name of presentation Month 2009
EXERCISE
A protein’s secondary structures may be visualized in 3D using various softwares like
DeepView. These structures are generated using different techniques, for e.g. x-ray
crystallography and published in various journals. All these structures are also available in
databases like RCSB protein data bank (see below). In this exercise, we shall look into the p53
protein structure and observe the mutations being introduced in this project.
Name of presentation Month 2009
I have previously downloaded the p53 homotetramer structure, PDB ID: A2HI. Open this
structure using the DeepView software. We can see the various control and visualization
panels below:
Name of presentation Month 2009
We can un-highlight a single p53 protein using the control panel (right).
Name of presentation Month 2009
Now, we can highlight a particular mutation, for e.g. R248G which is a change from Arginine
to Glycine at the 248th position, using the control panel. In this structure we can observe
Arginine at the 248th position. In the PCR we did yesterday, one set of primers introduced a
change in this particular Arginine and changed it into a Glycine.
Name of presentation Month 2009
FUTURE EXERCISE
You can also study how this would affect various changes like DNA binding effects (if the
mutation is in the DBD domain), electrostatic potentials, creation and deletion of different
types of bonds, etc using this software. A pdf manual for this software is available for your
perusal in your downtime. You may also find other p53 structures with other mutations in the
RCSB database. Some examples are below:
1) V143A - http://www.rcsb.org/pdb/explore.do?structureId=2j1w
PDB ID code 2J1W
2) R248G - http://www.nature.com/onc/journal/v26/n15/fig_tab/1210291f1.html#figure-title
PDB ID code 2AHI
3) R175H - http://www.nature.com/onc/journal/v26/n15/fig_tab/1210291f1.html#figure-title
PDB ID code 2AHI
Name of presentation Month 2009
Related documents