In my experience using Blogger over the past several months, I have begun to appreciate more its double editing mode: Compose and Edit HTML. The Compose mode is a simple WYSIWYG editor, convenient for most common tasks. Once in a while, however, I get stuck with some nasty formatting issues that could drive one crazy to fix. This is where the Edit HTML mode comes in handy, which allows for full flexibility in editing raw HTML.
In principle, I am pretty competent with HTML and could use the Edit HTML mode directly. However, raw HTML is verbose and thus not that convenient. So I normally start a blog post with the Compose mode for most of the content, and switch to the Edit HTML mode only when necessary. Blogger makes the switch between the two modes a simple button click. This is in contrast to the single editing mode in phpBB3 (BBCode) used by the 3DNA forum, which does not allow for direct access to HTML.
Ideally, a software tool should be both flexible and convenient. In reality, however, not that many software could strike a balance between the two factors. Blogger is a nice example. Interesting, in composing this post, I switched back and forth between the two editing modes a couple of occasions: one after copy-and-pasting Edit HTML to stop red coloring of the following text, and the other to qualify the 3DNA forum link text.
Saturday, October 10, 2009
Friday, October 9, 2009
Fiber models in 3DNA make it easy to build regular DNA helices
3DNA contains 55 fiber models compiled from literature. To the best of my knowledge, this is the most comprehensive collection of its kind (detailed list given below), including:
Once 3DNA in properly installed, the command-line interface is the most versatile and convenient way to generate, e.g., a regular double strand DNA (mostly, B-DNA) of arbitrary sequence. Moreover, the w3DNA and 3DART web-interfaces to 3DNA make it easy to generate a regular DNA model, especially for occasional use or educational purposes.
Theoretically, there is nothing to worth showing off in 3DNA's fiber model generation functionality. However, it serves as a clear example of the differences between a "proof of concept" and a practical software application. I initially decided to work on this issue simply for my own convenience. At that time, I had access to A-DNA and B-DNA fiber model generators, each as a separate program. Moreover, the constructed models did not comply to the PDB format in atom naming and other subtitles.
I started with the Chandrasekaran & Arnott fiber models which I had a copy of data files. Nevertheless, there are many details to work out, typos to correct, etc to put them in a consistent framework. For other models, I read each original publication, and typed into computer raw atomic cylindrical coordinates for each model. Again, quite a few inconsistencies popped up between the different publications with a time span over decades.
Overall, it was a quite tedious undertaking, requiring great attention to details. I am glad that I did that: I learned so much from the process, and more importantly, others can benefit from my effort. As I put in the 3DNA Nature Protocol paper (BOX 6 | FIBER-DIFFRACTION MODELS),
For those who want to understand what's going on under the hood, there is no better way than to try to reproduce the process using, e.g., fiber B-DNA as an example.
From the very beginning, I had expected the fiber functionality to be easily reachable and thus beneficial to all those who are interested in nucleic acid structures, especially to build a regular DNA duplex of chosen sequence. In my sense, the fiber program has not been widely used as it should have been, probably due to the fact that many people are not aware of its existence or capability. Hopefully, this blog post would help to get my message across.
PS. Given below is the content of the README file for fiber models in 3DNA
- Chandrasekaran & Arnott (from #1 to #43) — the most well-known set of fiber models
- Alexeev et al. (#44-#45)
- van Dam & Levitt (#46-#47)
- Premilat & Albiser (#48-#55)
Once 3DNA in properly installed, the command-line interface is the most versatile and convenient way to generate, e.g., a regular double strand DNA (mostly, B-DNA) of arbitrary sequence. Moreover, the w3DNA and 3DART web-interfaces to 3DNA make it easy to generate a regular DNA model, especially for occasional use or educational purposes.
Theoretically, there is nothing to worth showing off in 3DNA's fiber model generation functionality. However, it serves as a clear example of the differences between a "proof of concept" and a practical software application. I initially decided to work on this issue simply for my own convenience. At that time, I had access to A-DNA and B-DNA fiber model generators, each as a separate program. Moreover, the constructed models did not comply to the PDB format in atom naming and other subtitles.
I started with the Chandrasekaran & Arnott fiber models which I had a copy of data files. Nevertheless, there are many details to work out, typos to correct, etc to put them in a consistent framework. For other models, I read each original publication, and typed into computer raw atomic cylindrical coordinates for each model. Again, quite a few inconsistencies popped up between the different publications with a time span over decades.
Overall, it was a quite tedious undertaking, requiring great attention to details. I am glad that I did that: I learned so much from the process, and more importantly, others can benefit from my effort. As I put in the 3DNA Nature Protocol paper (BOX 6 | FIBER-DIFFRACTION MODELS),
In preparing this set of fiber models, we have taken great care to ensure the accuracy and consistency of the models. For completeness and user verification, 3DNA includes, in addition to 3DNA-processed files, the original coordinates collected from the literature.
For those who want to understand what's going on under the hood, there is no better way than to try to reproduce the process using, e.g., fiber B-DNA as an example.
From the very beginning, I had expected the fiber functionality to be easily reachable and thus beneficial to all those who are interested in nucleic acid structures, especially to build a regular DNA duplex of chosen sequence. In my sense, the fiber program has not been widely used as it should have been, probably due to the fact that many people are not aware of its existence or capability. Hopefully, this blog post would help to get my message across.
PS. Given below is the content of the README file for fiber models in 3DNA
1. The repeating units of each fiber structure are mostly based on the
work of Chandrasekaran & Arnott (from #1 to #43). More recent fiber
models are based on Alexeev et al. (#44-#45), van Dam & Levitt (#46
-#47) and Premilat & Albiser (#48-#55).
2. Clean up of each residue
a. currently ignore hydrogen atoms [can be easily added]
b. change ME/C7 group of thymine to C5M
c. re-assign O3' atom to be attached with C3'
d. change distance unit from nm to A [most of the entries]
e. re-ordering atoms according to the NDB convention
3. Fix up of problem structures.
a. str#8 has no N9 atom for guanine
b. str#10 is not available from the disk, manually input
c. str#14 C5M atom was named C5 for Thymine, resulting two C5 atoms
d. str#17 has wrong assignment of O3' atom on Guanine
e. str#33 has wrong C6 position in U3
f. str#37 to #str41 were typed in manually following Arnott's
new list as given in "Oxford Handbook of Nucleic Acid Structure"
edited by S. Neidle (Oxford Press, 1999)
g. str#38 coordinates for N6(A) and N3(T) are WRONG as given in the
original literature
h. str#39 and #40 have the same O3' coordinates for the 2nd strand
4. str#44 & 45 have fixed strand II residues (T)
5. str#46 & 47 have +z-axis upwards (based on BI.pdb & BII.pdb)
6. str#48 to 55 have +z-axis upwards
List of 55 fiber structures
id# Twist Rise Structure description
(dgrees) (A)
-------------------------------------------------------------------------------
1 32.7 2.548 A-DNA (calf thymus; generic sequence: A, C, G and T)
2 65.5 5.095 A-DNA poly d(ABr5U) : poly d(ABr5U)
3 0.0 28.030 A-DNA (calf thymus) poly d(A1T2C3G4G5A6A7T8G9G10T11) :
poly d(A1C2C3A4T5T6C7C8G9A10T11)
4 36.0 3.375 B-DNA (calf thymus; generic sequence: A, C, G and T)
5 72.0 6.720 B-DNA poly d(CG) : poly d(CG)
6 180.0 16.864 B-DNA (calf thymus) poly d(C1C2C3C4C5) : poly d(G6G7G8G9G10)
7 38.6 3.310 C-DNA (calf thymus; generic sequence: A, C, G and T)
8 40.0 3.312 C-DNA poly d(GGT) : poly d(ACC)
9 120.0 9.937 C-DNA poly d(G1G2T3) : poly d(A4C5C6)
10 80.0 6.467 C-DNA poly d(AG) : poly d(CT)
11 80.0 6.467 C-DNA poly d(A1G2) : poly d(C3T4)
12 45.0 3.013 D-DNA poly d(AAT) : poly d(ATT)
13 90.0 6.125 D-DNA poly d(CI) : poly d(CI)
14 -90.0 18.500 D-DNA poly d(A1T2A3T4A5T6) : poly d(A1T2A3T4A5T6)
15 -60.0 7.250 Z-DNA poly d(GC) : poly d(GC)
16 -51.4 7.571 Z-DNA poly d(As4T) : poly d(As4T)
17 0.0 10.200 L-DNA (calf thymus) poly d(GC) : poly d(GC)
18 36.0 3.230 B'-DNA alpha poly d(A) : poly d(T) (H-DNA)
19 36.0 3.233 B'-DNA beta2 poly d(A) : poly d(T) (H-DNA beta)
20 32.7 2.812 A-RNA poly (A) : poly (U)
21 30.0 3.000 A'-RNA poly (I) : poly (C)
22 32.7 2.560 Hybrid poly (A) : poly d(T)
23 32.0 2.780 Hybrid poly d(G) : poly (C)
24 36.0 3.130 Hybrid poly d(I) : poly (C)
25 32.7 3.060 Hybrid poly d(A) : poly (U)
26 36.0 3.010 10-fold poly (X) : poly (X)
27 32.7 2.518 11-fold poly (X) : poly (X)
28 32.7 2.596 Poly (s2U) : poly (s2U) (symmetric base-pair)
29 32.7 2.596 Poly (s2U) : poly (s2U) (asymmetric base-pair)
30 32.7 3.160 Poly d(C) : poly d(I) : poly d(C)
31 30.0 3.260 Poly d(T) : poly d(A) : poly d(T)
32 32.7 3.040 Poly (U) : poly (A) : poly(U) (11-fold)
33 30.0 3.040 Poly (U) : poly (A) : poly(U) (12-fold)
34 30.0 3.290 Poly (I) : poly (A) : poly(I)
35 31.3 3.410 Poly (I) : poly (I) : poly(I) : poly(I)
36 60.0 3.155 Poly (C) or poly (mC) or poly (eC)
37 36.0 3.200 B'-DNA beta2 Poly d(A) : poly d(U)
38 36.0 3.240 B'-DNA beta1 Poly d(A) : poly d(T)
39 72.0 6.480 B'-DNA beta2 Poly d(AI) : poly d(CT)
40 72.0 6.460 B'-DNA beta1 Poly d(AI) : poly d(CT)
41 144.0 13.540 B'-DNA Poly d(AATT) : poly d(AATT)
42 32.7 3.040 Poly(U) : poly d(A) : poly(U) [cf. #32]
43 36.0 3.200 Beta Poly d(A) : Poly d(U) [cf. #37]
44 36.0 3.233 Poly d(A) : poly d(T) (Ca salt)
45 36.0 3.233 Poly d(A) : poly d(T) (Na salt)
46 36.0 3.38 B-DNA (BI-type nucleotides; generic sequence: A, C, G and T)
47 40.0 3.32 C-DNA (BII-type nucleotides; generic sequence: A, C, G and T)
48 87.8 6.02 D(A)-DNA ploy d(AT) : ploy d(AT) (right-handed)
49 60.0 7.20 S-DNA ploy d(CG) : poly d(CG) (C_BG_A, right-handed)
50 60.0 7.20 S-DNA ploy d(GC) : poly d(GC) (C_AG_B, right-handed)
51 31.6 3.22 B*-DNA poly d(A) : poly d(T)
52 90.0 6.06 D(B)-DNA poly d(AT) : poly d(AT) [cf. #48]
53 -38.7 3.29 C-DNA (generic sequence: A, C, G and T) (depreciated)
54 32.73 2.56 A-DNA (generic sequence: A, C, G and T) [cf. #1]
55 36.0 3.39 B-DNA (generic sequence: A, C, G and T) [cf. #4]
-------------------------------------------------------------------------------
List 1-41 based on Struther Arnott: ``Polynucleotide secondary structures:
an historical perspective'', pp. 1-38 in ``Oxford Handbook of Nucleic
Acid Structure'' edited by Stephen Neidle (Oxford Press, 1999).
#42 and #43 are from Chandrasekaran & Arnott: "The Structures of DNA
and RNA Helices in Oriented Fibers", pp 31-170 in "Landolt-Bornstein
Numerical Data and Functional Relationships in Science and Technology"
edited by W. Saenger (Springer-Verlag, 1990).
#44-#45 based on Alexeev et al., ``The structure of poly(dA) . poly(dT)
as revealed by an X-ray fiber diffraction''. J. Biomol. Str. Dyn, 4,
pp. 989-1011, 1987.
#46-#47 based on van Dam & Levitt, ``BII nucleotides in the B and C forms
of natural-sequence polymeric DNA: a new model for the C form of DNA''.
J. Mol. Biol., 304, pp. 541-561, 2000.
#48-#55 based on Premilat & Albiser, ``A new D-DNA form of poly(dA-dT) .
poly(dA-dT): an A-DNA type structure with reversed Hoogsteen Pairing''.
Eur. Biophys. J., 30, pp. 404-410, 2001 (and several other publications).
Sunday, October 4, 2009
3DNA in molecular dynamics simulations
While updating 3DNA citations last Friday, I came across the paper "Flexibility of Short-Strand RNA in Aqueous Solution as Revealed by Molecular Dynamics Simulation: Are A-RNA and A′-RNA Distinct Conformational Structures?" by Ouyang et al. [Aust. J. Chem. 2009, 62, 1054–1061]. Through molecular dynamics (MD) simulations over a 30 ns period, the authors found that "the identification of distinct A-RNA and A′-RNA structures ... may not be generally relevant in the context of RNA in the aqueous phase." Overall, the paper is nicely written.
I have never performed MD simulations in my research experience, so normally I only read abstracts of such publications just to get a general idea of the main conclusions. What had attracted my attention to this work was its simultaneous citations to three earlier papers:
This is quite unusual, since nowadays it is far more common to only cite 3DNA itself (mostly 2003 NAR and/or 2008 NP). So I decided to have a look of the whole paper. It turned out that the authors used the extra citations to justify their choice of using 3DNA instead of Curves to calculate the helical parameters, including "the three main structural descriptors commonly used to differentiate between the two forms of RNA – namely major groove width, inclination and the number of base pairs in a helical twist [turn]".
I communicated to the authors about the availability of Curves+. Specifically, using one case (413D, one of the three structures in their Table 1, "Comparison of different results by CURVES and 3DNA programs"), I've tried to illustrate the point that Curves+ and 3DNA now give directly comparable parameters.
Of course, I am glad to see 3DNA being applied to molecular dynamics simulations of nucleic acid structures. Hopefully, more such applications will show up in the future, and I am willing to offer my help in ways that make sense to me.
I have never performed MD simulations in my research experience, so normally I only read abstracts of such publications just to get a general idea of the main conclusions. What had attracted my attention to this work was its simultaneous citations to three earlier papers:
This is quite unusual, since nowadays it is far more common to only cite 3DNA itself (mostly 2003 NAR and/or 2008 NP). So I decided to have a look of the whole paper. It turned out that the authors used the extra citations to justify their choice of using 3DNA instead of Curves to calculate the helical parameters, including "the three main structural descriptors commonly used to differentiate between the two forms of RNA – namely major groove width, inclination and the number of base pairs in a helical twist [turn]".
I communicated to the authors about the availability of Curves+. Specifically, using one case (413D, one of the three structures in their Table 1, "Comparison of different results by CURVES and 3DNA programs"), I've tried to illustrate the point that Curves+ and 3DNA now give directly comparable parameters.
Of course, I am glad to see 3DNA being applied to molecular dynamics simulations of nucleic acid structures. Hopefully, more such applications will show up in the future, and I am willing to offer my help in ways that make sense to me.
Saturday, October 3, 2009
Whenever in doubt, check with the author
Once in a while, I send emails to authors of papers I am interested in, sometimes simply to ask for PDF reprints, mostly to request for clarifications of points I cannot understand fully. Of course, the responses I have received vary significantly: some authors are responsive and are able to answer my questions concretely; while others respond less professionally; in no small percentage, I get no feedback at all. Whatever the case, though, sending querying emails is convenient, and the responses I get (even no response at all) are informative. Naturally, I would take more seriously the papers whose authors are responsive. On the other hand, in my memory, I have never ignored a reader's question of my publications.
Seeking clarification on a scientific software from author(s) or maintainer(s) is even more important due to inherent subtleties of (undocumented) details, as is common in (bio)informatics. In supporting 3DNA over the years, I've experienced quite a few cases where authors of some articles are misinformed in making judgment about 3DNA's functionality. In one case, I read in a paper claiming 3DNA cannot handle Hoogsteen base-pairs while Curves can. A few email exchanges with the corresponding author (who was very responsive and professional) turned out that an internally modified version of Curves was used. More recently, I found a paper claiming that find_pair from 3DNA failed to identify some base pairs in DNA-protein complexes where a new method succeeded. I asked for the missing list, and immediately noticed that simply relaxing some criteria recovered virtually all of the pairs. Thus, to make a convincing comparison of scientific software, it is crucial to check with the original authors to avoid misunderstandings. Serious scientific software developers and maintainers always welcome users' feedback. Why not ask for clarifications if one really wants to make a (strong) point in comparison? Of course, it is another story for unsupported software.
The Internet age has provided unprecedented convenience for scientific communication. It would be a pity not to take full advantage of it. One simple and important thing to do is: whenever in doubt, ask for clarification from corresponding author of a publication or maintainer of a software.
Seeking clarification on a scientific software from author(s) or maintainer(s) is even more important due to inherent subtleties of (undocumented) details, as is common in (bio)informatics. In supporting 3DNA over the years, I've experienced quite a few cases where authors of some articles are misinformed in making judgment about 3DNA's functionality. In one case, I read in a paper claiming 3DNA cannot handle Hoogsteen base-pairs while Curves can. A few email exchanges with the corresponding author (who was very responsive and professional) turned out that an internally modified version of Curves was used. More recently, I found a paper claiming that find_pair from 3DNA failed to identify some base pairs in DNA-protein complexes where a new method succeeded. I asked for the missing list, and immediately noticed that simply relaxing some criteria recovered virtually all of the pairs. Thus, to make a convincing comparison of scientific software, it is crucial to check with the original authors to avoid misunderstandings. Serious scientific software developers and maintainers always welcome users' feedback. Why not ask for clarifications if one really wants to make a (strong) point in comparison? Of course, it is another story for unsupported software.
The Internet age has provided unprecedented convenience for scientific communication. It would be a pity not to take full advantage of it. One simple and important thing to do is: whenever in doubt, ask for clarification from corresponding author of a publication or maintainer of a software.
Sunday, September 27, 2009
On reproducibility of scientific publications
In the September 25, 2009 issue of Science (Vol. 325, pp.1622-3), I read with interest the letter from Osterweil et al. "Forecast for Reproducible Data: Partly Cloudy" and the response from Nelson. This exchange of views highlights the difficulty/importance for one research team to precisely reproduce results from another when elaborate computation is involved. As is well known, subtle differences in computer hardware and software, different versions of the same software, or even different options of the same version, could all play a role. Without specifying those details, it is virtually impossible to repeat a publication exactly.
This reminds me of a recent paper "Repeatability of published microarray gene expression analyses" by Ioannidis et al. [Nat Genet. 2009, 41(2):149-55]. In the abstract, the authors summarized their findings:
Specifically, please note that:
In my experience and understanding, the methods section in journal articles is not, and should not aim to be, detailed enough for exact duplication by a qualified reader. Instead, most such reproducibility issues would be gone if journals require that authors provide raw data, detailed procedures used to process the data, software version and options used to generate the figures and tables reported in the publication. Such information could be made available in (journal or authors) websites. This is an effective way to solve the problem, especially for computational, informatics-related articles. Over the years, for papers I am the first author or I have made major contributions, I've always kept a folder for each article to include every detail (data files, scripts etc) so that the published tables and figures can be repeated precisely. This has turned out to be extremely helpful when I want to refer back to early publications, or when I was asked by readers for further details.
As noted by Osterweil et al., "repeatability, reproducibility, and transparency are the hallmarks of the scientific enterprise." To really achieve the goal, every scientist needs to pay more attention to details and be responsive. Do not be fool around by the impressive introduction or extensive discussions (which are important, of course) in a paper: to get the bottom of something, it is usually the details that count.
This reminds me of a recent paper "Repeatability of published microarray gene expression analyses" by Ioannidis et al. [Nat Genet. 2009, 41(2):149-55]. In the abstract, the authors summarized their findings:
Here we evaluated the replication of data analyses in 18 articles on microarray-based gene expression profiling published in Nature Genetics in 2005-2006. One table or figure from each article was independently evaluated by two teams of analysts. We reproduced two analyses in principle and six partially or with some discrepancies; ten could not be reproduced. The main reason for failure to reproduce was data unavailability, and discrepancies were mostly due to incomplete data annotation or specification of data processing and analysis.
Specifically, please note that:
- The authors are experts on microarray analysis, not occasional application software users.
- These 18 articles surveyed were published in Nature Genetics, one of the top journals of its field.
- Not a single analysis could be reproduced exactly: two were reproduced in principle, six only partially, and the other ten not at all.
In my experience and understanding, the methods section in journal articles is not, and should not aim to be, detailed enough for exact duplication by a qualified reader. Instead, most such reproducibility issues would be gone if journals require that authors provide raw data, detailed procedures used to process the data, software version and options used to generate the figures and tables reported in the publication. Such information could be made available in (journal or authors) websites. This is an effective way to solve the problem, especially for computational, informatics-related articles. Over the years, for papers I am the first author or I have made major contributions, I've always kept a folder for each article to include every detail (data files, scripts etc) so that the published tables and figures can be repeated precisely. This has turned out to be extremely helpful when I want to refer back to early publications, or when I was asked by readers for further details.
As noted by Osterweil et al., "repeatability, reproducibility, and transparency are the hallmarks of the scientific enterprise." To really achieve the goal, every scientist needs to pay more attention to details and be responsive. Do not be fool around by the impressive introduction or extensive discussions (which are important, of course) in a paper: to get the bottom of something, it is usually the details that count.
Saturday, September 26, 2009
RiboClub 10th Annual Meeting -- it was a good one!
I attended the RiboClub 10th Annual Meeting held during September 21 to 23 at Hotel Cheribourg, Orford, Quebec. Everyday, the schedule was fulfilled from early morning to late night. Indeed, it was so tight and intensive that I could not even find a time to walk around in the beautiful season. Overall, though, the meeting was well-organized and had a fantastic scientific program.
The RiboClub was founded ten years ago by researchers at the University of Sherbrooke, Quebec. Initially, it was local, and then the club was extended to nearby areas, and the whole Canada. As time goes, its influence has also passed the border so at its 10th anniversary, several leading RNA experts (including Phillip Sharp, Jack Szostak, Tim Nilsen, Tom Steitz et al.) from the USA also participated.
I noticed the RiboClub meeting early this year when I was writing up a manuscript on an RNA structural motif. As hinted in another post, "Does 3DNA work for RNA?", I've recently been attracted to the field of RNA structures. Using 3DNA, we have uncovered a simple RNA-specific interaction that is biologically relevant, yet virtually ignored by the community. So attending a RNA meeting would allow me to learn more about RNA and to pass our message across to a wider audience.
The meeting was mostly on various aspects of RNA biology. Of which, 3-dimensional structure is an integral part, yet no specific session was devoted to it. The same applied to bioinformatics tools and applications. Thus, for example, Tom Steitz's talk on the ribosome structures was under the session titled "Translation: targets and impact".
The organizers obviously paid attention to arrange meeting participants at the dinning table. So for Tuesday (Sept. 22) night, I sit next to Dr. Andrew MacMillan from University of Alberta. It was a nice surprise to know that Dr. MacMillan works on "structural and functional characterization of splicesome assembly and activation." I took this opportunity to read the abstracts from his lab and to talk to him about my findings that are related to pre-mRNA splicing. He visited my post the next day (Wednesday, Sept 23) and we discussed it in more details. On the Gala dinner on Wednesday, I was arranged to sit between Dr. Paul Griffiths, "a philosopher of science with a focus on biology and psychology", and Dr. W. Ford Doolittle, a leading scientist in Comparative Genomics. It was a valuable experience to hear them and others around the table talking on politics and science-related issues.
It is worth noting that Wednesday's dinner speaker was Alexander Rich. Wearing his tie from the famous RNA Tie Club, Dr. Rich talked about "The era of RNA awakening: structural biology of RNA in the early years." At the age of 85, he still spoke clearly and logically. His story telling style was very effective and his talk was well-received by the audience. For those who are interested in knowing more about Dr. Rich's work, I would strongly recommend his article "The excitement of discovery."
The "Neo-Traditional Quebec Music" show on Wednesday night was exciting and relaxing, following and in contrast to the three-day long intensive scientific program. Although I did not understand the music that well, I stayed until the very end.
The RiboClub was founded ten years ago by researchers at the University of Sherbrooke, Quebec. Initially, it was local, and then the club was extended to nearby areas, and the whole Canada. As time goes, its influence has also passed the border so at its 10th anniversary, several leading RNA experts (including Phillip Sharp, Jack Szostak, Tim Nilsen, Tom Steitz et al.) from the USA also participated.
I noticed the RiboClub meeting early this year when I was writing up a manuscript on an RNA structural motif. As hinted in another post, "Does 3DNA work for RNA?", I've recently been attracted to the field of RNA structures. Using 3DNA, we have uncovered a simple RNA-specific interaction that is biologically relevant, yet virtually ignored by the community. So attending a RNA meeting would allow me to learn more about RNA and to pass our message across to a wider audience.
The meeting was mostly on various aspects of RNA biology. Of which, 3-dimensional structure is an integral part, yet no specific session was devoted to it. The same applied to bioinformatics tools and applications. Thus, for example, Tom Steitz's talk on the ribosome structures was under the session titled "Translation: targets and impact".
The organizers obviously paid attention to arrange meeting participants at the dinning table. So for Tuesday (Sept. 22) night, I sit next to Dr. Andrew MacMillan from University of Alberta. It was a nice surprise to know that Dr. MacMillan works on "structural and functional characterization of splicesome assembly and activation." I took this opportunity to read the abstracts from his lab and to talk to him about my findings that are related to pre-mRNA splicing. He visited my post the next day (Wednesday, Sept 23) and we discussed it in more details. On the Gala dinner on Wednesday, I was arranged to sit between Dr. Paul Griffiths, "a philosopher of science with a focus on biology and psychology", and Dr. W. Ford Doolittle, a leading scientist in Comparative Genomics. It was a valuable experience to hear them and others around the table talking on politics and science-related issues.
It is worth noting that Wednesday's dinner speaker was Alexander Rich. Wearing his tie from the famous RNA Tie Club, Dr. Rich talked about "The era of RNA awakening: structural biology of RNA in the early years." At the age of 85, he still spoke clearly and logically. His story telling style was very effective and his talk was well-received by the audience. For those who are interested in knowing more about Dr. Rich's work, I would strongly recommend his article "The excitement of discovery."
The "Neo-Traditional Quebec Music" show on Wednesday night was exciting and relaxing, following and in contrast to the three-day long intensive scientific program. Although I did not understand the music that well, I stayed until the very end.
Saturday, September 5, 2009
Double helix groove width parameters from 3DNA
In the 3DNA output (from the analyze program) for a DNA/RNA duplex structure, there is a section on "Minor and major groove widths: direct P-P distances and refined P-P distances which take into account the directions of the sugar-phosphate backbones". The underlying algorithm is that of El Hassan and Calladine (1998). ``Two Distinct Modes of Protein-induced Bending in DNA.'' J. Mol. Biol., v282, pp331-343. Note that the P-P distances need to be subtracted by 5.8 Å to take account of the vdw radii of the phosphate groups (2.9 Å), and for comparisons with NewHelix/FreeHelix and Curves.
Using 3DNA fiber models #1 for A-DNA (calf thymus) and #4 for B-DNA (calf thymus), the groove widths are as follows:
One of the key structural differences between A- and B-DNA is their opposite groove dimensions: for B-DNA, the major groove width (~17 Å) is about 5 Å wider than the minor groove width (~12 Å); whereas for A-DNA, the major groove width (~11 Å) is narrower than the minor groove width (~17 Å) by a similar amount. Since the grooves provide binding sites, the difference between A- and B-DNA grooves has important implications in DNA (groove) recognitions by ligands or proteins.
In retrospect, I implemented the El Hassan and Calladine algorithm for calculating the groove widths mainly because of its simplicity: I can understand clearly how it works visually. The algorithm is described in a two-page appendix of the above cited paper. For those who are interest in DNA structures in general and how groove widths are defined in particular, I would strongly recommend them to read the appendix carefully and try to implement it: there is no substitute for first hand experience. For an idealized cases, as the above for fiber A- and B-DNA, the implementation should be straightforward. To be more realistic, an implementation should account for missing phosphate groups in some structures (for testing purpose, simply delete one P atom from a structure), for example.
As is obvious, 3DNA does not calculate groove depths. Over the years, I have actually been approached with requests/suggestions to provide such parameters to complement groove widths. However, for various reasons, none of the algorithms fits with 3DNA. As a general principle, I do not add new functionality to 3DNA simply for the seek of it. I must understand a new piece clearly in order to integrate it with the rest and to be able to respond concretely to possible questions from users.
Using 3DNA fiber models #1 for A-DNA (calf thymus) and #4 for B-DNA (calf thymus), the groove widths are as follows:
Minor Groove Major GrooveFrom the above table, it is clearly that for A-DNA, the minor and major groove widths for the refined set are smaller than their corresponding non-refined counterparts (i.e., those based on direct P-P distances). For B-DNA, there are no changes between the two sets. It should be noted that in real structures (i.e., non-perfectly regular, as in X-ray crystal structures in the NDB/PDB), there are nearly always some differences between the refined vs. direct P-P distances. As a general rule, the refined set should be used.
P-P Refined P-P Refined
-----------------------------------------------------
A-DNA (#1) 18.5 16.7 15.2 11.1
B-DNA (#4) 11.7 11.7 17.2 17.2
-----------------------------------------------------
One of the key structural differences between A- and B-DNA is their opposite groove dimensions: for B-DNA, the major groove width (~17 Å) is about 5 Å wider than the minor groove width (~12 Å); whereas for A-DNA, the major groove width (~11 Å) is narrower than the minor groove width (~17 Å) by a similar amount. Since the grooves provide binding sites, the difference between A- and B-DNA grooves has important implications in DNA (groove) recognitions by ligands or proteins.
In retrospect, I implemented the El Hassan and Calladine algorithm for calculating the groove widths mainly because of its simplicity: I can understand clearly how it works visually. The algorithm is described in a two-page appendix of the above cited paper. For those who are interest in DNA structures in general and how groove widths are defined in particular, I would strongly recommend them to read the appendix carefully and try to implement it: there is no substitute for first hand experience. For an idealized cases, as the above for fiber A- and B-DNA, the implementation should be straightforward. To be more realistic, an implementation should account for missing phosphate groups in some structures (for testing purpose, simply delete one P atom from a structure), for example.
As is obvious, 3DNA does not calculate groove depths. Over the years, I have actually been approached with requests/suggestions to provide such parameters to complement groove widths. However, for various reasons, none of the algorithms fits with 3DNA. As a general principle, I do not add new functionality to 3DNA simply for the seek of it. I must understand a new piece clearly in order to integrate it with the rest and to be able to respond concretely to possible questions from users.
Some emacs tricks
As a Linux/Unix fan, I am very familiar with vi and use it for quick and simple text editing purposes. Over the years, however, I have been using emacs the most: I like its color coding and programming language-specific editing mode.
Emacs is well-known for its extensibility, so there are many ways to customize it to suit one's taste. In my experience, I have found the following settings convenient:
Emacs is well-known for its extensibility, so there are many ways to customize it to suit one's taste. In my experience, I have found the following settings convenient:
- Highlight the current line with a background color (here "greenyellow")
(require 'highlight-current-line)
(highlight-current-line-on t)
(highlight-current-line-set-bg-color "greenyellow") - Set transient mark mode on, so that selected text become more obvious
(setq transient-mark-mode t)
- Show line number and column number
(setq line-number-mode t)
(setq column-number-mode t)
Sunday, August 30, 2009
JMB celebrates 50 years of protein structure determination
In the September 11, 2009 issue (vol.392, issue 1) of JMB, there are a series of three articles on the determination of the first two protein structures (myoglobin and haemoglobin), an achievement accomplished by Perutz and Kendrew and their colleagues at the Cambridge MRC laboratory in 1950s. Of special interest of this series is the three authors — Bror Strandberg, Richard Dickerson and Michael Rossmann — leading scientists in structure biology, then postdocs actively involved in the late stage of the structure determination.
These reviews are vividly written and provide interesting background information and some technical details on X-ray crystal structure determination (especially on phase angle), and have the following titles:
I have known Dickerson's work on nucleic acid structures for a while, firstly through the famous Drew-Dickerson dodecamer (CGCGAATTCGCG), and I am intimately familiar with his NewHelix/FreeHelix programs. Nevertheless, it is only after reading his above article do I become aware of his initial protein experience. I like Dickerson's writing a lot. For example, on commenting the different styles of Kendrew and Perutz, he wrote: "John was the mentor, guide, and organizer. …… In contrast, Max was a hands-on bench biochemist whose center of gravity was always the laboratory itself. …… Both styles had their merits: one learned from John, but one learned with Max."
An interesting point from Rossmann's article is his description of a secret he kept to himself for many decades: "I [Rossmann] had been privileged to work on the haemoglobin project with Max, but it was also a project that Max had given his whole life to develop. In my enthusiasm to look at the results, I stole the final discovery from Max. …… With the realization of what I had done, all desire to explore further was completely gone." While I vaguely remembered this story from reading the book "Max Perutz and the Secret of Life" by Georgina Ferry several months ago, Rossmann's personal account would make it unforgettable.
It is worth noting that Perutz and Rossmann were among those few who initiated the PDB in 1971 at a Cold Spring Harbor meeting. Finally, given the expertise of the three authors, it is not surprising to read in the Epilogue that "Indeed, structural biology has become the unifying factor of just about every aspect of biology."
These reviews are vividly written and provide interesting background information and some technical details on X-ray crystal structure determination (especially on phase angle), and have the following titles:
- "Building the Ground for the First Two Protein Structures: Myoglobin and Haemoglobin" by Bror Strandberg
- "Myoglobin: A Whale of a Structure!" by Richard Dickerson
- "Recollection of the Events Leading to the Discovery of the Structure of haemoglobin" by Michael Rossmann
I have known Dickerson's work on nucleic acid structures for a while, firstly through the famous Drew-Dickerson dodecamer (CGCGAATTCGCG), and I am intimately familiar with his NewHelix/FreeHelix programs. Nevertheless, it is only after reading his above article do I become aware of his initial protein experience. I like Dickerson's writing a lot. For example, on commenting the different styles of Kendrew and Perutz, he wrote: "John was the mentor, guide, and organizer. …… In contrast, Max was a hands-on bench biochemist whose center of gravity was always the laboratory itself. …… Both styles had their merits: one learned from John, but one learned with Max."
An interesting point from Rossmann's article is his description of a secret he kept to himself for many decades: "I [Rossmann] had been privileged to work on the haemoglobin project with Max, but it was also a project that Max had given his whole life to develop. In my enthusiasm to look at the results, I stole the final discovery from Max. …… With the realization of what I had done, all desire to explore further was completely gone." While I vaguely remembered this story from reading the book "Max Perutz and the Secret of Life" by Georgina Ferry several months ago, Rossmann's personal account would make it unforgettable.
It is worth noting that Perutz and Rossmann were among those few who initiated the PDB in 1971 at a Cold Spring Harbor meeting. Finally, given the expertise of the three authors, it is not surprising to read in the Epilogue that "Indeed, structural biology has become the unifying factor of just about every aspect of biology."
Subscribe to:
Posts (Atom)
