Sunday, March 20, 2011

3DNA citations reach over 500

On Friday, June 5, 2009, I blogged on the topic titled "3DNA citations reach over 300". At that time, I wrote (towards the end):
I still remember that the number of citations to 3DNA was less than 150 nearly two years ago [~ summer 2007], when I started to wrote the first draft of our 2008 Nature Protocols paper. Now it is more than doubled! I would blog on this topic again when the number reaches 500.
When I checked Google scholar for 3DNA citations right now, the citation number is already over 500 for the initial 2003 3DNA NAR paper alone. Combined with the two direct follow-ups – the 2008 Nature Protocols paper and the 2009 NAR web server paper – the three 3DNA publications have been cited a total of 550 times.

Again, as noted in that blog post,
In my opinion, some of 3DNA features are still (heavily) underused. Now that we have a sizable user community, 3DNA could only become better and would be more widely used. I have every reason to believe that in the not-so-distant-future, the citations to 3DNA would reach over 1000.
A decade after its initial humber release, 3DNA has been successfully applied to many real-world problems. As spare time permits, I have actively maintained and continuously refined 3DNA based largely on users' feedbacks. Over the time, I also see clearly that 3DNA can be moved to the next level both in functionality and usability to enjoy an even larger/broader impact.

Now more than half-way through, it won't be long when citations to 3DNA reach 1000, and then beyond.

Sunday, March 13, 2011

Review article on NMR analysis of protein–DNA interactions by Milon et al

Through Google scholar, I became aware of a recent review article by Milon et al., titled "Nuclear magnetic resonance analysis of protein–DNA interactions" in the journal J. R. Soc. Interface:
This review focuses on the experimental strategies currently employed to solve structures of protein–DNA complexes and to analyse their dynamics. It highlights how these approaches can help in understanding detailed molecular mechanisms of target recognition.
I browsed through the text to get myself more familiar with NMR the methodology and its applications in protein-DNA recognition. I was surprised that 3DNA was cited in the article, especially with respect to its unique analyze/rebuild complementarity:
In addition, several software programs have been developed to model DNA bending such as the 3DNA program, which allows analysis of DNA structural parameters and enables it to be rebuilt with customized DNA models [76]. Several Web servers have been created recently and provide interesting tools to analyse and rebuild DNA models [77,78].
I am only wishing that 3DNA's neat features could be more widely recognized; hopefully I'd have the opportunity to further refine 3DNA and move it to the next level.

Saturday, March 5, 2011

Retraction of scientific publications

Once in a while, I come across retraction notices of scientific publications in leading journals/magazines. Even for cases not directly related to my research areas, I normally browse through them.

In the March 3, 2011 issue of Nature, there is a retraction of the Letter "Mediation of pathogen resistance by exudation of antimicrobials from roots" [Nature 434, 217–221 (2005)]. I am intrigued by the first sentence of the note:
The authors wish to retract this Letter after a key reference by Walker et al. (ref. 9 in this Letter) was retracted from the scientific literature.

It turns out that the 2003 Walter et al. J. Agric. Food Chem. paper (withdrawn in October 2009) and the 2005 Nature Letter were from the same group. Overall, it took ~6 years each for the two papers to be retracted. As of today, they have been cited 76 and 84 times respectively accordingly to Google scholar.

Sunday, February 27, 2011

Evidences for transient Hoogsteen base pairs in canonical DNA duplex

In the February 24, 2011 issue of Nature, there is an interesting article by Nikolova et al., titled "Transient Hoogsteen base pairs in canonical duplex DNA". Its main discovery is succinctly summarized in the abstract:
By using nuclear magnetic resonance relaxation dispersion spectroscopy in concert with steered molecular dynamics simulations, we have observed transient sequence-specific excursions away from Watson–Crick base-pairing at CA and TA steps inside canonical duplex DNA towards low-populated and short-lived A•T and G•C Hoogsteen base pairs. The observation of Hoogsteen base pairs in DNA duplexes specifically bound to transcription factors and in damaged DNA sites implies that the DNA double helix intrinsically codes for excited state Hoogsteen base pairs as a means of expanding its structural complexity beyond that which can be achieved based on Watson–Crick base-pairing.
Geometrically, the Hoogsteen base pair is related to the Watson-Crick base pair by a 180-degree rotation about the glycosidic bond (N9–C1'). While the A•T Hoogsteen base pair is classic, the similar G•C+ Hoogsteen pair (with protonation of cytosine N3) is equally possible. The A•T and G•C Hoogsteen base pairs have two perfect H-bonds, so they are energetically stable. As for their existence in DNA duplex, the most direct evidence comes from the "trap" experiments (see Fig.3 of the paper). In the News & Views section, Honig and Rohs provide a nice recap of the main point and implications of this work.

As also observed in another recent publication, "Replication infidelity via a mismatch with Watson–Crick geometry", the base sequence has a subtle role in influencing the base-pairing schemes, three-dimensional structures and biological functions of DNA. However, we should not forget that only the Watson-Crick base pairs, and to a less extent, the G-U wobble pair, have the correct symmetry to ensure a "regular" double helical structure.

Sunday, February 20, 2011

Canned responses in gmail make it easy to send common messages

Through Gary Rosenzweig's MacMost Now video #509 "Gmail Labs" (January 28, 2011), I first heard of "Canned Responses" in Gmail Labs:
Email for the truly lazy. Save and then send your common messages using a button next to the compose form.
This is a truly handy feature that I have long been waiting for! Yet even though I am aware of Gmail Labs and enabled quite a few experimental features a while ago, I've not been searching Gmail Labs for new features ever since.

Over the past few weeks, I have found "Canned Responses" increasingly indispensable in my support of the 3DNA forum (as a sideline project):
  1. When I activate a new 3DNA forum registration, I've always included a "standard" message to "make the forum policy upfront and explicit, in order to avoid misunderstandings or surprises." Previously, I had to copy-and-paste, e.g., from a specifically created text file or elsewhere. Surely, this worked, but I had felt intuitively that there must be a better way to get the job done. Well, that's exactly where "Canned Responses" fit in!
  2. Over the past few months, I have been ever more bothered by spam registrations. So as a further filter, I have been sending the following enquiry message to each suspicious registration for activation:
    Thanks for your registration at the 3DNA forum. Please tell me a little bit about yourself and elaborate on how 3DNA could be useful to your project; we would like to make the forum spam-free.

    See also "Further notes on forum registration and posting" -- you may not need to register.
    Here once again, the "Canned Responses" feature makes my life much easier! Moreover, this step turns out to be extremely effective; a large percentage of registrations is filtered out at this final stage.
Are you using gmail? If so, you may also want to give "Canned Responses" a try.

Sunday, February 13, 2011

Making data maximally available?

In the February 11, 2011 issue of Science (Vol. 331 no. 6018 p. 649), there is an editorial, titled "Making Data Maximally Available". Indeed, the issue contains a special section on "Dealing with Data".
Science is driven by data. New technologies have vastly increased the ease of data collection and consequently the amount of data collected, while also enabling data to be independently mined and reanalyzed by others ... It is obvious that making data widely available is an essential element of scientific research.
Especially, I like the following two (proposed) new policies:
  1. To extended data access requirement "to include computer codes involved in the creation or analysis of data." If properly implemented/enforced, this policy could significantly increase the repeatability and assessment of published results. In my experience, I have observed too many times that secrets are hidden in the seemingly "little" subtle details.
  2. "To produce a single list that combines references from the main paper and the SOM" (supporting online material) to "provide credit and reveal data sources more clearly". Potentially, this will also increase the citation of method papers.
Hopefully, other journals will follow Science's lead to make data maximally available, and to present data more transparently.

Sunday, February 6, 2011

A G-T mismatch with perfect Watson-Crick geometry

In the February 1, 2011 issue of PNAS [108(5)], there is an interesting article "Replication infidelity via a mismatch with Watson–Crick geometry" by Bebenek et al. They solved the Pol λ DL ternary complex (PDB id: 3PML) which has a G-T nascent mispair in "perfect" Watson-Crick geometry (see their Fig 4, linked below).


From an H-bonding (energitic) point of view, the G-T mispair (with three "H-bonds") can only be possible if G or T is in the rare enol tautomeric state, or is ionized. The pH dependence of single nucleotide disincorporation seems to be consistent with an ionized base pair. Note the G-T mispair is different from a Wobble pair in which G and T have a relative sheared motion (see also Fig. 4C above).

I am glad to find that 3DNA was used in deriving the parameters. By design, 3DNA should be able to identify such "unusual" mispair as easily as for a normal Watson-Crick pair. As noted in our 2008 3DNA Nature Protocols paper,
By taking advantage of the standard base reference frame and selected geometric features, the find_pair program within 3DNA can identify all possible nucleic-acid base pairs, whether they are canonical Watson–Crick or noncanonical pairs and are made up of normal or modified bases, in any tautomeric form or protonation state. (p1217)
Moreover, 3DNA does notice and signify (wit a * instead of the normal -) the atypical H-bonding feature of the G-T mispair to draw further attentions.
5 T-*---g  [3]  O2 - N2  3.06  N3 * N1  2.97  O4 * O6  2.66
Hopefully, this example helps illustrate some of 3DNA's unique features that would hopefully be more widely recognized and applied.

Sunday, January 30, 2011

Trial of scientific papers by blogs and tweets?

In the January 20, 2011 issue of Nature, there is an interesting News Feature article by Apoorva Mandavilli, titled "Peer review: Trial by Twitter":
Blogs and tweets are ripping papers apart within days of publication, leaving researchers unsure how to react.
Specifically, two widely publicized papers in Science are singled out: one is about the longevity genes identified through genome-wide association study (GWAS), published last July; another is a more recent one on arsenic bacteria, published last December. In each case, while "the popular media was trumpeting the finding, other researchers were taking to the web to criticize the paper’s methodology." Yet, the authors failed to hold up their claims in the papers.

With great interest, I've been following the story on arsenic bacteria. I first noticed this work through Science Podcast, and found the topic of "arsenic" life intriguing. So for general knowledge, I read carefully the abstract and browses through the text. One week later came the Nature editorial "Response required", and from which, I followed the link to Rosie Redfield's blog post "Arsenic-associated bacteria (NASA's claims)". I read the post, and many of the comments therein; while I do not understand many of the technical details, I had no difficulty in following her argument. The lead author of the arsenic bacteria paper, Dr. Felisa Wolfe-Simon, did respond to comments on December 16, 2010. Science also published an interview with Wolfe-Simon, titled "Discoverer Asks for Time, Patience Over Arsenic Bacteria Controversy". On the same day, Redfield was quick to write another blog post "Comments on Dr. Wolfe-Simon's Response", which again has received many comments. So far, the story is still unfolding, and Science has promised to publish technical comments and responses in early 2011.

As a related topic, throughout the CCP4bb, I noticed the letter to editor titled "Is too ‘creative’ language acceptable in crystallography?" by Alexander Wlodawer et al. I agree fully with the authors that "While figures of speech are often useful and even educational, flashy titles combined with hyperbolae and imprecise language can mislead or deceive nonspecialist readers and should therefore be avoided."

In the Internet Age, bloggers and tweeters clearly have an important role to play in the assessment of research findings. In scientific publications, what counts is not how much one claims, but to what extent one can hold up such claims. Solid work holds up over time and the scrutiny of peers.

Saturday, January 22, 2011

Three structural biology papers in the latest issue of NAR cite 3DNA

While browsing the latest 39(2) January 2011 issue of Nucleic Acids Research (NAR), I found, to my great surprise, three papers that cite 3DNA. These papers, all under the "structural biology" section, are of interest to me from their titles and abstracts, so I downloaded the PDF versions and read through each of them.

For this blog post, #100 by incidence, it would be intriguing to look into the context to see how 3DNA is cited.


"Asymmetric DNA recognition by the OkrAI endonuclease, an isoschizomer of BamHI" by Vanamee et al. (Mount Sinai School of Medicine, and New England Biolabs):

Analysis of the stereochemical quality of the protein model and assignment of secondary structure were conducted with PROCHECK (13). DNA analysis was performed with 3DNA (14). Solvent-accessible surface areas were calculated in CNS with the algorithm of Lee and Richards employing a 1.4-Å probe(15). Figures were prepared using PyMOL (www.pymol.org). [p713, from bottom left to middle right]


"DNA intercalation without flipping in the specific ThaI–DNA complex" by Firczuk et al. (Poland, Germany and UK):
An oligoduplex with the correct sequence in standard B-DNA geometry was generated with the program 3DNA (44), and manually adjusted to fit the highly distorted DNA in the structure. ... The programs COOT (45), REFMAC (46) and CNS (47) were used for refinement. [p747, top left]
Analysis with the 3DNA software (44) shows that the intercalation increases the rise between base pairs to about 7 Å or approximately twice its usual value (Figure 5B). Phosphorus–phosphorus (Pn–Pn+1) distances in the DNA backbone are only mildly altered (values range from 5.6 to 7.0 Å). Instead, the extra height of the two CG steps comes at the expense of the twist, which is reduced from its usual value of about 36° (360°/10) to between 10 and 15°. A view toward the major groove shows that the inner base pairs of the recognition sequence are strongly tilted (Figure 5). According to the 3DNA software (44), the first CG step has a negative tilt of about ~12°, which results in the oblique orientation of the following base pairs. The central GC step is characterized by a tilt close to 0°, reflecting the nearly parallel arrangement of the middle bases. Finally, the second CG step has a positive tilt of about 15° which restores the standard orientation of the downstream base pairs. A side view of the DNA indicates a bend at the center of the recognition sequence which is primarily due to the positive ~12° roll of the central GC step into the major groove (Table 1). The 3DNA program also indicates that the propeller twist is positive for the specifically recognized sequence, and (as expected for the standard B-DNA) negative for most of the flanking base pairs. [p749, top right]
Table 1. DNA distortion in complex with ThaI restriction endonuclease: all parameters were calculated with the 3DNA software (44). [p750, middle left]


"On the molecular basis of uracil recognition in DNA: comparative study of T-A versus U-A structure, dynamics and open base pair kinetics" by Fadda and Pomès (Ireland and Canada):
MD simulations were run with versions 3.3.3 up to 4.0.4 of the GROMACS software package (47,48).

Structural parameters were determined with the 3DNA software package (51,52). The pymol (www .pymol.org) software package was used to generate figures. [p769, bottom right]

Established in 1974 and currently with an impact factor of 7.479, NAR has also been chosen by the Special Libraries Association as one of the top 100 most influential journals in medicine and biology over the last 100 years. The citations by the three papers in the latest issue of NAR illustrate unambiguously 3DNA's big impact in structural biology.