Saturday, July 25, 2009

PDB id vs NDB id

For nucleic-acid-containing structures, PDB and NDB are the two most widely used databases (databanks). Both PDB and NDB are maintained at Rutgers University. Among the two, PDB is primary, of which NDB is essentially a subset with extra derived parameters regarding base-pair geometry.

As is always the case, each entry is uniquely identified by an id in a database. Interestingly, PDB and NDB have adopted radically different approaches in picking up their ids.
  • PDB id is (currently) 4 characters long: the first character is a numeral in the range 1-9, while the rest can be either numerals or letters. Early PDB entries could be acronyms. For example, 1bna for the famous Dickerson-Drew B-DNA dodecamer with sequence CGCGAATTCGCG, the first full turn B-DNA duplex; and 1mbn for myoglobin, the first solved protein structure. Recently, due to the quick increase of deposited macromolecular structures, the PDB ids "are automatically assigned and do not have any meaning." (page 9)
  • NDB id by design seems to contain more information, even though detailed specifications cannot be located from online search. For examples, A-DNA, B-DNA and Z-DNA ids start with AD, BD, and ZD, respectively; protein-DNA complexes start with PD; and ribosomal RNAs start with RR etc. Furthermore, the third letter also has a meaning in the NDB code. E.g., L in BDL084 means 12 since it is the 12th letter in English alphabet, thus we know BDL084 is a B-DNA dodecamer. Similarly, the H in ADH026 means 8, thus ADH026 is an A-DNA octomer.
Since NDB (1992) appeared much late than PDB (1971) and was developed as a better database for macromolecular structures than the PDB (at that time), it is conceivable that its id scheme was part of the initial NDB design. However, even though the NDB id serves its purpose well (up to now, and in a broad sense), users need to be aware of one fundamental flaw inherent in the literal meaning of the NDB ids. As a concrete example, for the Ng et al. (1999) crystal structure of an A/B-DNA intermediate, PDB assigned it an id of 1dc0 -- no intuitive meaning or misleading, just an identifier. In contrast, NDB assigned it an id of BD0026, meaning B-DNA, following the pattern noted above, which is clearly misleading. Moreover, the structure is actually more similar to A-DNA than to B-DNA, as far as the characteristic parameters distinguishing A- and B-DNA -- slide, chi torsion angle, and sugar conformation -- are concerned.

I have no idea of how many such mis-picked ids exist in the NDB. What is clear is that as more and more weird structures (especially RNA) are deposited (or extracted from the PDB), it would be even harder to pick up an id in its 'canonical' sense. Inconsistency will then become a big issue. In contrast, PDB ids do not have such a problem by design, whether an id is an acronym or a random, automatic pick by a software program.

Over the years, NDB has served me no other purposes than as a pre-selected subset of PDB entries containing nucleic acid structures. It has become clear to me that starting directly from PDB would be a better choice, if nothing but to reduce a level of redundancy, and to avoid possible mis-leading ids.

Friday, July 17, 2009

Does open access increase citation?

In the July 17, 2009 issue of Science, there are several letters discussing the brevia titled "Open Access and Global Participation in Science" by Evans and Reimer who reported that:
The influence of OA [open access] is more modest than many have proposed, at ~8% for recently published research, but our work provides clear support for its ability to widen the global circle of those who can participate in science and benefit from it.
On one hand, Philip Davis from Cornell University argued that "Open Access: Increased Citations Not Guaranteed" in the title of his letter. On the other hand, Michael Eisen (HHMI and UC Berkeley) and Steven Salzberg (University of Maryland) stated that "Open Access: The Sooner the Better" -- In their opinion, "the 8% statistic that Evans and Reimer highlight is misleading. .... In particular, when articles were made freely available within 2 years of publication, their citations increased by almost 20%." Interesting, the same issue also published Evens' response, addressing comments and criticisms from other scientists.

Another very interesting point: Eisen and Salzberg also expressed their concern about the unavailability of the raw citation information used in the Evans and Reimer report, saying this "is an astonishing violation of the norms of science, and the explicitly stated publication policies of Science." In response to this point, Science's Editor noted that "[Science] do not preclude our authors from obtaining data from commercial sources when those are the only sources of the data and when those data are available to the scientific community."

Overall, the discrepancies about the influence of open access on citations, and how to possibly resolve them in relation to the availability of the original data, are typical in science. Presumably, this is a relatively simple case. Yet, there are still many variables in data selections and interpretations etc. Even if the "raw data" are made available, which certainly would be a big help, I still doubt that the discrepancies could be resolved. On the other hand, the "raw data" are only secondary in the sense that they were collected by Evans and Reimer using some specific criteria. If the detailed steps are made available such that the reported figures and tables can be reproduced by those who have access to the commercial sources, then things would become clear. In other words, it is not just the data nor the numbers (in published figures and tables), but the exact procedures, of how the numbers have been produced from the original data, that could provide a convincing resolution (if there is one).

Saturday, July 11, 2009

On maintaining the 3DNA forum

Over the past few years, maintaining the 3DNA forum (i.e., answering questions, performing administrative tasks) has taken up a significant amount of my spare time. Sometimes it could be quite demanding, especially because I need to pay great attention to details. Overall, though, it is a valuable experience, and I feel that the time is well-spent: 3DNA has been continuously refined and more widely used; my knowledge of nucleic acid structures (especially RNA) has been significantly sharpened; I have stayed aware of progress in related research fields and see more of the world; and I feel great pleasure in being of help to the community.

Some basic facts/statistics:
  • I tell everyone who sends me an email to ask a 3DNA-related question to register and re-post in the 3DNA forum, but less than 50% actually do this. However, if you want to be helped, you have to follow the rules.

  • Most of the forum registrations (over 90% at times) are spam. To forestall this, I must continually update phpBB3 to its latest version.

  • Over 50% of legitimate registrations end up being deleted instead of being activated due to the users' failure to send me an email for activation, as required. Among those whose send me an email, most use a subject line of "Re: 3DNA forum registration 'user-id'", as suggested. Only a few volunteer to share with me their real name, address, etc. (e.g., via a signature). Furthermore, some do not post back in the forum after their account is activated.

  • With very few exceptions, questions posted in the forum are normally addressed within a couple of days, or even sooner.

  • Except for one case, communicating with users has mostly been a pleasant experience. Some (though not many) users even posted back a summary and/or a thank-you note.

  • I would like to compliment esguerra, yrxin and tgaillar for sharing their tips and tricks in the section Users' contributions. Apart from me, ghzheng has contributed the most in the forum.

  • Overall, the forum is low volume (which is just fine) and spam-free (which is very important).

To make the forum policy upfront and explicit, in order to avoid misunderstandings or surprises, I have been enclosing the following note in each new registration confirmation message:
Your 3DNA forum registration has been activated — welcome aboard! See
http://xiang-jun.blogspot.com/2009/07/on-maintaining-3dna-forum.html

---------------------------------------------------------------------
I am so pleased that you have come thus far! To make the 3DNA forum a
more pleasant virtual community for all of us to learn from and
contribute to, please be considerate and practice good netiquette
(http://www.albion.com/netiquette/). More specifically, I would like
to reemphasize the following:

0. Do your homework; read the FAQ and browse the forum.

1. Ask your questions in the 3DNA forum instead of sending me emails.

2. Be specific with your questions; provide a minimal, reproducible
example if possible; use attachments where appropriate.

3. Do not ask for or expect immediate responses to your questions.
Lower your expectations and you will more likely end up feeling
happier.

4. Respond to requests for clarifications.

5. Summarize the solution to your problem(s) from a user's
perspective by providing details, for the benefit of other users.

6+ Contribute back to 3DNA if you can:
o Report bugs — including typos
o Make constructive suggestions — anything to make 3DNA better
o Answer other users' questions
o Share your use cases in the "Users' contributions" section

In a nutshell, you are welcome to participate and should not hesitate
to ask questions, but remember to play nice and preferably share what
you've learned!
---------------------------------------------------------------------
Thus, intentionally or otherwise, the forum has also acted as a filter to make my life easier. Whenever possible, though, I have tried my best to reward those who follow the simple, common sense rules. After all, nothing should be taken for granted, and no one likes to be taken advantage of. I am glad that through my contributions and user involvement, the forum has survived and 3DNA has thrived (evident from citations, numerous web links to its homepage, other services/tools — including NDB and PDB — taking advantage of parts of its functionality, and more recently, two dedicated web-interfaces), serving as a valuable resource to the community.



PS: Two related posts in the 3DNA forum:

Friday, July 10, 2009

Does 3DNA work for RNA?

At the C2B2 party this afternoon, I was asked the question: "Does 3DNA work for RNA?" Well, a good question, indeed. The short answer is definitely, YES. However, a detailed explanation is needed to address the underlying intuitive assumption: 3DNA is only for DNA.
  1. The name 3DNA was due to Dr. Olson, after we struggled quite a while. Initially, we played with NuStar (which was actually cited once by Richard Dickerson et al), and Carnival etc. I still remember the day when Dr. Olson asked me "How about 3DNA?" We immediately reached an agreement: that's it -- what a cute name! Another advantage (as it becomes clear later): since 3DNA starts with '3', it (mostly) shows up right at the top of many on-line lists of bioinformatics tools.
  2. Interpreted literally, 3DNA could mean 3-DNA, i.e., the three most common types of DNA: A-, B- and Z-form. That may be one of the reasons where the misconception that 3DNA is only for 3DNA comes from. Another reason could be that structural work on DNA is what the Olson lab best known for.
  3. The number '3' in 3DNA should also be associated with its three key components: analysis, rebuilding and visualization. In a sense, this is my favorite.
  4. Of course, 3DNA stands for 3D-NA, 3-Dimensional Nucleic Acids, as expressed explicitly in the titles of our two 3DNA papers (2003 NAR and 2008 NP).
The applications of 3DNA to RNA structures can be broadly categorized as follows:
  1. Automatically detect all existing base-pairs, Watson-Crick (A-U, G-C, wobble G-U) or non-canonical, using a set of simple geometric criteria. Furthermore, it has a unique base-pair classification system based on the six numerical structural parameters, suitable for database storage and search.
  2. Automatically detect all triplets or higher-order base-associations.
  3. Automatically detect double helical regions, regardless of backbone connection, thus ideal for finding pseudo-continuous coaxial stacking.
  4. The above three features are seamlessly integrated with the visualization component to allow for easy generation of publication quality images. See the 3DNA 2008 NP paper for detailed examples.
As further examples, the following two RNA publications take advantage of find_pair from 3DNA:
It is well worth noting that the base-pair detecting algorithm in RNAView is based on an earlier version of find_pair, a basic fact ignored in the RNAView publication.

In summary, 3DNA works for RNA as well as for DNA, and more.

Sunday, July 5, 2009

Errors in PDB entries

In the June 24, 2009 issue of Nature (v459, pp.1038-1039), there is an news item titled "New protein structures replace the old" by Katharine Sanderson, on a 'Dutch software to weed out errors in Protein Data Bank'.

In my experience with software development and using the PDB/NDB, it is certainly not a surprise that there are errors of various types in the macromolecular databases: whenever I apply an algorithm consistently to all the entries in the NDB (which is part of PDB, consisting of only nucleic-acid containing structures), I always notice some inconsistencies. As a more concrete example, blocview, a visualization tool initially developed as a by-product of another project while I was still at Rutgers (and partially involved with the NDB), was once used for correcting errors in the NDB as well.

Pure 're-refinement' of existing structures with software is surely helpful in catching obvious, systemic errors. However, it is impossible to catch all problems, no matter how sophisticated the software could be. Moreover, as put in the comment by yet another phd: "i have hard enough time getting the RCSB to change four atoms in a structure for me." and "scientists must remember that when they click the re-refine button, you read the paper where the structure was reported."

Errors will always be there -- that's just a basic fact of life. It is thus crucial for those who perform structural analysis to draw their conclusions based on not one or just a few purposely selected structures, but on a more objective and extensive ground.

Two web-interfaces to 3DNA, and more

As a nice surprise, I found in the 2009 web-server issue of NAR published on July 1, two articles on web-interface to 3DNA back-to-back:
  1. "3D-DART: a DNA structure modelling server" by van Dijk and Bonvin from Utrecht University, The Netherlands. Excerpt from the abstract:
    As a response to the demand for 3D-structural models reflecting the intrinsic plasticity of DNA we present the 3D-DART server (3DNA-Driven DNA Analysis and Rebuilding Tool). The server provides an easy interface to a powerful collection of tools for the generation of DNA-structural models in custom conformations. The computational engine beyond the server makes use of the 3DNA software suite together with a collection of home-written python scripts.
  2. "Web 3DNA—a web server for the analysis, reconstruction, and visualization of three-dimensional nucleic-acid structures" by Zheng, me and Olson. Excerpt from the abstract:
    The w3DNA (web 3DNA) server is a user-friendly web-based interface to the 3DNA suite of programs for the analysis, reconstruction, and visualization of three-dimensional (3D) nucleic-acid-containing structures, including their complexes with proteins and other ligands.
While I was aware of 3D-DART prior to its publication, I certainly did not expect it to appear in the same issue as w3DNA of which I am a co-author. There is nothing more compelling to illustrate 3DNA's value to the community than a third-party web-interface to it! Combined together, these two web-servers make 3DNA much more accessible to even wider audience. Specifically, they could well serve for education purposes, e.g., to conveniently build a DNA-model of A-, B- or C-form, with user-supplied sequences. Users would be glad to have a choice that better fits their needs. In the long run, the one which provides best user-support will survive.

It is also worthy noting the in the same 2009 NAR web-server issue, another paper titled "SARA: a server for function annotation of RNA structures" by Capriotti and Marti-Renom from Spain also makes use of 3DNA. This serves to emphasize the point that 3DNA is not just for DNA, but for RNA as well -- 3DNA has unique features for RNA that are not found in other currently available software tools that I am aware of.

Friday, June 19, 2009

RasMol 2.6.4 and 3DNA base rectangular block schematic representation

Over the years, I have been using RasMol v2.6.4, the last release in December 1998 by Roger Sayle, the original author of this molecular graphics visualization tool. According to the README file, RasMol v2.6.4 "includes mainly minor bug fixes and portability enhancements" to the version Sayle used internally.

Back in 1990s and early 2000s, RasMol was the most popular tool of its kind. However, even though RasMol v2.7.x series is actively maintained, in recent years JMol, PyMol and other similar tools have surpassed RasMol in popularity. So why still RasMol v2.6.4? To me, RasMol is a light-weight, robust tool for molecular structure visualization, just like ls, grep Unix commands for my daily computing life. Moreover, RasMol v2.6.4 is connected to 3DNA with the visualization of Alchemy format for rectangular base (or base-pair) block representation.

It is straightforward to install RasMol v2.6.4 in a modern computer system with a C compiler. You may also want to customize it, and more importantly, get to know how to use it effectively.
  • RasMol v2.6.4 can be downloaded from http://www.dcs.ed.ac.uk/home/rasmol/v2.6beta/RasMol2.tar.gz
  • In Linux, type "xmkmf; make" to compile, and this gives executable rasmol. By default, this is for 8-bit graphics. I normally renamed it to rasmol8 by "mv rasmol rasmol8" because the 8-bit version is useful for generating gif images (for web graphics, for example). To make a 32-bit version as default, edit Makefile macro DEPTHDEF: "# DEPTHDEF = -DEIGHTBIT" to "DEPTHDEF = -DTHIRTYTWOBIT", then "make clean; make". Move "rasmol8" and "rasmol" to a directory in your command line search path (e.g., ~/bin), copy "rasmol.hlp" to some location and set up environmental variable RASMOLPATH accordingly, you are all set.
  • By default, rasmol binary does not allow for outputting graphics in (automatic) scripting mode, this can be overwritten by changing AllowWrite = False; at line 1163 of "rasmol.c" (in function InitDefaultValues()) to AllowWrite = True;. While you are at it, you may also want to change other defaults. Just remember to "make clean; make".
  • It is worthwhile to study the help file to get family with its contents, especially the syntax for selections. The real power of RasMol comes with its scripting capability. It can be used to automate an otherwise tedious operation.
In relation to 3DNA (and SCHNArP for that matter), RasMol v2.6.4 can be used to visualize the Calladine-Drew style rectangular representation of bases (base-pairs) in Alchemy format. As a matter of fact, I initially picked up Alchemy format in SCHNArP (and later in 3DNA) primarily because it is one of the structure file formats supported by RasMol 2.6.x. JMol does not support 3DNA generated Alchemy format file until recently when I made a request, and PyMol does not support Alchemy format the last time I tried it.

I once upgraded to v2.7.x series, but noticed immediately a problem with visualization of Alchemy files from 3DNA. Initially, I thought it was a bug in 3DNA introduced in the source code somewhere. After some late night debugging, I re-tried to display the Alchemy in RasMol v2.6.4, and found that the fault was due to v2.7.x! It is actually quite easy to verify this v2.7.x bug: simply load a PDB structure, write to alchemy, then re-load the alchemy file, and you will see it. I reported this bug to the maintainer of v2.7.x, but was told that the bug was in v2.6.4. That's the reason I have stayed with v2.6.4 thereafter, and made the point that v2.7.x does not work with 3DNA generated schematic. As my programming experience grows, I realize more clearly the subtlety of a software product, and would normally trust the original author unless proved otherwise. It is far too easy to introduce a bug somewhere while fixing/improving other parts, unless you know every detail and/or have a through regression test suite for quality control.

In visualizing 3DNA generated Alchemy files with RasMol v2.6.4, it is important to remember to use command line options "-alchemy -noconnect". While the former option is obvious, the later one is to ensure that RasMol respects the connections explicitly specified by 3DNA generated Alchemy files, which are, after all, schematic, not real atomic structures. Sometimes (not always), this option would make a difference.

Saturday, June 13, 2009

blocview: a simple, effective visualization tool for nucleic acid structures

The 'blocview' Perl script in 3DNA was designed as a simple tool that nevertheless illustrates key features of nucleic acid structures effectively. Specifically, the 'bloc' part of the name means 'block', i.e., the base rectangular block in the Calladine-Drew style to show clearly the size (larger purine vs. smaller pyrimidine), identity (by color: red for A, yellow for C, green for G, and blue for T), and the groove (with minor groove edge filled in black); and the 'view' part means the most extended view (as defined by the principal axes of inertia). 'blocview' calls several utility programs of 3DNA and MolScript (for protein ribbons and nucleic acid backbone rods) to prepare the scenes, and then uses Raster3D (more specifically 'render') or PyMol (as of 3DNA v2.0) to generate a PNG image.

It has been my pleasure to notice that 'blocview' generated images are used in the NDB for each and every structure. See for example, an arbitrarily picked link on X-ray Drug-DNA Complexes. A nice thing about these images is that they can be generated automatically: using PDB id 1z8v as an example, the following command
blocview -i=1z8v.png 1z8v.pdb
would produce the image (named 1z8v.png above) shown in the right. In this representation, one can see clearly that there are two unpaired Gs (green) at each of the 5' end of the two DNA chains (red and yellow rods), and a drug molecule (ball-and-stick) binds in the minor groove (black edge of the rectangular blocks). Moreover, the first C-G pair has pronounced negative propeller twist.

It was a nice surprise when I found (several years ago) that such simple images have also be adopted by the PDB, prominently at the summary page, for every nucleic acid containing structure. There are so many sophisticated molecular graphics tools out there, why 'blocview'? After all, it is just a small utility tool in the 3DNA suite of programs, and I have never written a paper on it. Clearly, 'blocview' fills a niche.

On the other hand, given the effectiveness of this simple representation and its adoption by the NDB and PDB, I am surprised to found that 'blocview' images are not used that much in the literature. Maybe the situation will change as 3DNA gets more widely used, and when people pay more attention to its versatile functionalities: certainly, there are more handy features in 3DNA!

More blocview command-line options are available to suit some other common needs: just type 'blocview -h' for details. Furthermore, many aspects of the images can be fine-tuned by configuration files. Astute viewers could have noticed subtle differences of the 1z8v images shown here (above right), and in the NDB and PDB websites. Since 'blocview' is written in Perl script, users can easily check the implementation details to figure out exactly how it works.

Friday, June 12, 2009

Review of scientific manuscripts

While reading an editorial of Maddox [Nature. 1995 Dec 7; 378(6557):521-3], I found the following comment quite interesting:

It's a good joke (which I have often used) that Watson and Crick's paper on the structure of DNA could not be published now. If is only necessary to imagine what people would say if it reached them in the mail: "It's all model-building, just speculation, and such data they have are not theirs but Rosalind Franklin's!" Some would complain that the sentence beginning, "It has not escaped our attention..." is certainly unsubstantiated, and must be an attempt to claim credit for developments in genetics that lies years ahead.

Maddox continued to imagine what could have happened behind the scene that led to the publication of Watson and Crick's paper, together with Franklin's crucial experimental work illustrating the helical signature of crystalline DNA.

In some sense (to my understanding), Watson and Crick were lucky, not only in that they solved the puzzle of DNA structure elegantly by connecting the various experimental pieces together, but also they worked in the Cavendish Laboratory at Cambridge University and had the support from Sir Lawrence Bragg. One could easily imagine what would have happened if they did this work in a not so famous lab, without support from a leading, powerful figure.

In a scientific field (e.g. RNA), what appears to be a common sense, a decade old issue, could have something significant when view from a novel perspective. Such unconventional results, however, are often hard to get across the review process, especially for new comers to a field and without backing of big names. In principle, the reviewers should be fair to the author, being specific with his/her comments, especially in case of negative ones. That would be far more convincing to the authors than some vague, general remarks. As a general rule, I won't accept to review a manuscript for a journal unless I have time to read it thorough carefully, and the expertise to understand the details, in order to provide some concrete comments.

In this regard, it is worth reading the essay by Keith Manchester titled "Historical Opinion: Erwin Chargaff and his 'rules' for the base composition of DNA: why did he fail to see the possibility of complementarity?" [Trends Biochem Sci. 2008 Feb;33(2):65-70]. For completeness, the abstract is quoted below:
Erwin Chargaff was one of the more interesting and colourful figures of the historic decade that heralded the proposal of the double helical structure of DNA by Watson and Crick in 1953. In describing Chargaff's important contribution to the study of DNA, particularly its base composition, this article seeks to suggest why, despite his substantial achievements, he failed to anticipate some of the key features of the Watson-Crick model, particularly complementarity between bases--a failure that left him deeply embittered for the rest of his life.