Friday, July 16, 2010

PDB fingerprints make salient features of PDB entries obvious at a glance

In the past couple of days, I have followed with great interest the thread "PDBprints - salient, at-a-glance info about PDB entries" in CCP4BB. The idea behind PDBprints (short for PDB finger prints) is very simple: now with over 66 thousands entries, no one could possibly be intimately familiar with many of the structures in PDB. Yet "it would be very useful to have a way of obtaining at-a-glance information about the salient features of one or more PDB entries." In this scheme, key features (e.g., experimental method, species, protein/NA/ligand components, published or not) are represented by stylized icons called PDBlogos; a set of which organized in a specific order gives PDBprints.

The initial announcement, posted on Thursday by Gerard Kleywegt (the current Director of PDBe), has been well-received by the community. I am very impressed by the quality of the (over a dozen) feedbacks on icons coloring for presence/absence of a feature, and the proper icons for species etc. On the other hand, I am surprised by the quick and concrete response Gerard posted today, addressing the various issues.

Over the years, I have taken PDB/NDB mostly as data repositories; I regularly download (update) nucleic-acid containing structures for local analyses. Among the three wwPDB sites (RCSB, PDBe and PDBj), I've only visited the RCSB site on a regular basis. Now I find more features are available from PDBe, which I will browse more in the future.

Saturday, July 3, 2010

In 3DNA, analyze and rebuild are two sides of the same coin

Recently, I noticed the following two papers:
They both cite 3DNA, where the combined usage of its analyze/rebuild components has played a significant role.

When I first had access to the CEHS scheme, I was immediately attracted by its mathematical rigor which allows for complete reversibility in DNA structural analysis and model rebuilding. Historically, CEHS was the initial seed of the SCHNAaP/SCHNArP programs, and analyze/rebuild in 3DNA were directly derived from them.

Frequently, I think of analyze/rebuild as two sides of the same coin: starting from a nucleic acid structure (e.g., a DNA double helix), the analyze program gives a set of base-pair (propeller, buckle etc) and step (roll, slide etc) parameters. The structural parameters can be used to rebuild the structure, which is virtually identical in base geometry (i.e., without taking consideration of the sugar-phosphate backbone) to the original structure. Conversely, analyzing the rebuilt structure again gives exactly the same set of structural parameters describing the relative base geometry.

Reversibility is essentially a simple concept, nevertheless a very powerful one. Over the years, I am glad to see more people are taking advantage of this 3DNA unique feature in addressing real-world problems related to nucleic acid structures. I am confident to see more such applications.

Sunday, June 27, 2010

Get all torsion angles in a nucleic acid structure

In the field of nucleic acid structure analysis, a commonly calculated set of parameters is the torsion angles. Specifically, they include the main chain and chi (χ) torsion angles defined as follows:
  • α: O3'(i-1)-P-O5'-C5'
  • β: P-O5'-C5'-C4'
  • γ: O5'-C5'-C4'-C3'
  • δ: C5'-C4'-C3'-O3'
  • ε: C4'-C3'-O3'-P(i+1)
  • ζ: C3'-O3'-P(i+1)-O5'(i+1)
  • χ: for pyrimidines (Y, i.e., T, U, C): O4'-C1'-N1-C2; for purines (R, i.e., A, G): O4'-C1'-N9-C4
A related set of parameters characterizes the sugar conformation:
  • ν0: C4'-O4'-C1'-C2'
  • ν1: O4'-C1'-C2'-C3'
  • ν2: C1'-C2'-C3'-C4'
  • ν3: C2'-C3'-C4'-O4'
  • ν4: C3'-C4'-O4'-C1'
  • tm: amplitude of pseudorotation of the sugar ring
  • P: phase angle of pseudorotation of the sugar ring
These torsion angles are clearly defined and are readily available from various informatics programs. Not surprisingly, 3DNA also provides a complete set of DNA/RNA backbone torsions, calculated robustly and efficiently. The key is the "-s" option of find_pair" program, which is nevertheless little used, mostly because it is not the default.

Using the Haloarcula marismortui 50S large ribosomal subunit as an example (1jj2), the output file 1jj2.outs from the following command contains all the above mentioned main chain and sugar conformational parameters:
find_pair -s 1jj2.pdb stdout | analyze
Please see the Jena "Nucleic acid backbone parameters" website for a diagram (based on Saenger's book) defining the various backbone torsion angles.

Sunday, June 20, 2010

Subscribing to mailing lists in daily digest mode

I am currently on quite a few mailing lists that are relevant to my research topics in a broad sense, including Jmol, pdb-l and CCP4. Over the years, I've found it very handy to keep informed of a field by following in its (major) mailing list. However, I could easily get lost with so many posts each day on some active lists. So subscribing to the daily digest mode, provided by many mailing list management software, has become a norm. Thus, I would receive only one (or a few) email(s) from a mailing list. I could then easily browse through the subject lines to decide if to read further of a specific topic.

Among the three lists mentioned above, Jmol is quite active with several core contributors, most impressively from Prof. Robert Hanson. The PDB mailing list is of unbelievably low volume, with less than a post per day. CCP4 bullet board is overall the most well-organized; I am surprised by the knowledge-base from the posts there – it is a great resource in structural biology.

AMBER mailing list and Computational Chemistry List (CCL) are the other two lists I subscribed to before. However, I could not get into the daily digest mode; soon my mail box was flooded by many unrelated posts so I had to un-subscribe from them.

Journal impact factor and individual researcher h-index

The June 17, 2010 issue of Nature (Vol. 465, No. 7300) has extensive discussions on assessing the impact (influence/significance) of a journal or an individual researcher using quantitative metrics. Not surprisingly, there is no consensus. Over the past 20 years, the field of bibliometrics (scientometrics) has seen a ten-fold explosion in publications. In fact, the title of the editorial is "Assessing assessment"; it boils down to whether we should use a quantitative metric, a combination of several such metrics, or which metrics to choose among the so many possibilities.

Among the various view points, I agree more with David Pendlebury, Citation Analyst of Thomson Reuters. "No one enjoys being measured — unless he or she comes out on top." "Importantly, publication-based metrics provide an objective counterweight ... to bias of many kinds." "there are dangers [to] put too much faith in them. A quantitative profile should always be used to foster discussion, rather than to end it."

Among the many currently available metrics, it seems fair to say:
  • impact factor (IF), introduced in 1963, is the most important one to measure the impact of a journal. Nowadays, it is common to see the IF of a journal at its home page. For example, currently Nucleic Acids Research has an IF of 6.878. However, as emphasized by Anthony van Raan, director of the Centre for Science and Technology Studies at Leiden University in the Netherlands: "You should never use the journal impact factor to evaluate research performance for an article or for an individual — that is a mortal sin." Indeed, "In 2005, 89% of Nature’s impact factor was
    generated by 25% of the articles."
  • h-index, introduced in 2005 by Hirsch, is currently the most influential metric to quantify the productivity and impact of an individual researcher. An h-index is "defined as the number of papers with citation number ≥h". It "gives an estimate of the importance, significance, and broad impact of a scientist’s cumulative research contributions." I found it very informative to read the original, 4-page long PNAS paper to understand clearly why h-index is defined that way and to be aware of some of its caveats.
  • number of citations is undoubtedly the most objective metric to measure the impact of a publication.
Overall, it is easy to come up with a quantitative measure; the point is what that number really means. Along the line, it is of crucial importance to know exactly how a metric is calculated. No metric could be perfect; however, a metric, defined transparently and applied consistently, is more objective and convincing than other means.

Friday, June 11, 2010

3DNA citations from mainland Chinese researchers

Over the past couple of weeks, I have been quite surprised to notice the following four papers, by researchers from mainland China, that cite 3DNA:
For more ten years, this is the first time (to the best of my knowledge) that 3DNA has been cited by scientists from China, and it so happens that four come in a row. I am so happy to see that structural biology is shaping up in China. Surely, I would expect more 3DNA-citing papers from China to come in the following years.

block_atom, a little used 3DNA utility script

Following my previous post, "How is a base-pair rectangular block defined in 3DNA?", I feel it is in order to mention a tiny Perl script named block_atom that seems to have been little used by the 3DNA user community.

The basic idea of block_atom is to get an ALCHEMY data file which combines both atomic and block representations of nucleic-acid-containing structure in PDB format. The ALCHEMY file can then be displayed interactively using RasMol (v2.6.4, by its original author) (or Jmol).

Again, the point is best illustrated with an example. Here I am using 355d, the famous Drew-Dickerson B-DNA dodecamer solved at a high resolution by Williams and colleagues. Let the input PDB file be 355d.pdb, then the output file is automatically named 355d.alc, which can be displayed using RasMol with the options '-alchemy -noconnect'.
block_atom 355d.pdb
rasmol -alchemy -noconnect 355d.alc
One view of RasMol rendered image is as below, which shows more clearly the minor-groove (colored blue) of the DNA duplex and the various intra-base-pair deformations (e.g., propeller and buckle):
I initially wrote this utility script mostly for verification purpose, i.e., to make sure that the base reference frames are defined properly. For those who want to know more about 3DNA, having a look of the script (which is tiny) and understanding how it works could be a good exercise.

Saturday, June 5, 2010

How is a base-pair rectangular block defined in 3DNA?

One of 3DNA's unique features is the base-pair (and base) rectangular blocks, as shown in the figure below. Since such a schematic was initialized by Calladine and Drew in their Understanding DNA book, I normally refer it as the Calladine-Drew style representation.


By default, a base-pair [BP, (a)] has a dimension of 10x4.5x0.5 Å; a purine [R, (b) left] 4.5x4.5x0.5 Å; a pyrimidine [Y, (b) right] 3x4.5x0.5 Å; and a mean base [M, (c)], which is exactly half of the base-pair, 5x4.5x0.5 Å.

The blocks are put into specific files: Block_BP.alc, Block_R.alc, and Block_Y.alc respectively. To use M for R and Y, one just needs to copy file Block_M.alc to overwrite Block_R.alc and Block_Y.alc in the current directory or the system directory ($X3DNA/config/). These blocks are used in the rebuilding and visualization components of 3DNA (see my blog post "blocview: a simple, effective visualization tool for nucleic acid structures").

The blocks are stored in the ALCHEMY format for easy specification of the nodes and edges, and since ALCHEMY is a format supported by RasMol (and Jmol). As an example, Block_BP.alc has the following content:
   12 ATOMS,    12 BONDS
1 N -2.2500 5.0000 0.2500
2 N -2.2500 0.5000 0.2500
3 N -2.2500 0.5000 -0.2500
4 N -2.2500 5.0000 -0.2500
5 C 2.2500 3.5000 0.2500
6 C 2.2500 0.5000 0.2500
7 C 2.2500 0.5000 -0.2500
8 C 2.2500 3.5000 -0.2500
9 C -2.2500 5.0000 0.2500
10 C -2.2500 0.5000 0.2500
11 C -2.2500 0.5000 -0.2500
12 C -2.2500 5.0000 -0.2500
1 1 2
2 2 3
3 3 4
4 4 1
5 5 6
6 6 7
7 7 8
8 5 8
9 9 5
10 10 6
11 11 7
12 12 8
Astute viewers may observe that nodes 1-4 are identified as N (nitrogen) and have exactly the same coordinates as 9-12 (C, carbon). This is a trick so RasMol can display the the minor groove edge in a different color than the other five sides of the rectangular, as shown in the following figure:

Note that the rectangular is preset in the standard base reference frame. Thus the nodes have y-coordinates of +5 Å and -5 Å along the long edge of the base pair, and x-coordinates of +2.25 Å and -2.25 Å along the short edge.

As an extra bonus of storing the blocks in external ALCHEMY text files, the dimensions of the blocks are be readily changed. For example, the depth of a block (z-coordinates) can be easily increased from 0.5 to 1.0 Å to make it thicker (see $X3DNA/config/Block_BP1.alc). Moreover, the blocks do not need to be rectangular either – see $X3DNA/config/block/Block_R_nr.alc for an example.

Friday, May 28, 2010

Naming conventions: RasMol, PyMOL and Jmol

Over the years, I've used the following three molecular graphics programs the most: RasMol, PyMOL and Jmol. It is interesting to note their different naming conventions: while their names all end with m-o-l (for molecule, I believe), the cases (used in convention) are different: Mol as in RasMol, MOL as in PyMOL, and mol as in Jmol.

For Jmol, there is actually an FAQ item "How do you write Jmol?" that reads "Capital J, lower case mol. Please do not write it any other way." For PyMOL, I am not aware of a specification on its name – I know that's the "official name" from its home page. Regarding RasMol, that's the most common way to name it, even though the title of its 1995 publication is "RASMOL: biomolecular graphics for all."

As for the prefix in the names, J in Jmol stands for Java; Py in PyMOL for Python; and Ras in RasMol could well be the initials of Roger A. Sayle, the original author of the program.