Pages

I'm Hiring!

I direct the Bioinformatics Core at the University of Virginia, and I'm hiring. Visit this link on the UVA Jobs website for more information. Here's the description:
The University of Virginia Bioinformatics Core is seeking a full-time position as a bioinformatics analyst. The analyst will work with other core staff on grant-funded and chargeback-based projects to manage and analyze large-scale datasets produced by next-generation sequencing. The analyst will identify opportunities and implement solutions for managing, visualizing, analyzing, and interpreting genomic data, including studies of gene expression (RNA-seq and microarrays), pathway analysis, protein-DNA binding (e.g. ChIP-seq), DNA methylation, and DNA variation, using Affymetrix, Illumina, Nimblegen, Agilent, Roche 454, Ion Torrent, and other high-throughput platforms in both human and model organisms. The analyst will work closely with the core director to assist in experimental design and provide expert consultation, technical, and scientific support for UVA investigators, and assist in outreach and training activities. The analyst will organize large-scale sequence data sets, manipulate and format data with perl, python, or other scripting languages, use established software to assess quality and analyze data, schedule and run jobs on a high-performance computing cluster, use Unix or a scripting language to extract meaningful results from output, use software or genome browsers for visualization, and use established databases and techniques for annotating genetic variants and results from expression/DNA-binding experiments. The successful candidate will have a demonstrated ability to translate biological questions into technical designs, and to identify, prioritize, and execute bioinformatics tasks to meet project goals and deadlines. An M.S. in Bioinformatics, Genomics, Biostatistics, or a related field is required for this position.  
I'm Hiring - Bioinformatics Analyst in the UVA Bioinformatics Core

Your Publications (with PMCID) as a PubMed Query

I'm updating my CV and biosketch for a few grant applications, and for some time now, NIH has required you to include the PubMed Central ID for each article you publish that arose from NIH support. I only have a dozen or so papers indexed in PubMed, but I still wanted a way to do this automatically. If you have scores of publications, looking up all the PMCIDs could easily become a hassle.

First, create an account at My NCBI. Under your bibliography, click "Manage My Bibliography." Then click "Add citation," then in the new window that comes up, select "Citation from PubMed" and hit the "Go To PubMed" button.

Now the trick here is constructing a PubMed query that will get your publications only. There are lots of Stephen D. Turner's out there, so I had to get creative. This query construction tip comes to me by way of my colleague here at UVA, Aaron Mackey:

For many people, simple PubMed author searches suffice, e.g. "Pearson WR[Author]". For some, such name-based searches get it mostly right, but may include a few spurious false hits. For these cases, it's easy enough to exclude those false hits explicitly (e.g. "Mackey AJ"[Author] NOT 9850730[PMID] NOT 10730495[PMID] gets rid of the two AJ Mackey publications that are not, in fact, mine). For others, simple author searches do not suffice at all, but usually adding an institution and/or departmental affiliation does narrow the results sufficiently (e.g. for Jeff Smith, Biochemistry: "Smith JS"[au] AND "University of Virginia"[Affiliation] AND "Biochemistry"[Affiliation] identifies the 16 articles for which Jeff Smith is the senior author; Jeff could also add a few collaborative publications by adding those pubmed IDs to the search, i.e. adding "OR 17482543[PMID]" to the end of his query.

When I did this for myself, I searched by author, AND (any of my institutional affiliations separated by OR's), but NOT (any of the PMIDs that were not mine, separated by OR's). Apparently there was once another Stephen D. Turner at UVA in the department of Urology. Here are the results returned by my unique query:

"Turner SD"[Au] AND ("James Madison"[Affiliation] OR Vanderbilt[Affiliation] OR Hawaii[Affiliation] OR "University of Virginia"[Affiliation]) NOT (11514333[PMID] OR 11058553[PMID])

The final step is clicking the "Send to" link at the top right, and sending the results of your query to My Bibliography.


Now, when you are back at My NCBI, you should see a list of all your publications, complete with both the PMID and PMCID, ready to go in your biosketch.


You can then export this bibliography as text, or simply copy/paste. Finally, you have the option of making your bibliography public (example).



Webinar: Genomic Networks - Resolving Biomarkers from a Cloud of Data

Kevin White from the University of Chicago will be giving a special guest lecture at NCI next week on systems biology approaches to mine genomics data for biomarkers and therapeutic targets. The lecture will be available online as a videocast.

Title: Genomic Networks in Development and Cancer: Resolving Biomarkers and Therapeutic Targets from a Cloud of Data

Speaker: Kevin White, University of Chicago

When: Tuesday February 14, 2012, 1:00pm EST

Summary

Systems level approaches to construct abstract molecular networks can lead to predictions about genetic and biochemical functions in cells, organisms and in disease states. I will show examples of this approach from work in my laboratory. In one example we used an integrated experimental and computational approach to construct a large scale functional network in Drosophila melanogaster built around key transcription factors involved in the process of embryonic segmentation. Our network model is based on a combination of gene expression, transcription factor DNA binding site mapping, automated literature mining and protein-protein interaction mapping. We provide a strategy for reducing the dimensionality of the massive networks that result from such integrated whole genome analyses. 

Using results from one factor in particular, we demonstrated that our approach can rapidly translate a finding in a model organism to the development of a therapeutic target in kidney cancer. In another example, we built a large scale network based on gene expression and genome-wide ChIP results for 40 transcription factors, including two dozen Nuclear Receptor (NR) class proteins. Using this NR network we identified novel prognostic signatures for breast cancer survival and recurrence, as well as new therapeutic leads. 

Finally, if time permits I will talk about how we are mining The Cancer Genome Atlas along with data from the Chicago Cancer Genomes Project using the Bionimbus Cloud in order to identify new tumor suppressors and panels of genetic markers capable of classifying cancer subtypes that correspond to patient outcome.

Hadley Wickham: ggplot2 Webinar (Today!)


Title: A Backstage Tour of ggplot2 with Hadley Wickham
Date: Wednesday, February 8, 2012
Time: 11:00AM - 12:00PM Pacific
Presenter: Hadley Wickham, Professor of Statistics, Rice University

Register here.

I used ggplot2 extensively a few years ago, but reverted back to base graphics when ggplot2 was too slow for a project I was working on. But ggplot2 and plyr have improved much in the last few years, and I'm starting to pick it back up again. This webinar will give an overview of ggplot2, a preview of some of ggplot2's forthcoming new features, and will discuss ggplot2's internals and development over the last few years and how ggplot2 development is becoming easier.

I received an email yesterday saying that the registration list is over 1000 long, so it's a good idea to sign into the webinar early to make sure you get a spot. Hit the link below to register and you'll get a link to the webinar.

A Backstage Tour of ggplot2 with Hadley Wickham

Joint Techs Netcast: Enhancing Infrastructure Support for Data Intensive Science

The winter Joint Techs meeting is next week in Baton Rouge. I'm not going, but I plan on participating via a netcast to see what's going on. Jim Bottum, Clemson's CIO, is moderating an entire day devoted to the topic Enhancing Infrastructure Support for Data Intensive Science. Of particular interest to me are the talks from 9:30-11am Tuesday January 24 from researchers and those supporting climatology, genomics, and the XSEDE projects. The afternoon of January 24 has some talks from academic and government labs who've successfully deployed methods to enhance their infrastructure support for data intensive science. Check out the full agenda for the day here. These sessions sound particularly relevant for those researching and supporting large-scale genomics and bioinformatics projects.

Joint Techs Meeting Netcast

Annotating limma Results with Gene Names for Affy Microarrays

Lately I've been using the limma package often for analyzing microarray data. When I read in Affy CEL files using ReadAffy(), the resulting ExpressionSet won't contain any featureData annotation. Consequentially, when I run topTable to get a list of differentially expressed genes, there's no annotation information other than the Affymetrix probeset IDs or transcript cluster IDs. There are other ways of annotating these results (INNER JOIN to a MySQL database, biomaRt, etc), but I would like to have the output from topTable already annotated with gene information. Ideally, I could annotate each probeset ID with a gene symbol, gene name, Ensembl ID, and have that Ensembl ID hyperlink out to the Ensembl genome browser. With some help from Gordon Smyth on the Bioconductor Mailing list, I found that annotating the ExpressionSet object results in the output from topTable also being annotated.

The results from topTable are pretty uninformative without annotation:


After annotation:


You can generate an HTML file with clickable links to the Ensembl Genome Browser for each gene:


Here's the R code to do it:

New Year's Resolution: Learn How to Code

Farhad Manjoo at Slate has a good article on why you need to learn how to program. Chances are, if you're reading this post here you're already fairly adept at some form of programming. But if you're not, you should give it some serious thought.

Gina Trapani, former editor of tech blog Lifehacker, is quoted in the article:
“Learning to code demystifies tech in a way that empowers and enlightens. When you start coding you realize that every digital tool you have ever used involved lines of code just like the ones you're writing, and that if you want to make an existing app better, you can do just that with the same foreach and if-then statements every coder has ever used.”
Farhad makes the point that programming is important even in traditionally non-computational fields: if you were a travel agent in the 90's and knew how to code, not only would you have been able to see the approaching inevitable collapse of your profession, but perhaps you would have been able to get in early on the dot-com travel industry boom.

Q&A sites for biologists are littered with questions from researchers asking for non-technical, code-free ways of doing a particular analysis. Your friendly bioinformatics or computational biology neighbor can often point to a resource or design a solution that can get you 90% of the way, but usually won't grok the biological problem as truly as you do. By learning even the smallest bit of programming, you can at least be equipped with the knowledge of what is programmatically possible, and collaborations with your bioinformatician can be more fruitful. As every field of biological research becomes more computational in nature, learning how to code is becoming more important than ever.

Where to start

Getting started really isn't that difficult. Grab a good text editor like Notepad++ for windows, TextMate or Macvim for Mac, or vim for Linux/Unix. What language should you start with? This can be a subject of intense debate, but in reality, it doesn't matter - just pick something that's relevant to what you're doing. If you know Perl or Java, you can pick up the basics of Ruby or C++ in a weekend. I started with Perl (using the Llama book), but for scientific computing and basic scripting/automation, I would recommend learning Python instead. While Perl lets you get away with sloppy coding, terse shortcuts, with the motto of "there's more than one way to do it," Python forces you to keep your code tidy, and has a model that there's probably one best way to do something, and that's the way you should use. Python has a huge following in the scientific community - chances are you'll find plenty of useful functionality in the BioPython and SciPy modules. I learned Python in an afternoon through watching videos and doing exercises in Google's Python Class, and the free book Dive Into Python is a great reference. If you're on Windows, you can get Python from ActiveState; if you're on Mac or Linux, you already have Python.

The Slate article also points to Code Year - a site that will send you interactive coding projects once a week throughout 2012 starting January 9. Code Year is from the creators of Code Academy - a site with a series of fun, interactive JavaScript tutorials. Lifehacker has a 5-part "Night School" series on the basics of programming. Once you have some basic programming chops, take a look at Stanford's free machine learningartificial intelligence, and Natural Language Processing classes to hone your scientific computing skills. Need a challenge? Try the Python Challenge for fun puzzles to hone your Python skills, or check out Project Euler if you want to tackle more math-oriented programming challenges with any language. The point is - there is no lack of free resources to help you get started or get better at programming.

Slate - You Need to Learn How to Program