Pages

LocusZoom: Plot regional association results from GWAS

Update Friday, May 14, 2010: See this newer post on LocusZoom.



If you caught Cristen Willer's seminar here a few weeks ago you saw several beautiful figures in the style of a manhattan plot, but zoomed in around a region of interest, with several other useful information overlays.

Click to take a look at this plot below, showing the APOE region for an Alzheimer's Disease GWAS:

It's a simple plot of the -log10(p-values) for SNPs in a given region, but it also shows:

1. LD information (based on HapMap) shown by color-coded points (not much LD here).
2. Recombination rates (the blue line running through the plot). Peaks are hotspots.
3. Spatial orientation of the SNPs you plotted (running across the top)
3. Genes! The overlay along the bottom shows UCSC genes in the region.

You can very easily take a PLINK output file (or any other format) and make an image like this for your data for any SNP, gene, or region of interest using a tool Cristen and others at Michigan developed called LocusZoom.  LocusZoom is written in R with a Python wrapper that works from an easy to use web interface.

All the program needs is a list of SNP names and their associated P-values. If you're using PLINK, your *.assoc or *.qassoc files have this information, but first you'll have to run a quick command to format them. Run this command I discussed in a previous post to convert your PLINK output into a comma delimited CSV file (PLINK's default is irregular whitespace delimited):

cat plink.assoc | sed -r 's/^\s+//g' | sed -r 's/\s+/,/g' > plink.assoc.csv

Next, you'll want to compress this file so that it doesn't take forever to upload.

gzip plink.assoc.csv

Now, upload your new file (plink.assoc.csv.gz) on the LocusZoom website.  Tell it that your p-value column is named "P" and your marker column is named "SNP" (or whatever they're called if you're not using PLINK). Change the delimiter type to "comma", then put in a region of interest. I chose APOE, but you could also use a SNP name (include the "rs" before the number). Now hit "Plot your Data," and it should take about a minute.

There are some other options below, but I've had bad luck using any of them. For instance, I can never get it to output a PNG properly - only PDF works the last time I tried it. I also could not successfully make a plot if I turn off the recombination rate overlay. I know this is a very early version, but hopefully they'll clean up some of the code and document some of its features very soon. I could see this being a very useful tool, especially once it's available for download for local use. (Update: some of these bugs have been fixed. See this newer post on LocusZoom).

LocusZoom: Plot regional association results from GWAS

Papers from Feb 8 2010 Journal Club

Here are the papers we talked about in yesterday's Journal Club:

PLoS Biol. 2010 Jan 26;8(1):e1000294.
Rare variants create synthetic genome-wide associations. See my previous coverage of this paper and the comments.
Dickson SP, Wang K, Krantz I, Hakonarson H, Goldstein DB.

PLoS Genet. 2010 Jan;6(1):e1000798. Epub 2010 Jan 8.
Modeling of environmental effects in genome-wide association studies identifies SLC2A2 and HP as novel loci influencing serum cholesterol levels.
Igl W, Johansson A, Wilson JF, Wild SH, Polasek O, Hayward C, Vitart V, Hastie N, Rudan P, Gnewuch C, Schmitz G, Meitinger T, Pramstaller PP, Hicks AA, Oostra BA, van Duijn CM, Rudan I, Wright A, Campbell H, Gyllensten U; EUROSPAN Consortium.

Science. 2010 Jan 7. [Epub ahead of print]
A Composite of Multiple Signals Distinguishes Causal Variants in Regions of Positive Selection.
Grossman SR, Shylakhter I, Karlsson EK, Byrne EH, Morales S, Frieden G, Hostetter E, Angelino E, Garber M, Zuk O, Lander ES, Schaffner SF, Sabeti PC.

Regression Modeling Strategies Course by Frank Harrell

Frank Harrell is teaching his 3-session short course on regression modeling strategies using R here at Vanderbilt next month. Frank is a professor and chair of the Vanderbilt Biostatistics Department, and the author of several massively popular R libraries, including Design, rms, and the indispensable Hmisc.  He has also written a book, covering many topics related to regression modeling (Amazon, $98). You can find more information about the course and registration at this link.

The course consists of three half days:
Wednesday, March 31 (8:00AM-12:00PM)
Thursday, April 1 (8:00AM-12:00PM)
Friday, April 2 (8:00AM-4:00PM)

Registration Fees:
VU and MMC Students and Post-docs $50
VU and MMC Faculty and Staff $200
Other Students $200
Other Members of Non-Profit Institutions $400
Members of For-Profit Institutions $600
No charge to Department of Biostatistics faculty/staff

Regression Modeling Strategies - 2010 Short Course

Computational Genomics Journal Club, Monday Feb 8

We're restarting the PCG Journal Club again Monday February 8 at 4pm in the CHGR conference room. Most of you who usually attend are familiar with the format, but if not, bring any papers you've read recently and give a brief (i.e. 2 minute) overview of the paper and why you thought it was interesting. No slides allowed. We're also going to discuss the previously mentioned Rare Variants Create Synthetic Associations paper recently published in PLoS Biology, which has generated a bit of controversy lately.  Oh yeah, and we'll have beer and snacks.  In case you miss it, Julia will be posting links to all the papers we discuss under the Journal Club tag. You can also view any of the papers from previous journal clubs under that link.

Research Symposium at HudsonAlpha Institute for Biotechnology

If you were here for any of the talks Rick Myers has given here at Vanderbilt over the last few years you'll remember all the interesting biomedical research going on at his company, HudsonAlpha.  Their spring symposium is March 30, 8am-6pm, at the HudsonAlpha institute in Huntsville, AL. It's FREE, and poster sessions are open to all students and postdocs.  Participating institutions include HudsonAlpha, UAB, Auburn University, Emory University, The University of Alabama, and of course Vanderbilt University. See the link below for registration.

HudsonAlpha Institute for Biotechnology: Spring Symposium

Qualifying exam study guide phase I

Vanderbilt 2nd year grad students: Here is the study guide I made for studying for my general knowledge phase I qualifying exam. I'd recommend making your own, but this may help you with a place to start. You can download it at the link below.

Update January 25, 2013: I've uploaded the link to Figshare for a more permanent home.

http://dx.doi.org/10.6084/m9.figshare.154339

GRAIL: Gene Relationships Across Implicated Loci


If you caught Soumya Raychaudhuri's seminar last week you heard a lot about the tool he developed at the broad called GRAIL - Gene Relationships Across Implicated Loci. You've got GWAS results and now you want to prioritize SNPs to follow up in replication or functional studies. Of course you're going to take your stellar hits at p<10e-8, but what about that fuzzy region between 10e-4 and 10e-8? Here's where a tool like GRAIL may come in handy.

In essence, you feed GRAIL a list of SNPs and it maps these SNPs to gene regions using LD. It then uses a simple text-mining algorithm to ascertain the degree of connectivity among the associated genes by looking at the similarity of vectors of words pulled from PubMed abstracts which mention your gene of interest. In their most recent paper they took a list of 370 GWAS hits, and narrowed this down to a list of 22 candidate SNPs to follow up.  And it turns out these SNPs replicated in an independent set at a much higher frequency than random SNPs from the subset of 370. In his talk, Soumya offered convincing evidence that using the results from GRAIL you have a much better shot at replicating associations than if you just looked at the p-value rankings alone.  After the talk he did mention that this approach has had mixed success depending on the phenotype. Here's the original GRAIL paper, and the OpenHelix blog has a nice 5-minute video screencast demonstration of GRAIL where they take a few SNPs from the GWAS catalog and run them through GRAIL.

GRAIL is a free web application (beta) available at the Broad's website below.

GRAIL: Gene Relationships Across Implicated Loci