Tuesday, 8 November 2011

[Bio-bwa-help] BWA 0.6.0 works with a genome > 4 GB

Pengcheng has posted mapping statistics for
maq version 0.7.1 and bwa 0.6.0
to align 1M PE100 reads to 7.3g reference genome

The run was the first working example of real life data for mapping > 4 gb long ref genomes

Sunday, 6 November 2011

SlideShare The Next, Next Generation of Sequencing - From Semiconductor to Single Molecule

'The Next, Next Generation of Sequencing - From Semiconductor to Single Molecule' to SlideShare. 

Take This Home• There are many challenges before we even get to picking a platform – Technical Expertise – Standards in Prep and Analysis With Great NGS Power Comes Great Responsibility

Resequencing ConclusionUsing appropriate aligners and variant callers we show bothplatforms have high accuracy, each with strengths and weaknesses…


Questions Twitter: @Bioinfojjohnson@edgebio.com

Multiplexing Target enrichment of exomes for cost savings in NGS and other papers


1.Multiplexed array-based and in-solution genomic enrichment for flexible and cost-effective targeted next-generation sequencing.
Harakalova M, Mokry M, Hrdlickova B, Renkens I, Duran K, van Roekel H, Lansu N, van Roosmalen M, de Bruijn E, Nijman IJ, Kloosterman WP, Cuppen E.
Nat Protoc. 2011 Nov 3;6(12):1870-86. doi: 10.1038/nprot.2011.396.
PMID: 22051800 [PubMed - in process]

Abstract

The unprecedented increase in the throughput of DNA sequencing driven by next-generation technologies now allows efficient analysis of the complete protein-coding regions of genomes (exomes) for multiple samples in a single sequencing run. However, sample preparation and targeted enrichment of multiple samples has become a rate-limiting and costly step in high-throughput genetic analysis. Here we present an efficient protocol for parallel library preparation and targeted enrichment of pooled multiplexed bar-coded samples. The procedure is compatible with microarray-based and solution-based capture approaches. The high flexibility of this method allows multiplexing of 3-5 samples for whole-exome experiments, 20 samples for targeted footprints of 5 Mb and 96 samples for targeted footprints of 0.4 Mb. From library preparation to post-enrichment amplification, including hybridization time, the protocol takes 5-6 d for array-based enrichment and 3-4 d for solution-based enrichment. Our method provides a cost-effective approach for a broad range of applications, including targeted resequencing of large sample collections (e.g., follow-up genome-wide association studies), and whole-exome or custom mini-genome sequencing projects. This protocol gives details for a single-tube procedure, but scaling to a manual or automated 96-well plate format is possible and discussed.

2.A computational index derived from whole-genome copy number analysis is a novel tool for prognosis in early stage lung squamous cell carcinoma.
Belvedere O, Berri S, Chalkley R, Conway C, Barbone F, Pisa F, Maclennan K, Daly C, Alsop M, Morgan J, Menis J, Tcherveniakov P, Papagiannopoulos K, Rabbitts P, Wood HM.
Genomics. 2011 Oct 25. [Epub ahead of print]
PMID: 22050995 [PubMed - as supplied by publisher]

Abstract

Squamous cell carcinoma of the lung is remarkable for the extent to which the same chromosomal abnormalities are detected in individual tumours. We have used next generation sequencing at low coverage to produce high resolution copy number karyograms of a series of 89 non-small cell lung tumours specifically of the squamous cell subtype. Because this methodology is able to create karyograms from formalin-fixed paraffin-embedded material, we were able to use archival stored samples for which survival data were available and correlate frequently occurring copy number changes with disease outcome. No single region of genomic change showed significant correlation with survival. However, adopting a whole-genome approach, we devised an algorithm that relates to total genomic damage, specifically the relative ratios of copy number states across the genome. This algorithm generated a novel index, which is an independent prognostic indicator in early stage squamous cell carcinoma of the lung.

3.The utility of gene expression in blood cells for diagnosing neuropsychiatric disorders.
Woelk CH, Singhania A, Pérez-Santiago J, Glatt SJ, Tsuang MT.
Int Rev Neurobiol. 2011;101:41-63.
PMID: 22050848 [PubMed - in process]

Abstract

Objective diagnostic tools are required for neuropsychiatric disorders. Gene expression in blood cells may provide such a tool and has already been used to construct classifiers capable of diagnosing many human diseases. This chapter discusses the use of microarray gene expression data to construct diagnostic classifiers for neuropsychiatric disorders. The potential pitfalls of microarray gene expression analysis and the experimental design and methods suitable for classifier construction are described in detail. A review of studies that have analyzed gene expression in blood cells from patients with neuropsychiatric disorders is presented with an emphasis on the feasibility of generating a diagnostic classifier for schizophrenia. Finally, the future directions of the field are discussed with respect to using blood gene expression to tailor antipsychotic medications to individual patients, applying microRNA expression for diagnostic purposes, as well as the implications of next-generation sequencing technologies for gene expression analysis.


Thursday, 3 November 2011

password protected RSS feed

I am curious what good is a RSS feed if I have to key in a password to see it.
That effectively disables all the RSS feed readers I use from accessing it.

Ion Community Home > Torrent Dev > Content
RSS feed of this list
http://lifetech-it.hosted.jivesoftware.com/community/feeds/documents?containerType=14&container=2005

Available references for Ion Torrent Server download

There are reference files (plain fasta files zipped up) available on http://updates.iontorrent.com/reference/ to use with your TS.

I would think it trivial for someone to write a shell script to curl / wget the reference and manually (or commandline ) create the reference index in the proper location.

But as of now, I can't find how to
1) manually create the reference index so that it shows in the realignment plugin as well.
2) reference files (there's only hg19 on the website)

2 isn't a big problem since the fasta requirements are simple enough.
no blank lines
no non IUPAC sequences

it's something u can include in the bash script with sed

Ok shall bother my PGM bioinfo FAS soon.

otherwise, I am stuck with the routine of downloading the sequences and zipping it and uploading on a win box (I do not have sliverlight on my Ubuntu)

employment ads do you have mad ngs / genomic analysis skills

EdgeBio Hiring #bioinformatics engineers (bit.ly/u4WXQR) & scientists (bit.ly/ujjBa7). Gotta have mad #ngs / #genomic analysis skills.

I like looking at recruitment ads :)
often they list skills that gets added to my todo /to learn list to keep current with research trends

not a plug for them but do check out the list of skills
$$ denotes shared skills Hmmm MATLAB  ... that's something I have no experience with, shall endeavour to find out it's relevance to NGS ... guessing it's a systems biology requirement ..

I am curious as to 'Gnome scale data analysis techniques'

anyone in the know to explain that to me ?

We are looking for a Bioinformatics Engineer to work in close collaboration with our informatics research and laboratory staff on site to provide sequencing core and data analysis infrastructure built atop a cutting edge data center and computation facility.

Required Skills:
• Proven experience with large scale data management, qc, and visualization
• Ability to leverage innovative 3rd party tools in published research and the public domain
• Gnome scale data analysis techniques (sequence similarity, assembly, annotation, etc)
• Comfortable in a Unix/Linux environment (you need to "get" dorky Linux jokes) $$
• Relational Database (Postgres, MySQL)
• Scripting languages (Perl, Python, shell) and associated tools
• Next generation sequencing (NGS) tools and techniques
• Statistics, mathematical analysis, R, MATLAB $$
• Ability to develop and implement data analysis algorithms

With Edge BioServ, you will find those "dare to be great" challenges and opportunities. We offer competitive salaries, excellent benefits, and a rewarding environment.


For Bioinformatics Scientists

Do you feel you have more to contribute to bioinformatics research? Are you looking for a flexible, start-up environment where you can take responsibility and apply novel computational methods to help analyze terabytes of high throughput sequence data? If so we are looking for talented, energetic professionals with proven experience in:

• In depth knowledge of bioinformatics algorithms and tools (alignment, expression profiling, etc)
• Next generation sequencing (NGS) tools and techniques
• High performance computing clusters, cloud computing
• Statistics, Mathematical Analysis, R, MATLAB $$
• Strong scientific and project management skills
• Comfortable in a Unix/Linux environment (you need to "get" dorky Linux jokes) $$



#disclaimer- not affliated with EdgeBio though it sounds like a place I might enjoy working in ;p

Video Tip of the Week: MizBee Synteny Browser


Video Tip of the Week: MizBee Synteny Browser
http://blog.openhelix.eu/?p=9634

mproved visualization strategies that offer increased flexibility about the features being viewed, and that scale up to the current data deluge, are both going to be necessary. So when I spotted this tweet about a group doing visualization I was intrigued:

RT @mcmahanl: RT @chlalanne: Check out Miriah Meyer's work in #dataviz for #bioinformatics, http://t.co/xo5Ei7K (via @FILWD)

The video tip this week is produced by the Meyer team, and you can click on that image to go to the page where you can see it.

Datanami, Woe be me