Tuesday, 18 May 2010

Book review:Programming Collective Intelligence

Programming Collective Intelligence: Building Smart Web 2.0 Applications by Toby Segaran

BGI to Sequence and Assemble 100 Vertebrates within Two Years for Genome 10K Project

Got to know of this news by www.genomeweb.com

BGI to sequence 100 vertebrate species for Genome 10K project. Andrea Anderson. May 14, 2010. GenomeWeb.

The http://www.genome10k.org/ project has lofty aims to capture ".. the genetic diversity of vertebrate species would create an unprecedented resource for the life sciences and for worldwide conservation efforts."

I do wonder however, if the selection will be biased towards life sciences or conservation... Never the twain shall meet? or am I wrong?

Friday, 14 May 2010

Lincoln Stein makes his case for moving genome informatics to the Cloud

Matthew Dublin summarizes Lincoln's paper in Making the Case for Cloud Computing & Genomics in genomeweb

excerpt "....
Stein walks the reader through an nice explanation of what exactly cloud computing is, the benefits of using a compute solution that grows and shrinks as needed, and makes an attempt at tackling the question of the cloud's economic viability when compared to purchasing and managing local compute resources.
The take away is that Moore's Law and its effect on sequencing technology will soon force researchers to analyze their mountains of sequencing data in a paradigm where the software comes to the data rather than the current, and opposite, approach. Stein says that this means now more than ever, cloud computing is a viable and attractive option..... "

Yet to read it (my weekend bedtime story) will post comments here.

The elusive Bioscope on cloud service.

Reference to my last post about Cloud enabled Bioscope. I have found new documentation at applied biosystems.

But the service appears to be not yet public.
Oh the suspense!

Tuesday, 11 May 2010

Solutions for Applying Automation to Next-Generation Sequencing Sample Prep

FYI might be useful for some.
disclaimer: I am not affliated with the companies involved.
http://www.genengnews.com/ngs


Solutions for Applying Automation to Next-Generation Sequencing Sample Prep

In recent years, next-generation sequencing (NGS) technologies have rapidly evolved to provide faster, better, cheaper and more reliable mapping of DNA and RNA sequences thus enabling a diverse set of genomic discoveries. This has been largely driven by innovations and improvisations at the technical end. However, challenges still remain with increasing sample throughput, enabling sample preparation, minimizing errors, and with improving data analysis.
This webinar provides the audience with an overview of the developments in NGS technologies, with an emphasis on the challenges that are routinely encountered at various stages in the sequencing and analysis of samples. It offers detailed information on the specific challenges associated with sample preparation and the benefits of using automation to alleviate some of the bottlenecks. The webinar features viewpoints expressed by three experts in the field, who share examples from their laboratories on how to effectively adopt and utilize automation for NGS applications.

What will be covered:

  • Overview of NGS technologies and potential challenges in their adoption and use
  • Tackling challenges associated specifically with sample prep for NGS
  • Effective use of automation for alleviating some of the bottlenecks in sequencing
  • Ways to increase and improve the sample throughput in sequencing
  • Use and creation of high-diversity sequencing libraries
  • Application of automated NGS platforms in areas like cancer genetics and CNS research
  • Overview of targeted resequencing applications including effective whole exon sequencing, indexed/barcoded resequencing of small regions, automated targeted resequencing

Who will benefit from attending:

  • Scientists involved in pharmaceutical/biotechnology R&D and clinical services
  • Scientists keen to learn more about the use and adoption of NGS technologies
  • Scientists and clinicians active in biomarker research
  • Scientists looking to use sequencing for oncology and CNS research

Panelists include

  • Shawn Levy, Ph.D., Faculty Investigator, HudsonAlpha Institute for Biotechnology
  • Brian Minie, Ph.D., Broad Institute of MIT and Harvard
  • Emily Leproust, Ph.D., Director, Applications and Chemistry R&D, Genomics, Agilent Technology
John Sterling, Editor in Chief, Genetic Engineering & Biotechnology News, will be the host of this webinar.   A live Q&A session will follow the presentations, offering you a chance to pose questions to our expert panelists.

Monday, 10 May 2010

A plethora of solid2fastq or csfasta convertors to fastq

I hadn't realised that there's an accumulation of prog/scripts to do the same task. Last count is 4 of these in my tool closet.


The C binary from bfast

solid2fastq 0.6.4a

Usage: solid2fastq [options]
        -c              produce no output.
        -n      INT     number of reads per file.
        -o      STRING  output prefix.
        -j              input files are bzip2 compressed.
        -z              input files are gzip compressed.
        -J              output files are bzip2 compressed.
        -Z              output files are gzip compressed.
        -t      INT     trim INT bases from the 3' end of the reads.
        -h              print this help message.

 send bugs to bfast-help@lists

solid2fastq.pl from bfast-0.6.4a
with notes in the script to refer to the above
# Author: Nils Homer
# Please see the C implementation of this script.


EDIT: THANKS to iceman for his reminder in the comments
"Make sure that you use the BWA's solid2fastq.pl if you are going to use BWA as it "double-encodes" the reads."

solid2fastq.pl from bwa-0.5.7
Usage: solid2fastq.pl

Note: is the string showed in the `# Title:' line of a
      ".csfasta" read file. Then F3.csfasta is read sequence
      file and F3_QV.qual is the quality file. If
      R3.csfasta is present, this script assumes reads are
      paired; otherwise reads will be regarded as single-end.

      The read name will be :panel_x_y/[12] with `1' for R3
      tag and `2' for F3. Usually you may want to use short
      to save diskspace. Long also causes troubles to maq.
 

# Author: lh3
# Note: Ideally, this script should be written in C. It is a bit slow at present.
# Also note that this script is different from the one contained in MAQ.

maq-0.7.1/scripts/solid2fastq.pl

Usage: solid2fastq.pl

Note: is the string showed in the `# Title:' line of a
      ".csfasta" read file. Then F3.csfasta is read sequence
      file and F3_QV.qual is the quality file. If
      R3.csfasta is present, this script assumes reads are
      paired; otherwise reads will be regarded as single-end.

      The read name will be :panel_x_y/[12] with `1' for F3
      tag and `2' for R3. Usually you may want to use short
      to save diskspace. Long also causes troubles to maq.


# Author: lh3
# Note: Ideally, this script should be written in C. It is a bit slow at present.

Friday, 7 May 2010

Yet another viewer for NGS data, MagicViewer

    MagicViewer: integrated solution for next-generation sequencing data visualization and genetic variation detection and annotation.
    Hou H, Zhao F, Zhou L, Zhu E, Teng H, Li X, Bao Q, Wu J, Sun Z.
    Nucleic Acids Res. 2010 May 5. [Epub ahead of print]
    PMID: 20444865 [PubMed - as supplied by publisher]

Abstract

New sequencing technologies, such as Roche 454, ABI SOLiD and Illumina, have been increasingly developed at an astounding pace with the advantages of high throughput, reduced time and cost. To satisfy the impending need for deciphering the large-scale data generated from next-generation sequencing, an integrated software MagicViewer is developed to easily visualize short read mapping, identify and annotate genetic variation based on the reference genome. MagicViewer provides a user-friendly environment in which large-scale short reads can be displayed in a zoomable interface under user-defined color scheme through an operating system-independent manner. Meanwhile, it also holds a versatile computational pipeline for genetic variation detection, filtration, annotation and visualization, providing details of search option, functional classification, subset selection, sequence association and primer design. In conclusion, MagicViewer is a sophisticated assembly visualization and genetic variation annotation tool for next-generation sequencing data, which can be widely used in a variety of sequencing-based researches, including genome re-sequencing and transcriptome studies. MagicViewer is freely available at http://bioinformatics.zj.cn/magicviewer/.
PMID: 20444865 [PubMed - as supplied by publisher]

Datanami, Woe be me