Showing posts with label AGBT. Show all posts
Showing posts with label AGBT. Show all posts

Monday, 25 February 2013

Michael Schatz:Assembling Crop Genomes With SMS

PDF of the presentation on Feb 22, 2013 AGBT, Marco Island, FL

http://schatzlab.cshl.edu/presentations/2013-02-20.AGBT.Assembling%20Crop%20Genomes.pdf

if you need an intro

"In a talk during the evening session, Mike Schatz, an assistant professor at Cold Spring Harbor Laboratory, spoke about “Assembling Crop Genomes with Single Molecule Sequencing.” Crops are important to sequence — 15 crops represent 90% of the world’s food, Schatz said — but are notoriously difficult to study because of their large genome size, high repeat content, and higher ploidy. Along with Sergey Koren and Adam Phillippy, he has built a pipeline to create hybrid genome assemblies using PacBio long reads combined with shorter-read sequence — either CCS reads from PacBio or data from another sequencing platform. In an example he offered of a rice strain, an attempted genome assembly using just Illumina reads yielded an N50 contig of 16Kb, but adding PacBio long reads to that boosted the N50 contig to 25Kb. Ultimately, Schatz said, he expects that as PacBio's readlength improves, this kind of approach could routinely generate megabase-size contigs or even pull plant chromosomes into single contigs.

For more information on Mike Schatz’s work using SMRT Sequencing, check out this case studydescribing an automated pipeline for genome finishing with PacBio long reads."

source: http://blog.pacificbiosciences.com/2013/02/notes-from-agbt-long-read-sequence-data.html

He includes a snippet of code to answer this question from twitter
'What's the longest single contig from a de Bruijn assembler without PE or a jumping library?'


$ perl -e 'print ">random\n"; @D=split //,"ACGT"; \for (1...100000000){print $D[int(rand(4))];} \print "\n"’ | fold > random.fa$ wgsim –r 0 -e 0 -N 50000000 -1 100 -2 1 \random.fa random.reads.fq /dev/null$ SOAPdenovo-63mer all –s random.cfg -K 63 -o random.63$ getlengths random.63.contig           1 99999990

Saturday, 18 February 2012

Oxford Nanopore megaton announcement: “Why do you need a machine?” – exclusive interview for this blog!


http://pathogenomics.bham.ac.uk/blog/2012/02/oxford-nanopore-megaton-announcement-why-do-you-need-a-machine-exclusive-interview-for-this-blog/

woke up this morning to see a whole bunch of excited tweets on Oxford Nanopore and I can totally understand why. This is the real democratization of DNA sequencing. Move over benchtop / desktop sequencers for 'laptop sequencers'!

Hmmm or a cluster of sequencers, on your compute cluster ... !

Using USB powered sequencers, and a pipette to put in the dsDNA and you might have your sequence read to FASTQ directly to your laptop. 
No known limit to read length. 
4% seq error (the good thing is that the form of error is known and therefore correctable)


Do read the url above for more info, here's the excerpted executive summary for the impatient

Executive Summary
  • Nanopore have announced a strand sequencing method, made possible by a heavily modified biological nanopore and an industrially-fabricated polymer
  • DNA passes through the nanopore and tri-nucleotides in contact with the pore are detected through electrochemistry
  • Demonstrated 2x50kb sense & anti-sense of same molecules (lambda phage) – no theoretical read length limit
  • Can sequence direct from blood without need for sample preparation
  • Two products announced:
    • MinIon – USB disposable sequencer for ~ $900 has 512 nanopores – target 150mb/hour
    • MinIon can run at 120-1000 bases/minute per pore for up to 6 hours
    • GridIon – two versions of rack-mountable sequencer with 2000 nanopores (2nd half 2012), 8000 nanopores (2013)
    • GridIons can be racked in parallel, 20 could do a whole human genome in 15 minutes
    • Each GridIon can do "tens of gigabases" over 24 hours
  • Both machines commercially available 2nd half 2012
  • Sequencing can be paused, sample recovered, replaced, started again
  • Accuracy is 96%, errors are deletions, error profile will improve through software


Check out Forbes interview with 454 / PGM inventor Jon Rothberg

"Rothberg noted that Ion Torrent’s new machine, the Proton, the company showed three completed human genomes yesterday at AGBT. More importantly, he had the machine – not a mock-up or a design – on the stage. “That’s where you need to be to ship mid-year,” he writes."


Over at Genomes Unzipped 
Oxford Nanopore CTO Clive Brown related how sequencing library prep is as simple as diluting rabbit's blood with water. Now that is impressive!




This post is getting too long because I keep updating it. 
Over at the BioITWorld, there's an interview with Clive Brown which cites other interesting info. 
First of which is the opening paragraph which is amusing in the light of ONT's rivals comments
"Clive Brown, vice president of development and informatics for Oxford Nanopore Technologies (ONT), a.k.a “the most honest guy in all of next-gen sequencing,” as dubbed by The Genome Center's David Dooling, is hoping to catch lightning in a bottle again. "


Oxford Nanopore has not yet revealed details of its future platform, but in early 2009, published a lovely paper in Nature Nanotechnology showing that its alpha-hemolysin nanopores can discriminate between the four bases of DNA (not to mention a fifth, methyl C)




Directly get methylation information from your sequencing sans complicated sample prep? That has to be another selling point. 


Not sure whether Nanopore is truly vaporware. However, gauging by the excitement over the blogosphere and the hit rates for the first to blog about it. I think Nanopore is upping the ante for the next IT sequencer. 
maybe we can only survive 2 more AGBT like this and AGBT might fizzle out as new sequencing technologies fade as our computation advances trails behind the ability to generate more data. 
Maybe you will see scientists start attending Big Data tech conferences or AGBT's  main draw will  fancy new software to assemble, align and make sense out of all the data being generated ... 




This picture tells quite a story (Wordle constructed from 3,386 tweets and retweets tagged #AGBT with @s removed).
No prizes for guessing the winner ... 

Wednesday, 9 February 2011

1st feedback from Ion Torrent at AGBT

Definitely exciting! Wished I was there instead of reading twitter and blog reports on the event. 
excerpted
During a Life Tech conference workshop, Kevin McKernan, vice president of advanced research R&D, said that PGM customers can expect a 10-fold increase in output about every six months.
Since Life Tech launched the system in December, it has announced its first chip upgrade, from the Ion 314 Chip to the Ion 316 Chip. The upgrade promises to increase the output from 10 to 100 megabases per run and will be available to early-access customers this quarter, and more generally in the second quarter (IS 1/11/2011).
The best internal run today has yielded 300,000 reads 100 base pairs in length of quality Q17, according to Maneesh Jain, Ion Torrent's vice president of marketing and business development, who spoke during a separate Ion Torrent conference workshop. He said that the company plans to "address" RNA-seq as an additional application "later in the year."
The system's read length is currently about 100 base pairs. According to information provided during the Life Tech and Ion Torrent workshops, read length is expected to increase to 200 base pairs in the fourth quarter and to 400 base pairs in 2012.
McKernan said that in a single run to sequence the E. coli genome, the system provided "uniform genome coverage regardless of GC content." Coverage of human genes has also been "very even" and has included areas that were missed by both SOLiD and Illumina sequencing, he said.
The PGM's per-base accuracy also continues to improve. According to McKernan, based on 50-base reads, it was about 98.7 percent at the end of 2010 and has since improved to 99.6 percent. In the second quarter of this year, it is projected to improve further, he added.
The company has also improved the accuracy for homopolymer regions, he said, largely based on better software. While the per base accuracy for a stretch of four identical bases was 94 percent at the end of last year, it has increased to 98 percent for five identical bases today, and is expected to go up to 99 percent during the second half of this year.

Datanami, Woe be me