Showing posts with label 454. Show all posts
Showing posts with label 454. Show all posts

Wednesday, 24 July 2013

Illumina produces 3k of 8500 bp reads on HiSeq using Moleculo Technology

Keith blogged about how super long read sequencing methods would be a threat to Illumina in Jan 2013. Today, Illumina can now openly acknowledge the shortcomings of their short reads for various applications like
  • assembly of complex genomes (polyploid, containing excessive long repeat regions, etc.), 
  • accurate transcript assembly, 
  • metagenomics of complex communities, 
  • and phasing of long haplotype blocks.


the reason?
This latest set of data released on BaseSpace
Read length distribution of synthetic long reads for a D. melanogaster library
The data set, available as a single project in BaseSpace, can be accessed here.

image source: http://blog.basespace.illumina.com/2013/07/22/first-data-set-from-fasttrack-long-reads-early-access-service/

with the integration of Moleculo they have managed to generate ~30 gb of raw sequence data. They have refrained from talking about 'key analysis metrics' that's available in the pdf report. Perhaps it's much easier to let the blogosphere and data scientists dissect the new data themselves.

Am wondering when the 454 versus Illumina Long Reads side-by-side comparison will pop up

UPDATE:

Can't find the 'key analysis metrics' in the pdf report files. Perhaps it's still being uploaded? *shrugs*
so please update me if you see it  otherwise I just have to run something on it


These are the files that I have now

total 512M
 259M Jul 18 01:01 mol-32-2832.fastq.gz
  44K Jul 24  2013 FastTrackLongReads_dmelanogaster_281c.pdf
 149K Jul 24  2013 mol-32-281c-scaffolds.txt
  44K Jul 24  2013 FastTrackLongReads_dmelanogaster_2832.pdf
 151K Jul 24  2013 mol-32-2832-scaffolds.txt
 253M Jul 24  2013 mol-32-281c.fastq.gz

md5sums
6845fc3a4da9f93efc3a52f288e2d7a0  FastTrackLongReads_dmelanogaster_281c.pdf
02f5de4f7e15bbcd96ada6e78f659fdb  FastTrackLongReads_dmelanogaster_2832.pdf
586599bb7fca3c20ba82a82921e8ba3f  mol-32-281c-scaffolds.txt
b25010e9e5e13dc7befc43b5dff8c3d6  mol-32-281c.fastq.gz
6822cfbd3eb2a535a38a5022c1d3c336  mol-32-2832-scaffolds.txt
873f09080cdf59ed37b3676cddcbe26f  mol-32-2832.fastq.gz


I have ran FastQC (FastQC v0.10.1) on both samples the images below are from 281c.
you can download the full HTML report here
https://www.dropbox.com/sh/5unu3zba9u21ywj/JT4HdkzfOP/mol-32-281c_fastqc.zip
https://www.dropbox.com/s/mpxa5wx51iqmiz3/mol-32-2832_fastqc.zip

Reading about the Moleculo sample prep method, it seems like it's just a rather ingenious way to stitch short reads which are barcoded to form a single long contig. if that is the case, then I am not sure if the base quality scores here are meaningful anymore since it's a mini-assembly. Also this takes out any quantitative value of the number of reads I presume. So accurate quantification of long RNA molecules or splice variants isn't possible. Nevertheless it's an interesting development on the Illumina platform. Looking forward to seeing more news about it.













Other links

Illumina Long-Read Sequencing Service
Moleculo technology: synthetic long reads for genome phasing, de novo sequencing
CoreGenomics: Genome partitioning: my moleculo-esque idea
Moleculo and Haplotype Phasing - The Next Generation TechnologistNext Generation Technologist
Abstract: Production Of Long (1.5kb – 15.0kb), Accurate, DNA Sequencing Reads Using An Illumina HiSeq2000 To Support De Novo Assembly Of The Blue Catfish Genome (Plant and Animal Genome XXI Conference)
http://www.moleculo.com/ (no info on this page though)
Illumina Announces Phasing Analysis Service for Human Whole-Genome Sequencing - MarketWatch
Patent information on the Long Read technology
https://docs.google.com/viewer?url=patentimages.storage.googleapis.com/pdfs/US20130079231.pdf









Tuesday, 15 May 2012

NATURE BIOTECHNOLOGY | Performance comparison of benchtop high-throughput sequencing platforms


Performance comparison of benchtop high-throughput sequencing platforms

Nature Biotechnology
 
30,
 
434–439
 
(2012)
 
doi:10.1038/nbt.2198
Received
 
Accepted
 
Published online
 
Corrected online
 

Abstract

Three benchtop high-throughput sequencing instruments are now available. The 454 GS Junior (Roche), MiSeq (Illumina) and Ion Torrent PGM (Life Technologies) are laser-printer sized and offer modest set-up and running costs. Each instrument can generate data required for a draft bacterial genome sequence in days, making them attractive for identifying and characterizing pathogens in the clinical setting. We compared the performance of these instruments by sequencing an isolate of Escherichia coli O104:H4, which caused an outbreak of food poisoning in Germany in 2011. The MiSeq had the highest throughput per run (1.6 Gb/run, 60 Mb/h) and lowest error rates. The 454 GS Junior generated the longest reads (up to 600 bases) and most contiguous assemblies but had the lowest throughput (70 Mb/run, 9 Mb/h). Run in 100-bp mode, the Ion Torrent PGM had the highest throughput (80–100 Mb/h). Unlike the MiSeq, the Ion Torrent PGM and 454 GS Junior both produced homopolymer-associated indel errors (1.5 and 0.38 errors per 100 bases, respectively).

Figures at a glance

Tuesday, 6 December 2011

Complete Khoisan and Bantu genomes from southern Africa : Article : Nature

http://www.nature.com/nature/journal/v463/n7283/full/nature08795.html

Just attended a very good lecture by Stephan Schuster, entitled "African Genomes: Charting Human Diversity"
He offered unbiased views / charts on the platform differences between 454, GAIIx, HiSeq, SOLiD for NGS sequencing coverage (which I think I should not repeat here). It points to the need to do sequencing on 2 different platforms to get a more accurate SNP list.

He also gave compelling reasons for getting a 20x coverage of Human genome done in 454 to complete the human genome (457 gaps in hg19). 

Yes, I often forget that the media / lay person thinks that the human genome is 'complete'. It's a often ignored 'secret' that actually it isn't. Maybe the next marketing ploy(i mean strategy) for emerging sequencing platforms would be to be THE ONE that actually finishes the human genome.


Saturday, 10 September 2011

Rapid and Efficient Human Mutation Detection Using a Bench-Top Next-Generation DNA Sequencer.

Hum Mutat. 2011 Sep 6. doi: 10.1002/humu.21602. [Epub ahead of print]
Click here to read

Rapid and Efficient Human Mutation Detection Using a Bench-Top Next-Generation DNA Sequencer.

Source

Center for Complex Disease Genomics, McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University School of Medicine, Baltimore, MD 21205, USA.

Abstract

Next-generation sequencing (NGS) technologies can be a boon to human mutation detection given their high throughput: consequently, many genes and samples may be simultaneously studied with high coverage for accurate detection of heterozygotes. In circumstances requiring the intensive study of a few genes, particularly in clinical applications, a rapid turn-around is another desirable goal. To this end, we assessed the performance of the bench-top 454 GS Junior platform as an optimized solution for mutation detection by amplicon sequencing of three type 3 semaphorin genes SEMA3A, SEMA3C and SEMA3D implicated in Hirschsprung disease (HSCR). We performed mutation detection on 39 PCR amplicons totaling 14,014bp in 47 samples studied in pools of 12 samples. Each 10-hour run was able to generate ∼75,000 reads and ∼28 million high-quality bases at an average read length of 371bp. The overall sequencing error was 0.26 changes per kb at a coverage depth of ≥20 reads. Altogether, 37 sequence variants were found in this study of which 10 were unique to HSCR patients. We identified five missense mutations in these three genes that may potentially be involved in the pathogenesis of HSCR and need to be studied in larger patient samples. ©2011 Wiley-Liss, Inc.
© 2011 Wiley-Liss, Inc.
PMID:
21898659
[PubMed - as supplied by publisher]

Wednesday, 13 April 2011

ZORRO is an hybrid sequencing technology assembler:tested with Solexa 454

Typos in the header aside... you have to love the name!

waiting for the name to become a verb... " I zorroed the NGS reads the other and i had a fantastic assembly!" lol..



Here goes: http://lge.ibi.unicamp.br/zorro/


Overview

ZORRO is an hybrid sequencing technology assembler. It takes 2 sets of pre-assembled contigs and merge them into a more contiguous and consistent assembly. We have already tested Zorro with Illumina Solexa and 454 from some of organisms varying from 3Mb to 100Mb. The main caracteristic of Zorro is the treatment before and after assembly to avoid errors.
The ZORRO project is maintained by Gustavo Lacerda, Ramon Vidal and Marcelo Carazzole and were first used in this Yeast assembly: Genome structure of a Saccharomyces cerevisiae strain widely used in bioethanol production
ZORRO needs to be better documented and has not undergone enough testing. If you want to discuss the pipeline you can join the mailing list: zorro-google group

Zorro: The Complete Series

p.s. the typo is here
"ZORROthe masked assember "

Saturday, 12 March 2011

Chimeric 16S rRNA sequence formation and detection in Sanger and 454-pyrosequenced PCR amplicons [RESOURCES]


Chimeric 16S rRNA sequence formation and detection in Sanger and 454-pyrosequenced PCR amplicons [RESOURCES]

Bacterial diversity among environmental samples is commonly assessed with PCR-amplified 16S rRNA gene (16S) sequences. Perceived diversity, however, can be influenced by sample preparation, primer selection, and formation of chimeric 16S amplification products. Chimeras are hybrid products between multiple parent sequences that can be falsely interpreted as novel organisms, thus inflating apparent diversity. We developed a new chimera detection tool called Chimera Slayer (CS). CS detects chimeras with greater sensitivity than previous methods, performs well on short sequences such as those produced by the 454 Life Sciences (Roche) Genome Sequencer, and can scale to large data sets. By benchmarking CS performance against sequences derived from a controlled DNA mixture of known organisms and a simulated chimera set, we provide insights into the factors that affect chimera formation such as sequence abundance, the extent of similarity between 16S genes, and PCR conditions. Chimeras were found to reproducibly form among independent amplifications and contributed to false perceptions of sample diversity and the false identification of novel taxa, with less-abundant species exhibiting chimera rates exceeding 70%. Shotgun metagenomic sequences of our mock community appear to be devoid of 16S chimeras, supporting a role for shotgun metagenomics in validating novel organisms discovered in targeted sequence surveys.

Friday, 4 March 2011

Guide/tutorial for the analysis of RNA-seq data

link in seqanswers

Excellent starting point for those confused about the RNA-seq data analysis procedure.

Hello,

I've written a guide to the analysis of RNA-seq data, for the purpose of differential expression analysis. It currently lives on our internal wiki that can't be viewed outside of our division, although printouts have been used at workshops. It is by no means perfect and very much a work in progress, but a number of people have found it helpful, so I thought it would useful to have it somewhere more publicly accessible.

I've attached a pdf version of the guide, although really what I was hoping was that someone here could suggest somewhere where it could be publicly hosted as a wiki. This area is so multifaceted and fast-moving that the only way such a guide can remain useful is if it can be constantly extended and updated.

If anyone has any suggestions about potential hosting, they can contact me at myoung @wehi.edu.au

Cheers

Matt

Update: I've put a few extra things on our local Wiki and seeing as people here seem to be finding this useful I thought I'd post an updated version. I'm also an author on a review paper on Differential Expression using RNA-seq which people who find the guide useful, might also find relevant...

RNA-seq Review

Friday, 24 December 2010

Exome sequencing reveals mutations in previously unconsidered candidate genes

Excerpted from Bio-IT World 

December 20, 2010 | Doctors at the Medical College in Wisconsin have published in the journal Genetics and Medicine the results of exome sequencing in a seriously ill boy with undiagnosed bowel disease. The study (using 454 sequencing) revealed mutations in a gene called XIAP, which interestingly was not previously considered among more than 2,000 candidate genes before the DNA sequencing was performed. ....read full article

Tuesday, 10 August 2010

PyroNoise:Accurate determination of microbial diversity from 454 pyrosequencing data

Using 454 to do microbial ecology / metagenomics of environmental / soil samples?
Then I think you should take a look at this paper.

Quince, C., Lanzén, A., Curtis, T., Davenport, R., Hall, N., Head, I., Read, L., & Sloan, W. (2009). Accurate determination of microbial diversity from 454 pyrosequencing data Nature Methods, 6 (9), 639-641 DOI: 10.1038/nmeth.1361

The Pathogens blog has a good summary post on it.

Wednesday, 14 July 2010

Shiny new tool to index NGS reads G-SQZ

This is a long over due tool for those trying to do non-typical analysis with your reads.
Finally you can index and compress your NGS reads

http://www.ncbi.nlm.nih.gov/pubmed/20605925

Bioinformatics. 2010 Jul 6. [Epub ahead of print]
G-SQZ: Compact Encoding of Genomic Sequence and Quality Data.

Tembe W, Lowey J, Suh E.

Translational Genomics Research Institute, 445 N 5th Street, Phoenix, AZ 85004, USA.
Abstract

SUMMARY: Large volumes of data generated by high-throughput sequencing instruments present non-trivial challenges in data storage, content access, and transfer. We present G-SQZ, a Huffman coding-based sequencing-reads specific representation scheme that compresses data without altering the relative order. G-SQZ has achieved from 65% to 81% compression on benchmark datasets, and it allows selective access without scanning and decoding from start. This paper focuses on describing the underlying encoding scheme and its software implementation, and a more theoretical problem of optimal compression is out of scope. The immediate practical benefits include reduced infrastructure and informatics costs in managing and analyzing large sequencing data. AVAILABILITY: http://public.tgen.org/sqz Academic/non-profit: Source: available at no cost under a non-open-source license by requesting from the web-site; Binary: available for direct download at no cost. For-Profit: Submit request for for-profit license from the web-site. CONTACT: Waibhav Tembe (wtembe@tgen.org).

read the discussion thread in seqanswers for more tips and benchmarks

I am not affliated with the author btw

Sunday, 30 May 2010

Cofactor genomics on the different NGS platforms

Original post here

They are a commercial company that offers NGS on ABI and Illumina platforms and since this is on their company page I guess its their official stand on what rocks on each platform

Excerpted.

Applied Biosystems SOLiD 3

The Applied Biosystems SOLiD 3 has the shortest but also the highest quantity of reads. The SOLiD produces up to 240 million 50bp reads per slide per end. As with the Illumina, Mate-Pairs produce double the output by duplicating the read length on each end, and the SOLiD supports a variety of insert lengths like the 454. The SOLiD can also run 2 slides at once to again double the output. SOLiD has the lowest *raw* base qualities but the highest processed base qualities when using a reference due to its 2-base encoding. Because of the number of reads and more advanced library types, we recommend the SOLiD for all RNA and bisulfite sequencing projects.

Solexa/Illumina

The Solexa/Illumina generates shorter reads at 36-75bp but produces up to 160 million reads per run.  All reads are of similar length.  The Illumina has the highest *raw* quality scores and its errors are mostly base substitutions. Paired-end reads with ~200 bp inserts are possible with high efficiency and double the output of the machine by duplicating the read length on each end. Paired-end Illumina reads are suitable for de novo assemblies, especially in combination with 454. The large number of reads makes the Illumina appropriate for de novo transcriptome studies with simultaneous discovery and quantification of RNAs at qRT-PCR accuracy.

Roche/454 FLX

The Roche/454 FLX with Titanium chemistry generates the longest reads (350-500bp) and the most contiguous assemblies, can phase SNPs or other features into blocks, and has the shortest run times. However, 454 also produces the fewest total reads (~1 million) at the highest cost per base. Read lengths are variable. Errors occur mostly at the ends of long same-nucleotide stretches. Libraries can be constructed with many insert sizes (8kb - 20kb) but at half of the read length for each end and with low efficiency.

Friday, 28 May 2010

Illumina: an alternative to 454 in metagenomics?

Check out this BMC Bioinformatics paper entitled "Short clones or long clones? A simulation study on the use of paired reads in metagenomics" 


"This paper addresses the problem of taxonomical analysis of paired reads. We describe a new feature of our metagenome analysis software MEGAN that allows one to process sequencing reads in pairs and makes assignments of such reads based on the combined bit scores of their matches to reference sequences. Using this new software in a simulation study, we investigate the use of Illumina paired-sequencing in taxonomical analysis and compare the performance of single reads, short clones and long clones. In addition, we also compare against simulated Roche-454 sequencing runs."

"Our study suggests that a higher percentage of Illumina paired reads than of Roche-454 single reads are correctly assigned to species."

"The gain of long-clone data (75 bp paired reads) over long single-read data (250 bp reads) is still significant at ≈ 4% (not shown)."

of course more importantly
"The authors declare that they have no competing interests."

I am not sure if such a program exists but I wonder if there is a aligner that takes into account the size between mate pairs and paired ends. Theoratically it should improve mapping. but by how much is unknown

Thursday, 27 May 2010

178 Microbial Reference Genomes Associated with the Human Body

Venter Institute Scientists, Along with Consortium Members of the NIH's Human Microbiome Project, Sequence 178 Microbial Reference Genomes Associated with the Human Body

Researchers from the J. Craig Venter Institute, a not-for-profit genomic research organization, have published (along with other members of the National Institutes of Health (NIH) Human Microbiome Jumpstart Reference Strains Consortium), a catalog of 178 microbial reference genomes isolated from the human body.  Other members of the Consortium are: Baylor College of Medicine Human Genome Sequencing Center, the Broad Institute, and the Genome Center at Washington University. The paper is being published in the May 21 issue of the journal Science
The human body is teeming with a variety of microbial species. This collective community is called the human microbiome. The role these microbes play in human health and disease is still relatively unknown but likely very important. The NIH Human Microbiome Project was launched in 2007, as part of the National Institutes of Health’s (NIH) Common Fund’s Roadmap for Medical Research. It is a $157 million, five-year effort that will implement a series of increasingly complicated studies that reveal the interactive role of the microbiome in human health. 

Venter Institute Press Release 20th May

Wednesday, 26 May 2010

A scientific spectator's guide to next-generation sequencing

ROFL
I love the title!

A scientific spectator's guide to next-generation sequencing

Dr Keith not only looks at next gen sequencing but also the emerging technologies of single molecule sequencing. Interesting read!

My fave parts of the review
"Finally, there is the cost per base, generally expressed in a cost per human genome sequenced at approximately 40X coverage. To show one example of how these trade off, the new PacBio machine has a great cost per sample (~U$100) and per run (you can run just one sample) but a poor cost per human genome – you’d need around 12,000 of those runs to sequence a human genome (~U$120K). In contrast, one can buy a human genome on the open market for U$50K and sub U$10K genomes will probably be generally available this year."


"Length is critical to genome sequencing and RNA-seq experiments, but really short reads in huge numbers are what counts for DGE/SAGE and many of the functional tag sequencing methods. Technologies with really long reads tend not to give as many, and with all of them you can always choose a much shorter run to enable the machine to be turned over to another job sooner – if your application doesn’t need long reads."

 

 

Coral Transcriptomics-a budget NGS approach?

Was surprised I didn't blog about this earlier.
Dr Mikhail Matz is a researcher in the field of coral genomics. His approach to doing de novo transcriptomics for an organism whose genome is unavailable.


his compute cluster is basically
"two Dell PowerEdge 1900 servers joined together with ROCKS clustering software v5.0. Each server had: two Intel Quad Core E5345 (2.33 Ghz, 1333 Mhz FSB, 2x4MB L2 Cache) CPU’s and 16 GB of 667 Mhz DDR2 RAM. The cluster had a combined total of 580 GB disk space."




Tools used are
- Blast executables from NCBI, including blast, blastcl3, and blastclust
- Washington University blast (Wu-blast)
- ESTate sequence clustering software
- Perl

He admits that the assembled transcriptome might be incomplete (~40,000 contigs with five-fold average sequencing see Figure 2 for the size distribution of the assembled contigs
But it is "good enough" to use as a reference transcriptome to align SOLiD reads accurately and to generate the coverage that 454 can't give for the same amount of grant money.

the results are published in BMC Genomics

Not sure if you have heard of just in time inventory. But I think "good enough" science takes a bit of dare to spend that money to ask those what-ifs.


Wednesday, 7 April 2010

Datanami, Woe be me