Tuesday, 7 February 2012

Google Refine is a power tool for working with messy data sets

http://google-refine.blogspot.com/

Google Refine is a power tool for working with messy data sets, including cleaning up inconsistencies, transforming them from one format into another, and extending them with new data from external web services or other databases. Version 2.0 introduces a new extensions architecture, a reconciliation framework for linking records to other databases (like Freebase), and a ton of new transformation commands and expressions.

Monday, 6 February 2012

Public Poll: What data storage solutions are you using now?

BooRoo

Just a public poll created curious as to what systems are out there and what is being used day to day. Results are immediately viewable. Order of choices is randomized. 


Which data storage platform are you using?

Create an online survey quiz or web poll
http://www.bluearc.com/ BlueArc is your Life Sciences Solution0%
http://www.isilon.com/ Isilon scale-out NAS, BIG data is no longer a challenge, it’s an opportunity.0%
http://www.panasas.com/ PANASAS SOLVES BIG DATA PROBLEMS0%
http://www.ddn.com/ DataDirect Networks’ (DDN) Big Data Storage Technology Powers More0%
http://aws.amazon.com/ Store data and build dependable backup solutions using AWS’s highly reliable, inexpensive data storage services.0%
Inhouse solutions using commodity hardware0%
http://www-03.ibm.com/systems/software/gpfs/ IBM GPFS0%
Other: (Please specify)0%
Create an online survey quiz or web poll

[circos] Circos v0.56 released


---------- Forwarded message ----------
From: "Martin"
Date: Feb 6, 2012 9:56 am
Subject: [circos] Circos v0.56 released

Dear Circos users!

I have made v0.56 available.

http://circos.ca/software/download/

Thank you to everyone who submitted bug reports and patiently waited
for me to respond.

Remember that now tutorials and the core code are kept separate. The
core circos-0.56.tgz archive contains an example directory (example/)
that creates a relatively complex image. Use this to ensure that
everything is working.

Most of the changes are bug fixes. Error handling is improved - many
errors have a gentler format and should be more explanatory.

*** IMPORTANT CHANGES AND REMINDERS ***

1. Import colors/fonts/patterns all at once

Use a single line to add color, font and pattern definitions in
circos.conf:

<<include etc/colors_fonts_patterns.conf>>

instead of

<colors>
<<include etc/colors.conf>>
...

Note that colors.conf now imports brewer.conf automatically.

Please see example/etc/circos.conf for ideas of how to make your
configuration more modular.

2. Make sure you're using etc/housekeeping.conf

A while back, I had centralized system parameters that used to appear
at the bottom of circos.conf into a separate file. For those of you
who still have these parameters in their circos.conf file, please
remove them.

<<include etc/housekeeping.conf>>

instead of

anglestep       = 0.5
minslicestep    = 10
beziersamples   = 40
debug           = no
warnings        = no
imagemap        = no
units_ok = bupr
units_nounit = n

3. Use debug_group

You'll notice that Circos' default debug output contains more text. By
default debugging that falls into the 'summary' category is created. A
complete list of debug categories is listed here

http://mkweb.bcgsc.ca/dev/circos/documentation/tutorials/configuration/debugging/

In particular, the following are very useful

  timer - show code timings
  io - loading and locating files
  textplace - reports whether a text in a text track has been
successfully placed

4. SVG text handling improved

I've revamped how SVG text labels are placed. You'll notice that fonts
now have font family definitions in etc/fonts.conf. This string is
used in the SVG file to specify the font. Once you have the font
installed on your system, labels won't all appear in the SVG app or
viewer default (e.g. Arial) font.

There's still more to do here, but it's getting better.

0.56 CHANGELOG

Parallel ideogram labels are now centered with respect to the
ideogram.

restrict_parameter_names now controls whether parameters are
restricted to a pre-defined list (e.g. color, thickness, etc). If you
have custom parameters (e.g. 'myspecialcode') then set
restrict_parameter_names=no. By default, this is always set to 'yes'.

Added link_orientation for text tracks. When set to "out" links from
text labels face out, rather than in.

Added font names to SVG files via font-family tag.

Removed -verbose. The -v flag now reports version.

Removed dependence on Graphics::ColorObject.

Added error handling framework.

Circos now requires Text::Format

Bug fix to heat map color mapping of last color.

Fixed bug which was causing line links to be drawn with a thickness
half of what was requested.

Fixed bug that prevented parameters made acceptable by the
restrict_parameter_names=no setting from being parsed.

Configuration file location is now guessed if guess_conf_location=yes
(see etc/housekeeping.conf)

Color file cache can now be static (color_cache_static) and dir/file
can be changed (color_cache_{file,dir})

Added 'placed' and 'not_placed' output for labels in text tracks to.
Use -debug_group text to see this.

  # create data file of labels that were not placed
  circos -conf ... -debug_group text | grep not_placed >
text.notplaced.txt

Parts of the code were made faster (unit checking) through Memoize.

Fixed a bug which prevented links with thickness=1 from being shown
with correct transparency.

Fixed bug in which color errors were produced when PNG file was not
asked for.

Fixed bug in which opacity of links in SVG files was missing in some
cases.

Fixed but which was assigning the wrong color to transparent links in
PNG files in some cases.

Centralized color configuration files colors.conf now includes
brewer.conf - no need to include brewer.conf separately.

Added -paranoid flag to exit on warnings.

Made error messages friendlier. Revamped internal error handling
mechanism.

Added data_path to allow adding to locations where files are searched
for.

Added data/ to prefix path of locations where files are searched for.
Now the default search locations are

 { CWD | CIRCOS_PATH } + { . | .. | ../.. } + { . | etc | data }

--
You received this message because you are subscribed to the Google Groups "Circos" group.

Sunday, 5 February 2012

Hive Plots - Linear Layout for Network Visualization - Visually Interpreting Network Structure and Content Made Possible

http://mkweb.bcgsc.ca/linnet/
Krzywinski M, Birol I, Jones S, Marra M (2011). Hive Plots — Rational Approach to Visualizing Networks. Briefings in Bioinformatics (early access 9 December 2011, doi: 10.1093/bib/bbr069). (download citation)

Hive plots are useful, geeky and simple to implement. Go nuts.  [ Hive Plots - Rational Network Visualization - A Simple, Informative and Pretty Linear Layout for Network Analytics - Martin Krzywinski ]

the hive plot is a perceptually uniform and scalable linear layout visualization for network visual analytics 


UNDERSTANDING NETWORK STRUCTURE WITH HIVE PLOTS.

(A) Normalized (top) and absolute (bottom) connectivity of E. coli gene regulatory network and Linux function call network (Yan et al.)

(B) Gene co-regulation networks in neuroblastoma samples.

(C) Network edges shown as ribbons creating circularly composited stacked bar plots (a periodic steamgraph).

(D) Syntenic network of three modern crucifer species to ancestral genome.

(E) Layered network correlation matrix. In each cell two layers u,v are depicted with u used to order axes and nodes while links for v are shown.


Hive plots — for the impatient

The hive plot is a rational visualization method for drawing networks. Nodes are mapped to and positioned on radially distributed linear axes — this mapping is based on network structural properties. Edges are drawn as curved links. Simple and interpretable.

The purpose of the hive plot is to establish a new baseline for visualization of large networks — a method that is both general and tunable and useful as a starting point in visually exploring network structure.

 [ Hive Plots - Rational Network Visualization - A Simple, Informative and Pretty Linear Layout for Network Analytics - Martin Krzywinski ]

Hive plots give the reader a passing chance to quantitatively understand important aspects of a network's structure. Unlike hairballs, hive plots are excellent at managing the visual complexity arising from large number of edges and exposing both trends and outlier patterns in network structure.



only for networks?

No.

Hive plots can be applied to data structures other than networks. The method requires that your data be mappable onto a set of pairwise relationships. For networks, this pairwise relationship is the edge between two nodes. In other circumstances, it can relate two spatial positions (where the axis corresponds to an object with a physical length scale) or two intervals (two axis segments are related, thereby creating a ratio comparison).


 [ Hive Plots - Rational Network Visualization - A Simple, Informative and Pretty Linear Layout for Network Analytics - Martin Krzywinski ] - Assessing genome assembly quality with a hive plot, which compares reads, assembly and reference.

This hive plot provides a visual recipe for assessing the quality of a genomic assembly. An assembly is composed of reads (bottom axis), which are assembled into contigs (right axis). Independently, a reference assembly (left axis) may exist and act as a comparator. Among others, this hive plot answers the following questions

  • what fraction of reads are unassembled? 20%
  • what fraction of reads are unaligned to reference? 30%
  • what fraction of reference has no read coverage? 2%
  • what fraction of reference has no contig coverage? 15%
  • what fraction of reference is constructed by contigs < 200kb? 60%
  • are there contigs > 200kb? no.
  • what fraction of contigs are unaligned to the reference? 20%
  • what fraction of the overall assembly is derived from k=27 assembly? 80%

The benefit of this stacked bar plot layout is that the circular layout is both periodic and has visual weight. This approach is similar to a parallel coordinate plot, except here the plot wraps around.

Multiple panels can be combined to display a very large number of ratios. Below are shown 3 x 3 x 3 (27) comparisons, each with 3 x 8 ratios, for a total of 648 ratios. By categorizing each ratio using a spectral color scheme, patterns can be quickly spotted and interpreted. The image below was created for the EMBO Journal 2011 cover contest.

Friday, 3 February 2012

R-SAP: a multi-threading computational pipeline for the characterization of high-throughput RNA-sequencing data

http://nar.oxfordjournals.org/content/early/2012/01/28/nar.gks047.long
The rapid expansion in the quantity and quality of RNA-Seq data requires the development of sophisticated high-performance bioinformatics tools capable of rapidly transforming this data into meaningful information that is easily interpretable by biologists. Currently available analysis tools are often not easily installed by the general biologist and most of them lack inherent parallel processing capabilities widely recognized as an essential feature of next-generation bioinformatics tools. We present here a user-friendly and fully automated RNA-Seq analysis pipeline (R-SAP) with built-in multi-threading capability to analyze and quantitate high-throughput RNA-Seq datasets. R-SAP follows a hierarchical decision making procedure to accurately characterize various classes of transcripts and achieves a near linear decrease in data processing time as a result of increased multi-threading. In addition, RNA expression level estimates obtained using R-SAP display high concordance with levels measured by microarrays.

To initiate analyses using R-SAP, the user provides two required inputs for the pipeline: the sequence alignment file and known transcripts’ coordinate file. Currently R-SAP accepts alignment files only in psl format that are generated by mapping RNA-Seq reads to the reference genome using BLAT (Blast like alignment tool) (21) or SSAHA2 (Sequence search and alignment by hashing algorithm) (22). RNA-Seq reads mapping to the genome may result in the alignments scattered across multiple exons separated by introns. We chose psl as the alignment format for the pipeline because the scattered alignments are precisely stitched together and reported as a large single alignment. As a result, for each sequencing read the most likely alignment and corresponding genomic locus can be readily found in the alignment files. Moreover, the psl format preserves the orientation of alignment blocks originating from the contiguous genomic loci enabling their accurate re-mapping to the annotated exons and determination of associated reference structural variants.
R-SAP is also configured to work with two of the currently available transcript assemblers: Cufflinks (23) and Scripture (24). Assembled transcripts can be supplied to R-SAP either in GTF (Gene Transfer Format) or in BED (Browser Extensible Data) format. GTF and BED are default output formats from Cufflinks and Scripture respectively. 

  1. Nucl. Acids Res. (2012) doi: 10.1093/nar/gks047
This article is Open Access

Performance comparison and evaluation of software tools for microRNA deep-sequencing data analysis

http://nar.oxfordjournals.org/content/early/2012/01/28/nar.gks043.long

With the development of next-generation sequencing (NGS) techniques, many software tools have emerged for the discovery of novel microRNAs (miRNAs) and for analyzing the miRNAs expression profiles. An overall evaluation of these diverse software tools is lacking. In this study, we evaluated eight software tools based on their common feature and key algorithms. Three deep-sequencing data sets were collected from different species and used to assess the computational time, sensitivity and accuracy of detecting known miRNAs as well as their capacity for predicting novel miRNAs. Our results provide useful information for researchers to facilitate their selection of the optimal software tools for miRNA analysis depending on their specific requirements, i.e. novel miRNAs discovery or miRNA expression profile analysis of sequencing data sets.


Are we slaves to the scientific publishing industry?

Found this fascinating analogy in the post http://conservationbytes.com/2012/01/29/knowledge-slavery/

...our slavery to the scientific publishing industry.
And 'slavery' is definitely the most appropriate term here, for how else would you describe a business where the product is produced by others for free1 (scientific results), is assessed for quality by others for free (reviewing), is commissioned, overviewed and selected by yet others for free (editing), and then sold back to the very same scientists and the rest of the world's consumers at exorbitant prices.
This isn't just a whinge about a specialised and economically irrelevant sector of the economy, we're talking about an industry worth hundreds of billions of dollars annually. In fact, Elsevier (agreed by many to be the leader in the greed-pack – see how some scientists are staging their protest; also here) made US$1.1 billion in 2010!

what are your thoughts? I don't think that making money is necessarily bad, but apparently rich is even a derogatory term now, with rich ppl (oops) preferring to be called high net worth individuals. But when it's larger entities, corporations, it might be easy to mud sling them.
However, I think the author has a point when making the analogy, my 1st shock was discovering that I have to read lengthy copyright info on what I can or cannot do with something that I researched / wrote  (see http://www.elsevier.com/wps/find/authorsview.authors/rights )

It's so much easier to understand from a photographer's point of view
http://www.photosecrets.com/photography-law-copyright

"Copyright" is "the right to copy." This right is a legal construct, designed for you — the artist — to support your artistic endeavors. Without copyright, people would be free to use your artistic work
You can negotiate a "license" to copy, and perhaps even get paid in real money. Hopefully this will give you more incentive to create art, and the world will be a better place.

Will People Steal My Work?

Generally no, as publishers live by copyright law and usually have established rates which they gladly pay. A more likely problem is that publishers may not know that you are the copyright owner, which goes back to that "©" symbol and digital watermark.

Hey wait a minute .. "the world will be a better place "  hmmm I thought scientists are the ones trying to help the world with advancing human knowledge where our incentive to churn out more scientific results?

hmmm food for thought ..
Post comments!

Datanami, Woe be me