Tuesday, 14 February 2012

FAQ Admin/Data Libraries/Uploading Library Files local instance - Galaxy Wiki

This is a FAQ that's best answered in the Wiki page
the most intuitive is to use the web GUI to upload but this creates unnecessary overhead and slowdown when you can point to the files 

Options for Uploading Files from the Admin Perspective

There are currently four options available to a Galaxy admin user for uploading files to a data library. Some of these same options are available to all regular users that have been granted permission to add items to a Data Library or folder, but this section describes the features from the Galaxy admin perspective which is accessed by clicking on the "Admin" link in the top Galaxy menu bar.
http://wiki.g2.bx.psu.edu/Admin/Data%20Libraries/Uploading%20Library%20Files

Monday, 13 February 2012

Get to Know Btrfs | Linux.com

https://www.linux.com/learn/tutorials/533112-weekend-project-get-to-know-btrfs

As expected, Storage Spaces will indeed be a feature of both desktop and server editions of the operating system.
If the feature does indeed ship in desktop Windows, it will overnight obsolete a range of SOHO-oriented storage systems; products like Drobo and ReadyNAS will find it hard to survive in a Windows 8 world.

Btrfs Features

  • RAID 0, 1, 10
  • COW
  • Incremental backup
  • Online defrag
  • gzip and LZO compression
  • Space-efficient packing of small files
  • Dynamic inode allocation
  • Checksums on data and metadata
  • Shrink and grow storage volumes
  • Extents
  • Snapshots
  • 16 EiB maximum file size
Planned features include RAID 5 and 6, deduplication, and a ready-for-primetime filesystem checker, btrfsck. You can try out btrfsck now because it is included in btrfsprogs. (Which of course Debian/Ubuntu/Mint etc. changes to btrfs-tools, and Fedora calls it btrfs-progs.) But it is not ready for production systems yet.
Putting the finishing touches on btrfsck is the last big step before Oracle makes it the default filesystem in their next Unbreakable Linux release. Fedora 16 Linux was supposed to default to Btrfs, but now they're aiming for Fedora 17 in May 2012.


also see https://btrfs.wiki.kernel.org/

Windows 8 Storage Spaces detailed: pooling redundant disk space for all

This will definitely change mainstream consumer ideas on backup and data storage .. maybe it will ripple down to labs who currently do not backup their data ..


Unlike RAID systems of old, but in common with other modern storage technologies such as Solaris' ZFS and Linux's btrfs, pools can use disks of different interface technologies—USB, SATA, Serial Attached SCSI—and different, mismatched sizes. New disks can be added to a pool at any time. Pools can also include one or more hot spares: drives allocated to a pool but kept in standby until another disk in the pool fails, at which point they spring into life.

As expected, Storage Spaces will indeed be a feature of both desktop and server editions of the operating system.
If the feature does indeed ship in desktop Windows, it will overnight obsolete a range of SOHO-oriented storage systems; products like Drobo and ReadyNAS will find it hard to survive in a Windows 8 world.

Thursday, 9 February 2012

Unix join on more than two files - Stack Overflow

http://stackoverflow.com/questions/9212893/unix-join-on-more-than-two-files

This has to be useful someday ...

ea-utils - FASTQ processing utilities - Google Project Hosting

http://code.google.com/p/ea-utils/

Suite of processing tools for sequencing output. Barcode demultiplexing, adapter trimming, etc.
Primarily written to support an Illumina based pipeline - but should work with any FASTQs.

Overview:

  • fastq-mcf
  • Scans a sequence file for adapters, and, based on a log-scaled threshold, determines a set of clipping parameters and performs clipping. Also does skewing detection and quality filtering.
  • fastq-multx
  • Demultiplexes a fastq. Capable of auto-determining barcode id's based on a master set fields. Keeps multiple reads in-sync during demultiplexing. Can verify that the reads are in-sync as well, and fail if they're not.
  • fastq-join
  • Similar to audy's stitch program, but in C, more efficient and supports some automatic benchmarking and tuning. It uses the same "squared distance for anchored alignment" as other tools.

Other Stuff:

  • sam-stats - Basic sam/bam stats. Like other tools, but produces what I want to look at, in a format suitable for passing to other programs. (Click for source)
  • fastq-stats - Basic fastq stats. Counts duplicates. Option for per-cycle stats, or not (irrelevant for many sequencers). (Click for source)
  • determine-phred - Returns the phred scale of the input file. Works with sams, fastq's or pileups and gzipped files.
  • Chrdex.pm - indexes a delimited file by chromosome start/stop. There are lots of tools for this. This one works pretty well if you're a perl user. It handles overlapping regions reasonably well. It uses RAM comparable to the size of the annotation file.
  • Sqldex.pm - just like Chrdex.pm, except uses a disk-based btree. Not as fast, but close, and uses very little RAM.
  • qsh - Runs a bash script file like a "cluster aware makefile"...only processing newer things, die'ing if things go wrong, and sending jobs to a queue manager if they're big. That way you don't have to write makefiles, or wrap things in "qsub" calls for every little program. Not really ready yet.
  • grun - Fast, lightweight grid queue software. Keeps the job queue on disk at all times. Very fast. Works well by now
  • gwrap - Bash wrapper shell that downloads all dependencies that are not the local system.... good for EC2 nodes. Linux only. Will use it if we ever go to EC2.

Tuesday, 7 February 2012

Amazon S3 Lowers Standard Storage Prices

Dear Amazon S3 Customer,

We are excited to announce that we have reduced the Amazon S3 standard storage prices in all regions. With this price change, all Amazon S3 standard storage customers will see a reduction in their storage costs. For instance, if you store 50 TB of data on average, you'll see a 12% reduction in costs. If you store 500 TB of data on average, you'll see a 13.5% reduction in costs. The price reduction applies to all of your standard storage- both existing storage and all new storage you add. Here is a summary of price changes for the US Standard region:

                          Old         New
First 1TB           $0.140    $0.125
Next 49TB         $0.125    $0.110
Next 450TB       $0.110    $0.095
Next 500TB       $0.095    $0.090
Next 4000TB     $0.080    $0.080 (no change)
Over 5000TB     $0.055    $0.055 (no change)

The new lower prices for all regions can be found on the Amazon S3 Detail Page. New prices are effective February 1st and will be applied to your bill for all storage on or after this date.

We are happy to pass along these savings to you as we continue to innovate and drive down our costs.

Sincerely,
The Amazon S3 Team

We hope you enjoyed receiving this message. If you wish to remove yourself from receiving future product announcements and the monthly AWS Newsletter, please update your communication preferences.

Amazon Web Services LLC is a subsidiary of Amazon.com, Inc. Amazon.com is a registered trademark of Amazon.com, Inc. This message produced and distributed by Amazon Web Services, LLC, 1918 8th Avenue, Seattle, WA 98101.

Datanami, Woe be me