Taylor, Ronald C.

doi:10.1186/1471-2105-11-S12-S1

Back to matches

Your institution may have access to this item. Find your institution then sign in to continue.

Title: An overview of the Hadoop/MapReduce/HBase framework and its current applications in bioinformatics.
Authors: Taylor, Ronald C.
Abstract: Background: Bioinformatics researchers are now confronted with analysis of ultra large-scale data sets, a problem that will only increase at an alarming rate in coming years. Recent developments in open source software, that is, the Hadoop project and associated software, provide a foundation for scaling to petabyte scale data warehouses on Linux clusters, providing fault-tolerant parallelized analysis on such data using a programming style named MapReduce. Description: An overview is given of the current usage within the bioinformatics community of Hadoop, a top-level Apache Software Foundation project, and of associated open source software projects. The concepts behind Hadoop and the associated HBase project are defined, and current bioinformatics software that employ Hadoop is described. The focus is on next-generation sequencing, as the leading application area to date. Conclusions: Hadoop and the MapReduce programming paradigm already have a substantial base in the bioinformatics community, especially in the field of next-generation sequencing analysis, and such use is increasing. This is due to the cost-effectiveness of Hadoop-based analysis on commodity Linux clusters, and in the cloud via data upload to cloud vendors who have implemented Hadoop/HBase; and due to the effectiveness and ease-of-use of the MapReduce method in parallelization of many data analysis algorithms.
Subjects: BIOINFORMATICS; OPEN source software; LINUX operating systems; ALGORITHMS; SOFTWARE sequencers
Publication: BMC Bioinformatics, 2010, Vol 11, p1
ISSN: 1471-2105
Publication type: Article
DOI: 10.1186/1471-2105-11-S12-S1

We found a match

An overview of the Hadoop/MapReduce/HBase framework and its current applications in bioinformatics.

Taylor, Ronald C.

BIOINFORMATICS; OPEN source software; LINUX operating systems; ALGORITHMS; SOFTWARE sequencers

BMC Bioinformatics, 2010, Vol 11, p1

1471-2105

Article

10.1186/1471-2105-11-S12-S1