Extending science gateway frameworks to support Big Data applications in the cloud
Ver/ Abrir
Registro completo
Mostrar el registro completo DCFecha
2016-12Derechos
Attribution 4.0 International
Publicado en
Journal of Grid Computing, 2016, 14(4), 589-601
Editorial
Springer Nature
Enlace a la publicación
Palabras clave
Big data
Hadoop
MapReduce
Science gateway
WS-PGRADE
Workflow
Resumen/Abstract
Cloud computing offers massive scalability and elasticity required by many scientific and commercial applications. Combining the computational and data handling capabilities of clouds with parallel processing also has the potential to tackle Big Data problems efficiently. Science gateway frameworks and workflow systems enable application developers to implement complex applications and make these available for end-users via simple graphical user interfaces. The integration of such frameworks with Big Data processing tools on the cloud opens new opportunities for application developers. This paper investigates how workflow systems and science gateways can be extended with Big Data processing capabilities. A generic approach based on infrastructure aware workflows is suggested and a proof of concept is implemented based on the WS-PGRADE/gUSE science gateway framework and its integration with the Hadoop parallel data processing solution based on the MapReduce paradigm in the cloud. The provided analysis demonstrates that the methods described to integrate Big Data processing with workflows and science gateways work well in different cloud infrastructures and application scenarios, and can be used to create massively parallel applications for scientific analysis of Big Data.
Colecciones a las que pertenece
- D20 Artículos [468]