• No se han encontrado resultados

Instrument technologies for making science observations have evolved considerably with data sets now in the petabyte range for many of the science disciplines that NASA and other researchers are working with. The scale and complexity of current and future science data changes the nature of data processing, management, sharing, analyzing, and archiving extremely large-scale science data sets. This becomes one of the grand challenges of the twenty-fi rst century. Cloud computing has emerged as one of the decade’s chief technologies that can help meet these new compute- and data-intensive science challenges.

Today, within science, there is a trend toward ensuring that high-capacity science data pipelines are established to produce science data products as effi ciently and cost effectively as possible. Pipelines are used to take data, acquired from scientifi c obser- vations, and generate data in structures that are useful for scientifi c analysis. These pipelines employ complex algorithms that can often be compute-intensive. Many of these pipelines produce intermediate products that must be captured. In addition, reprocessing and rerunning data through pipelines generates newer data. Because of the importance of capturing the provenance of the data, multiple versions of science data are routinely kept. The end result is that massive amounts of data and computa- tion must be made available for these science disciplines. Due to the disparate nature of how these projects are funded, they are often housed and developed in isolation of one another. Local laboratories and ground systems both acquire dedicated systems to support these compute-intensive jobs and data storage needs.

Capturing and archiving data for the long term is also critical for preserving scientifi c knowledge and distributing that knowledge to the worldwide science com- munity. In many scientifi c disciplines, archives have been doubling in size about every 5 years. Local laboratories and institutions that have routinely instantiated their own storage to archive data need to scale up to construct petabyte scale archives. In addition, addressing issues of data redundancy, integrity, and security is critical to providing a robust storage infrastructure.

These two scenarios are well suited for the cloud environment. Many of these systems are highly distributed where both storage and computation can be resourced externally. However, the movement of massive data, the provisioning of reliable services, and the policies and security related to access to data are issues that must be resolved when integrating cloud services and technologies into what we are now building as an emerging scientifi c data ecosystem of distributed software services and data.

One of the most attractive aspects of cloud computing is the elastic nature which allows for services to be provisioned “on demand.” Many scientifi c projects have differing requirements and needs relative to the use of operational capabilities the cloud can offer. Instantiating computation and data internally by each project often leaves quite a bit of untapped computing power. Using Earth scientifi c research as an example, often data processing and image data tiling for effi cient retrieval,

involves advanced computational algorithms and parallelism that demand high- performance computing resources such as supercomputers and computer clusters that are expensive to establish and maintain. Many of these jobs require processing systems that approach the terafl ops region to support the computational demands. Within airborne science missions, many of the compute-intensive jobs are only performed on demand since the airborne campaigns occur infrequently. Sharing of computational resources among airborne science missions is an effective way to achieve better economies of scale. The cloud computing model is ideal for such jobs. Likewise, many scientifi c disciplines have similar needs that can achieve better economies of scale through cloud computing.

In addition to lowering cost and scalable capacity, science data systems can benefi t from other services with cloud computing. The cloud provides a virtual environment that allows users to share both science data and services without knowledge, expertise, nor control over the technology infrastructure that supports them. This promotes collaboration of scientifi c research by making it possible for researchers across laboratories, institutions, and disciplines globally to work together on common projects. Researchers need access to data that can be inte- grated to generate science products for multidisciplinary experiments. This is espe- cially important in data-intensive science, where the power of discovery lies in applying computational approaches to collection of large data sets that need to be sharable in advancing research.

On the other hand, there are barriers to the success of integration of cloud and science. There are a number of challenges that inhibit the leveraging of the full potential that cloud computing promises. These challenges are related to security, data movement, and reliability. Virtualization, cyber, and data security have always been a key concern to most cloud enterprises, and they are just starting to grasp and not fully understand all the issues. This contributes to skepticism of data integrity and privacy since users often do not trust cloud providers’ effectiveness of their security (e.g., hacker attack) and privacy controls (e.g., poor encryption key man- agement, user management).

Another challenge that directly affects science is the ability to move massive data from data providers to the cloud and from the cloud to the users. Science applications (e.g., biomedical research) require frequently upload to or download very large amounts of data to and from the cloud. Often, there are data transfer bottlenecks affecting performance and reliability because of physical networking bandwidth limitation, and no single data movement protocol can support all data movement tasks. Despite the exponential growth in network and hardware speeds, data movement continues to exceed the capability. It is one of the broad areas where research and development are needed to minimize the time spent on data transfer and to bolster the ability to recover from failures and incomplete transfers.

Recent outages at Amazon , Google and Microsoft have shown that cloud services still have glitches and raise reliability concerns. In addition to service disruption, some data were permanently lost. This possibility of data loss and losing access to the data is unsettling for mission-critical and science data, as well as processing

algorithms and research results. It is critical for science data pipelines to take extra measures in order to provide reliable services.

Cloud computing promises to provide more effi ciency and reduced expenditure on IT services to providers of science services. It presents potential opportunities for integration of science and the cloud as well as improving science data pipelines even though there are still some challenges to fostering this new science paradigm on the cloud.

2.3

Scaling Scientifi c Data Systems with Cloud Computing

Documento similar