B) Factores de Inestabilidad de los Peligros Geológicos
2.2.3. Zonificación de la amenaza y susceptibilidad por MM
Autonomic failover is a unique enterprise class feature of the HP B6200 StoreOnce Backup system.
When integrated with various backup applications it makes it possible for the backup process to continue even if a node within a B6200 couplet fails. ISV scripts are usually required to complete this process. The failover process is best visualized by watching the video on:
http://www.youtube.com/watch?v=p9A3Ql1-BBs
What happens during autonomic failover?
At a logical level, all the virtual devices (VTL, NAS and replication) associated with the failing node are transferred by the B6200 operating system onto the paired healthy node of the couplet. The use of Virtual IP addresses for Ethernet and NPIV virtualization on the Fibre Channel ports are the key technology enablers that allow this to happen without manual intervention.
NAS target failover is via the Virtual IP system used in the HP B6200 Backup System – the service set simply presents the failed node Virtual IP address on the remaining node.
FC (VTL device) failover relies on the customer’s fabric switches supporting NPIV, and NPIV being enabled and the zones set up correctly. Here the situation is more complex as several permutations are possible. For more details see Appendix A – FC failover supported configurations.
Note: To prevent data corruption, the system must confirm that the failing node is shutdown before the other node starts writing to disk. This can be seen in the video where the “service set” is stopping.
At a hardware level the active cluster manager is sending a shutdown command via the dedicated iLO3 port on the failing node. Email alerts and SNMP traps are also sent on node failure.
The HP B6200 Backup System failover process can take approximately 15 minutes to complete. The following figure illustrates the failover timeline.
Figure 27: Failover timeline
55
Failover support with backup applications
Backup applications do not have an awareness of advanced features such as autonomic failover because they are designed for use with physical tape libraries and NAS storage. From the
perspective of the backup application, when failover occurs, the virtual tape libraries and the NAS shares on the HP B6200 Backup System go offline and after a period of time they come back online again. This is similar to a scenario where the backup device has been powered off and powered on again.
Each backup application deals with backup devices going offline differently. In some cases, once a backup device goes offline the backup application will keep retrying until the target backup device comes back online and the backup job can be completed. In other cases, once a backup device goes offline it must be brought back online again manually within the backup application before it can be used to retry the failed backups.
In this section we shall briefly describe three popular backup applications and their integration with the autonomic failover feature. Information for additional backup applications will be published on the B6200 support documentation pages when it is available.
HP Data Protector 6.21: job retries are currently supported by using a post-exec script.
Download from B6200 support documentation.
Symantec NetBackup 7.x: job retries are automatic, but after a period without a response from the backup device the software marks the devices as “down”. Once failover has completed and the backup device is responding again the software does not
automatically mark the device as “up” again. A script is available from HP that continually checks Symantec device status and ensures that backup devices are marked as “up”. With this script deployed on the NetBackup media server, the HP B6200 Backup System failover works seamlessly. Download from B6200 support documentation. NetBackup can go back to the last checkpoint and carry on from there, if checkpointing has been enabled in the backup job. So, all the data backed up prior to failover is preserved and the job does not have to go right back to the beginning and start again.
EMC Networker 7.x:
VTL: Job retries are automatically enabled for scheduled backup jobs. No additional scripts or configuration are required in order to achieve seamless integration with the HP B6200 Backup System. In the event of a failover scenario, the backup jobs are automatically retried once the HP B6200 Backup System has completed the failover process. EMC Networker also has a checkpoint facility that can be enabled. This allows failed backup jobs to be restarted from the most recent checkpoint.
NAS: The combination of Networker and NAS is not supported with autonomic failover and use could cause irrecoverable data loss.
It is strongly recommended that all backup jobs to all nodes be configured to restart (if any action to do this is required) because there is no guarantee which nodes are more likely to fail than others. It is best to cover all eventualities by ensuring all backups to all nodes have restart capability enabled, if required.
Whilst the failover process is autonomic, the failback process is manual because the replacement or repaired node must be brought back on line before failback can happen. Failback can be
implemented either from the CLI or the GUI interface.
56
Restores are generally a manual process and restore jobs are typically not automatically retried because they are rarely scheduled.Designing for failover
One node is effectively doing the work of two nodes in the failed over condition. There is some performance degradation but the backup jobs will continue after the autonomic failover.
The following best practices apply when designing for autonomic failover support:
The customer must choose whether SLAs will remain the same after failover as they did before failover. If they do, the solution must be sized in advance to only use up to 50% of the available performance. This is to ensure that there is sufficient headroom in system resources so that in the case of failover there is no appreciable degradation in performance after failover and the SLAs are still met.
For customers who are more price-conscious and where failover is an “exception condition”
the solution can be sized for cost effectiveness. Here most of the available throughput is utilized on the nodes. In this case when failover happens there will be a degradation in performance. The amount of degradation observed will depend on the relative “imbalance”
of throughput requirements between the two nodes. This is another reason for keeping both nodes in a couplet as evenly loaded as possible.
Ensure the correct ISV patches/scripts are applied and do a dry run to test the solution. In some cases a post execution script must be added to each and every backup job/policy. The customer can configure which jobs will retry in the event of failover (which is a temporary condition) in order to limit the load on the single remaining node in the couplet by:-
o Only putting the post execution script to retry the job in the most urgent and important jobs, not all jobs. This is the method for HP Data Protector.
o Modifying the “bring device back on line scripts” to only apply to certain drives and robots – those used by the most urgent and important jobs. This is the method for Symantec NetBackup.
Remember replication is also considered as a virtual device within a service set and replication fails over as well as backup devices
For replication failover there are two scenarios:
o Replication was not running – that is, failover occurred outside the replication window, in which case replication will start when the replication windows is next open.
o If replication was in progress when failover occurred, after failover has completed replication will start again from the last known good checkpoint (about every 10MB of replicated data).
Failback (via CLI or GUI) is a manual process and should be scheduled to occur during a period of inactivity.
Remember all failover related events are recorded in the Event Logs.