Scenario
Example: node B has taken over node A. A takeover can keep data available, but the reason must be understood before giveback.
Before you change production
Commands and screens can differ by release and platform. Replace example names and documentation IP addresses. Check prerequisites, impact, current health and rollback with your change owner.
1. Establish the state
- Capture node health, HA state, current aggregate ownership, LIF placement, affected applications, EMS timeline and planned maintenance.
- Determine whether the takeover was operator-initiated, automatic due to a fault, or incomplete. Check power, interconnect, storage and network alerts.
Read-only investigation
cluster show
storage failover show
system health alert show
event log show -severity ERROR2. Decide the next action
- If partner hardware is faulted or aggregates are not healthy, escalate and repair before giveback. Do not force a giveback to clear an alert.
- If the cause is resolved, follow the exact platform and ONTAP giveback procedure in an approved window. Monitor client paths and LIFs during the transition.
Verify
Partner eligible, underlying fault resolved, and recovery procedure approved.
3. Validate recovery
- Check both nodes and HA readiness, aggregate and LIF placement, volumes, SAN paths, application transactions, and new EMS messages.
- Record the initiating event and preventive change.
Verify
Both nodes healthy and HA ready; client access and replication normal.
If validation fails
- A takeover that succeeded does not prove the failed node is safe to return. Check the initiating hardware and storage events first.
- If giveback is vetoed, record the exact veto and use the vendor procedure for that cause. Forcing it can reintroduce a fault.
- After giveback, compare client access and SAN paths with the pre-incident baseline, then observe for repeated alerts.
Verify
Re-run the original validation and record the observed result, exact error, time, and corrective action.