Scenario
MetroCluster has its own IP/FC topology and recovery procedures. This checklist is for assessment and coordination; use the matching NetApp procedure for the exact configuration.
Before you change production
Commands and screens can differ by release and platform. Replace example names and documentation IP addresses. Check prerequisites, impact, current health and rollback with your change owner.
1. Identify configuration and failure
- Confirm MetroCluster IP or FC, node count, mediator/tiebreaker, local HA, intersite connectivity, storage health and the exact failed components.
- For planned maintenance, perform the supported simulation and prechecks. For a real disaster, establish whether the remote site can still write and follow the forced/negotiated switchover decision tree.
Verify
An approved site state and switchover type are recorded.
2. Switchover and service validation
- Execute the platform-specific switchover procedure, monitoring each stage. Never run forced switchover merely because an alert appeared.
- Validate SVMs, LIF placement, SAN paths, client access, application startup order and data consistency at the surviving site.
Verify
Production is served by the expected site and no conflicting active copy exists.
3. Heal and switch back
- Repair the failed site and follow the documented healing sequence. Run the supported readiness/simulation checks before switchback.
- Switch back in an approved window. Confirm both sites, aggregates, LIFs and application access return to expected states. Escalate if LIF placement or storage state is wrong.
Verify
MetroCluster is back to normal and all health checks pass.
If validation fails
- If the configuration is not ready for switchover, read the simulated check and health alerts; do not bypass a veto without the documented recovery procedure.
- If LIFs land on unexpected ports after switchover, pause client cutover and compare network topology with the planned DR design.
- If healing or switchback fails, preserve status and logs, keep the surviving site stable and escalate according to the platform guide.
Verify
Re-run the original validation and record the observed result, exact error, time, and corrective action.