Storage Admin HubConfiguration · Protection · Recovery
Operations · ONTAP 9

Daily and weekly health operations

An auditable routine for hardware, capacity, protocols, and protection.

Scenario

Run this at a consistent time and save outputs with the cluster name and timestamp. Compare against baseline, not merely whether a command returned output.

Before you change production

Commands and screens can differ by release and platform. Replace example names and documentation IP addresses. Check prerequisites, impact, current health and rollback with your change owner.

Daily: health and service

  1. Review unresolved system health and EMS alerts, node and HA state, aggregate and volume headroom, data LIF status, SnapMirror health and lag.
  2. Check application tickets and any AutoSupport or monitoring alarms. Investigate new critical errors rather than clearing them to make the dashboard green.
Read-only daily checklist
cluster show
storage failover show
system health alert show
event log show -severity ERROR
storage aggregate show
volume show
network interface show
snapmirror show

Weekly: trends and recoverability

  1. Compare capacity and latency to last week. Review rapidly growing volumes, snapshots, replication lag, failed transfers, and a sample of host path health.
  2. Validate a recovery point by restoring a test file or isolated clone. Review firmware, ONTAP support notices, expiring certificates, and pending changes.
Verify

No unresolved priority alert, capacity trend has an owner, and the recovery test is documented.

Escalation record

  1. For any incident capture exact time, client/host, SVM, volume or LUN, error, recent changes, commands and outputs, impacted users, current workaround, and next owner.
  2. Avoid attaching secrets, private client data, or full logs to a public site.
Verify

A second administrator can reproduce the investigation from the incident record.

If validation fails

  1. A new EMS alert with no obvious client impact still needs a named owner and deadline; compare to the preceding day and recent maintenance.
  2. A healthy SnapMirror status can hide excessive lag. Compare actual lag against the application RPO and review transfer history.
  3. A test restore failure is an operational incident. Record the recovery point, error and alternate copy, then retest after correction.
Verify

Re-run the original validation and record the observed result, exact error, time, and corrective action.

NetApp references