Storage Admin HubConfiguration · Protection · Recovery
Troubleshooting · ONTAP 9

Investigate high storage latency

Correlate workload, node, volume, host, and network metrics.

Scenario

Example: database writes rose from 3 ms to 40 ms at 14:10. A useful diagnosis compares the same interval on host, storage, and fabric.

Before you change production

Commands and screens can differ by release and platform. Replace example names and documentation IP addresses. Check prerequisites, impact, current health and rollback with your change owner.

1. Capture a baseline comparison

  1. Identify affected application, protocol, host, volume/LUN, start time, IOPS, throughput, latency and errors. Compare a healthy workload on the same cluster and the same workload before the spike.
  2. Check for recent changes: backup, snapshot, failover, new workload, switch change, host rescan, or space pressure.
Verify

Time-aligned metrics from application, host and ONTAP.

2. Separate components

  1. Use System Manager or Unified Manager to inspect latency, operations, throughput and busy resources for the affected workload, node and aggregate.
  2. For NAS, review client network errors, DNS/name-service delay, protocol retransmits and mount behavior. For SAN, inspect host multipath, HBA/NIC errors and fabric congestion.
  3. Correlate with EMS and health alerts. A high total latency value alone does not prove the storage subsystem is the cause.
Read-only context
system health alert show
event log show -severity ERROR
storage aggregate show
network port show

3. Validate the hypothesis

  1. Change only one suspected bottleneck under change control. Compare the same workload and interval after the change.
  2. Escalate with a timeline, affected object IDs, metrics and a support bundle when the bottleneck is not isolated.
Verify

Measured improvement against baseline, or a documented unresolved hypothesis with evidence.

If validation fails

  1. If only one host is slow, compare its network errors, multipath and application queue with a healthy host on the same volume or LUN.
  2. If all workloads slow together, inspect shared node, aggregate, network and background operations over the same time window.
  3. If capacity is near full, include physical space and snapshot churn in the performance hypothesis before tuning protocol settings.
Verify

Re-run the original validation and record the observed result, exact error, time, and corrective action.