跳转到主内容

引用已删除后端的 Trident 孤立卷导致工作节点 CPU 使用率过高

Views:
94
Visibility:
Public
Votes:
0
Category:
trident-openshift
Specialty:
snapx
Last Updated:

适用于

  • NetApp Astra Trident CSI 驱动程序
  • Kubernetes / OpenShift 环境
  • ONTAP SAN/NAS 后端

问题

运行 Trident CSI 控制器 Pod 的工作节点表现出持续的高 CPU 利用率(高达 100%)。 

导致:

  • 集群范围的性能下降
  • 跨多个 Pod 的应用程序中断
  • 配置和删除工作流中的延迟增加
观察到的日志模式:

从Trident控制器日志:

level=error msg="Invalid backend state." expectedState=online/deleting state=failed workflow="core=node_reconcile"
msg="Unable to delete snapshot from backend."
error="backend <backend-name> is not Online or Deleting"
msg="Unable to delete volume from backend."
error="backend <backend-name> is not Online or Deleting"
  • 这些错误连续重复(高频重试)
  • 控制器进入持久协调循环

Sign in to view the entire content of this KB article.

New to NetApp?

Learn more about our award-winning Support

NetApp provides no representations or warranties regarding the accuracy or reliability or serviceability of any information or recommendations provided in this publication or with respect to any results that may be obtained by the use of the information or observance of any recommendations provided herein. The information in this document is distributed AS IS and the use of this information or the implementation of any recommendations or techniques herein is a customer's responsibility and depends on the customer's ability to evaluate and integrate them into the customer's operational environment. This document and the information contained herein may be used solely in connection with the NetApp products discussed in this document.