EC 重新平衡失败,服务不可用。联系 EC Job Manager 时出错
适用于
问题描述
EC-rebalance 报告: Service unavailable. Error contacting EC Job Manager. 一个或多个节点可能在 GUI 中处于未知状态,但实际上不会报告任何错误。节点最终将自行恢复在线状态。
Bycast.log 来自 EC 负责人的报告:
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740693 ECJM ???? 2025-02-04T12:29:51.840953| NOTICE 0040 54375572ef20a272 ECJM: Starting job 1581323391038607979: Site Rebalance - Group ID 10.
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740693 ECJM ^RDY 2025-02-04T12:29:51.841797| NOTICE 0106 54375572ef20a272 ECJM: Job status of job 1581323391038607979 is JOBSTATUS_IN_PROGRESS
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740693 ECJM ^RDY 2025-02-04T12:29:51.841824| NOTICE 0112 54375572ef20a272 ECJM: Resuming job 1581323391038607979
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740693 ECJM ^RDY 2025-02-04T12:29:51.841833| NOTICE 0219 54375572ef20a272 ECJM: 1581323391038607979(rebalance 10): Resuming
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740693 ECJM %MDW 2025-02-04T12:29:51.850301| WARNING 1005 54375572ef20a272 ECJM: Volume Info request timed out.
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740693 ECJM %MDW 2025-02-04T12:29:51.850336| NOTICE 0994 54375572ef20a272 ECJM: 1581323391038607979(rebalance 10): Cannot determine if there are Offline Volumes in the Grid.
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740693 ECJM %MDW 2025-02-04T12:29:51.850347| NOTICE 1271 54375572ef20a272 ECJM: 1581323391038607979(rebalance 10): saving state. status: JOBSTATUS_PAUSED
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740693 ECJM ^RDY 2025-02-04T12:29:51.852021| NOTICE 1057 54375572ef20a272 ECJM: 1581323391038607979(rebalance 10): Stopping child jobs.
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740693 ECJM ^RDY 2025-02-04T12:29:51.852132| WARNING 0062 54375572ef20a272 ECJM: Caught exception 'Failed to ensure all volumes are online. pausing job...' when running job 1581323391038607979: Site Rebalance - Group ID 10.
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740691 ECJM _DON 2025-02-04T12:29:51.852204| NOTICE 0934 54375572ef20a272 ECJM: Received job completion message.
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740691 ECJM _DON 2025-02-04T12:29:51.852230| NOTICE 0940 54375572ef20a272 ECJM: Job 1581323391038607979 completed with result GERR.
Feb 4 12:29:51 <nodename> ADE: |21835099 0110740693 ECJM ^RDY 2025-02-04T12:29:51.852232| ERROR 1081 54375572ef20a272 PROC: Exception: Dynamic exception type: std::runtime_error#012std::exception::what: Failed to ensure all volumes are online. pausing job...#012
bycast-err.log 可能会报告:ERROR Internal server error. The server encountered an error and could not complete your request. Try again. If the problem persists, contact support. EC job manager unavailable. (MgmtApi::LocalizedRuntimeError)ERROR /usr/local/lib/site_ruby/mgmt-api/rest-client/resource.rb:833:in `handle_errors!'ERROR /usr/local/lib/site_ruby/mgmt-api/rest-client/resource.rb:486:in `all'ERROR /usr/local/lib/site_ruby/mgmt-api/data-recovery/data-recovery.rb:696:in `get_status'ERROR Failed to retrieve erasure-coded repair status, EC job manager might be down.ERROR Failed to retry erasure-coded repair due to EC job manager unavailable.ERROR DataRecoveryManager failed to retry EC repairERROR Failed to retry repair: EC job manager unavailable. (MgmtApi::LocalizedRuntimeError)