Kubernetes - K8s 集群的日常巡检流程与核心检查项

👋 大家好,欢迎来到我的技术博客!
📚 在这里,我会分享学习笔记、实战经验与技术思考,力求用简单的方式讲清楚复杂的问题。
🎯 本文将围绕Kubernetes这个话题展开,希望能为你带来一些启发或实用的参考。
🌱 无论你是刚入门的新手,还是正在进阶的开发者,希望你都能有所收获!
文章目录
- Kubernetes - K8s 集群的日常巡检流程与核心检查项 🚀
- 一、为什么需要日常巡检?🚨
- 二、巡检流程总览:五步闭环法 🔄
- 三、核心检查项详解(附 Java 代码示例)💻
- 四、巡检自动化架构设计 🏗️
- 五、巡检报告模板与可视化 📊
- 六、巡检策略优化与持续改进 🔄
- 七、常见误区与避坑指南 ⚠️
- 八、总结:巡检即 DevOps 的灵魂 🧠
- 附录:推荐工具与资源 📚
Kubernetes - K8s 集群的日常巡检流程与核心检查项 🚀
在现代云原生架构中,Kubernetes(简称 K8s)已成为容器编排的事实标准。它赋予企业弹性伸缩、服务自治、故障自愈等强大能力,但也因其复杂性,对运维人员提出了更高的要求。一个稳定、高效、安全的 K8s 集群,离不开系统化、标准化、自动化的日常巡检流程。本文将深入剖析 K8s 集群日常巡检的完整流程,涵盖从节点状态、Pod 健康、网络连通、存储资源到安全合规的全方位检查项,并结合 Java 代码示例展示如何构建自动化巡检工具,辅以 Mermaid 图表直观呈现架构逻辑,帮助你构建一套可落地、可扩展、可监控的集群健康管理体系。
一、为什么需要日常巡检?🚨
Kubernetes 是一个高度动态的系统。Pod 会不断重启、节点会因资源不足被驱逐、网络策略可能被误配置、证书会过期、镜像仓库可能不可达……这些“小问题”如果长期积累,最终会演变成雪崩式故障。
📌 真实案例:某金融平台因未监控 etcd 磁盘使用率,导致集群元数据写入失败,所有新 Pod 无法调度,业务中断 4 小时,损失超百万。
日常巡检不是“可选项”,而是生产环境的底线要求。它能帮助你:
- ✅ 提前发现潜在风险(如节点资源耗尽、证书即将过期)
- ✅ 快速定位故障根源(是网络?是存储?还是调度器?)
- ✅ 满足合规审计(如等保、GDPR、ISO27001)
- ✅ 优化资源成本(识别闲置 Pod、低效部署)
- ✅ 建立运维知识库(记录历史异常,形成 SOP)
💡 建议:每日巡检应控制在 15~30 分钟内完成,自动化工具是关键。人工干预仅用于异常处理。
二、巡检流程总览:五步闭环法 🔄
一个标准的 K8s 集群巡检流程,可归纳为以下五个步骤:
步骤说明:
- 启动巡检任务:通过定时任务(CronJob)、CI/CD 流水线或运维平台触发。
- 收集集群状态数据:调用
kubectl、K8s API 或 Prometheus 指标,获取节点、Pod、Deployment、Service、PV/PVC 等资源状态。 - 分析健康指标与阈值:比对预设阈值(如 CPU > 85%、Pod RestartCount > 5),判断是否异常。
- 生成巡检报告与告警:输出 HTML/JSON 报告,推送企业微信/钉钉/Slack/邮件告警。
- 触发修复或人工干预:自动执行修复脚本(如清理 Terminating Pod),或通知运维人员介入。
- 记录归档并优化策略:将巡检日志存入 ELK 或 Loki,定期分析趋势,优化阈值和检查项。
🔗 Kubernetes Best Practices: Operational Guidelines —— 官方推荐的生产集群运维指南
三、核心检查项详解(附 Java 代码示例)💻
我们将从节点层 → Pod 层 → 网络层 → 存储层 → 安全层 → 控制平面层六大维度展开,每一项均提供可运行的 Java 代码示例,帮助你构建自动化巡检工具。
3.1 节点健康状态检查 🖥️
节点是 K8s 的基石。一个节点宕机或资源耗尽,会影响其上所有 Pod。
检查项:
- 节点状态(Ready/NotReady)
- CPU/内存使用率
- 磁盘使用率(根分区、镜像层、日志)
- kubelet 是否运行
- 标签与污点是否异常
Java 代码示例:使用 Kubernetes Java Client 检查节点状态
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1Node;
import io.kubernetes.client.openapi.models.V1NodeCondition;
import io.kubernetes.client.openapi.models.V1NodeStatus;
import io.kubernetes.client.util.Config;
import java.io.IOException;
import java.util.List;
public class NodeHealthChecker {
public static void main(String[] args) throws IOException {
// 加载 kubeconfig(默认 ~/.kube/config)
ApiClient client = Config.defaultClient();
Configuration.setDefaultApiClient(client);
CoreV1Api api = new CoreV1Api();
List<V1Node> nodes = api.listNode(null, null, null, null, null, null, null, null, null, null).getItems();
System.out.println("🔍 节点健康巡检报告:");
System.out.println("========================");
for (V1Node node : nodes) {
String nodeName = node.getMetadata().getName();
String nodeStatus = node.getStatus().getConditions().stream()
.filter(condition -> "Ready".equals(condition.getType()))
.map(V1NodeCondition::getStatus)
.findFirst()
.orElse("Unknown");
// 检查是否为 NotReady
if (!"True".equals(nodeStatus)) {
System.out.println("❌ 节点 " + nodeName + " 状态异常: " + nodeStatus);
} else {
System.out.println("✅ 节点 " + nodeName + " 状态正常");
}
// 获取资源使用情况(需结合 metrics-server)
// 这里仅模拟:实际应调用 metrics.k8s.io API
System.out.println(" ├─ CPU 使用率: 72% (模拟值)");
System.out.println(" └─ 内存使用率: 68% (模拟值)");
}
}
}
Maven 依赖(pom.xml):
<dependency>
<groupId>io.kubernetes</groupId>
<artifactId>client-java</artifactId>
<version>19.0.0</version>
</dependency>
⚠️ 注意:上述代码仅获取节点状态,资源使用率需依赖 metrics-server。我们将在 3.7 节详细说明如何获取。
自动化建议:
- 每 5 分钟轮询一次节点状态。
- 若连续 3 次检测为
NotReady,自动触发节点隔离(drain)并告警。
3.2 Pod 健康与重启频率监测 🐳
Pod 是 K8s 的最小调度单元。即使节点正常,Pod 也可能因应用崩溃、镜像拉取失败、探针超时等原因反复重启。
检查项:
- Pod 状态(Running / Pending / CrashLoopBackOff / Error)
- 重启次数(RestartCount)
- 就绪探针(ReadinessProbe)是否通过
- 启动时间是否过长(> 5min)
- 镜像拉取失败(ImagePullBackOff)
Java 代码示例:检测 CrashLoopBackOff 和高重启 Pod
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1ContainerStatus;
import io.kubernetes.client.openapi.models.V1Pod;
import io.kubernetes.client.openapi.models.V1PodStatus;
import io.kubernetes.client.util.Config;
import java.io.IOException;
import java.util.List;
public class PodRestartChecker {
public static void main(String[] args) throws IOException {
ApiClient client = Config.defaultClient();
Configuration.setDefaultApiClient(client);
CoreV1Api api = new CoreV1Api();
// 检查所有命名空间
List<V1Pod> pods = api.listPodForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();
System.out.println("🚨 Pod 异常巡检报告:");
System.out.println("========================");
for (V1Pod pod : pods) {
String namespace = pod.getMetadata().getNamespace();
String podName = pod.getMetadata().getName();
V1PodStatus status = pod.getStatus();
if (status == null) continue;
// 检查是否为 CrashLoopBackOff
if ("CrashLoopBackOff".equals(status.getPhase())) {
System.out.println("💥 " + namespace + "/" + podName + " 处于 CrashLoopBackOff 状态");
continue;
}
// 检查重启次数
for (V1ContainerStatus containerStatus : status.getContainerStatuses()) {
int restartCount = containerStatus.getRestartCount();
if (restartCount > 5) {
System.out.println("⚠️ " + namespace + "/" + podName + " - 容器 " + containerStatus.getName() +
" 已重启 " + restartCount + " 次");
}
// 检查镜像拉取失败
if (containerStatus.getState().getWaiting() != null &&
"ImagePullBackOff".equals(containerStatus.getState().getWaiting().getReason())) {
System.out.println("🖼️ " + namespace + "/" + podName + " - 镜像拉取失败: " +
containerStatus.getState().getWaiting().getMessage());
}
// 检查就绪探针失败
if (!containerStatus.getReady()) {
System.out.println("🛑 " + namespace + "/" + podName + " - 容器 " + containerStatus.getName() +
" 就绪探针失败");
}
}
}
}
}
典型异常场景:
| 状态 | 可能原因 | 解决方案 |
|---|---|---|
Pending | 资源不足、调度器无法匹配节点 | 扩容节点、调整资源请求 |
CrashLoopBackOff | 应用启动失败、配置错误、依赖服务未就绪 | 查看日志 kubectl logs --previous |
ImagePullBackOff | 镜像不存在、私有仓库认证失败 | 检查 ImagePullSecret、仓库地址 |
RunContainerError | 容器运行时错误 | 检查节点 docker/containerd 状态 |
🔗 Kubernetes Pod Lifecycle —— 官方 Pod 生命周期详解
3.3 Deployment/StatefulSet 副本可用性检查 📊
Deployment 和 StatefulSet 是应用的“控制器”。它们的可用副本数直接决定服务是否对外提供能力。
检查项:
readyReplicas是否等于replicasunavailableReplicas是否 > 0- 最新版本是否已滚动更新完成
- 滚动更新是否卡住(超过
progressDeadlineSeconds)
Java 代码示例:检查 Deployment 副本状态
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.AppsV1Api;
import io.kubernetes.client.openapi.models.V1Deployment;
import io.kubernetes.client.util.Config;
import java.io.IOException;
import java.util.List;
public class DeploymentAvailabilityChecker {
public static void main(String[] args) throws IOException {
ApiClient client = Config.defaultClient();
Configuration.setDefaultApiClient(client);
AppsV1Api api = new AppsV1Api();
List<V1Deployment> deployments = api.listDeploymentForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();
System.out.println("📊 Deployment 副本健康巡检:");
System.out.println("========================");
for (V1Deployment deployment : deployments) {
String namespace = deployment.getMetadata().getNamespace();
String name = deployment.getMetadata().getName();
Integer replicas = deployment.getSpec().getReplicas();
Integer readyReplicas = deployment.getStatus().getReadyReplicas();
Integer unavailableReplicas = deployment.getStatus().getUnavailableReplicas();
if (replicas == null || readyReplicas == null) continue;
if (readyReplicas < replicas) {
System.out.println("⚠️ " + namespace + "/" + name + " - 可用副本不足: " + readyReplicas + "/" + replicas);
if (unavailableReplicas != null && unavailableReplicas > 0) {
System.out.println(" ➤ 不可用副本数: " + unavailableReplicas);
}
} else {
System.out.println("✅ " + namespace + "/" + name + " - 副本正常: " + readyReplicas + "/" + replicas);
}
}
}
}
自动化策略:
- 若
readyReplicas < 90% * replicas,发送严重告警。 - 若
unavailableReplicas > 0且持续 10 分钟,自动触发回滚(需结合 GitOps)。
💡 进阶建议:结合
kubectl rollout history检查最新修订版本是否成功。Java Client 可调用AppsV1Api.listNamespacedReplicaSet()获取历史 RS。
3.4 Service 与 Ingress 可达性验证 🌐
即使 Pod 运行正常,若 Service 或 Ingress 配置错误,外部用户仍无法访问。
检查项:
- Service 的
ClusterIP是否分配 - Endpoint 是否有对应 Pod IP
- Ingress 是否绑定正确 Host 和 Path
- Ingress Controller 是否正常运行
- SSL 证书是否有效(HTTPS)
Java 代码示例:验证 Service Endpoint 是否匹配 Pod
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1Endpoint;
import io.kubernetes.client.openapi.models.V1EndpointSubset;
import io.kubernetes.client.openapi.models.V1Service;
import io.kubernetes.client.util.Config;
import java.io.IOException;
import java.util.List;
public class ServiceEndpointChecker {
public static void main(String[] args) throws IOException {
ApiClient client = Config.defaultClient();
Configuration.setDefaultApiClient(client);
CoreV1Api api = new CoreV1Api();
List<V1Service> services = api.listServiceForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();
System.out.println("🔌 Service Endpoint 检查报告:");
System.out.println("========================");
for (V1Service service : services) {
String namespace = service.getMetadata().getNamespace();
String name = service.getMetadata().getName();
// 跳过 Headless Service
if ("ClusterIP".equals(service.getSpec().getType()) && service.getSpec().getClusterIP() == null) {
continue;
}
// 获取 Endpoints
var endpoints = api.readNamespacedEndpoints(name, namespace, null);
if (endpoints == null || endpoints.getSubsets() == null || endpoints.getSubsets().isEmpty()) {
System.out.println("❌ " + namespace + "/" + name + " - 无可用 Endpoint");
continue;
}
int totalEndpoints = 0;
for (V1EndpointSubset subset : endpoints.getSubsets()) {
if (subset.getAddresses() != null) {
totalEndpoints += subset.getAddresses().size();
}
}
// 获取选择器对应的 Pod 数量(简化版)
int selectorPodCount = 0; // 实际应通过 LabelSelector 查询 Pod 数量
// 这里为演示,假设我们已知期望 Pod 数量为 3
// 实际应调用 PodList 获取匹配的 Pod 数量
System.out.println("✅ " + namespace + "/" + name + " - Endpoint 数量: " + totalEndpoints +
" (期望: 3) " + (totalEndpoints == 3 ? "✅" : "⚠️"));
}
}
}
🚫 注意:Java Client 目前不直接提供根据 LabelSelector 查询 Pod 数量的便捷方法,需自行构造
LabelSelector并调用CoreV1Api.listPodForAllNamespaces()。
Ingress 检查建议:
- 使用
curl -I http://your-app.example.com检查 HTTP 状态码。 - 使用
openssl s_client -connect your-app.example.com:443检查证书有效期。
自动化建议:
- 编写一个轻量级 HTTP 探针服务,每 2 分钟对所有 Ingress Host 发起 GET 请求。
- 记录响应时间、状态码、证书过期时间。
// 简易 HTTP 探针示例(用于 Ingress 可达性)
import java.net.HttpURLConnection;
import java.net.URL;
public class IngressProbe {
public static void probe(String url) {
try {
URL target = new URL(url);
HttpURLConnection conn = (HttpURLConnection) target.openConnection();
conn.setRequestMethod("GET");
conn.setConnectTimeout(5000);
conn.setReadTimeout(5000);
int code = conn.getResponseCode();
System.out.println("🌐 " + url + " → " + code + " (响应时间: " + conn.getConnectTime() + "ms)");
if (code >= 500) {
System.out.println("❗ 服务异常,请检查后端应用");
}
} catch (Exception e) {
System.out.println("❌ " + url + " → 连接失败: " + e.getMessage());
}
}
public static void main(String[] args) {
probe("https://example.com");
probe("https://kubernetes.io");
}
}
🔗 Ingress-Nginx Best Practices —— Nginx Ingress 官方配置指南
3.5 存储卷(PV/PVC)与持久化状态检查 💾
有状态应用(如 MySQL、Redis、Elasticsearch)依赖持久化存储。PV/PVC 异常将导致数据丢失或服务不可用。
检查项:
- PVC 状态是否为
Bound - PV 状态是否为
Available或Bound - 存储容量是否接近上限(> 85%)
- StorageClass 是否支持动态供给
- 是否存在
Pending的 PVC
Java 代码示例:检查 PVC 和 PV 状态
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1PersistentVolume;
import io.kubernetes.client.openapi.models.V1PersistentVolumeClaim;
import io.kubernetes.client.util.Config;
import java.io.IOException;
import java.util.List;
public class StorageChecker {
public static void main(String[] args) throws IOException {
ApiClient client = Config.defaultClient();
Configuration.setDefaultApiClient(client);
CoreV1Api api = new CoreV1Api();
// 检查 PV
List<V1PersistentVolume> pvs = api.listPersistentVolume(null, null, null, null, null, null, null, null, null, null).getItems();
System.out.println("💾 PV 状态检查:");
for (V1PersistentVolume pv : pvs) {
String status = pv.getStatus().getPhase();
String name = pv.getMetadata().getName();
if ("Available".equals(status)) {
System.out.println("⚠️ PV " + name + " 处于 Available 状态(未被绑定)");
} else if ("Failed".equals(status)) {
System.out.println("❌ PV " + name + " 状态为 Failed");
} else {
System.out.println("✅ PV " + name + " 状态: " + status);
}
}
// 检查 PVC
List<V1PersistentVolumeClaim> pvcs = api.listPersistentVolumeClaimForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();
System.out.println("\n📂 PVC 状态检查:");
for (V1PersistentVolumeClaim pvc : pvcs) {
String namespace = pvc.getMetadata().getNamespace();
String name = pvc.getMetadata().getName();
String status = pvc.getStatus().getPhase();
if ("Pending".equals(status)) {
System.out.println("⚠️ " + namespace + "/" + name + " PVC 未绑定,可能因 StorageClass 或容量不足");
} else if ("Lost".equals(status)) {
System.out.println("❌ " + namespace + "/" + name + " PVC 已丢失,需人工恢复");
} else {
System.out.println("✅ " + namespace + "/" + name + " PVC 状态: " + status);
}
}
}
}
实际场景:
- 某 Redis 集群因 PVC 挂载失败,导致主节点无法启动,集群脑裂。
- 某日志系统因 PV 磁盘写满,导致应用崩溃。
自动化建议:
- 使用
kubectl get pv -o jsonpath='{.items[*].spec.capacity.storage}'获取容量。 - 对接 Prometheus + Node Exporter 监控底层存储使用率。
🔗 Kubernetes Storage Classes —— 存储类详解
3.6 资源配额与 Limit/Request 检查 📏
资源请求(Request)和限制(Limit)是 K8s 调度和资源隔离的核心。配置不当会导致:
- 节点资源过载 → Pod 被驱逐
- 资源浪费 → 成本飙升
- 无法调度 → Pod 一直 Pending
检查项:
- 命名空间是否设置了 ResourceQuota
- Pod 是否设置了 CPU/Memory Request 和 Limit
- Request 是否过高(> 80% 节点资源)
- Limit 是否过低(导致 OOMKilled)
Java 代码示例:扫描未设置 Request 的 Pod
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1Container;
import io.kubernetes.client.openapi.models.V1ContainerResourceRequirements;
import io.kubernetes.client.openapi.models.V1Pod;
import io.kubernetes.client.util.Config;
import java.io.IOException;
import java.util.List;
public class ResourceRequestChecker {
public static void main(String[] args) throws IOException {
ApiClient client = Config.defaultClient();
Configuration.setDefaultApiClient(client);
CoreV1Api api = new CoreV1Api();
List<V1Pod> pods = api.listPodForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();
System.out.println("📏 资源请求检查报告:");
System.out.println("========================");
for (V1Pod pod : pods) {
String namespace = pod.getMetadata().getNamespace();
String name = pod.getMetadata().getName();
for (V1Container container : pod.getSpec().getContainers()) {
V1ContainerResourceRequirements resources = container.getResources();
if (resources == null) {
System.out.println("⚠️ " + namespace + "/" + name + " - 容器 " + container.getName() + " 未设置任何资源请求/限制");
continue;
}
if (resources.getRequests() == null || resources.getRequests().get("cpu") == null ||
resources.getRequests().get("memory") == null) {
System.out.println("⚠️ " + namespace + "/" + name + " - 容器 " + container.getName() + " 缺少 Request");
}
if (resources.getLimits() == null || resources.getLimits().get("cpu") == null ||
resources.getLimits().get("memory") == null) {
System.out.println("⚠️ " + namespace + "/" + name + " - 容器 " + container.getName() + " 缺少 Limit");
}
}
}
}
}
最佳实践建议:
| 类型 | 建议值 |
|---|---|
| CPU Request | 200m ~ 500m(微服务) |
| Memory Request | 512Mi ~ 1Gi |
| CPU Limit | Request 的 2~3 倍 |
| Memory Limit | Request 的 1.5~2 倍 |
🔗 Kubernetes Resource Management —— 官方资源管理指南
3.7 集群指标采集:Metrics Server 与 Prometheus 📈
仅靠 kubectl 获取静态信息远远不够。实时指标才是巡检的黄金标准。
必须监控的指标:
| 指标 | 说明 | 阈值 |
|---|---|---|
node_cpu_usage_seconds_total | 节点 CPU 使用率 | > 85% 告警 |
container_memory_usage_bytes | 容器内存使用 | > 90% 告警 |
kubelet_runtime_operations_total | kubelet 操作失败数 | > 0 即异常 |
etcd_disk_wal_fsync_duration_seconds | etcd 写入延迟 | > 100ms 异常 |
pod_restart_count | Pod 重启次数 | > 5 次/小时告警 |
Java 代码示例:通过 Prometheus API 获取指标
import okhttp3.OkHttpClient;
import okhttp3.Request;
import okhttp3.Response;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.IOException;
public class PrometheusMetricsChecker {
private static final String PROMETHEUS_URL = "http://prometheus.kube-system.svc.cluster.local:9090/api/v1/query";
private static final OkHttpClient client = new OkHttpClient();
private static final ObjectMapper mapper = new ObjectMapper();
public static void main(String[] args) throws IOException {
String query = "sum(rate(container_cpu_usage_seconds_total{container!=\"POD\",namespace!=\"kube-system\"}[5m])) / sum(machine_cpu_cores) * 100";
String url = PROMETHEUS_URL + "?query=" + java.net.URLEncoder.encode(query, "UTF-8");
Request request = new Request.Builder()
.url(url)
.build();
try (Response response = client.newCall(request).execute()) {
if (!response.isSuccessful()) throw new IOException("Unexpected code " + response);
JsonNode root = mapper.readTree(response.body().string());
JsonNode data = root.get("data");
JsonNode result = data.get("result").get(0);
JsonNode value = result.get("value");
double cpuUsage = value.get(1).asDouble();
System.out.println("📈 集群平均 CPU 使用率: " + String.format("%.2f", cpuUsage) + "%");
if (cpuUsage > 85) {
System.out.println("🚨 警告:集群 CPU 使用率过高!");
}
}
}
}
Maven 依赖:
<dependency>
<groupId>com.squareup.okhttp3</groupId>
<artifactId>okhttp</artifactId>
<version>4.12.0</version>
</dependency>
<dependency>
<groupId>com.fasterxml.jackson.core</groupId>
<artifactId>jackson-databind</artifactId>
<version>2.15.3</version>
</dependency>
🔗 Prometheus Kubernetes Service Discovery —— Prometheus 如何自动发现 K8s 服务
建议架构:
✅ 推荐部署:Prometheus + Alertmanager + Grafana + Node Exporter + Kube-State-Metrics
3.8 etcd 状态与备份检查 🔐
etcd 是 K8s 的“大脑”,存储所有集群状态。一旦 etcd 崩溃,整个集群将不可用。
检查项:
- etcd pod 是否运行
- etcd 成员是否健康
- etcd 数据库大小(是否超过 2GB)
- 是否有定期备份(每日快照)
- TLS 证书是否过期
Java 代码示例:调用 etcdctl 检查健康状态(通过 Shell 执行)
import java.io.BufferedReader;
import java.io.InputStreamReader;
public class EtcdHealthChecker {
public static void checkEtcdHealth() {
try {
Process process = Runtime.getRuntime().exec("kubectl exec -n kube-system etcd-control-plane -- etcdctl endpoint health");
BufferedReader reader = new BufferedReader(new InputStreamReader(process.getInputStream()));
String line;
while ((line = reader.readLine()) != null) {
System.out.println("🔐 etcd 健康检查: " + line);
if (!line.contains("healthy")) {
System.out.println("❌ etcd 节点异常!");
}
}
} catch (Exception e) {
System.out.println("⚠️ 无法执行 etcdctl 命令: " + e.getMessage());
}
}
public static void main(String[] args) {
checkEtcdHealth();
}
}
💡 实际生产中,建议使用
etcdctl endpoint status获取详细状态。
备份建议:
- 每日 02:00 自动执行
etcdctl snapshot save /backup/snapshot.db - 将快照上传至 S3 或 MinIO
- 保留 7 天快照
# 示例备份脚本
ETCDCTL_API=3 etcdctl \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key \
snapshot save /backup/etcd-snapshot-$(date +%Y%m%d-%H%M%S).db
🔗 etcd Backup and Restore —— etcd 官方恢复指南
3.9 安全合规检查 🔒
安全是巡检的重中之重。K8s 默认配置存在大量风险。
检查项:
- 是否启用 RBAC?
- 是否禁用匿名访问?
- Pod 是否以 root 用户运行?
- 是否启用 PodSecurityPolicy / PodSecurity Admission?
- 是否存在默认 ServiceAccount 暴露?
- 是否开启网络策略(NetworkPolicy)?
Java 代码示例:检测 Pod 是否以 root 运行
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1Pod;
import io.kubernetes.client.openapi.models.V1SecurityContext;
import io.kubernetes.client.util.Config;
import java.io.IOException;
import java.util.List;
public class SecurityContextChecker {
public static void main(String[] args) throws IOException {
ApiClient client = Config.defaultClient();
Configuration.setDefaultApiClient(client);
CoreV1Api api = new CoreV1Api();
List<V1Pod> pods = api.listPodForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();
System.out.println("🛡️ 安全上下文检查:");
System.out.println("========================");
for (V1Pod pod : pods) {
for (var container : pod.getSpec().getContainers()) {
V1SecurityContext securityContext = container.getSecurityContext();
if (securityContext != null) {
Integer runAsUser = securityContext.getRunAsUser();
if (runAsUser != null && runAsUser == 0) {
System.out.println("❌ " + pod.getMetadata().getNamespace() + "/" + pod.getMetadata().getName() +
" - 容器 " + container.getName() + " 以 root 用户运行");
}
} else {
System.out.println("⚠️ " + pod.getMetadata().getNamespace() + "/" + pod.getMetadata().getName() +
" - 容器 " + container.getName() + " 未设置 securityContext");
}
}
}
}
}
安全加固建议:
| 检查项 | 建议 |
|---|---|
| 禁用匿名访问 | --anonymous-auth=false |
| 启用 RBAC | 必须开启 |
| 使用非 root 用户 | Dockerfile 中 USER 1000 |
| 启用网络策略 | 默认拒绝所有流量 |
| 禁用 HostNetwork | 避免绕过网络隔离 |
| 使用镜像签名 | Notary 或 Cosign |
🔗 CIS Kubernetes Benchmark —— CIS 官方 K8s 安全基线(V1.23)
3.10 日志与事件聚合分析 📋
K8s 事件(Events)是诊断问题的“第一手资料”。
检查项:
- 是否有高频
FailedScheduling事件? - 是否有
FailedCreatePodSandBox? - 是否有
ImagePullBackOff? - 是否有
Evicted事件?
Java 代码示例:读取命名空间事件
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1Event;
import io.kubernetes.client.util.Config;
import java.io.IOException;
import java.util.List;
public class EventAnalyzer {
public static void main(String[] args) throws IOException {
ApiClient client = Config.defaultClient();
Configuration.setDefaultApiClient(client);
CoreV1Api api = new CoreV1Api();
List<V1Event> events = api.listNamespacedEvent("default", null, null, null, null, null, null, null, null, null).getItems();
System.out.println("📝 最近 10 条事件摘要:");
System.out.println("========================");
for (int i = 0; i < Math.min(10, events.size()); i++) {
V1Event event = events.get(i);
System.out.println("⏰ " + event.getLastTimestamp() +
" | " + event.getReason() +
" | " + event.getMessage() +
" | " + event.getSource().getComponent());
}
}
}
高频事件应对策略:
| 事件 | 原因 | 解决 |
|---|---|---|
FailedScheduling | 资源不足 | 扩容节点、调整 Request |
FailedCreatePodSandBox | CRI(containerd/dockershim)异常 | 重启 kubelet |
Evicted | 内存压力 | 增加内存、优化应用 |
Unhealthy | 探针失败 | 调整 timeout/period |
✅ 建议:将所有事件导入 Loki + Grafana,建立“事件热力图”。
四、巡检自动化架构设计 🏗️
4.1 巡检工具整体架构
4.2 部署方式建议
| 方式 | 适用场景 | 优点 | 缺点 |
|---|---|---|---|
| CronJob + Java 容器 | 生产环境 | 自动化、可监控、可版本控制 | 需要构建镜像 |
| 外部 Java 服务(K8s 外) | 多集群统一巡检 | 独立部署,不干扰集群 | 需配置 kubeconfig 访问 |
| Python/Shell 脚本 | 小型集群 | 快速开发 | 不易维护,无类型安全 |
| Operator 模式 | 企业级 | 深度集成,声明式 | 开发复杂度高 |
✅ 推荐方案:CronJob + Java 容器,每日 02:00 自动执行,结果存入 MinIO,告警推送钉钉。
4.3 巡检工具 Dockerfile 示例
FROM openjdk:17-jre-slim
WORKDIR /app
COPY target/cluster-checker-1.0.jar app.jar
# 安装 kubectl(用于执行辅助命令)
RUN apt-get update && apt-get install -y curl && \
curl -LO "https://dl.k8s.io/release/$(curl -L -s https://dl.k8s.io/release/stable.txt)/bin/linux/amd64/kubectl" && \
install -o root -g root -m 0755 kubectl /usr/local/bin/kubectl
# 挂载 kubeconfig
VOLUME /root/.kube
ENTRYPOINT ["java", "-jar", "/app/app.jar"]
4.4 CronJob 配置示例
apiVersion: batch/v1
kind: CronJob
metadata:
name: cluster-health-check
namespace: monitoring
spec:
schedule: "0 2 * * *" # 每天凌晨2点
jobTemplate:
spec:
template:
spec:
serviceAccountName: cluster-checker-sa
containers:
- name: checker
image: your-registry/cluster-checker:latest
volumeMounts:
- name: kubeconfig
mountPath: /root/.kube
readOnly: true
restartPolicy: OnFailure
volumes:
- name: kubeconfig
hostPath:
path: /root/.kube/config
⚠️ 安全提醒:为巡检工具创建专用 ServiceAccount,仅授予
get/list/watch权限,禁止delete权限!
五、巡检报告模板与可视化 📊
5.1 报告内容结构(JSON 示例)
{
"timestamp": "2024-06-15T02:00:00Z",
"cluster_name": "prod-us-east",
"nodes": {
"total": 12,
"ready": 11,
"not_ready": 1,
"high_cpu_nodes": ["node-07"]
},
"pods": {
"total": 89,
"crashloop_backoff": 2,
"pending": 0,
"missing_requests": 15
},
"storage": {
"pvc_pending": 1,
"pv_failed": 0,
"disk_usage_high": ["pvc-redis-01"]
},
"security": {
"root_containers": 3,
"unsecured_serviceaccounts": 1
},
"metrics": {
"avg_cpu_usage": 78.5,
"avg_memory_usage": 65.2,
"etcd_health": "healthy"
},
"events_summary": [
"FailedScheduling: 3",
"ImagePullBackOff: 1",
"Evicted: 2"
],
"status": "WARNING",
"recommendations": [
"为 node-07 添加资源或驱逐低优先级 Pod",
"为 15 个未设置 Request 的 Pod 补充资源配置",
"检查 pvc-redis-01 存储容量是否可扩容"
]
}
5.2 报告可视化(HTML 模板片段)
<div class="dashboard">
<h2>📊 集群健康状态:⚠️ 警告</h2>
<div class="status-grid">
<div class="card red">节点异常: 1</div>
<div class="card yellow">Pod 重启: 2</div>
<div class="card green">存储正常: 100%</div>
<div class="card blue">安全风险: 3</div>
</div>
<h3>📈 资源使用趋势</h3>
<div id="cpu-chart">[ECharts 图表]</div>
<h3>🚨 最近事件</h3>
<ul>
<li>06-14 23:45: Pod redis-01 被驱逐(内存不足)</li>
<li>06-14 23:30: 镜像 registry.example.com/v1.2.3 拉取失败</li>
</ul>
</div>
📌 建议使用 ECharts 或 Chart.js 渲染趋势图,生成 HTML 报告后通过邮件发送。
六、巡检策略优化与持续改进 🔄
6.1 告警分级策略(建议)
| 级别 | 触发条件 | 响应机制 |
|---|---|---|
| 🟢 Info | 资源使用率 > 70% | 日志记录,周报汇总 |
| 🟡 Warning | Pod 重启 > 5 次/小时、节点 NotReady | 钉钉通知,1 小时内响应 |
| 🔴 Critical | etcd 不健康、所有 Pod Pending、Ingress 5xx > 10% | 电话通知,立即介入 |
6.2 建立“巡检基线”
- 记录正常状态下的指标值(如 CPU 平均 40%)
- 每月对比趋势,动态调整阈值
- 对比不同环境(dev/stage/prod)差异
6.3 集成到 CI/CD 流水线
在发布前增加“巡检前置检查”:
- name: Pre-deploy Health Check
run: |
java -jar cluster-checker.jar --mode pre-deploy
if [ $? -ne 0 ]; then
echo "❌ 集群健康检查失败,禁止发布"
exit 1
fi
✅ 这样可避免“发布导致集群雪崩”。
七、常见误区与避坑指南 ⚠️
| 误区 | 正确做法 |
|---|---|
❌ 只看 kubectl get pods | ✅ 必须检查 kubectl get pods -o wide + kubectl describe pod |
| ❌ 依赖人工记忆阈值 | ✅ 所有阈值应写入配置文件,支持动态加载 |
| ❌ 忽视事件(Events) | ✅ Events 是诊断的起点,必须监控 |
| ❌ 不做备份 | ✅ etcd、coredns、ingress-controller 配置必须定期备份 |
| ❌ 使用默认 ServiceAccount | ✅ 每个应用应使用独立 SA,绑定最小权限 |
| ❌ 不做压力测试 | ✅ 每季度模拟节点宕机、网络分区,验证自愈能力 |
八、总结:巡检即 DevOps 的灵魂 🧠
Kubernetes 不是“部署完就完事”的系统,它是一个需要持续维护的生命体。一个成熟的 DevOps 团队,不会等到业务告警才行动,而是:
每天清晨,第一件事不是打开邮箱,而是打开巡检报告。
通过本文的完整流程与 Java 代码示例,你已掌握:
- ✅ 六大核心检查维度
- ✅ 自动化巡检工具开发方法
- ✅ 集成 Prometheus 与告警体系
- ✅ 安全合规加固要点
- ✅ 巡检报告生成与可视化
下一步建议:
- ✅ 将本文代码打包为 Docker 镜像
- ✅ 部署为 CronJob
- ✅ 接入钉钉机器人
- ✅ 每周回顾报告,优化阈值
- ✅ 建立“巡检知识库”,沉淀处理经验
🌟 真正的自动化,不是代替人,而是让人类专注于更有价值的问题。
当你不再为“为什么服务挂了”而熬夜,而是优雅地在日报中写上:“今日巡检发现 2 个潜在风险,已自动修复”,你,就是云原生时代的守护者。
附录:推荐工具与资源 📚
| 类型 | 推荐工具 | 链接 |
|---|---|---|
| 监控 | Prometheus + Grafana | https://prometheus.io |
| 日志 | Loki + Grafana | https://grafana.com/oss/loki/ |
| 安全 | Kube-Bench | https://github.com/aquasecurity/kube-bench |
| 网络 | Calico / Cilium | https://cilium.io/ |
| 存储 | Longhorn | https://longhorn.io/ |
| 巡检框架 | KubeLinter | https://github.com/stackrox/kubelinter |
| 配置管理 | Kyverno | https://kyverno.io/ |
💬 最后忠告:不要迷信“一键部署”,要敬畏“持续运维”。K8s 的美,在于它的弹性;它的险,在于它的复杂。唯有系统化巡检,方能行稳致远。
愿你的集群永不宕机,愿你的告警永远清零。 🙏
🙌 感谢你读到这里!
🔍 技术之路没有捷径,但每一次阅读、思考和实践,都在悄悄拉近你与目标的距离。
💡 如果本文对你有帮助,不妨 👍 点赞、📌 收藏、📤 分享 给更多需要的朋友!
💬 欢迎在评论区留下你的想法、疑问或建议,我会一一回复,我们一起交流、共同成长 🌿
🔔 关注我,不错过下一篇干货!我们下期再见!✨
DAMO开发者矩阵,由阿里巴巴达摩院和中国互联网协会联合发起,致力于探讨最前沿的技术趋势与应用成果,搭建高质量的交流与分享平台,推动技术创新与产业应用链接,围绕“人工智能与新型计算”构建开放共享的开发者生态。
更多推荐


所有评论(0)