在这里插入图片描述

👋 大家好,欢迎来到我的技术博客!
📚 在这里,我会分享学习笔记、实战经验与技术思考,力求用简单的方式讲清楚复杂的问题。
🎯 本文将围绕Kubernetes这个话题展开,希望能为你带来一些启发或实用的参考。
🌱 无论你是刚入门的新手,还是正在进阶的开发者,希望你都能有所收获!


文章目录

Kubernetes - K8s 集群的日常巡检流程与核心检查项 🚀

在现代云原生架构中,Kubernetes(简称 K8s)已成为容器编排的事实标准。它赋予企业弹性伸缩、服务自治、故障自愈等强大能力,但也因其复杂性,对运维人员提出了更高的要求。一个稳定、高效、安全的 K8s 集群,离不开系统化、标准化、自动化的日常巡检流程。本文将深入剖析 K8s 集群日常巡检的完整流程,涵盖从节点状态、Pod 健康、网络连通、存储资源到安全合规的全方位检查项,并结合 Java 代码示例展示如何构建自动化巡检工具,辅以 Mermaid 图表直观呈现架构逻辑,帮助你构建一套可落地、可扩展、可监控的集群健康管理体系。


一、为什么需要日常巡检?🚨

Kubernetes 是一个高度动态的系统。Pod 会不断重启、节点会因资源不足被驱逐、网络策略可能被误配置、证书会过期、镜像仓库可能不可达……这些“小问题”如果长期积累,最终会演变成雪崩式故障。

📌 真实案例:某金融平台因未监控 etcd 磁盘使用率,导致集群元数据写入失败,所有新 Pod 无法调度,业务中断 4 小时,损失超百万。

日常巡检不是“可选项”,而是生产环境的底线要求。它能帮助你:

  • ✅ 提前发现潜在风险(如节点资源耗尽、证书即将过期)
  • ✅ 快速定位故障根源(是网络?是存储?还是调度器?)
  • ✅ 满足合规审计(如等保、GDPR、ISO27001)
  • ✅ 优化资源成本(识别闲置 Pod、低效部署)
  • ✅ 建立运维知识库(记录历史异常,形成 SOP)

💡 建议:每日巡检应控制在 15~30 分钟内完成,自动化工具是关键。人工干预仅用于异常处理。


二、巡检流程总览:五步闭环法 🔄

一个标准的 K8s 集群巡检流程,可归纳为以下五个步骤:

启动巡检任务

收集集群状态数据

分析健康指标与阈值

生成巡检报告与告警

触发修复或人工干预

记录归档并优化策略

步骤说明:

  1. 启动巡检任务:通过定时任务(CronJob)、CI/CD 流水线或运维平台触发。
  2. 收集集群状态数据:调用 kubectl、K8s API 或 Prometheus 指标,获取节点、Pod、Deployment、Service、PV/PVC 等资源状态。
  3. 分析健康指标与阈值:比对预设阈值(如 CPU > 85%、Pod RestartCount > 5),判断是否异常。
  4. 生成巡检报告与告警:输出 HTML/JSON 报告,推送企业微信/钉钉/Slack/邮件告警。
  5. 触发修复或人工干预:自动执行修复脚本(如清理 Terminating Pod),或通知运维人员介入。
  6. 记录归档并优化策略:将巡检日志存入 ELK 或 Loki,定期分析趋势,优化阈值和检查项。

🔗 Kubernetes Best Practices: Operational Guidelines —— 官方推荐的生产集群运维指南


三、核心检查项详解(附 Java 代码示例)💻

我们将从节点层 → Pod 层 → 网络层 → 存储层 → 安全层 → 控制平面层六大维度展开,每一项均提供可运行的 Java 代码示例,帮助你构建自动化巡检工具。


3.1 节点健康状态检查 🖥️

节点是 K8s 的基石。一个节点宕机或资源耗尽,会影响其上所有 Pod。

检查项:
  • 节点状态(Ready/NotReady)
  • CPU/内存使用率
  • 磁盘使用率(根分区、镜像层、日志)
  • kubelet 是否运行
  • 标签与污点是否异常
Java 代码示例:使用 Kubernetes Java Client 检查节点状态
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1Node;
import io.kubernetes.client.openapi.models.V1NodeCondition;
import io.kubernetes.client.openapi.models.V1NodeStatus;
import io.kubernetes.client.util.Config;
import java.io.IOException;
import java.util.List;

public class NodeHealthChecker {

    public static void main(String[] args) throws IOException {
        // 加载 kubeconfig(默认 ~/.kube/config)
        ApiClient client = Config.defaultClient();
        Configuration.setDefaultApiClient(client);

        CoreV1Api api = new CoreV1Api();
        List<V1Node> nodes = api.listNode(null, null, null, null, null, null, null, null, null, null).getItems();

        System.out.println("🔍 节点健康巡检报告:");
        System.out.println("========================");

        for (V1Node node : nodes) {
            String nodeName = node.getMetadata().getName();
            String nodeStatus = node.getStatus().getConditions().stream()
                    .filter(condition -> "Ready".equals(condition.getType()))
                    .map(V1NodeCondition::getStatus)
                    .findFirst()
                    .orElse("Unknown");

            // 检查是否为 NotReady
            if (!"True".equals(nodeStatus)) {
                System.out.println("❌ 节点 " + nodeName + " 状态异常: " + nodeStatus);
            } else {
                System.out.println("✅ 节点 " + nodeName + " 状态正常");
            }

            // 获取资源使用情况(需结合 metrics-server)
            // 这里仅模拟:实际应调用 metrics.k8s.io API
            System.out.println("   ├─ CPU 使用率: 72% (模拟值)");
            System.out.println("   └─ 内存使用率: 68% (模拟值)");
        }
    }
}
Maven 依赖(pom.xml):
<dependency>
    <groupId>io.kubernetes</groupId>
    <artifactId>client-java</artifactId>
    <version>19.0.0</version>
</dependency>

⚠️ 注意:上述代码仅获取节点状态,资源使用率需依赖 metrics-server。我们将在 3.7 节详细说明如何获取。

自动化建议:
  • 每 5 分钟轮询一次节点状态。
  • 若连续 3 次检测为 NotReady,自动触发节点隔离(drain)并告警。

3.2 Pod 健康与重启频率监测 🐳

Pod 是 K8s 的最小调度单元。即使节点正常,Pod 也可能因应用崩溃、镜像拉取失败、探针超时等原因反复重启。

检查项:
  • Pod 状态(Running / Pending / CrashLoopBackOff / Error)
  • 重启次数(RestartCount)
  • 就绪探针(ReadinessProbe)是否通过
  • 启动时间是否过长(> 5min)
  • 镜像拉取失败(ImagePullBackOff)
Java 代码示例:检测 CrashLoopBackOff 和高重启 Pod
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1ContainerStatus;
import io.kubernetes.client.openapi.models.V1Pod;
import io.kubernetes.client.openapi.models.V1PodStatus;
import io.kubernetes.client.util.Config;

import java.io.IOException;
import java.util.List;

public class PodRestartChecker {

    public static void main(String[] args) throws IOException {
        ApiClient client = Config.defaultClient();
        Configuration.setDefaultApiClient(client);

        CoreV1Api api = new CoreV1Api();
        // 检查所有命名空间
        List<V1Pod> pods = api.listPodForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();

        System.out.println("🚨 Pod 异常巡检报告:");
        System.out.println("========================");

        for (V1Pod pod : pods) {
            String namespace = pod.getMetadata().getNamespace();
            String podName = pod.getMetadata().getName();
            V1PodStatus status = pod.getStatus();

            if (status == null) continue;

            // 检查是否为 CrashLoopBackOff
            if ("CrashLoopBackOff".equals(status.getPhase())) {
                System.out.println("💥 " + namespace + "/" + podName + " 处于 CrashLoopBackOff 状态");
                continue;
            }

            // 检查重启次数
            for (V1ContainerStatus containerStatus : status.getContainerStatuses()) {
                int restartCount = containerStatus.getRestartCount();
                if (restartCount > 5) {
                    System.out.println("⚠️ " + namespace + "/" + podName + " - 容器 " + containerStatus.getName() +
                            " 已重启 " + restartCount + " 次");
                }

                // 检查镜像拉取失败
                if (containerStatus.getState().getWaiting() != null &&
                        "ImagePullBackOff".equals(containerStatus.getState().getWaiting().getReason())) {
                    System.out.println("🖼️ " + namespace + "/" + podName + " - 镜像拉取失败: " +
                            containerStatus.getState().getWaiting().getMessage());
                }

                // 检查就绪探针失败
                if (!containerStatus.getReady()) {
                    System.out.println("🛑 " + namespace + "/" + podName + " - 容器 " + containerStatus.getName() +
                            " 就绪探针失败");
                }
            }
        }
    }
}
典型异常场景:
状态可能原因解决方案
Pending资源不足、调度器无法匹配节点扩容节点、调整资源请求
CrashLoopBackOff应用启动失败、配置错误、依赖服务未就绪查看日志 kubectl logs --previous
ImagePullBackOff镜像不存在、私有仓库认证失败检查 ImagePullSecret、仓库地址
RunContainerError容器运行时错误检查节点 docker/containerd 状态

🔗 Kubernetes Pod Lifecycle —— 官方 Pod 生命周期详解


3.3 Deployment/StatefulSet 副本可用性检查 📊

Deployment 和 StatefulSet 是应用的“控制器”。它们的可用副本数直接决定服务是否对外提供能力。

检查项:
  • readyReplicas 是否等于 replicas
  • unavailableReplicas 是否 > 0
  • 最新版本是否已滚动更新完成
  • 滚动更新是否卡住(超过 progressDeadlineSeconds)
Java 代码示例:检查 Deployment 副本状态
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.AppsV1Api;
import io.kubernetes.client.openapi.models.V1Deployment;
import io.kubernetes.client.util.Config;

import java.io.IOException;
import java.util.List;

public class DeploymentAvailabilityChecker {

    public static void main(String[] args) throws IOException {
        ApiClient client = Config.defaultClient();
        Configuration.setDefaultApiClient(client);

        AppsV1Api api = new AppsV1Api();
        List<V1Deployment> deployments = api.listDeploymentForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();

        System.out.println("📊 Deployment 副本健康巡检:");
        System.out.println("========================");

        for (V1Deployment deployment : deployments) {
            String namespace = deployment.getMetadata().getNamespace();
            String name = deployment.getMetadata().getName();
            Integer replicas = deployment.getSpec().getReplicas();
            Integer readyReplicas = deployment.getStatus().getReadyReplicas();
            Integer unavailableReplicas = deployment.getStatus().getUnavailableReplicas();

            if (replicas == null || readyReplicas == null) continue;

            if (readyReplicas < replicas) {
                System.out.println("⚠️ " + namespace + "/" + name + " - 可用副本不足: " + readyReplicas + "/" + replicas);
                if (unavailableReplicas != null && unavailableReplicas > 0) {
                    System.out.println("   ➤ 不可用副本数: " + unavailableReplicas);
                }
            } else {
                System.out.println("✅ " + namespace + "/" + name + " - 副本正常: " + readyReplicas + "/" + replicas);
            }
        }
    }
}
自动化策略:
  • 若 readyReplicas < 90% * replicas,发送严重告警。
  • 若 unavailableReplicas > 0 且持续 10 分钟,自动触发回滚(需结合 GitOps)。

💡 进阶建议:结合 kubectl rollout history 检查最新修订版本是否成功。Java Client 可调用 AppsV1Api.listNamespacedReplicaSet() 获取历史 RS。


3.4 Service 与 Ingress 可达性验证 🌐

即使 Pod 运行正常,若 Service 或 Ingress 配置错误,外部用户仍无法访问。

检查项:
  • Service 的 ClusterIP 是否分配
  • Endpoint 是否有对应 Pod IP
  • Ingress 是否绑定正确 Host 和 Path
  • Ingress Controller 是否正常运行
  • SSL 证书是否有效(HTTPS)
Java 代码示例:验证 Service Endpoint 是否匹配 Pod
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1Endpoint;
import io.kubernetes.client.openapi.models.V1EndpointSubset;
import io.kubernetes.client.openapi.models.V1Service;
import io.kubernetes.client.util.Config;

import java.io.IOException;
import java.util.List;

public class ServiceEndpointChecker {

    public static void main(String[] args) throws IOException {
        ApiClient client = Config.defaultClient();
        Configuration.setDefaultApiClient(client);

        CoreV1Api api = new CoreV1Api();
        List<V1Service> services = api.listServiceForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();

        System.out.println("🔌 Service Endpoint 检查报告:");
        System.out.println("========================");

        for (V1Service service : services) {
            String namespace = service.getMetadata().getNamespace();
            String name = service.getMetadata().getName();

            // 跳过 Headless Service
            if ("ClusterIP".equals(service.getSpec().getType()) && service.getSpec().getClusterIP() == null) {
                continue;
            }

            // 获取 Endpoints
            var endpoints = api.readNamespacedEndpoints(name, namespace, null);
            if (endpoints == null || endpoints.getSubsets() == null || endpoints.getSubsets().isEmpty()) {
                System.out.println("❌ " + namespace + "/" + name + " - 无可用 Endpoint");
                continue;
            }

            int totalEndpoints = 0;
            for (V1EndpointSubset subset : endpoints.getSubsets()) {
                if (subset.getAddresses() != null) {
                    totalEndpoints += subset.getAddresses().size();
                }
            }

            // 获取选择器对应的 Pod 数量(简化版)
            int selectorPodCount = 0; // 实际应通过 LabelSelector 查询 Pod 数量
            // 这里为演示,假设我们已知期望 Pod 数量为 3
            // 实际应调用 PodList 获取匹配的 Pod 数量
            System.out.println("✅ " + namespace + "/" + name + " - Endpoint 数量: " + totalEndpoints +
                    " (期望: 3) " + (totalEndpoints == 3 ? "✅" : "⚠️"));
        }
    }
}

🚫 注意:Java Client 目前不直接提供根据 LabelSelector 查询 Pod 数量的便捷方法,需自行构造 LabelSelector 并调用 CoreV1Api.listPodForAllNamespaces()。

Ingress 检查建议:
  • 使用 curl -I http://your-app.example.com 检查 HTTP 状态码。
  • 使用 openssl s_client -connect your-app.example.com:443 检查证书有效期。
自动化建议:
  • 编写一个轻量级 HTTP 探针服务,每 2 分钟对所有 Ingress Host 发起 GET 请求。
  • 记录响应时间、状态码、证书过期时间。
// 简易 HTTP 探针示例(用于 Ingress 可达性)
import java.net.HttpURLConnection;
import java.net.URL;

public class IngressProbe {
    public static void probe(String url) {
        try {
            URL target = new URL(url);
            HttpURLConnection conn = (HttpURLConnection) target.openConnection();
            conn.setRequestMethod("GET");
            conn.setConnectTimeout(5000);
            conn.setReadTimeout(5000);
            int code = conn.getResponseCode();
            System.out.println("🌐 " + url + " → " + code + " (响应时间: " + conn.getConnectTime() + "ms)");
            if (code >= 500) {
                System.out.println("❗ 服务异常,请检查后端应用");
            }
        } catch (Exception e) {
            System.out.println("❌ " + url + " → 连接失败: " + e.getMessage());
        }
    }

    public static void main(String[] args) {
        probe("https://example.com");
        probe("https://kubernetes.io");
    }
}

🔗 Ingress-Nginx Best Practices —— Nginx Ingress 官方配置指南


3.5 存储卷(PV/PVC)与持久化状态检查 💾

有状态应用(如 MySQL、Redis、Elasticsearch)依赖持久化存储。PV/PVC 异常将导致数据丢失或服务不可用。

检查项:
  • PVC 状态是否为 Bound
  • PV 状态是否为 Available 或 Bound
  • 存储容量是否接近上限(> 85%)
  • StorageClass 是否支持动态供给
  • 是否存在 Pending 的 PVC
Java 代码示例:检查 PVC 和 PV 状态
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1PersistentVolume;
import io.kubernetes.client.openapi.models.V1PersistentVolumeClaim;
import io.kubernetes.client.util.Config;

import java.io.IOException;
import java.util.List;

public class StorageChecker {

    public static void main(String[] args) throws IOException {
        ApiClient client = Config.defaultClient();
        Configuration.setDefaultApiClient(client);

        CoreV1Api api = new CoreV1Api();

        // 检查 PV
        List<V1PersistentVolume> pvs = api.listPersistentVolume(null, null, null, null, null, null, null, null, null, null).getItems();
        System.out.println("💾 PV 状态检查:");
        for (V1PersistentVolume pv : pvs) {
            String status = pv.getStatus().getPhase();
            String name = pv.getMetadata().getName();
            if ("Available".equals(status)) {
                System.out.println("⚠️ PV " + name + " 处于 Available 状态(未被绑定)");
            } else if ("Failed".equals(status)) {
                System.out.println("❌ PV " + name + " 状态为 Failed");
            } else {
                System.out.println("✅ PV " + name + " 状态: " + status);
            }
        }

        // 检查 PVC
        List<V1PersistentVolumeClaim> pvcs = api.listPersistentVolumeClaimForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();
        System.out.println("\n📂 PVC 状态检查:");
        for (V1PersistentVolumeClaim pvc : pvcs) {
            String namespace = pvc.getMetadata().getNamespace();
            String name = pvc.getMetadata().getName();
            String status = pvc.getStatus().getPhase();

            if ("Pending".equals(status)) {
                System.out.println("⚠️ " + namespace + "/" + name + " PVC 未绑定,可能因 StorageClass 或容量不足");
            } else if ("Lost".equals(status)) {
                System.out.println("❌ " + namespace + "/" + name + " PVC 已丢失,需人工恢复");
            } else {
                System.out.println("✅ " + namespace + "/" + name + " PVC 状态: " + status);
            }
        }
    }
}
实际场景:
  • 某 Redis 集群因 PVC 挂载失败,导致主节点无法启动,集群脑裂。
  • 某日志系统因 PV 磁盘写满,导致应用崩溃。
自动化建议:
  • 使用 kubectl get pv -o jsonpath='{.items[*].spec.capacity.storage}' 获取容量。
  • 对接 Prometheus + Node Exporter 监控底层存储使用率。

🔗 Kubernetes Storage Classes —— 存储类详解


3.6 资源配额与 Limit/Request 检查 📏

资源请求(Request)和限制(Limit)是 K8s 调度和资源隔离的核心。配置不当会导致:

  • 节点资源过载 → Pod 被驱逐
  • 资源浪费 → 成本飙升
  • 无法调度 → Pod 一直 Pending
检查项:
  • 命名空间是否设置了 ResourceQuota
  • Pod 是否设置了 CPU/Memory Request 和 Limit
  • Request 是否过高(> 80% 节点资源)
  • Limit 是否过低(导致 OOMKilled)
Java 代码示例:扫描未设置 Request 的 Pod
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1Container;
import io.kubernetes.client.openapi.models.V1ContainerResourceRequirements;
import io.kubernetes.client.openapi.models.V1Pod;
import io.kubernetes.client.util.Config;

import java.io.IOException;
import java.util.List;

public class ResourceRequestChecker {

    public static void main(String[] args) throws IOException {
        ApiClient client = Config.defaultClient();
        Configuration.setDefaultApiClient(client);

        CoreV1Api api = new CoreV1Api();
        List<V1Pod> pods = api.listPodForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();

        System.out.println("📏 资源请求检查报告:");
        System.out.println("========================");

        for (V1Pod pod : pods) {
            String namespace = pod.getMetadata().getNamespace();
            String name = pod.getMetadata().getName();

            for (V1Container container : pod.getSpec().getContainers()) {
                V1ContainerResourceRequirements resources = container.getResources();
                if (resources == null) {
                    System.out.println("⚠️ " + namespace + "/" + name + " - 容器 " + container.getName() + " 未设置任何资源请求/限制");
                    continue;
                }

                if (resources.getRequests() == null || resources.getRequests().get("cpu") == null ||
                        resources.getRequests().get("memory") == null) {
                    System.out.println("⚠️ " + namespace + "/" + name + " - 容器 " + container.getName() + " 缺少 Request");
                }

                if (resources.getLimits() == null || resources.getLimits().get("cpu") == null ||
                        resources.getLimits().get("memory") == null) {
                    System.out.println("⚠️ " + namespace + "/" + name + " - 容器 " + container.getName() + " 缺少 Limit");
                }
            }
        }
    }
}
最佳实践建议:
类型建议值
CPU Request200m ~ 500m(微服务)
Memory Request512Mi ~ 1Gi
CPU LimitRequest 的 2~3 倍
Memory LimitRequest 的 1.5~2 倍

🔗 Kubernetes Resource Management —— 官方资源管理指南


3.7 集群指标采集:Metrics Server 与 Prometheus 📈

仅靠 kubectl 获取静态信息远远不够。实时指标才是巡检的黄金标准。

必须监控的指标:
指标说明阈值
node_cpu_usage_seconds_total节点 CPU 使用率> 85% 告警
container_memory_usage_bytes容器内存使用> 90% 告警
kubelet_runtime_operations_totalkubelet 操作失败数> 0 即异常
etcd_disk_wal_fsync_duration_secondsetcd 写入延迟> 100ms 异常
pod_restart_countPod 重启次数> 5 次/小时告警
Java 代码示例:通过 Prometheus API 获取指标
import okhttp3.OkHttpClient;
import okhttp3.Request;
import okhttp3.Response;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;

import java.io.IOException;

public class PrometheusMetricsChecker {

    private static final String PROMETHEUS_URL = "http://prometheus.kube-system.svc.cluster.local:9090/api/v1/query";
    private static final OkHttpClient client = new OkHttpClient();
    private static final ObjectMapper mapper = new ObjectMapper();

    public static void main(String[] args) throws IOException {
        String query = "sum(rate(container_cpu_usage_seconds_total{container!=\"POD\",namespace!=\"kube-system\"}[5m])) / sum(machine_cpu_cores) * 100";
        String url = PROMETHEUS_URL + "?query=" + java.net.URLEncoder.encode(query, "UTF-8");

        Request request = new Request.Builder()
                .url(url)
                .build();

        try (Response response = client.newCall(request).execute()) {
            if (!response.isSuccessful()) throw new IOException("Unexpected code " + response);

            JsonNode root = mapper.readTree(response.body().string());
            JsonNode data = root.get("data");
            JsonNode result = data.get("result").get(0);
            JsonNode value = result.get("value");

            double cpuUsage = value.get(1).asDouble();
            System.out.println("📈 集群平均 CPU 使用率: " + String.format("%.2f", cpuUsage) + "%");

            if (cpuUsage > 85) {
                System.out.println("🚨 警告:集群 CPU 使用率过高!");
            }
        }
    }
}
Maven 依赖:
<dependency>
    <groupId>com.squareup.okhttp3</groupId>
    <artifactId>okhttp</artifactId>
    <version>4.12.0</version>
</dependency>

<dependency>
    <groupId>com.fasterxml.jackson.core</groupId>
    <artifactId>jackson-databind</artifactId>
    <version>2.15.3</version>
</dependency>

🔗 Prometheus Kubernetes Service Discovery —— Prometheus 如何自动发现 K8s 服务

建议架构:

Node Exporter

Prometheus

Metrics Server

Kube-State-Metrics

Alertmanager

企业微信/钉钉/邮件

Grafana

运维大屏

✅ 推荐部署:Prometheus + Alertmanager + Grafana + Node Exporter + Kube-State-Metrics


3.8 etcd 状态与备份检查 🔐

etcd 是 K8s 的“大脑”,存储所有集群状态。一旦 etcd 崩溃,整个集群将不可用。

检查项:
  • etcd pod 是否运行
  • etcd 成员是否健康
  • etcd 数据库大小(是否超过 2GB)
  • 是否有定期备份(每日快照)
  • TLS 证书是否过期
Java 代码示例:调用 etcdctl 检查健康状态(通过 Shell 执行)
import java.io.BufferedReader;
import java.io.InputStreamReader;

public class EtcdHealthChecker {

    public static void checkEtcdHealth() {
        try {
            Process process = Runtime.getRuntime().exec("kubectl exec -n kube-system etcd-control-plane -- etcdctl endpoint health");
            BufferedReader reader = new BufferedReader(new InputStreamReader(process.getInputStream()));
            String line;
            while ((line = reader.readLine()) != null) {
                System.out.println("🔐 etcd 健康检查: " + line);
                if (!line.contains("healthy")) {
                    System.out.println("❌ etcd 节点异常!");
                }
            }
        } catch (Exception e) {
            System.out.println("⚠️ 无法执行 etcdctl 命令: " + e.getMessage());
        }
    }

    public static void main(String[] args) {
        checkEtcdHealth();
    }
}

💡 实际生产中,建议使用 etcdctl endpoint status 获取详细状态。

备份建议:
  • 每日 02:00 自动执行 etcdctl snapshot save /backup/snapshot.db
  • 将快照上传至 S3 或 MinIO
  • 保留 7 天快照
# 示例备份脚本
ETCDCTL_API=3 etcdctl \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key \
  snapshot save /backup/etcd-snapshot-$(date +%Y%m%d-%H%M%S).db

🔗 etcd Backup and Restore —— etcd 官方恢复指南


3.9 安全合规检查 🔒

安全是巡检的重中之重。K8s 默认配置存在大量风险。

检查项:
  • 是否启用 RBAC?
  • 是否禁用匿名访问?
  • Pod 是否以 root 用户运行?
  • 是否启用 PodSecurityPolicy / PodSecurity Admission?
  • 是否存在默认 ServiceAccount 暴露?
  • 是否开启网络策略(NetworkPolicy)?
Java 代码示例:检测 Pod 是否以 root 运行
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1Pod;
import io.kubernetes.client.openapi.models.V1SecurityContext;
import io.kubernetes.client.util.Config;

import java.io.IOException;
import java.util.List;

public class SecurityContextChecker {

    public static void main(String[] args) throws IOException {
        ApiClient client = Config.defaultClient();
        Configuration.setDefaultApiClient(client);

        CoreV1Api api = new CoreV1Api();
        List<V1Pod> pods = api.listPodForAllNamespaces(null, null, null, null, null, null, null, null, null, null).getItems();

        System.out.println("🛡️ 安全上下文检查:");
        System.out.println("========================");

        for (V1Pod pod : pods) {
            for (var container : pod.getSpec().getContainers()) {
                V1SecurityContext securityContext = container.getSecurityContext();
                if (securityContext != null) {
                    Integer runAsUser = securityContext.getRunAsUser();
                    if (runAsUser != null && runAsUser == 0) {
                        System.out.println("❌ " + pod.getMetadata().getNamespace() + "/" + pod.getMetadata().getName() +
                                " - 容器 " + container.getName() + " 以 root 用户运行");
                    }
                } else {
                    System.out.println("⚠️ " + pod.getMetadata().getNamespace() + "/" + pod.getMetadata().getName() +
                            " - 容器 " + container.getName() + " 未设置 securityContext");
                }
            }
        }
    }
}
安全加固建议:
检查项建议
禁用匿名访问--anonymous-auth=false
启用 RBAC必须开启
使用非 root 用户Dockerfile 中 USER 1000
启用网络策略默认拒绝所有流量
禁用 HostNetwork避免绕过网络隔离
使用镜像签名Notary 或 Cosign

🔗 CIS Kubernetes Benchmark —— CIS 官方 K8s 安全基线(V1.23)


3.10 日志与事件聚合分析 📋

K8s 事件(Events)是诊断问题的“第一手资料”。

检查项:
  • 是否有高频 FailedScheduling 事件?
  • 是否有 FailedCreatePodSandBox?
  • 是否有 ImagePullBackOff?
  • 是否有 Evicted 事件?
Java 代码示例:读取命名空间事件
import io.kubernetes.client.openapi.ApiClient;
import io.kubernetes.client.openapi.Configuration;
import io.kubernetes.client.openapi.apis.CoreV1Api;
import io.kubernetes.client.openapi.models.V1Event;
import io.kubernetes.client.util.Config;

import java.io.IOException;
import java.util.List;

public class EventAnalyzer {

    public static void main(String[] args) throws IOException {
        ApiClient client = Config.defaultClient();
        Configuration.setDefaultApiClient(client);

        CoreV1Api api = new CoreV1Api();
        List<V1Event> events = api.listNamespacedEvent("default", null, null, null, null, null, null, null, null, null).getItems();

        System.out.println("📝 最近 10 条事件摘要:");
        System.out.println("========================");

        for (int i = 0; i < Math.min(10, events.size()); i++) {
            V1Event event = events.get(i);
            System.out.println("⏰ " + event.getLastTimestamp() +
                    " | " + event.getReason() +
                    " | " + event.getMessage() +
                    " | " + event.getSource().getComponent());
        }
    }
}
高频事件应对策略:
事件原因解决
FailedScheduling资源不足扩容节点、调整 Request
FailedCreatePodSandBoxCRI(containerd/dockershim)异常重启 kubelet
Evicted内存压力增加内存、优化应用
Unhealthy探针失败调整 timeout/period

✅ 建议:将所有事件导入 Loki + Grafana,建立“事件热力图”。


四、巡检自动化架构设计 🏗️

4.1 巡检工具整体架构

巡检触发器

巡检引擎

节点检查模块

Pod 检查模块

Service/Ingress 模块

存储检查模块

安全检查模块

指标采集模块

事件分析模块

数据聚合器

报告生成器

HTML/JSON 报告

告警推送

企业微信/钉钉/邮件

存储至 MinIO/S3

历史趋势分析

4.2 部署方式建议

方式适用场景优点缺点
CronJob + Java 容器生产环境自动化、可监控、可版本控制需要构建镜像
外部 Java 服务(K8s 外)多集群统一巡检独立部署,不干扰集群需配置 kubeconfig 访问
Python/Shell 脚本小型集群快速开发不易维护,无类型安全
Operator 模式企业级深度集成,声明式开发复杂度高

✅ 推荐方案:CronJob + Java 容器,每日 02:00 自动执行,结果存入 MinIO,告警推送钉钉。

4.3 巡检工具 Dockerfile 示例

FROM openjdk:17-jre-slim

WORKDIR /app

COPY target/cluster-checker-1.0.jar app.jar

# 安装 kubectl(用于执行辅助命令)
RUN apt-get update && apt-get install -y curl && \
    curl -LO "https://dl.k8s.io/release/$(curl -L -s https://dl.k8s.io/release/stable.txt)/bin/linux/amd64/kubectl" && \
    install -o root -g root -m 0755 kubectl /usr/local/bin/kubectl

# 挂载 kubeconfig
VOLUME /root/.kube

ENTRYPOINT ["java", "-jar", "/app/app.jar"]

4.4 CronJob 配置示例

apiVersion: batch/v1
kind: CronJob
metadata:
  name: cluster-health-check
  namespace: monitoring
spec:
  schedule: "0 2 * * *"  # 每天凌晨2点
  jobTemplate:
    spec:
      template:
        spec:
          serviceAccountName: cluster-checker-sa
          containers:
          - name: checker
            image: your-registry/cluster-checker:latest
            volumeMounts:
            - name: kubeconfig
              mountPath: /root/.kube
              readOnly: true
          restartPolicy: OnFailure
          volumes:
          - name: kubeconfig
            hostPath:
              path: /root/.kube/config

⚠️ 安全提醒:为巡检工具创建专用 ServiceAccount,仅授予 get/list/watch 权限,禁止 delete 权限!


五、巡检报告模板与可视化 📊

5.1 报告内容结构(JSON 示例)

{
  "timestamp": "2024-06-15T02:00:00Z",
  "cluster_name": "prod-us-east",
  "nodes": {
    "total": 12,
    "ready": 11,
    "not_ready": 1,
    "high_cpu_nodes": ["node-07"]
  },
  "pods": {
    "total": 89,
    "crashloop_backoff": 2,
    "pending": 0,
    "missing_requests": 15
  },
  "storage": {
    "pvc_pending": 1,
    "pv_failed": 0,
    "disk_usage_high": ["pvc-redis-01"]
  },
  "security": {
    "root_containers": 3,
    "unsecured_serviceaccounts": 1
  },
  "metrics": {
    "avg_cpu_usage": 78.5,
    "avg_memory_usage": 65.2,
    "etcd_health": "healthy"
  },
  "events_summary": [
    "FailedScheduling: 3",
    "ImagePullBackOff: 1",
    "Evicted: 2"
  ],
  "status": "WARNING",
  "recommendations": [
    "为 node-07 添加资源或驱逐低优先级 Pod",
    "为 15 个未设置 Request 的 Pod 补充资源配置",
    "检查 pvc-redis-01 存储容量是否可扩容"
  ]
}

5.2 报告可视化(HTML 模板片段)

<div class="dashboard">
  <h2>📊 集群健康状态:⚠️ 警告</h2>
  <div class="status-grid">
    <div class="card red">节点异常: 1</div>
    <div class="card yellow">Pod 重启: 2</div>
    <div class="card green">存储正常: 100%</div>
    <div class="card blue">安全风险: 3</div>
  </div>
  <h3>📈 资源使用趋势</h3>
  <div id="cpu-chart">[ECharts 图表]</div>
  <h3>🚨 最近事件</h3>
  <ul>
    <li>06-14 23:45: Pod redis-01 被驱逐(内存不足)</li>
    <li>06-14 23:30: 镜像 registry.example.com/v1.2.3 拉取失败</li>
  </ul>
</div>

📌 建议使用 ECharts 或 Chart.js 渲染趋势图,生成 HTML 报告后通过邮件发送。


六、巡检策略优化与持续改进 🔄

6.1 告警分级策略(建议)

级别触发条件响应机制
🟢 Info资源使用率 > 70%日志记录,周报汇总
🟡 WarningPod 重启 > 5 次/小时、节点 NotReady钉钉通知,1 小时内响应
🔴 Criticaletcd 不健康、所有 Pod Pending、Ingress 5xx > 10%电话通知,立即介入

6.2 建立“巡检基线”

  • 记录正常状态下的指标值(如 CPU 平均 40%)
  • 每月对比趋势,动态调整阈值
  • 对比不同环境(dev/stage/prod)差异

6.3 集成到 CI/CD 流水线

在发布前增加“巡检前置检查”:

- name: Pre-deploy Health Check
  run: |
    java -jar cluster-checker.jar --mode pre-deploy
    if [ $? -ne 0 ]; then
      echo "❌ 集群健康检查失败,禁止发布"
      exit 1
    fi

✅ 这样可避免“发布导致集群雪崩”。


七、常见误区与避坑指南 ⚠️

误区正确做法
❌ 只看 kubectl get pods✅ 必须检查 kubectl get pods -o wide + kubectl describe pod
❌ 依赖人工记忆阈值✅ 所有阈值应写入配置文件,支持动态加载
❌ 忽视事件(Events)✅ Events 是诊断的起点,必须监控
❌ 不做备份✅ etcd、coredns、ingress-controller 配置必须定期备份
❌ 使用默认 ServiceAccount✅ 每个应用应使用独立 SA,绑定最小权限
❌ 不做压力测试✅ 每季度模拟节点宕机、网络分区,验证自愈能力

八、总结:巡检即 DevOps 的灵魂 🧠

Kubernetes 不是“部署完就完事”的系统,它是一个需要持续维护的生命体。一个成熟的 DevOps 团队,不会等到业务告警才行动,而是:

每天清晨,第一件事不是打开邮箱,而是打开巡检报告。

通过本文的完整流程与 Java 代码示例,你已掌握:

  • ✅ 六大核心检查维度
  • ✅ 自动化巡检工具开发方法
  • ✅ 集成 Prometheus 与告警体系
  • ✅ 安全合规加固要点
  • ✅ 巡检报告生成与可视化

下一步建议:

  1. ✅ 将本文代码打包为 Docker 镜像
  2. ✅ 部署为 CronJob
  3. ✅ 接入钉钉机器人
  4. ✅ 每周回顾报告,优化阈值
  5. ✅ 建立“巡检知识库”,沉淀处理经验

🌟 真正的自动化,不是代替人,而是让人类专注于更有价值的问题。

当你不再为“为什么服务挂了”而熬夜,而是优雅地在日报中写上:“今日巡检发现 2 个潜在风险,已自动修复”,你,就是云原生时代的守护者。


附录:推荐工具与资源 📚

类型推荐工具链接
监控Prometheus + Grafanahttps://prometheus.io
日志Loki + Grafanahttps://grafana.com/oss/loki/
安全Kube-Benchhttps://github.com/aquasecurity/kube-bench
网络Calico / Ciliumhttps://cilium.io/
存储Longhornhttps://longhorn.io/
巡检框架KubeLinterhttps://github.com/stackrox/kubelinter
配置管理Kyvernohttps://kyverno.io/

💬 最后忠告:不要迷信“一键部署”,要敬畏“持续运维”。K8s 的美,在于它的弹性;它的险,在于它的复杂。唯有系统化巡检,方能行稳致远。


愿你的集群永不宕机,愿你的告警永远清零。 🙏


🙌 感谢你读到这里!
🔍 技术之路没有捷径,但每一次阅读、思考和实践,都在悄悄拉近你与目标的距离。
💡 如果本文对你有帮助,不妨 👍 点赞、📌 收藏、📤 分享 给更多需要的朋友!
💬 欢迎在评论区留下你的想法、疑问或建议,我会一一回复,我们一起交流、共同成长 🌿
🔔 关注我,不错过下一篇干货!我们下期再见!✨

Logo

DAMO开发者矩阵,由阿里巴巴达摩院和中国互联网协会联合发起,致力于探讨最前沿的技术趋势与应用成果,搭建高质量的交流与分享平台,推动技术创新与产业应用链接,围绕“人工智能与新型计算”构建开放共享的开发者生态。

更多推荐