k8s-cluster-api

v2026.09.24

Kubernetes Cluster API v1.12. Covers clusterctl CLI, ClusterClass, GitOps integration. Scripts for health checks, backup, migration, linting. Templates: clusters, DR, Prometheus. Use when provisioning, upgrading, or operating Kubernetes clusters with CAPI, or running clusterctl and ClusterClass workflows. Keywords: CAPI, clusterctl, kubeadm, cluster lifecycle.

GitHub
安装命令
npx skhub add itechmeat/k8s-cluster-api
Markdown
SKILL.md

Kubernetes Cluster API

Kubernetes Cluster API (CAPI) is a Kubernetes sub-project focused on providing declarative APIs and tooling to simplify provisioning, upgrading, and operating multiple Kubernetes clusters.

Overview

Started by SIG Cluster Lifecycle, Cluster API uses Kubernetes-style APIs and patterns to automate cluster lifecycle management. The infrastructure (VMs, networks, load balancers, VPCs) and Kubernetes configuration are defined declaratively, enabling consistent and repeatable cluster deployments across environments.

Why Cluster API?

While kubeadm reduces installation complexity, it doesn't address day-to-day cluster management:

  • How to consistently provision infrastructure across providers and locations?
  • How to automate cluster lifecycle (upgrades, deletion)?
  • How to scale processes to manage any number of clusters?

Cluster API addresses these gaps with declarative, Kubernetes-style APIs that automate cluster creation, configuration, and management.

Goals

  • Manage lifecycle (create, scale, upgrade, destroy) of Kubernetes-conformant clusters via declarative API
  • Work in different environments (on-premises and cloud)
  • Define common operations with swappable implementations
  • Reuse existing ecosystem components (cluster-autoscaler, node-problem-detector)
  • Provide transition path for existing tools to adopt incrementally

Non-Goals

  • Add APIs to Kubernetes core
  • Manage infrastructure unrelated to Kubernetes clusters
  • Force all lifecycle products to use these APIs
  • Manage non-CAPI provisioned clusters
  • Manage single cluster spanning multiple providers
  • Configure machines after create/upgrade

Quick Navigation

TopicReference
Getting Startedgetting-started.md
Concepts & Architectureconcepts.md
Certificatescertificates.md
Bootstrap (Kubeadm/MicroK8s)bootstrap.md
Cluster Operationscluster-operations.md
Experimental Featuresexperimental.md
clusterctl CLIclusterctl.md
Developer Guidedeveloper.md
Troubleshootingtroubleshooting.md
API Reference & Providersapi-reference.md
Security & PSSsecurity.md
Controllerscontrollers.md
Version Migrationsmigrations.md
FAQfaq.md
Best Practicesbest-practices.md

When to Use

  • Provisioning Kubernetes clusters across multiple infrastructure providers
  • Managing cluster lifecycle (create, scale, upgrade, destroy)
  • Automating cluster operations with declarative APIs
  • Implementing GitOps workflows for cluster management
  • Building custom infrastructure providers

Core Concepts

Architecture

┌─────────────────────────────────────────┐
│         Management Cluster              │
│  ┌─────────────┐  ┌─────────────────┐   │
│  │ CAPI Core   │  │ Infrastructure  │   │
│  │ Controllers │  │ Provider        │   │
│  └─────────────┘  └─────────────────┘   │
│  ┌─────────────┐  ┌─────────────────┐   │
│  │  Bootstrap  │  │  Control Plane  │   │
│  │  Provider   │  │  Provider       │   │
│  └─────────────┘  └─────────────────┘   │
└─────────────────────┬───────────────────┘
                      │ manages
          ┌───────────┴───────────┐
          ▼                       ▼
┌─────────────────┐     ┌─────────────────┐
│ Workload        │     │ Workload        │
│ Cluster 1       │     │ Cluster N       │
└─────────────────┘     └─────────────────┘

Key Components

ComponentPurpose
Management ClusterHosts CAPI controllers, manages workloads
Workload ClusterUser clusters managed by CAPI
Infrastructure ProviderProvisions VMs, networks, load balancers
Bootstrap ProviderGenerates cloud-init/ignition configs
Control Plane ProviderManages control plane nodes lifecycle

Core Resources

ResourceDescription
ClusterRepresents a Kubernetes cluster
MachineRepresents a single node/VM
MachineSetManages replicas of Machines
MachineDeploymentDeclarative updates for MachineSets
MachineHealthCheckAutomatic remediation of unhealthy nodes

Quick Start

# Install clusterctl
curl -L https://github.com/kubernetes-sigs/cluster-api/releases/download/v1.12.0/clusterctl-linux-amd64 -o clusterctl
chmod +x clusterctl
sudo mv clusterctl /usr/local/bin/

# Initialize management cluster
clusterctl init --infrastructure docker

# Create workload cluster
clusterctl generate cluster my-cluster --kubernetes-version v1.32.0 --control-plane-machine-count 1 --worker-machine-count 3 | kubectl apply -f -

# Get cluster kubeconfig
clusterctl get kubeconfig my-cluster > my-cluster.kubeconfig

# Delete cluster
kubectl delete cluster my-cluster

Common Workflows

Cluster Lifecycle

# Create cluster from template
clusterctl generate cluster prod-cluster \
  --infrastructure aws \
  --kubernetes-version v1.32.0 \
  --control-plane-machine-count 3 \
  --worker-machine-count 5 \
  | kubectl apply -f -

# Scale workers
kubectl scale machinedeployment prod-cluster-md-0 --replicas=10

# Upgrade Kubernetes version
kubectl patch cluster prod-cluster --type merge -p '{"spec":{"topology":{"version":"v1.33.0"}}}'

# Move cluster to new management cluster
clusterctl move --to-kubeconfig target-mgmt.kubeconfig

Health Monitoring

apiVersion: cluster.x-k8s.io/v1beta1
kind: MachineHealthCheck
metadata:
  name: my-cluster-mhc
spec:
  clusterName: my-cluster
  maxUnhealthy: 40%
  nodeStartupTimeout: 10m
  selector:
    matchLabels:
      cluster.x-k8s.io/cluster-name: my-cluster
  unhealthyConditions:
    - type: Ready
      status: "False"
      timeout: 5m
    - type: Ready
      status: Unknown
      timeout: 5m

Critical Prohibitions

  • Do NOT modify management cluster directly without proper backup
  • Do NOT delete Machine objects directly (use MachineDeployment scale)
  • Do NOT mix provider versions without checking compatibility
  • Do NOT skip cluster upgrade steps (control plane before workers)
  • Do NOT ignore MachineHealthCheck alerts

Release Highlights (1.14.x)

  • Kubernetes compatibility moves to management clusters v1.33.x -> v1.37.x and workload clusters v1.31.x -> v1.37.x by the 1.14.1 line.
  • API types move into a dedicated Golang module: stronger compatibility guarantees and a much smaller, tightly controlled dependency tree that reduces CVE exposure from transitive dependencies.
  • Kubeadm control plane robustness: safe joining of worker nodes on older Kubernetes versions, improved forward etcd leadership during control plane machine deletion, and remediation of unhealthy machines during intermediate steps of chained upgrades.
  • Observability: the upgrade plan is surfaced in cluster status, aggregated machine versions are exposed in status, and runtime extension errors appear in cluster conditions.
  • Scale/performance: ClusterCache clients expose cache tuning options for a smaller memory footprint, and managedFields interning reduces stored state.
  • Deprecation warning: the v1beta1 API is on track to be unserved in CAPI v1.16 (migrate to v1beta2), and Docker* resources will be removed in v1.15 (migrate to Dev* resources). An experimental clusterctl convert command is available.

Release Highlights (1.13.x)

  • Kubernetes compatibility moves to management clusters v1.32.x -> v1.36.x and workload clusters v1.30.x -> v1.36.x by the 1.13.2 line.
  • v1alpha3 and v1alpha4 API versions are now removed; providers should keep moving toward the v1beta2 contract because v1beta1 remains on the path to becoming unserved in a later release.
  • Cluster topology can now drive rolloutAfter for both control plane and MachineDeployment resources.
  • KubeadmControlPlane improves remediation tolerance for multiple failures and better surfaces common join/remediation symptoms.
  • PriorityQueue and ReconcilerRateLimiting are now beta defaults in the 1.13 line, which can change reconciliation behavior under load.

Scripts

Go-based tools in scripts/. Run via go run ./tool-name from the scripts directory.

ToolPurpose
validate-manifestsValidate YAML manifests against CRD schemas
run-clusterctl-diagnoseRun clusterctl describe and save diagnostic report
migration-checkerCheck v1beta1→v1beta2 migration readiness
check-cluster-healthAnalyze conditions across all cluster objects
analyze-conditionsParse and report False/Unknown conditions
scaffold-providerGenerate new provider directory structure
generate-cluster-templateGenerate templates from ClusterClass
export-cluster-stateExport cluster state for backup/move
audit-securityCheck PSS compliance and security posture
timeline-eventsBuild provisioning event timeline
compare-versionsCompare CAPI version specs and API changes
check-provider-contractVerify provider CRD compliance with contracts
lint-cluster-templatesLint and validate CAPI manifests

Assets

Reusable templates in assets/:

  • Cluster templates: cluster-minimal.yaml, cluster-production.yaml, cluster-clusterclass.yaml, clusterclass-example.yaml
  • Provider configs: docker-quickstart.yaml, aws-credentials.yaml, azure-credentials.yaml, provider-matrix.md
  • Operations: upgrade-checklist.md, migration-v1beta2.md, troubleshooting-flow.md, security-audit-report.md, dr-backup-restore.md, etcd-backup.yaml
  • GitOps: argocd-cluster-app.yaml, flux-kustomization.yaml, gitops-rbac.yaml
  • Monitoring: prometheus-alerts.yaml

Links

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

MIT

源路径

skills/k8s-cluster-api

默认分支

master

最新提交

7ae8a00

Tree SHA

47f5439