← Home

Let's build a k8s cluster with Ansible

As part of my HPC curricular internship at Cineca — infrastructure, lifecycle management and security hardening for HPC services on Kubernetes — I built a digital twin of a production-style cluster on my own homelab hardware, fully provisioned with Ansible. Goal: a highly-available Kubernetes cluster (3 control plane + 2 workers), replicated storage, and a service reachable through a single stable IP even if a node dies — all reproducible from scratch, never clicked by hand.

Infrastructure

Host: a Dell R720 — 2× Intel Xeon E5-2650 v2 (16 cores / 32 threads total), 144 GB RAM. Storage: 10× 1 TB SAS disks in a ZFS raidz2 pool (tank, tolerates 2 disk failures), a 250 GB HDD dedicated to Proxmox itself.

Five VMs cloned from a Debian 13 cloud-init template onto tank, each 2 vCPU / 2 GB RAM / 1 TB thin disk, hostname nodeN, fixed MAC → fixed IP via DHCP reservation:

NodesRoleIP
node1–3control plane (cp1–3)<CP_IP_1>–<CP_IP_3>
node4–5workers (w1–2)<WORKER_IP_1>–<WORKER_IP_2>

Software: Kubernetes v1.34 (kubeadm), containerd 1.7, Calico v3.30 as CNI, Longhorn v1.12 for storage, MetalLB v0.16 for service IPs.

VM creation (Proxmox CLI, for N = 1..5):

qm clone 9000 10N --name nodeN --full 1 --storage tank
qm set 10N --net0 virtio=BC:24:11:AA:00:0N,bridge=vmbr0 --onboot 1
qm resize 10N scsi0 1T
qm start 10N

qm clone --full makes an independent copy of the template's disk, so the template can be deleted later without breaking the VMs. qm set --net0 pins the interface's MAC address to a per-node value (...:01 through ...:05), which is what the router's DHCP reservation matches to hand out the same IP every time. qm resize grows the thin-provisioned disk to 1 TB — thin, so tank only allocates space as data is actually written. qm start boots the VM and hands off to cloud-init, which sets the hostname, installs the SSH key and brings up networking.

Node preparation

The whole project is a handful of Ansible files that run in a fixed order, plus one file of shared variables. Each one below is walked through piece by piece — comments stripped from the snippets, explained in the text instead — with a link at the end to download the real file. Quick map, same order everything runs in:

FilePurpose
inventory.inihosts and groups
updateAPT.ymlapt update/upgrade, one node at a time
prep-k8s.ymlcontainerd + kubeadm prerequisites
group_vars/all.ymlshared variables (VIP, versions, pools)
lb.ymlVIP + load balancer for the API server
bootstrap-k8s.ymlinit, join, CNI
longhorn.ymlreplicated storage
test-web.ymlMetalLB, test service, resilience test

inventory.ini

What it does: lists every node, grouped by role, and how Ansible connects to them.

Before any playbook, Ansible needs to know which hosts exist and how to reach them. This is what hosts: k8s in every playbook actually refers to.

[control_plane]
cp1 ansible_host=node1.pve.homelab.nicologiuliani.site
cp2 ansible_host=node2.pve.homelab.nicologiuliani.site
cp3 ansible_host=node3.pve.homelab.nicologiuliani.site

[workers]
w1 ansible_host=node4.pve.homelab.nicologiuliani.site
w2 ansible_host=node5.pve.homelab.nicologiuliani.site

[k8s:children]
control_plane
workers

[k8s:vars]
ansible_user=debian
ansible_ssh_private_key_file=ssh_key/id_ed25519
ansible_ssh_common_args='-o StrictHostKeyChecking=accept-new'

Each line in [control_plane] and [workers] gives a short alias (cp1, w1...) to an actual node, resolved through ansible_host to its DNS name on the router. The alias is what later playbooks and group_vars files use — cp1 instead of node1.pve.homelab.nicologiuliani.site — and it's also what makes a node's role explicit: moving a host between [control_plane] and [workers] is a one-line edit, not a rename.

[k8s:children] merges both groups into one, k8s, which is the group every playbook in this project actually targets (hosts: k8s) — no play needs to list control plane and workers separately just to run against all five nodes. [k8s:vars] applies to every host in that merged group: the SSH user cloud-init created (debian), the path to the private half of the key pair generated for this cluster, and StrictHostKeyChecking=accept-new so Ansible accepts a new node's host key on first contact instead of stopping to ask.

Download inventory.ini → · ↑ file map

updateAPT.yml

What it does: updates and upgrades every node via apt, rebooting one node at a time only if the update actually requires it.

- name: Aggiorna i nodi con APT
  hosts: k8s
  become: true
  serial: 1

serial: 1 runs this whole play on one node at a time instead of all five in parallel — the point is to never reboot every node together.

  tasks:
    - name: Aggiorna la cache apt
      ansible.builtin.apt:
        update_cache: true
        cache_valid_time: 3600

    - name: Upgrade di tutti i pacchetti
      ansible.builtin.apt:
        upgrade: dist
        autoremove: true
        autoclean: true

Refreshes the apt cache (reusing it if it's under an hour old), then a full dist-upgrade, dropping packages that are no longer needed and clearing the local package cache afterwards.

    - name: Controlla se serve il reboot
      ansible.builtin.stat:
        path: /var/run/reboot-required
      register: reboot_required

    - name: Riavvia se necessario
      ansible.builtin.reboot:
        reboot_timeout: 600
      when: reboot_required.stat.exists

Debian itself drops /var/run/reboot-required when an update (typically the kernel) needs a reboot to take effect. This checks for that file and only reboots when it's actually there, waiting up to 10 minutes for the node to come back — combined with serial: 1 above, that's what keeps a kernel update from taking down the whole cluster's worth of nodes at once.

Download updateAPT.yml → · ↑ file map

prep-k8s.yml

What it does: turns a bare Debian node into one kubeadm can use — kernel modules, sysctl, containerd configured correctly, and the kubelet/kubeadm/kubectl packages installed and held.

- name: Preparazione nodi Kubernetes
  hosts: k8s
  become: true
  vars:
    k8s_version: "v1.34"

The Kubernetes version is a variable, set once here and reused later for the apt repository URL — bumping the cluster version later is a one-line change. Check the latest stable release on kubernetes.io before reusing this on a new cluster.

  tasks:
    - name: Disabilita swap a runtime
      ansible.builtin.command: swapoff -a
      when: ansible_facts.swaptotal_mb > 0

    - name: Commenta swap in fstab
      ansible.builtin.replace:
        path: /etc/fstab
        regexp: '^([^#].*\sswap\s.*)$'
        replace: '# \1'

kubelet refuses to start with swap active. swapoff -a turns it off right now, but only runs if there's swap to turn off; commenting out the swap line in fstab stops it from being remounted on the next reboot.

    - name: Moduli kernel al boot
      ansible.builtin.copy:
        dest: /etc/modules-load.d/k8s.conf
        content: |
          overlay
          br_netfilter

    - name: Carica i moduli ora
      ansible.builtin.command: "modprobe {{ item }}"
      loop: [overlay, br_netfilter]
      changed_when: false

overlay is the filesystem containerd uses for container images; br_netfilter makes bridged traffic visible to iptables, which pod networking depends on. The first task makes both load on every future boot; the second loads them immediately, so a reboot isn't needed now. changed_when: false just stops Ansible from reporting a "change" for a modprobe that's a no-op if the module is already loaded.

    - name: Sysctl per Kubernetes
      ansible.builtin.copy:
        dest: /etc/sysctl.d/99-k8s.conf
        content: |
          net.bridge.bridge-nf-call-iptables = 1
          net.bridge.bridge-nf-call-ip6tables = 1
          net.ipv4.ip_forward = 1
      register: sysctl_conf

    - name: Applica sysctl
      ansible.builtin.command: sysctl --system
      when: sysctl_conf.changed

Enables IP forwarding and lets iptables see bridged traffic — both required for pod-to-pod networking. register + when: sysctl_conf.changed means sysctl --system only reruns when this file actually changed, not on every single playbook run.

    - name: Pacchetti base e containerd
      ansible.builtin.apt:
        name: [apt-transport-https, ca-certificates, curl, gpg, containerd]
        state: present
        update_cache: true

Installs containerd itself, plus the tools needed a few tasks later to add the Kubernetes apt repository over HTTPS with a signed key (apt-transport-https, ca-certificates, curl, gpg).

    - name: Directory config containerd
      ansible.builtin.file:
        path: /etc/containerd
        state: directory
        mode: "0755"

    - name: Controlla se config containerd e' completa
      ansible.builtin.command: grep -q SystemdCgroup /etc/containerd/config.toml
      register: containerd_cfg
      failed_when: false
      changed_when: false

    - name: Genera config di default containerd
      ansible.builtin.shell: containerd config default > /etc/containerd/config.toml
      when: containerd_cfg.rc != 0

    - name: Abilita SystemdCgroup
      ansible.builtin.replace:
        path: /etc/containerd/config.toml
        regexp: 'SystemdCgroup = false'
        replace: 'SystemdCgroup = true'
      notify: Restart containerd

Debian's containerd package ships a minimal config.toml that doesn't even mention SystemdCgroup, so a plain "does this file exist" check isn't enough — hence the grep, with failed_when: false so a no-match isn't treated as an error. If the setting isn't in there, the full default config is generated first. Then SystemdCgroup is flipped to true: containerd defaults to the cgroupfs cgroup driver, kubelet expects systemd, and that mismatch is fatal to the control plane. The notify queues a containerd restart, but only if this task actually changed the file.

    - name: Directory plugin CNI
      ansible.builtin.replace:
        path: /etc/containerd/config.toml
        regexp: 'bin_dir = "/usr/lib/cni"'
        replace: 'bin_dir = "/opt/cni/bin"'
      notify: Restart containerd

containerd's default CNI plugin directory on Debian is /usr/lib/cni, but Calico installs its plugins into /opt/cni/bin. Without this, pods stay in ContainerCreating with failed to find plugin "calico". Same restart-on-change pattern as above.

    - name: Keyring directory
      ansible.builtin.file:
        path: /etc/apt/keyrings
        state: directory
        mode: "0755"

    - name: Chiave repo Kubernetes
      ansible.builtin.get_url:
        url: "https://pkgs.k8s.io/core:/stable:/{{ k8s_version }}/deb/Release.key"
        dest: /etc/apt/keyrings/kubernetes-apt-keyring.asc
        mode: "0644"

    - name: Repo Kubernetes
      ansible.builtin.copy:
        dest: /etc/apt/sources.list.d/kubernetes.list
        content: "deb [signed-by=/etc/apt/keyrings/kubernetes-apt-keyring.asc] https://pkgs.k8s.io/core:/stable:/{{ k8s_version }}/deb/ /\n"

Standard signed-by apt repository setup for the official Kubernetes packages: fetch the signing key, drop it in /etc/apt/keyrings, then point a sources.list.d entry at it — both URLs built from the same k8s_version variable set at the top of the file.

    - name: Installa kubelet kubeadm kubectl
      ansible.builtin.apt:
        name: [kubelet, kubeadm, kubectl]
        state: present
        update_cache: true

    - name: Hold dei pacchetti k8s
      ansible.builtin.dpkg_selections:
        name: "{{ item }}"
        selection: hold
      loop: [kubelet, kubeadm, kubectl]

    - name: Abilita servizi
      ansible.builtin.systemd:
        name: "{{ item }}"
        enabled: true
        state: started
      loop: [containerd, kubelet]

  handlers:
    - name: Restart containerd
      ansible.builtin.systemd:
        name: containerd
        state: restarted

Installs the three Kubernetes binaries, then holds them at the dpkg level — without this, the updateAPT.yml dist-upgrade above would be free to bump kubelet out from under a running cluster. Last task enables and starts containerd and kubelet. The handler at the bottom is what actually restarts containerd; it only fires because one of the two notify: Restart containerd tasks earlier changed the config, and Ansible handlers run once, after all tasks, no matter how many tasks notified them.

Download prep-k8s.yml → · ↑ file map

group_vars/all.yml

What it does: one file of shared values — VIP, versions, storage pools — that every playbook in the project can read.

Ansible loads anything under group_vars/ automatically, no include needed; the values below become variables usable as {{ vip }} etc. in any playbook. The real file also carries a vrrp_pass value (the shared secret the three keepalived instances use to authenticate each other over VRRP) — redacted to a placeholder here and in the downloadable copy, since this repo is public.

vip: "<VIP>"
vip_port: 8443
vrrp_pass: "CHANGE_ME"

pod_cidr: "10.244.0.0/16"
calico_version: "v3.30.3"

longhorn_version: "v1.12.1"
longhorn_replicas: 2

metallb_version: "v0.16.1"
metallb_pool: "<METALLB_POOL_START>-<METALLB_POOL_END>"
web_ip: "<WEB_IP>"

vip/vip_port are the address lb.yml sets up next. pod_cidr and calico_version feed the CNI install in bootstrap-k8s.yml. longhorn_replicas is 2 because there are 2 workers right now — the comment in the real file is a reminder to raise it to 3 once a third worker exists. metallb_pool and web_ip are consumed by test-web.yml; the pool sits outside both the DHCP range and the five nodes' addresses, so nothing else on the LAN can be handed one of those IPs by accident.

Download group_vars/all.yml → · ↑ file map

lb.yml

What it does: puts keepalived and haproxy on the three control-plane nodes, so the Kubernetes API stays reachable on one address even if one of them goes down.

- name: Load balancer HA sui control plane (keepalived + haproxy)
  hosts: control_plane
  become: true
  tasks:
    - name: Installa keepalived e haproxy
      ansible.builtin.apt:
        name: [keepalived, haproxy]
        state: present
        update_cache: true

hosts: control_plane — this play only touches the three control-plane nodes, not the workers. keepalived will hold the shared VIP; haproxy will sit behind it and spread API requests across the three apiservers.

    - name: Configurazione haproxy
      ansible.builtin.copy:
        dest: /etc/haproxy/haproxy.cfg
        content: |
          global
              log /dev/log local0
              maxconn 2000
          defaults
              mode tcp
              log global
              option tcplog
              timeout connect 5s
              timeout client 30s
              timeout server 30s
          frontend k8s_api
              bind *:{{ vip_port }}
              default_backend k8s_api_be
          backend k8s_api_be
              option httpchk GET /healthz
              http-check expect status 200
              balance roundrobin
              default-server inter 3s fall 3 rise 2
          {% for h in groups['control_plane'] %}
              server {{ h }} {{ hostvars[h].ansible_facts.default_ipv4.address }}:6443 check check-ssl verify none
          {% endfor %}
      notify: Restart haproxy

Plain TCP load balancing (mode tcp — haproxy isn't trying to read the encrypted API traffic, just forward it). It listens on {{ vip_port }} (8443, from group_vars/all.yml) and round-robins across a backend built with a Jinja loop over groups['control_plane'] — so adding a fourth control-plane node to the inventory is enough for it to show up here too, no manual edit. Each backend server is health-checked against the apiserver's own /healthz endpoint on port 6443; a server that fails 3 checks in a row is pulled out of rotation, and needs 2 good checks to go back in.

    - name: Configurazione keepalived
      ansible.builtin.copy:
        dest: /etc/keepalived/keepalived.conf
        content: |
          global_defs {
              router_id {{ inventory_hostname }}
              enable_script_security
              script_user root
          }
          vrrp_script check_haproxy {
              script "/usr/bin/pgrep -x haproxy"
              interval 2
              fall 2
              rise 2
          }
          vrrp_instance VI_1 {
              state BACKUP
              interface {{ ansible_facts.default_ipv4.interface }}
              virtual_router_id 51
              priority {{ 150 - 10 * groups['control_plane'].index(inventory_hostname) }}
              advert_int 1
              authentication {
                  auth_type PASS
                  auth_pass {{ vrrp_pass }}
              }
              virtual_ipaddress {
                  {{ vip }}
              }
              track_script {
                  check_haproxy
              }
          }
      notify: Restart keepalived

Every node runs the same vrrp_script: check every 2 seconds that haproxy's process is actually alive on this node (pgrep -x haproxy) — a node with a dead haproxy has no business holding the VIP, even if it's otherwise healthy. All three nodes start as BACKUP — nobody is hardcoded as master — and priority is derived from each node's position in the control_plane group (150, 140, 130 for cp1, cp2, cp3): highest priority with a passing health check wins the VIP. auth_pass is the vrrp_pass shared secret from group_vars/all.yml, so a rogue VRRP announcement on the LAN can't steal the VIP. virtual_router_id 51 just has to be unique among any other VRRP setups on this LAN segment (this network doesn't have another one).

    - name: Abilita e avvia i servizi
      ansible.builtin.systemd:
        name: "{{ item }}"
        enabled: true
        state: started
      loop: [haproxy, keepalived]

  handlers:
    - name: Restart haproxy
      ansible.builtin.systemd:
        name: haproxy
        state: restarted

    - name: Restart keepalived
      ansible.builtin.systemd:
        name: keepalived
        state: restarted

Enables and starts both services; the two handlers only fire if their matching config file was actually rewritten by the tasks above.

Download lb.yml → · ↑ file map

Cluster bootstrap

bootstrap-k8s.yml

What it does: runs kubeadm init on the first control plane, joins the other four nodes, installs Calico, and waits until every node and pod is actually Ready.

- name: Init del primo control plane
  hosts: cp1
  become: true
  tags: [init]
  tasks:
    - name: kubeadm init
      ansible.builtin.command: >
        kubeadm init
        --control-plane-endpoint {{ vip }}:{{ vip_port }}
        --pod-network-cidr {{ pod_cidr }}
      args:
        creates: /etc/kubernetes/admin.conf

Runs only on cp1. The important part is --control-plane-endpoint {{ vip }}:{{ vip_port }}: it tells kubeadm the cluster's "official" address is the load balancer's VIP, not cp1 itself — every other node and every kubeconfig will point at the VIP from here on, so any control-plane node can disappear without the endpoint changing. --pod-network-cidr reserves the pod IP range Calico will hand out from. creates: /etc/kubernetes/admin.conf makes the whole task a no-op if the node has already been initialized — that file is one of the first things kubeadm init writes.

    - name: Comando di join
      ansible.builtin.command: kubeadm token create --print-join-command
      register: join_out
      changed_when: false

    - name: Chiave certificati per i join dei control plane
      ansible.builtin.command: kubeadm init phase upload-certs --upload-certs
      register: cert_out
      changed_when: false

    - name: Salva i valori per gli altri nodi
      ansible.builtin.set_fact:
        join_cmd: "{{ join_out.stdout }}"
        cert_key: "{{ cert_out.stdout_lines[-1] }}"

kubeadm token create --print-join-command mints a fresh join token and prints the full command any node needs to join as a worker. Control-plane certificates (etcd, apiserver...) are private and don't travel in that plain token, so upload-certs encrypts them into a temporary Secret in the cluster and hands back a decryption key — only cp2 and cp3 need it, not the workers. set_fact saves both outputs as join_cmd and cert_key, which the next two plays read straight out of hostvars['cp1']. Both command tasks are marked changed_when: false because printing a join command doesn't change anything on cp1 itself.

- name: Join degli altri control plane
  hosts: cp2,cp3
  become: true
  serial: 1
  tags: [join_cp]
  tasks:
    - name: kubeadm join come control plane
      ansible.builtin.command: >
        {{ hostvars['cp1'].join_cmd }} --control-plane
        --certificate-key {{ hostvars['cp1'].cert_key }}
      args:
        creates: /etc/kubernetes/kubelet.conf

- name: Join dei worker
  hosts: workers
  become: true
  tags: [join_workers]
  tasks:
    - name: kubeadm join come worker
      ansible.builtin.command: "{{ hostvars['cp1'].join_cmd }}"
      args:
        creates: /etc/kubernetes/kubelet.conf

cp2 and cp3 reuse cp1's join command with --control-plane and the certificate key added, joining the etcd cluster — serial: 1 so they don't both join at once and race each other. The workers just run the plain join command, and since they don't need certificate-key at all, Ansible runs both in parallel. Same creates: /etc/kubernetes/kubelet.conf guard on both, so re-running the playbook after a partial failure doesn't try to join an already-joined node.

- name: CNI Calico e kubeconfig
  hosts: cp1
  become: true
  tags: [cni]
  environment:
    KUBECONFIG: /etc/kubernetes/admin.conf
  tasks:
    - name: Operator Calico
      ansible.builtin.command: >
        kubectl apply --server-side -f
        https://raw.githubusercontent.com/projectcalico/calico/{{ calico_version }}/manifests/tigera-operator.yaml
      changed_when: false

Back on cp1 only, with KUBECONFIG set so every kubectl call in this play just works without --kubeconfig. First installs the Calico operator itself, applied server-side (lets the API server, not kubectl, merge the object — the manifest is large).

    - name: Configurazione Calico
      ansible.builtin.copy:
        dest: /root/calico-cr.yaml
        content: |
          apiVersion: operator.tigera.io/v1
          kind: Installation
          metadata:
            name: default
          spec:
            calicoNetwork:
              ipPools:
                - cidr: {{ pod_cidr }}
                  encapsulation: VXLAN
                  natOutgoing: Enabled
                  nodeSelector: all()
          ---
          apiVersion: operator.tigera.io/v1
          kind: APIServer
          metadata:
            name: default
          spec: {}

    - name: Applica Calico (riprova finché le CRD sono pronte)
      ansible.builtin.command: kubectl apply -f /root/calico-cr.yaml
      register: calico_apply
      until: calico_apply.rc == 0
      retries: 10
      delay: 10
      changed_when: false

The operator only defines the CRDs; this second manifest is the actual configuration — "use this pod CIDR, wrap inter-node traffic in a VXLAN tunnel, NAT outgoing traffic to the internet, apply to every node". It's applied with retries because the operator's CRDs may not be registered yet right after the previous task — until keeps trying every 10 seconds, up to 10 times, instead of failing on a race condition.

    - name: Directory .kube per l'utente
      ansible.builtin.file:
        path: "/home/{{ ansible_user }}/.kube"
        state: directory
        owner: "{{ ansible_user }}"
        mode: "0755"

    - name: Copia kubeconfig
      ansible.builtin.copy:
        src: /etc/kubernetes/admin.conf
        dest: "/home/{{ ansible_user }}/.kube/config"
        remote_src: true
        owner: "{{ ansible_user }}"
        mode: "0600"

Gives the debian user their own working kubectl setup — a copy of the admin kubeconfig in ~/.kube/config. remote_src: true means this is a local copy on cp1, not a file pushed from the control machine. mode: "0600" matters: that file holds full cluster-admin credentials, so it's readable by nobody but that user.

    - name: Attendi che tutti i nodi siano Ready
      ansible.builtin.command: kubectl wait --for=condition=Ready nodes --all --timeout=600s
      changed_when: false

    - name: Attendi che tutti i pod siano Ready
      ansible.builtin.command: kubectl wait --for=condition=Ready pod --all -A --timeout=600s
      changed_when: false

Two separate waits on purpose: nodes going Ready only means kubelet is up and talking to the API server — if the CNI plugin is broken, pods (CoreDNS, calico's own controllers...) can still sit stuck in ContainerCreating forever while every node looks fine. The second wait is what actually catches that.

Result of the whole playbook: 5 Ready nodes, 3 of them control plane, every pod Running. To start over from scratch on any node: kubeadm reset -f; rm -rf /etc/cni/net.d /var/lib/etcd.

Download bootstrap-k8s.yml → · ↑ file map

Replicated storage

longhorn.yml

What it does: installs Longhorn, Kubernetes' replicated block storage, and waits for it to be fully up before handing control back.

- name: Prerequisiti Longhorn sui worker
  hosts: workers
  become: true
  tags: [prereq]
  tasks:
    - name: Pacchetti iSCSI e NFS
      ansible.builtin.apt:
        name: [open-iscsi, nfs-common]
        state: present
        update_cache: true

    - name: Modulo kernel iscsi_tcp al boot
      ansible.builtin.copy:
        dest: /etc/modules-load.d/longhorn.conf
        content: "iscsi_tcp\n"

    - name: Carica il modulo ora
      ansible.builtin.command: modprobe iscsi_tcp
      changed_when: false

    - name: Abilita e avvia iscsid
      ansible.builtin.systemd:
        name: iscsid
        enabled: true
        state: started

Only hosts: workers — Longhorn exposes volumes to pods over iSCSI (and NFS for RWX), and pods only run on workers, so only workers need the packages, the kernel module loaded at boot, and iscsid running.

- name: Installa Longhorn
  hosts: cp1
  become: true
  tags: [install]
  environment:
    KUBECONFIG: /etc/kubernetes/admin.conf
  tasks:
    - name: Scarica il manifest ufficiale
      ansible.builtin.get_url:
        url: "https://raw.githubusercontent.com/longhorn/longhorn/{{ longhorn_version }}/deploy/longhorn.yaml"
        dest: /root/longhorn.yaml
        mode: "0644"

    - name: Numero repliche della StorageClass
      ansible.builtin.replace:
        path: /root/longhorn.yaml
        regexp: 'numberOfReplicas: "\d+"'
        replace: 'numberOfReplicas: "{{ longhorn_replicas }}"'

    - name: Applica Longhorn
      ansible.builtin.command: kubectl apply -f /root/longhorn.yaml
      changed_when: false

Installation itself only runs on cp1, since it's just kubectl talking to the API server — the actual Longhorn manager pods start themselves on the workers (control planes carry a NoSchedule taint). The official manifest defaults to 3 replicas, but there are only 2 workers here, so that number is patched to {{ longhorn_replicas }} (2, from group_vars/all.yml) before applying — with fewer replicas than workers, volumes would sit degraded forever. Bump longhorn_replicas in group_vars once a third worker exists.

    - name: Attendi i manager Longhorn (uno per worker)
      ansible.builtin.command: kubectl -n longhorn-system rollout status ds/longhorn-manager --timeout=600s
      changed_when: false

    - name: Attendi il driver CSI
      ansible.builtin.command: kubectl -n longhorn-system rollout status deploy/longhorn-driver-deployer --timeout=600s
      changed_when: false

    - name: Attendi che tutti i pod Longhorn siano Ready
      ansible.builtin.command: kubectl -n longhorn-system wait --for=condition=Ready pod --all --timeout=60s
      register: lh_ready
      until: lh_ready.rc == 0
      retries: 10
      delay: 10
      changed_when: false

    - name: Verifica StorageClass longhorn
      ansible.builtin.command: kubectl get storageclass longhorn
      changed_when: false

Three separate waits, each for a different piece: the manager DaemonSet (one pod per worker), then the CSI driver deployer, then — with retries, since the CSI driver spawns further pods after it's up — every remaining pod in the namespace. Only once all of that is Ready does the last task confirm the longhorn StorageClass actually exists and is usable.

Download longhorn.yml → · ↑ file map

One IP for the service, and a resilience test

test-web.yml

What it does: installs MetalLB, gives it an IP pool, then deploys a real nginx + Longhorn test workload behind a single LoadBalancer IP to prove the whole stack works together.

- name: MetalLB
  hosts: cp1
  become: true
  tags: [metallb]
  environment:
    KUBECONFIG: /etc/kubernetes/admin.conf
  tasks:
    - name: Scarica il manifest ufficiale
      ansible.builtin.get_url:
        url: "https://raw.githubusercontent.com/metallb/metallb/{{ metallb_version }}/config/manifests/metallb-native.yaml"
        dest: /root/metallb-native.yaml
        mode: "0644"

    - name: Applica MetalLB
      ansible.builtin.command: kubectl apply -f /root/metallb-native.yaml
      changed_when: false

    - name: Attendi il controller
      ansible.builtin.command: kubectl -n metallb-system rollout status deploy/controller --timeout=300s
      changed_when: false

    - name: Attendi gli speaker (uno per nodo)
      ansible.builtin.command: kubectl -n metallb-system rollout status ds/speaker --timeout=300s
      changed_when: false

A NodePort would expose a service on every node's own IP — not one address. MetalLB fixes that: the controller deployment assigns IPs to LoadBalancer Services from a pool, and the speaker DaemonSet (one pod per node) is what actually announces an assigned IP on the LAN.

    - name: Pool di IP e annuncio L2
      ansible.builtin.copy:
        dest: /root/metallb-pool.yaml
        content: |
          apiVersion: metallb.io/v1beta1
          kind: IPAddressPool
          metadata:
            name: pool
            namespace: metallb-system
          spec:
            addresses:
              - {{ metallb_pool }}
          ---
          apiVersion: metallb.io/v1beta1
          kind: L2Advertisement
          metadata:
            name: l2
            namespace: metallb-system
          spec:
            ipAddressPools: [pool]

    - name: Applica il pool (riprova finche' il webhook di MetalLB e' pronto)
      ansible.builtin.command: kubectl apply -f /root/metallb-pool.yaml
      register: pool_apply
      until: pool_apply.rc == 0
      retries: 10
      delay: 6
      changed_when: false

IPAddressPool is the range MetalLB is allowed to hand out ({{ metallb_pool }}, <METALLB_POOL_START>-<METALLB_POOL_END>); L2Advertisement is what tells it to announce those IPs over plain ARP — mode L2, meaning one node holds and announces a given IP at a time, and another takes over if it disappears, the same idea as keepalived one layer up. Applying the pool right after installing MetalLB can race the admission webhook still starting up, hence the retry loop.

- name: Test - server web nginx con volume Longhorn
  hosts: cp1
  become: true
  tags: [deploy]
  environment:
    KUBECONFIG: /etc/kubernetes/admin.conf
  tasks:
    - name: Scrivi il manifest
      ansible.builtin.copy:
        dest: /root/test-web.yaml
        content: |
          apiVersion: v1
          kind: Namespace
          metadata:
            name: test-web
          ---
          apiVersion: v1
          kind: PersistentVolumeClaim
          metadata:
            name: web-data
            namespace: test-web
          spec:
            accessModes: [ReadWriteOnce]
            storageClassName: longhorn
            resources:
              requests:
                storage: 1Gi

A dedicated namespace, and a 1 GiB PVC against the longhorn StorageClass — this is what actually exercises the storage layer, not just the network one.

          ---
          apiVersion: apps/v1
          kind: Deployment
          metadata:
            name: web
            namespace: test-web
          spec:
            replicas: 1
            strategy:
              type: Recreate
            selector:
              matchLabels: {app: web}
            template:
              metadata:
                labels: {app: web}
              spec:
                initContainers:
                  - name: init-page
                    image: busybox:1.37
                    command:
                      - sh
                      - -c
                      - '[ -f /data/index.html ] || echo "<h1>Longhorn OK</h1><p>pagina creata il $(date) dal pod $(hostname)</p>" > /data/index.html'
                    volumeMounts:
                      - {name: data, mountPath: /data}
                containers:
                  - name: nginx
                    image: nginx:stable-alpine
                    ports:
                      - containerPort: 80
                    volumeMounts:
                      - {name: data, mountPath: /usr/share/nginx/html}
                volumes:
                  - name: data
                    persistentVolumeClaim:
                      claimName: web-data

strategy: Recreate because the PVC is RWO (one writer at a time): with the default rolling update, Kubernetes would try to start the new pod before the old one lets go of the volume, and it would just hang. The initContainer writes index.html with a timestamp and the pod's hostname, but only if that file doesn't already exist — which is exactly the point: if the page survives a pod moving to a different node, its creation time won't change, proving the data itself moved with it rather than being recreated from scratch.

          ---
          apiVersion: v1
          kind: Service
          metadata:
            name: web
            namespace: test-web
            annotations:
              metallb.io/loadBalancerIPs: "{{ web_ip }}"
          spec:
            type: LoadBalancer
            selector: {app: web}
            ports:
              - port: 80
                targetPort: 80

A LoadBalancer Service pinned to a fixed address ({{ web_ip }}, <WEB_IP>) via the MetalLB-specific annotation — without it, MetalLB would just hand out the next free IP from the pool instead of always the same one.

    - name: Attendi che la PVC sia Bound
      ansible.builtin.command:
        argv: [kubectl, -n, test-web, wait, "--for=jsonpath={.status.phase}=Bound", pvc/web-data, --timeout=120s]
      changed_when: false

    - name: Attendi che il pod nginx sia pronto
      ansible.builtin.command: kubectl -n test-web rollout status deploy/web --timeout=300s
      changed_when: false

    - name: Richiedi la pagina sull'IP unico
      ansible.builtin.uri:
        url: "http://{{ web_ip }}"
        return_content: true
      register: web_page
      until: web_page.status == 200
      retries: 10
      delay: 5
      changed_when: false

Three checks in sequence: the PVC actually gets a Longhorn volume and goes Bound, the Deployment's rollout finishes, and finally an HTTP request against web_ip itself — retried for almost a minute, since MetalLB announcing a fresh IP over ARP isn't instant. Getting a real 200 back here is the end-to-end proof: scheduling, pod networking, DNS, storage and the Service all worked together. A separate play, gated behind --tags cleanup (never run by default), tears the whole namespace back down.

Download test-web.yml → · ↑ file map

Resilience test

None of this proves anything survives a node dying until it's actually tested. With the test service up, draining the worker the pod is running on:

kubectl drain <node> --ignore-daemonsets --delete-emptydir-data
kubectl uncordon <node>

drain marks the node unschedulable and evicts every pod on it except DaemonSets, forcing the nginx pod to be rescheduled on the other worker — that's the actual failure being simulated. uncordon makes the node schedulable again once the test is done; it does not move anything back by itself.

Results: