This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
IMPORTANT: This repository allows direct commits and pushes to master branch.
- Direct push to master is ALLOWED and PREFERRED
- Feature branches are optional
- Pull requests are optional
- This overrides global CLAUDE.md Git workflow rules
CRITICAL: ALWAYS use the homelab kubectl context when working with this cluster.
# All kubectl commands MUST use --context homelab
kubectl --context homelab get pods
kubectl --context homelab apply --filename manifest.yaml
# Or set context for session
kubectl config use-context homelab- Context name:
homelab - API server: https://172.16.101.4:6443 (k8s-storage-01)
- Credentials: admin via client certificate
Never run kubectl commands without specifying context — this prevents accidental operations on wrong clusters.
This is a Kubernetes cluster configuration for a single-node home cluster (1x x86_64 storage node) managed via GitOps with ArgoCD. The repository contains all Kubernetes manifests, Helm values, and ArgoCD application definitions for a production home cluster.
Previously the cluster ran 3 ARM64 Raspberry Pi nodes as HA control plane, but these were retired on 2026-04-18 due to recurring macb silent NIC death (LP #2133877). The amd64 storage node now serves as the sole control-plane and workload node.
The repository uses ArgoCD's App-of-Apps pattern:
argocd/meta/meta.yaml- Root application that manages all other ArgoCD applications- ArgoCD applications are organized by category in
argocd/subdirectories:meta/- ArgoCD self-management and projectsinfra/- Infrastructure components (networking, storage, ingress)monitoring/- Monitoring stack componentsworkloads/- User applicationssmarthome/- Smart home related applicationsdefault/- Default namespace applications
- Disabled/experimental applications are kept in
argocd-disabled/
-
argocd/- ArgoCD Application manifests defining what to deploy- Each Application references either Helm charts or directories in
manifests/ - Applications use
automatedsync policy withselfHeal: trueandprune: true - Organized by ArgoCD Projects (meta, infra, monitoring, workloads, smarthome, default)
- Each Application references either Helm charts or directories in
-
manifests/- Raw Kubernetes YAML manifests for applications- Each subdirectory contains manifests for a specific application
- Referenced by ArgoCD Applications for deployment
-
values/- Helm chart values files for infrastructure componentsargocd.yaml- ArgoCD configuration including Crossplane health checks, server.insecure for Gateway TLS termination, and HTTPRoute configurationcoredns.yaml- CoreDNS configuration with custom cluster domain k8s.home.lex.la (also containstemplate IN AAAA . rcode NOERRORto suppress AAAA since the node has IPv6 ULA but no default v6 route — Happy Eyeballs would time out otherwise)cilium.yaml- Cilium CNI configuration (native routing, kube-proxy replacement, L2 announcements, Gateway API, Hubble enabled)
-
charts/- Reserved for vendored upstream Helm charts when consuming them directly is not possible. Currently empty. History:charts/obico/(byte-for-byte copy from the obico-serverreleasebranch) was vendored 2026-07-10 because the upstream repo was too heavy for ArgoCD to clone and no published chart existed; retired 2026-07-11 once upstream started publishing the chart as a signed OCI artifact (oci://ghcr.io/thespaghettidetective/charts/obico, images baked per-commit).charts/authelia/(0.11.3 + localsecret_nameschema fix) was retired 2026-04-25 once upstream 0.11.4 landed an equivalent fix; Authelia consumesoci://ghcr.io/authelia/chartrepodirectly. -
secrets/- Bootstrap secrets (GPG encrypted)openbao-seal-key.yaml.asc- OpenBao auto-unseal key (required before OpenBao starts)authelia.yaml.asc- Authelia OIDC secrets (references OpenBao paths)- All other secrets migrated to OpenBao and managed via External Secrets Operator
The cluster runs on K3s with these core components (in deployment order):
- DNS: CoreDNS with custom cluster domain k8s.home.lex.la
- Networking: Cilium CNI with native routing on 10.42.0.0/16
- Control Plane: Single-node k3s server on storage-01 (no HA). extractedprism per-node TCP LB (127.0.0.1:7445) still deployed as thin abstraction for CNI/kubelet
- GitOps: ArgoCD (self-managed via the meta application)
- Storage: OpenEBS ZFS LocalPV on pool/* datasets (primary); Longhorn still installed but with ZFS data locality (mostly unused after CP migration)
- Load Balancing: Cilium L2 Announcements (LB IPAM) with dedicated IP pools
- Gateway API: Cilium Gateway API for internal routing + Cloudflare Tunnel for public routing
- Certificate Management: cert-manager with automatic Gateway API integration
- External DNS: external-dns with Gateway API HTTPRoute support
- Monitoring: Grafana operator, node-exporter, metrics-server
- Secrets Management: OpenBao (Vault fork) with auto-unseal
- Secrets Sync: External Secrets Operator (ESO) for OpenBao → K8s sync
- Authentication: Authelia for SSO/OIDC (integrated with OpenBao)
- Workflow Automation: Argo Workflows for Kubernetes-native workflow orchestration
- Event-Driven Automation: Argo Events for event-driven workflow triggers
- Node uplink:
ens3f0— HP 560SFP+ (Intel 82599ES,ixgbe) in PCI-E Slot 3, 10G, carries the172.16.101.4reservation at netplanroute-metric: 100. Onboardeno2(1G) is the fallback at metric 200.- Predictable names encode the PCI slot, so the same card is
ens2f0in slot 2 andens3f0in slot 3. Netplan (ansible/roles/node-prep/files/90-all-ethernets.yaml) matches on driver plus port suffix rather than name so a reseat does not leave the port unconfigured. - Slot 2 is x8 mechanically but x1 electrically (PCH root port), which caps a card there at 4 Gb/s.
dmidecodereports the connector type, not the lane count — checkLnkCapon the root port. - Changing the active uplink requires three places to agree: k3s
node-ip, Ciliumdevices, and the netplan metrics. Disagreement breaks either etcd member identity or pod egress masquerading.
- Predictable names encode the PCI slot, so the same card is
- extractedprism: Per-node TCP load balancer on 127.0.0.1:7445, retained for single-CP architecture
- DaemonSet with hostNetwork, static bootstrap endpoint (172.16.101.4:6443) + dynamic EndpointSlice discovery
- Cilium k8sServiceHost points to extractedprism (127.0.0.1:7445) instead of direct apiserver IP
- Reduced value for single CP (no failover to balance) — kept as a thin stable abstraction should CP ever grow again
- Pod network: 10.42.0.0/16 (Cilium tunnel mode with VXLAN)
- Cilium kube-proxy replacement for service load balancing and NodePort
- Cilium L2 Announcements for LoadBalancer IP allocation:
- Internal Gateway pool: 172.16.100.250
- Syslog pool: 172.16.100.249
- Transmission pool: 172.16.100.252
- Minecraft pool: 172.16.100.253
- ArgoCD API pool: 172.16.100.254
- Default pool: 172.16.100.101-110
- Dual Gateway setup for HTTP/HTTPS routing:
- Public Gateway (cloudflare-tunnel): via Cloudflare Tunnel
- GatewayClass:
cloudflare-tunnel(managed by cloudflare-tunnel-gateway-controller) - Traffic routed through Cloudflare Tunnel (no direct IP/port-forwarding)
- Tunnel ID: 59db961d-3851-468a-bf80-4d39f942b9e0
- Public hostnames: eta.lex.la, job.lex.la, map.lex.la, aleksei.sviridk.in, abs.lex.la, wish.lex.la, test.lex.la, authelia.lex.la
- DDoS protection via Cloudflare
- GatewayClass:
- Internal Gateway (cilium-gateway-internal): 172.16.100.250
- GatewayClass:
cilium(managed by Cilium) - NOT port-forwarded - accessible only from local network
- Uses wildcard listener (*.home.lex.la) for simplified internal routing
- Cloudflare DNS-only mode (grey cloud) - no proxy
- Automatic HTTP→HTTPS redirect (301) via dedicated HTTPRoute
- GatewayClass:
- TLS certificates automatically managed by cert-manager via Gateway API integration
- Supports wildcard certificates for *.lex.la, *.home.lex.la, *.k8s.home.lex.la, *.sviridk.in
- Public Gateway (cloudflare-tunnel): via Cloudflare Tunnel
- External DNS automatically creates DNS records from Gateway annotations
- Internal Gateway: external-dns.alpha.kubernetes.io/target: "172.16.100.250" (DNS-only)
- Public DNS records managed by Cloudflare Tunnel automatically
- Cluster domain:
k8s.home.lex.la(configured in K3s and CoreDNS) - Hubble: Enabled for network observability
- Web UI: https://hubble.home.lex.la (internal Gateway)
- CLI:
hubble --server 172.16.100.101:4245 observe - Relay LoadBalancer: 172.16.100.101 (pinned via annotation)
- CRITICAL:
peerService.clusterDomain: k8s.home.lex.larequired due to custom cluster domain
Bootstrap the cluster (run on management machine after K3s installation):
# Add Helm repositories
helm repo add coredns https://coredns.github.io/helm
helm repo add cilium https://helm.cilium.io/
helm repo add argo https://argoproj.github.io/argo-helm
helm repo update
# Install core components in order
helm install coredns coredns/coredns --namespace kube-system --values values/coredns.yaml
helm install cilium cilium/cilium --namespace kube-system --values values/cilium.yaml
helm install argocd argo/argo-cd --namespace argocd --values values/argocd.yaml --create-namespace
# Apply Cilium LB IP pools, L2 announcement policy, and Gateway
kubectl apply --filename manifests/cilium/
# Deploy meta application (deploys all other applications via GitOps)
# This will deploy extractedprism (per-node API LB) and all other apps
kubectl apply --filename argocd/meta/meta.yamlIMPORTANT: Always use ArgoCD CLI for operations, not kubectl commands directly.
# Login to ArgoCD
argocd login api.argocd.home.lex.la:443 --username admin --plaintext
# Check all ArgoCD applications
argocd app list
# View application details
argocd app get APP_NAME
# Sync a specific application (after GitOps changes)
# 1. Make changes in git repository
# 2. Commit and push
# 3. Sync meta app first (it manages all other apps)
argocd app sync argocd/meta
# 4. Sync target application
argocd app sync argocd/APP_NAME
# Force hard refresh (reconcile from git)
argocd app sync argocd/APP_NAME --prune --forceGitOps Workflow for Updates:
- Make changes in git repository (manifests, values, argocd definitions)
- Commit and push to master
- CRITICAL: When changing Application definitions (argocd/CATEGORY/*.yaml):
- First sync meta app:
argocd app sync argocd/metaorkubectl annotate application meta --namespace argocd argocd.argoproj.io/refresh=normal --overwrite - Meta app will update Application CRDs in cluster
- Then sync target app:
argocd app sync argocd/TARGET_APP
- First sync meta app:
- When changing only manifests/values (NOT Application definitions):
- Sync target application directly:
argocd app sync argocd/TARGET_APP
- Sync target application directly:
- Verify:
argocd app get argocd/TARGET_APP
Before committing:
# Validate Kubernetes YAML syntax
kubectl apply --dry-run=client --filename manifests/APP_NAME/
# Validate ArgoCD application
kubectl apply --dry-run=client --filename argocd/CATEGORY/APP_NAME.yaml
# Check Helm values render correctly (for components using Helm)
helm template TEST_NAME CHART_NAME --values values/COMPONENT.yamlArgo Workflows provides Kubernetes-native workflow orchestration. Argo Events enables event-driven automation.
# View workflows
argo list --namespace argo-events
# View workflow details
argo get WORKFLOW_NAME --namespace argo-events
# View workflow logs
argo logs WORKFLOW_NAME --namespace argo-events
argo logs @latest --namespace argo-events
# Manual workflow submission from template
argo submit --from workflowtemplate/system-upgrade --namespace argo-events
# Check EventSource status
kubectl get eventsource --namespace argo-events
# Check Sensor status
kubectl get sensor --namespace argo-events
# Check EventBus status
kubectl get eventbus --namespace argo-events
# Web UI
# https://argo-workflows.home.lex.laEvent-Driven Architecture:
- EventSource (calendar, resource, webhook, GitHub, etc.) → EventBus (NATS JetStream) → Sensor → Workflow
- Calendar EventSource triggers workflows on schedule (e.g., weekly system upgrades)
- Resource EventSource can trigger on Kubernetes events (new nodes, pod failures, etc.)
- Sensor watches EventBus and submits WorkflowTemplates when events match
Example Use Cases:
- Weekly system upgrade plan application via calendar trigger
- New node provisioning automation via resource events
- Automated recovery workflows on pod failures
- GitHub webhook-triggered deployments
Secrets are managed via OpenBao (Vault fork) and External Secrets Operator (ESO):
-
OpenBao: Central secrets store deployed in
securitynamespace- UI: https://openbao.home.lex.la
- Auto-unseal via static seal with key from
secrets/openbao-seal-key.yaml.asc - KV v2 secrets engine at
secret/ - Kubernetes auth method for ESO
-
External Secrets Operator: Syncs secrets from OpenBao to Kubernetes
ClusterSecretStorenamedopenbaoconnects to OpenBaoExternalSecretresources in each namespace define what to sync- Secrets refresh every 1 hour
-
Bootstrap secrets in
secrets/(GPG encrypted):openbao-seal-key.yaml.asc- Required BEFORE OpenBao startsauthelia.yaml.asc- Uses path references to OpenBao secrets
# Check ExternalSecrets status
kubectl get externalsecrets --all-namespaces
# Check ClusterSecretStore
kubectl get clustersecretstore openbao
# View secret in OpenBao (requires root token or appropriate policy)
kubectl exec -it openbao-0 -n security -- bao kv get secret/PATHThe ansible/ directory contains Ansible playbooks for node management. Always use Ansible for node operations instead of ad-hoc SSH commands.
# IMPORTANT: All ansible commands must be run from the ansible/ directory
cd ansible/
# Upgrade all nodes with automatic reboot if required (TRUSTED - use this by default)
ansible-playbook playbooks/upgrade-nodes.yaml
# Upgrade without automatic reboot (only if you need manual control)
ansible-playbook playbooks/upgrade-nodes.yaml --extra-vars "auto_reboot=false"
# Upgrade only control plane
ansible-playbook playbooks/upgrade-nodes.yaml --limit server
# Upgrade only workers
ansible-playbook playbooks/upgrade-nodes.yaml --limit agent
# Upgrade specific node
ansible-playbook playbooks/upgrade-nodes.yaml --limit k8s-cp-01
# Ad-hoc commands (must be from ansible/ directory for inventory and become)
ansible k3s_cluster --module-name shell --args "uname -r"
ansible k3s_cluster --module-name apt --args "name=PACKAGE state=latest"Configuration:
- Inventory:
ansible/inventory/production.yaml - SSH user:
ansible(with dedicated key~/.ssh/ansible_ed25519) - Become: enabled by default (passwordless sudo)
- Serial execution: nodes upgraded one at a time to maintain cluster availability
Trust auto_reboot:
- The
upgrade-nodes.yamlplaybook handles reboots safely - Nodes are upgraded sequentially (serial: 1)
- Reboot only happens if
/var/run/reboot-requiredexists - Post-reboot delay ensures node is ready before proceeding
- Use
auto_reboot=true(default) for routine upgrades
K3s installation and upgrades use the k3s.orchestration collection.
# IMPORTANT: All ansible commands must be run from the ansible/ directory
cd ansible/
# Install/upgrade K3s cluster (full deployment)
ansible-playbook k3s.orchestration.site
# Upgrade K3s only (sequential server upgrade, then agents)
ansible-playbook k3s.orchestration.upgrade
# Reset K3s cluster (DESTRUCTIVE)
ansible-playbook k3s.orchestration.reset
# Reboot all nodes
ansible-playbook k3s.orchestration.rebootK3s Version:
- Defined in
ansible/inventory/production.yaml→k3s_version - Renovate auto-updates version via GitHub releases datasource
- Collection playbooks use FQCN format:
k3s.orchestration.<playbook>
Collection Installation:
cd ansible/
ansible-galaxy collection install --requirements-file requirements.yaml- Create manifests in
manifests/NEW_APP/ - Create ArgoCD Application in appropriate
argocd/CATEGORY/directory - Reference the manifests directory in the Application spec
- Choose appropriate Gateway based on service accessibility:
- Public services (accessible from internet): Use
cloudflare-tunnelGateway - Internal services (home network only): Use
cilium-gateway-internal
- Public services (accessible from internet): Use
- Create HTTPRoute resource:
# Example for PUBLIC service (via Cloudflare Tunnel) apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: name: your-app namespace: your-namespace spec: parentRefs: - name: cloudflare-tunnel namespace: cloudflare-tunnel-system sectionName: https hostnames: - your-app.lex.la rules: - matches: - path: type: PathPrefix value: / backendRefs: - name: your-service port: 80
# Example for INTERNAL service apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: name: your-app namespace: your-namespace spec: parentRefs: - name: cilium-gateway-internal namespace: kube-system sectionName: https-home-lex-la # Wildcard listener for *.home.lex.la hostnames: - your-app.home.lex.la rules: - matches: - path: type: PathPrefix value: / backendRefs: - name: your-service port: 80
- For PUBLIC services: Configure DNS in Cloudflare to point to the tunnel (automatic via cloudflare-tunnel-gateway-controller)
- For INTERNAL services: Use existing wildcard listener
https-home-lex-la(no Gateway changes needed) - Add your namespace to ReferenceGrant in
manifests/cilium/reference-grant.yaml(for internal Gateway routes) - Commit and push - ArgoCD meta app will auto-deploy
- Update Helm values in
values/COMPONENT.yaml - For components with HTTPRoute support (e.g., ArgoCD), use built-in Helm chart configuration instead of manual manifests
- Commit changes
- ArgoCD will detect and sync changes automatically (selfHeal: true)
Move ArgoCD Application manifest from argocd/CATEGORY/ to argocd-disabled/
- ZFS pool on storage-01 is scrubbed monthly via
zfs-scrub-monthly@pool.timer(ansible-managed inplaybooks/zfs.yaml);zfs-zedruns and emails alerts on pool/disk errors - K3s cluster with specific components disabled (local-storage, servicelb, metrics-server, coredns, kube-proxy, flannel, traefik)
- Custom CoreDNS and Cilium replace default K3s networking
- Designed for ARM64 architecture (Raspberry Pi)
- Node access: Only network (SSH/Talos API) and UART available — no HDMI/display access to nodes
- All changes deploy automatically via ArgoCD (selfHeal: true, prune: true)
- Domain references throughout use
lex.la- must be updated for different domains - Public services use Cloudflare Tunnel via cloudflare-tunnel-gateway-controller (GatewayClass: cloudflare-tunnel)
- cert-manager automatically manages certificates for Gateway API listeners
- TLS certificates issued via DNS-01 ACME challenge with Cloudflare API
- Gateway API provides modern, role-oriented routing instead of legacy Ingress
- SECURITY: Public Gateway uses Cloudflare Tunnel — traffic never touches local IP directly
- SECURITY: Internal Gateway uses wildcard (*.home.lex.la) since it's not exposed to internet
- SECURITY: Internal Gateway (172.16.100.250) MUST NOT be port-forwarded - local network access only
- Cluster domain: k8s.home.lex.la is configured in both K3s and CoreDNS, must be consistent across all components
- extractedprism: Per-node TCP LB DaemonSet (127.0.0.1:7445 → all control plane endpoints)
- Cilium k8sServiceHost points to extractedprism (cannot be empty due to kube-proxy replacement)
- Static bootstrap endpoints + dynamic Kubernetes EndpointSlice discovery
- No CNI dependency at boot (hostNetwork + static endpoints)
- ArgoCD server.insecure MUST be true when behind Gateway with TLS termination
- ArgoCD HTTPRoute is managed via Helm chart values (server.httproute section), not manual manifests
- ArgoCD dual-access architecture:
- WebUI: https://argocd.home.lex.la (via Gateway API with TLS termination)
- CLI/gRPC: api.argocd.home.lex.la:443 (via dedicated LoadBalancer 172.16.100.254, plaintext)
- Gateway uses HTTP/1.1 (breaks gRPC), LoadBalancer allows direct cmux access for CLI
- ArgoCD LoadBalancer uses dedicated IP pool (argocd-api-pool) with external-dns integration
- Argo Workflows: Deployed in argo-events namespace for workflow orchestration
- WorkflowTemplates define reusable workflow specifications
- ServiceAccount argo-workflow has ClusterRole for kubectl operations
- Web UI at https://argo-workflows.home.lex.la (internal Gateway)
- Argo Events: Event-driven automation framework deployed in argo-events namespace
- EventBus uses NATS JetStream with Longhorn persistence (1Gi)
- EventSources define event triggers (calendar, resource, webhook, GitHub, etc.)
- Sensors watch EventBus and submit Workflows when events match
- Current setup: Calendar EventSource triggers weekly system-plan.yaml application
- ZFS Storage: k8s-storage-01 (172.16.101.4) provides local ZFS storage via OpenEBS ZFS LocalPV
- StorageClass
zfs-localpv: native ZFS datasets, compression, thin provisioning - Static PVs for existing datasets: pool/lex, pool/dump, pool/transmission, pool/papermc-data
- Dynamic PVs for workloads: OpenBao, Loki, VMSingle, Transmission config
- Node-local only — pods using zfs-localpv are pinned to k8s-storage-01 via nodeAffinity
- StorageClass
- Hubble configuration constraints:
hubble.peerService.clusterDomainMUST match cluster domain (k8s.home.lex.la) - Relay DNS resolution fails otherwise- Hubble Relay IP pinning: Helm chart doesn't support service annotations, use
kubectl annotate svc hubble-relay --namespace kube-system io.cilium/lb-ipam-ips=IP(ArgoCD won't overwrite since Helm doesn't set annotations)
- OpenBao: Deployed in
securitynamespace with static auto-unseal- Seal key stored in
secrets/openbao-seal-key.yaml.asc(must exist before OpenBao starts) - Uses
zfs-localpvStorageClass on k8s-storage-01
- Seal key stored in
- External Secrets Operator: Deployed in
external-secretsnamespace- ClusterSecretStore
openbaoauthenticates via Kubernetes auth method - All application secrets synced from OpenBao KV v2 (
secret/path)
- ClusterSecretStore
- Authelia: OIDC provider in
securitynamespace- Configured clients: Grafana, ArgoCD, OpenBao
- OIDC secrets stored in OpenBao, referenced via path in authelia config
- Uses
claims_policiesto include groups in ID token (Authelia 4.38+ feature) - Dex OIDC connector requires
insecureEnableGroups: trueto read groups from ID token (name refers to stale data on refresh, not a security issue)
CRITICAL: All services MUST use OIDC authentication via Authelia. Local/password authentication is FORBIDDEN.
- ArgoCD: Local admin login disabled (
admin.enabled: false), OIDC via Dex → Authelia - Grafana: Login form disabled (
disable_login_form: true), OIDC via Authelia - OpenBao: OIDC via Authelia
- When adding new services that support authentication, ALWAYS configure OIDC with Authelia
- NEVER enable local password authentication for any service
- User groups are managed in Authelia
users_database.yml - RBAC policies reference Authelia groups (e.g.,
adminsgroup → admin role)
Renovate bot is configured with:
- Semantic commits enabled
- Automerge enabled
- Version pinning for all dependencies
- Separate labels for major/minor/argocd/github-actions updates
- Staging periods: 5 days for major, 3 days for minor
- PR limits: 5 hourly, 10 concurrent
- Timezone: Asia/Tbilisi