Workload cluster release aws-35.2.0 for CAPA

Changes compared to v35.1.1

Components

  • cluster-aws from v10.3.1 to v10.4.1
  • cluster from v8.3.1 to v8.4.2
  • Flatcar from v4593.2.5 to v4757.2.1
  • Kubernetes from v1.35.8 to v1.35.9
  • os-tooling from v1.34.0 to v1.36.0

cluster-aws v10.3.1…v10.4.1

Added

  • Add global.nodePools.<pool>.kubeletVolume to place /var/lib/kubelet on the instance-store disks of a node pool’s nodes, so that pod emptyDir volumes use local disks instead of the EBS lib volume. This rolls all nodes.
  • Add global.providerSpecific.iam.ecr.allowedRepositories so the read-only Amazon ECR permissions of the control plane and worker node IAM roles can be restricted to selected repositories. Leaving the map empty or undefined keeps allowing all repositories (backwards-compatible).

cluster v8.3.1…v8.4.2

Added

  • Add filesTemplateName hook under providerIntegration.workers.kubeadmConfig. It names a provider template that renders a YAML list of files, once per node pool, so that a provider can deploy files to selected node pools only. A node pool for which the template renders nothing keeps its KubeadmConfig spec, and therefore its spec hash, unchanged.
  • Enable mergeDefaultEvictionSettings to keep defaults for eviction like nodefs.available and nodefs.inodesFree which would otherwise be set to 0. This rolls all nodes.

Changed

  • kubeadm: Exclude /etc/.systemd-confext from restorecon.
  • Cilium: Replace the catch-all - operator: Exists toleration on cilium-operator with an explicit list.
  • Stop deleting and recreating /etc/ssl/certs on nodes for fixing SELinux labeling, so any custom certificates placed there directly are preserved.

Fixed

  • The helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.

Apps

  • aws-ebs-csi-driver from v4.3.1 to v4.3.2
  • cert-exporter from v2.12.0 to v2.12.1
  • cilium from v1.5.1 to v1.5.2
  • coredns from v1.32.0 to v1.34.0
  • etcd-defrag from v1.2.10 to v1.2.12
  • external-dns from v3.5.0 to v3.6.0
  • karpenter from v2.4.1 to v2.5.0
  • net-exporter from v1.24.0 to v1.24.1
  • network-policies from v0.2.0 to v0.3.2
  • node-exporter from v1.20.13 to v1.21.0
  • observability-bundle from v3.3.1 to v3.5.0
  • prometheus-blackbox-exporter from v0.9.0 to v0.10.0
  • security-bundle from v2.3.0 to v2.4.0
  • teleport-kube-agent from v0.11.1 to v0.12.0

aws-ebs-csi-driver v4.3.1…v4.3.2

Changed

  • Conditionally depend on Kyverno and VPA CRDs since security-bundle isn’t yet implemented for cluster-eks and customers may disable that bundle for cluster adoption cases

cert-exporter v2.12.0…v2.12.1

Fixed

  • A cert file that cannot be read no longer aborts the scan of its whole cert path. Previously one unreadable file (such as a root-only 0600 ca.crt on an AKS node, where the exporter runs as an unprivileged user) stopped the walk, silently dropping every file sorting after it from the metrics. Unreadable files are now logged and skipped individually.

cilium v1.5.1…v1.5.2

Changed

coredns v1.32.0…v1.34.0

Added

  • Chart metadata: add io.giantswarm.application.managed annotation ("true").
  • Chart metadata: add keywords.

Changed

  • Update coredns image to 1.14.7.
  • Update coredns image to 1.14.6.
  • Run the E2E test suites automatically on release PRs by adding .github/release-pr-body.md.

Fixed

  • The helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.
  • Honor the deprecated configmap.log, loadbalancePolicy and configmap.cache again. Since 1.31.0 coredns.<zone>.log, coredns.<zone>.loadbalance and coredns.<zone>.cache.success.ttl shipped defaults that shadowed them, so the old keys were silently ignored. They are now unset by default, restoring the documented fallback chain. Rendering with default values is unchanged.

etcd-defrag v1.2.10…v1.2.12

Changed

  • Chart: Update dependency ahrtr/etcd-defrag to v0.45.0. (#136)
  • Chart: Update dependency ahrtr/etcd-defrag to v0.44.0. (#129)

external-dns v3.5.0…v3.6.0

Added

  • Chart metadata: add io.giantswarm.application.managed annotation ("true").
  • Chart metadata: add keywords.

Changed

  • Run the E2E test suites automatically on release PRs by adding .github/release-pr-body.md.
  • Upgrade external-dns to v0.22.0.
  • Sync to upstream helm chart 1.22.0.
    • Add replicaCount value to scale the deployment down to 0 or back to 1.
    • Add service.enabled value to skip creating the Service.
    • Add hostAliases value to inject entries into the pod’s /etc/hosts.
    • Add crd as a valid registry value, with the matching RBAC on dnsrecords.
    • Install the new DNSRecord CRD.
    • Narrow the RBAC verbs on dnsendpoints/status from * to update.
    • policy is now a required value. Our default of sync is unchanged.
    • Pin annotationPrefix to external-dns.alpha.kubernetes.io/, keeping the previous annotation prefix after upstream changed the default.

Fixed

  • The helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.

karpenter v2.4.1…v2.5.0

Added

  • Add iam:CreateServiceLinkedRole permission to the Karpenter IAM role, scoped to the AWSServiceRoleForEC2Spot service-linked role, so Karpenter can create such role when launching spot instances in accounts where it does not exist yet.
  • Support karpenter on EKS clusters
    • Add eks:DescribeCluster permission to discover the cluster endpoint
    • Conditionally depend on Kyverno and pod logs CRDs since security-bundle isn’t yet implemented for cluster-eks and customers may disable that bundle and also observability-bundle for cluster adoption cases

Fixed

  • Strip -, _, and . from the helm.sh/chart label to support dev versions

net-exporter v1.24.0…v1.24.1

Added

  • Add keywords to Chart.yaml.
  • Add the io.giantswarm.application.audience (all) and io.giantswarm.application.managed

Changed

  • Replace interface{} with any and use for-range over integers (Go modernization).
  • Move the team annotation from the legacy application.giantswarm.io/team key to
  • Set Chart.yaml apiVersion to v2. The chart declares no dependencies and had no

Fixed

  • The helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.

network-policies v0.2.0…v0.3.2

Added

  • Add keywords to Chart.yaml.
  • Declare the io.giantswarm.application.audience (all) and
  • Add optional denyEgressToIMDS policy denying pod egress to the instance metadata service. Disabled by default.
  • Add apptest-framework e2e test suite.

Changed

  • Convert the allow-ingress-from-konnectivity CiliumNetworkPolicy into a CiliumClusterwideNetworkPolicy named
  • Always exclude the karpenter and aws-load-balancer-controller namespaces from denyEgressToIMDS.
  • Move the team annotation from the legacy application.giantswarm.io/team key to

node-exporter v1.20.13…v1.21.0

Added

  • disableSystemdCollector value to turn the systemd collector off, mirroring the existing disableConntrackCollector and disableNvmeCollector toggles. The collector needs a D-Bus connection to the host, which is refused on nodes where AppArmor mediates D-Bus (such as AKS Ubuntu nodes running under the default containerd profile), making it fail on every scrape. Defaults to false, so behaviour is unchanged.

observability-bundle v3.3.1…v3.5.0

Added

  • Point kube-prometheus-stack’s control-plane ServiceMonitors at the alloy-metrics token Secret.

Changed

  • Update alloy apps to 0.23.1 (Alloy v1.19.2).
  • Update prometheus-operator-crd to 24.0.0 (Prometheus Operator CRDs v0.94.0).
  • Update kube-prometheus-stack to 24.0.0 (chart 91.2.3, Prometheus Operator v0.94.0).
  • Values: Update Prometheus Operator CRD and Kube Prometheus Stack to v23.0.0.

Fixed

  • KSM custom resource state: Set the Gateway API TCPRoute and UDPRoute collectors to v1, which is the version the API server serves.

prometheus-blackbox-exporter v0.9.0…v0.10.0

Added

  • Probe github.com, gsoci.azurecr.io and grafana.com from every node (targets egress-github, egress-registry, egress-grafana), so an allowlist-style firewall change blocking a single domain becomes visible. The new serviceMonitor.externalTargets value is a map so per-installation, per-region or per-customer overrides can disable, change or add individual entries without copying the whole list. A separate serviceMonitor.additionalExternalTargets key takes regional/customer additions, structurally separated from the Giant Swarm defaults. Adds the http_2xx_or_401 module for registry endpoints that answer unauthenticated requests with 401. See giantswarm/giantswarm#33409.
  • Add the http_2xx_egress module, used by the internet egress targets. It carries a 15s timeout so probe_success reports whether an endpoint is reachable rather than whether it is fast. It is a separate module rather than a longer timeout on http_2xx because several installations pin http_2xx in their own custom values, which would silently revert the change there.

Removed

  • Remove the inert instance metric relabeling from the ServiceMonitor template. It interpolated a url field that no target defines, so it rendered empty and Prometheus fell back to its $1 default, leaving instance unchanged.

Fixed

  • Give the internet egress targets a probe deadline above the cross-border baseline: scrapeTimeout: 20s on http-giantswarm, egress-github, egress-registry and egress-grafana, and a 15s timeout on http_2xx_or_401. The exporter applies min(module timeout, scrapeTimeout - 0.5s offset), so the 5s serviceMonitor.defaults.scrapeTimeout capped every probe at a 4.5s deadline and raising a module timeout alone had no effect. Installations whose baseline latency is a large fraction of that deadline crossed it on endpoints that were still returning HTTP 200, making probe_success report latency rather than reachability.
  • Point the dns-tcp-internal and dns-udp-internal ServiceMonitors at the dns_*_internal modules. They referenced the _external modules, so both probed www.prometheus.io and in-cluster DNS resolution was never monitored.

security-bundle v2.3.0…v2.4.0

Added

  • Add e2e scenarios covering trivy-operator VulnerabilityReport creation, starboard-exporter metrics for that report, kyverno restricted PSS enforcement, and kyverno-policy-operator PolicyException translation.

Changed

  • Update exception-recommender (app) to v0.3.0.
  • Update falco (app) to v0.13.0.
  • Update jiralert (app) to v0.1.4.
  • Update kubescape (app) to v0.1.1.
  • Update kyverno-policies (app) to v0.27.1.
  • Update policy-api (app) to v0.0.12.
  • Update starboard-exporter (app) to v1.2.15.
  • Update trivy (app) to v0.18.0.
  • Update trivy-operator (app) to v0.15.0.

Removed

  • Remove gel (app).

Fixed

  • Give kubescape a 15m install and upgrade timeout. It does not finish installing within Flux’s 5m default, so it failed with context deadline exceeded and then retried indefinitely.
  • Set createNamespace on every app in the bundle, so each one creates its target namespace instead of relying on another app to have created it first. Previously kubescape failed with namespaces "kubescape" not found, and the apps targeting security-bundle could only install after kyverno-policy-operator had created it.

teleport-kube-agent v0.11.1…v0.12.0

Changed

  • Updated teleport-kube-agent to upstream version v18.10.7.