Workload cluster release cloud-director-35.2.0 for CAPVCD

Changes compared to v35.1.1

Components

  • cluster-cloud-director from v7.3.1 to v7.4.1
  • cluster from v8.3.1 to v8.4.2
  • Flatcar from v4593.2.5 to v4757.2.1
  • Kubernetes from v1.35.8 to v1.35.9
  • os-tooling from v1.34.0 to v1.36.0

cluster-cloud-director v7.3.1…v7.4.1

Changed

  • Enable blackbox-exporter by default.
  • The helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.

cluster v8.3.1…v8.4.2

Added

  • Add filesTemplateName hook under providerIntegration.workers.kubeadmConfig. It names a provider template that renders a YAML list of files, once per node pool, so that a provider can deploy files to selected node pools only. A node pool for which the template renders nothing keeps its KubeadmConfig spec, and therefore its spec hash, unchanged.
  • Enable mergeDefaultEvictionSettings to keep defaults for eviction like nodefs.available and nodefs.inodesFree which would otherwise be set to 0. This rolls all nodes.

Changed

  • kubeadm: Exclude /etc/.systemd-confext from restorecon.
  • Cilium: Replace the catch-all - operator: Exists toleration on cilium-operator with an explicit list.
  • Stop deleting and recreating /etc/ssl/certs on nodes for fixing SELinux labeling, so any custom certificates placed there directly are preserved.

Fixed

  • The helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.

Apps

  • cert-exporter from v2.12.0 to v2.12.1
  • cilium from v1.5.1 to v1.5.2
  • cloud-provider-cloud-director from v0.5.2 to v0.6.1
  • coredns from v1.32.0 to v1.34.0
  • etcd-defrag from v1.2.10 to v1.2.12
  • net-exporter from v1.24.0 to v1.24.1
  • network-policies from v0.2.0 to v0.3.2
  • node-exporter from v1.20.13 to v1.21.0
  • observability-bundle from v3.3.1 to v3.5.0
  • Added prometheus-blackbox-exporter v0.10.0
  • security-bundle from v2.3.0 to v2.4.0
  • teleport-kube-agent from v0.11.1 to v0.12.0

cert-exporter v2.12.0…v2.12.1

Fixed

  • A cert file that cannot be read no longer aborts the scan of its whole cert path. Previously one unreadable file (such as a root-only 0600 ca.crt on an AKS node, where the exporter runs as an unprivileged user) stopped the walk, silently dropping every file sorting after it from the metrics. Unreadable files are now logged and skipped individually.

cilium v1.5.1…v1.5.2

Changed

cloud-provider-cloud-director v0.5.2…v0.6.1

Changed

  • Values: Downgrade images to compatible versions.
  • Update architect to v10.10.0 (giantswarm/cloud-provider-cloud-director-app#182)
  • Update architect to v10.11.1 (giantswarm/cloud-provider-cloud-director-app#183)
  • Update architect to v10.12.0 (giantswarm/cloud-provider-cloud-director-app#185)
  • Update architect to v10.12.1 (giantswarm/cloud-provider-cloud-director-app#186)
  • chore(deps): update gsoci.azurecr.io/giantswarm/cloud-director-named-disk-csi-driver docker tag to v1.6.1
  • chore(deps): update projects.registry.vmware.com/vmware-cloud-director/cloud-provider-for-cloud-director docker tag to v1.6.2
  • chore(deps): update registry.k8s.io/sig-storage/csi-node-driver-registrar docker tag to v2.18.0
  • chore(deps): update registry.k8s.io/sig-storage/csi-attacher docker tag to v4
  • chore(deps): update registry.k8s.io/sig-storage/csi-provisioner docker tag to v6

Fixed

  • The helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.

coredns v1.32.0…v1.34.0

Added

  • Chart metadata: add io.giantswarm.application.managed annotation ("true").
  • Chart metadata: add keywords.

Changed

  • Update coredns image to 1.14.7.
  • Update coredns image to 1.14.6.
  • Run the E2E test suites automatically on release PRs by adding .github/release-pr-body.md.

Fixed

  • The helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.
  • Honor the deprecated configmap.log, loadbalancePolicy and configmap.cache again. Since 1.31.0 coredns.<zone>.log, coredns.<zone>.loadbalance and coredns.<zone>.cache.success.ttl shipped defaults that shadowed them, so the old keys were silently ignored. They are now unset by default, restoring the documented fallback chain. Rendering with default values is unchanged.

etcd-defrag v1.2.10…v1.2.12

Changed

  • Chart: Update dependency ahrtr/etcd-defrag to v0.45.0. (#136)
  • Chart: Update dependency ahrtr/etcd-defrag to v0.44.0. (#129)

net-exporter v1.24.0…v1.24.1

Added

  • Add keywords to Chart.yaml.
  • Add the io.giantswarm.application.audience (all) and io.giantswarm.application.managed

Changed

  • Replace interface{} with any and use for-range over integers (Go modernization).
  • Move the team annotation from the legacy application.giantswarm.io/team key to
  • Set Chart.yaml apiVersion to v2. The chart declares no dependencies and had no

Fixed

  • The helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.

network-policies v0.2.0…v0.3.2

Added

  • Add keywords to Chart.yaml.
  • Declare the io.giantswarm.application.audience (all) and
  • Add optional denyEgressToIMDS policy denying pod egress to the instance metadata service. Disabled by default.
  • Add apptest-framework e2e test suite.

Changed

  • Convert the allow-ingress-from-konnectivity CiliumNetworkPolicy into a CiliumClusterwideNetworkPolicy named
  • Always exclude the karpenter and aws-load-balancer-controller namespaces from denyEgressToIMDS.
  • Move the team annotation from the legacy application.giantswarm.io/team key to

node-exporter v1.20.13…v1.21.0

Added

  • disableSystemdCollector value to turn the systemd collector off, mirroring the existing disableConntrackCollector and disableNvmeCollector toggles. The collector needs a D-Bus connection to the host, which is refused on nodes where AppArmor mediates D-Bus (such as AKS Ubuntu nodes running under the default containerd profile), making it fail on every scrape. Defaults to false, so behaviour is unchanged.

observability-bundle v3.3.1…v3.5.0

Added

  • Point kube-prometheus-stack’s control-plane ServiceMonitors at the alloy-metrics token Secret.

Changed

  • Update alloy apps to 0.23.1 (Alloy v1.19.2).
  • Update prometheus-operator-crd to 24.0.0 (Prometheus Operator CRDs v0.94.0).
  • Update kube-prometheus-stack to 24.0.0 (chart 91.2.3, Prometheus Operator v0.94.0).
  • Values: Update Prometheus Operator CRD and Kube Prometheus Stack to v23.0.0.

Fixed

  • KSM custom resource state: Set the Gateway API TCPRoute and UDPRoute collectors to v1, which is the version the API server serves.

prometheus-blackbox-exporter v0.10.0

Added

  • Probe github.com, gsoci.azurecr.io and grafana.com from every node (targets egress-github, egress-registry, egress-grafana), so an allowlist-style firewall change blocking a single domain becomes visible. The new serviceMonitor.externalTargets value is a map so per-installation, per-region or per-customer overrides can disable, change or add individual entries without copying the whole list. A separate serviceMonitor.additionalExternalTargets key takes regional/customer additions, structurally separated from the Giant Swarm defaults. Adds the http_2xx_or_401 module for registry endpoints that answer unauthenticated requests with 401. See giantswarm/giantswarm#33409.
  • Add the http_2xx_egress module, used by the internet egress targets. It carries a 15s timeout so probe_success reports whether an endpoint is reachable rather than whether it is fast. It is a separate module rather than a longer timeout on http_2xx because several installations pin http_2xx in their own custom values, which would silently revert the change there.

Removed

  • Remove the inert instance metric relabeling from the ServiceMonitor template. It interpolated a url field that no target defines, so it rendered empty and Prometheus fell back to its $1 default, leaving instance unchanged.

Fixed

  • Give the internet egress targets a probe deadline above the cross-border baseline: scrapeTimeout: 20s on http-giantswarm, egress-github, egress-registry and egress-grafana, and a 15s timeout on http_2xx_or_401. The exporter applies min(module timeout, scrapeTimeout - 0.5s offset), so the 5s serviceMonitor.defaults.scrapeTimeout capped every probe at a 4.5s deadline and raising a module timeout alone had no effect. Installations whose baseline latency is a large fraction of that deadline crossed it on endpoints that were still returning HTTP 200, making probe_success report latency rather than reachability.
  • Point the dns-tcp-internal and dns-udp-internal ServiceMonitors at the dns_*_internal modules. They referenced the _external modules, so both probed www.prometheus.io and in-cluster DNS resolution was never monitored.

security-bundle v2.3.0…v2.4.0

Added

  • Add e2e scenarios covering trivy-operator VulnerabilityReport creation, starboard-exporter metrics for that report, kyverno restricted PSS enforcement, and kyverno-policy-operator PolicyException translation.

Changed

  • Update exception-recommender (app) to v0.3.0.
  • Update falco (app) to v0.13.0.
  • Update jiralert (app) to v0.1.4.
  • Update kubescape (app) to v0.1.1.
  • Update kyverno-policies (app) to v0.27.1.
  • Update policy-api (app) to v0.0.12.
  • Update starboard-exporter (app) to v1.2.15.
  • Update trivy (app) to v0.18.0.
  • Update trivy-operator (app) to v0.15.0.

Removed

  • Remove gel (app).

Fixed

  • Give kubescape a 15m install and upgrade timeout. It does not finish installing within Flux’s 5m default, so it failed with context deadline exceeded and then retried indefinitely.
  • Set createNamespace on every app in the bundle, so each one creates its target namespace instead of relying on another app to have created it first. Previously kubescape failed with namespaces "kubescape" not found, and the apps targeting security-bundle could only install after kyverno-policy-operator had created it.

teleport-kube-agent v0.11.1…v0.12.0

Changed

  • Updated teleport-kube-agent to upstream version v18.10.7.