Workload cluster release azure-35.2.0 for CAPZ
Changes compared to v35.1.1
Components
- cluster-azure from v9.3.1 to v9.4.1
- cluster from v8.3.1 to v8.4.2
- Flatcar from v4593.2.4 to v4757.2.1
- Kubernetes from v1.35.8 to v1.35.9
- os-tooling from v1.34.0 to v1.36.0
Added
- Add
filesTemplateName hook under providerIntegration.workers.kubeadmConfig. It names a provider template that renders a YAML list of files, once per node pool, so that a provider can deploy files to selected node pools only. A node pool for which the template renders nothing keeps its KubeadmConfig spec, and therefore its spec hash, unchanged. - Enable
mergeDefaultEvictionSettings to keep defaults for eviction like nodefs.available and nodefs.inodesFree which would otherwise be set to 0. This rolls all nodes.
Changed
kubeadm: Exclude /etc/.systemd-confext from restorecon.- Cilium: Replace the catch-all
- operator: Exists toleration on cilium-operator with an explicit list. - Stop deleting and recreating
/etc/ssl/certs on nodes for fixing SELinux labeling, so any custom certificates placed there directly are preserved.
Fixed
- The
helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.
Apps
- azuredisk-csi-driver from v2.1.0 to v2.2.0
- azurefile-csi-driver from v2.0.0 to v2.1.0
- cert-exporter from v2.12.0 to v2.12.1
- cilium from v1.5.1 to v1.5.2
- coredns from v1.32.0 to v1.34.0
- etcd-defrag from v1.2.10 to v1.2.12
- external-dns from v3.5.0 to v3.6.0
- net-exporter from v1.24.0 to v1.24.1
- network-policies from v0.2.0 to v0.3.2
- node-exporter from v1.20.13 to v1.21.0
- observability-bundle from v3.3.1 to v3.5.0
- prometheus-blackbox-exporter from v0.9.0 to v0.10.0
- security-bundle from v2.3.0 to v2.4.0
- teleport-kube-agent from v0.11.1 to v0.12.0
Changed
- Migrate to App Build Suite (ABS).
- Chart: Update to upstream v1.34.5.
Removed
- Removed
PodSecurityPolicy. - Removed
global.podSecurityStandards.enforced helm value.
Changed
- Migrate to App Build Suite (ABS).
- Chart: Update to upstream v1.35.7.
Removed
- Removed
PodSecurityPolicy. - Removed
global.podSecurityStandards.enforced helm value.
Fixed
- A cert file that cannot be read no longer aborts the scan of its whole cert path. Previously one unreadable file (such as a root-only
0600 ca.crt on an AKS node, where the exporter runs as an unprivileged user) stopped the walk, silently dropping every file sorting after it from the metrics. Unreadable files are now logged and skipped individually.
Changed
Added
- Chart metadata: add
io.giantswarm.application.managed annotation ("true"). - Chart metadata: add
keywords.
Changed
- Update
coredns image to 1.14.7. - Update
coredns image to 1.14.6. - Run the E2E test suites automatically on release PRs by adding
.github/release-pr-body.md.
Fixed
- The
helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _. - Honor the deprecated
configmap.log, loadbalancePolicy and configmap.cache again. Since 1.31.0 coredns.<zone>.log, coredns.<zone>.loadbalance and coredns.<zone>.cache.success.ttl shipped defaults that shadowed them, so the old keys were silently ignored. They are now unset by default, restoring the documented fallback chain. Rendering with default values is unchanged.
Changed
- Chart: Update dependency ahrtr/etcd-defrag to v0.45.0. (#136)
- Chart: Update dependency ahrtr/etcd-defrag to v0.44.0. (#129)
Added
- Chart metadata: add
io.giantswarm.application.managed annotation ("true"). - Chart metadata: add
keywords.
Changed
- Run the E2E test suites automatically on release PRs by adding
.github/release-pr-body.md. - Upgrade external-dns to v0.22.0.
- Sync to upstream helm chart 1.22.0.
- Add
replicaCount value to scale the deployment down to 0 or back to 1. - Add
service.enabled value to skip creating the Service. - Add
hostAliases value to inject entries into the pod’s /etc/hosts. - Add
crd as a valid registry value, with the matching RBAC on dnsrecords. - Install the new
DNSRecord CRD. - Narrow the RBAC verbs on
dnsendpoints/status from * to update. policy is now a required value. Our default of sync is unchanged.- Pin
annotationPrefix to external-dns.alpha.kubernetes.io/, keeping the previous annotation prefix after upstream changed the default.
Fixed
- The
helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.
Added
- Add
keywords to Chart.yaml. - Add the
io.giantswarm.application.audience (all) and io.giantswarm.application.managed
Changed
- Replace
interface{} with any and use for-range over integers (Go modernization). - Move the team annotation from the legacy
application.giantswarm.io/team key to - Set
Chart.yaml apiVersion to v2. The chart declares no dependencies and had no
Fixed
- The
helm.sh/chart label is valid for long chart versions: the 63-character cut trims the whole trailing run of -, . and _.
Added
- Add
keywords to Chart.yaml. - Declare the
io.giantswarm.application.audience (all) and - Add optional
denyEgressToIMDS policy denying pod egress to the instance metadata service. Disabled by default. - Add apptest-framework e2e test suite.
Changed
- Convert the
allow-ingress-from-konnectivity CiliumNetworkPolicy into a CiliumClusterwideNetworkPolicy named - Always exclude the
karpenter and aws-load-balancer-controller namespaces from denyEgressToIMDS. - Move the team annotation from the legacy
application.giantswarm.io/team key to
Added
disableSystemdCollector value to turn the systemd collector off, mirroring the existing disableConntrackCollector and disableNvmeCollector toggles. The collector needs a D-Bus connection to the host, which is refused on nodes where AppArmor mediates D-Bus (such as AKS Ubuntu nodes running under the default containerd profile), making it fail on every scrape. Defaults to false, so behaviour is unchanged.
Added
- Point kube-prometheus-stack’s control-plane ServiceMonitors at the alloy-metrics token Secret.
Changed
- Update
alloy apps to 0.23.1 (Alloy v1.19.2). - Update
prometheus-operator-crd to 24.0.0 (Prometheus Operator CRDs v0.94.0). - Update
kube-prometheus-stack to 24.0.0 (chart 91.2.3, Prometheus Operator v0.94.0). - Values: Update Prometheus Operator CRD and Kube Prometheus Stack to v23.0.0.
Fixed
- KSM custom resource state: Set the Gateway API
TCPRoute and UDPRoute collectors to v1, which is the version the API server serves.
Added
- Probe
github.com, gsoci.azurecr.io and grafana.com from every node (targets egress-github, egress-registry, egress-grafana), so an allowlist-style firewall change blocking a single domain becomes visible. The new serviceMonitor.externalTargets value is a map so per-installation, per-region or per-customer overrides can disable, change or add individual entries without copying the whole list. A separate serviceMonitor.additionalExternalTargets key takes regional/customer additions, structurally separated from the Giant Swarm defaults. Adds the http_2xx_or_401 module for registry endpoints that answer unauthenticated requests with 401. See giantswarm/giantswarm#33409. - Add the
http_2xx_egress module, used by the internet egress targets. It carries a 15s timeout so probe_success reports whether an endpoint is reachable rather than whether it is fast. It is a separate module rather than a longer timeout on http_2xx because several installations pin http_2xx in their own custom values, which would silently revert the change there.
Removed
- Remove the inert
instance metric relabeling from the ServiceMonitor template. It interpolated a url field that no target defines, so it rendered empty and Prometheus fell back to its $1 default, leaving instance unchanged.
Fixed
- Give the internet egress targets a probe deadline above the cross-border baseline:
scrapeTimeout: 20s on http-giantswarm, egress-github, egress-registry and egress-grafana, and a 15s timeout on http_2xx_or_401. The exporter applies min(module timeout, scrapeTimeout - 0.5s offset), so the 5s serviceMonitor.defaults.scrapeTimeout capped every probe at a 4.5s deadline and raising a module timeout alone had no effect. Installations whose baseline latency is a large fraction of that deadline crossed it on endpoints that were still returning HTTP 200, making probe_success report latency rather than reachability. - Point the
dns-tcp-internal and dns-udp-internal ServiceMonitors at the dns_*_internal modules. They referenced the _external modules, so both probed www.prometheus.io and in-cluster DNS resolution was never monitored.
Added
- Add e2e scenarios covering trivy-operator
VulnerabilityReport creation, starboard-exporter metrics for that report, kyverno restricted PSS enforcement, and kyverno-policy-operator PolicyException translation.
Changed
- Update
exception-recommender (app) to v0.3.0. - Update
falco (app) to v0.13.0. - Update
jiralert (app) to v0.1.4. - Update
kubescape (app) to v0.1.1. - Update
kyverno-policies (app) to v0.27.1. - Update
policy-api (app) to v0.0.12. - Update
starboard-exporter (app) to v1.2.15. - Update
trivy (app) to v0.18.0. - Update
trivy-operator (app) to v0.15.0.
Removed
Fixed
- Give
kubescape a 15m install and upgrade timeout. It does not finish installing within Flux’s 5m default, so it failed with context deadline exceeded and then retried indefinitely. - Set
createNamespace on every app in the bundle, so each one creates its target namespace instead of relying on another app to have created it first. Previously kubescape failed with namespaces "kubescape" not found, and the apps targeting security-bundle could only install after kyverno-policy-operator had created it.
Changed
- Updated
teleport-kube-agent to upstream version v18.10.7.