Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

dynamic-config on Kubernetes

Annotate a pod; an agent appears in it that renders configuration from a remote store to a file the application watches. The agent-injector shape, for configuration.

metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "consul"
    dynamic-config.rs/endpoint: "http://consul:8500"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/path: "/config/rendered.toml"

The application needs no store client, no store credential, and no code change: it reads a file, and if it reads it with any dynamic-config binding it also reloads on every re-render — the agent writes atomically, exactly the whole-file event a watcher wants.

When not to use this. The engine runs in-process everywhere; a Rust, Python or Node service that can hold a store credential should usually use its own binding's remote support and skip the sidecar entirely. This integration exists for the pods that want files rendered for them: Java services reading .properties, anything that must not carry store credentials in-process, and fleets standardising one injection pattern.

How a document REACHES a workload — a live file, real environment variables, or a native Kubernetes Secret — is its own decision, with a map and two honest comparison tables (Vault Agent Injector, External Secrets Operator) on The Four Deliveries.

The four pieces, staged

pieceships intoday
agent0.3.0all nine stores — the blocking six since 0.1.0, etcd/nats/s3 on the async path since 0.1.1 — watched rather than polled since 0.2.0, and since 0.3.0: last-known-good with a startup policy, dynamic secrets with lease renewal, a readiness probe that means there is a document, drift, history, acknowledgement and canary cohorts
webhook0.3.0golden-tested; the annotation contract is v1 and is now a registry a test checks this book against; three TLS modes incl. selfRotate; admission warnings and a validate CLI since 0.3.0
operator0.3.0Render → ConfigMap reconciler shipped, Class watch wired, e2e-gated, leader-elected since 0.2.0; deletionPolicy, observedGeneration and three refusal reasons since 0.3.0
node agent0.3.0one process per node instead of one beside every render, delivering as a CSI volume; pods sharing a document share a watch. Off by default — it holds the credentials of every pod on its node, so the sidecar stays the shape to reach for

Install

# From the OCI registry (ArtifactHub lists the same chart):
helm install dynamic-config \
  oci://ghcr.io/dynamic-config-rs/charts/dynamic-config --version 0.3.0

# Or from a checkout:
helm install dynamic-config deploy/helm

Every image the chart deploys is on ghcr and mirrored to Docker Hub (docker.io/ctolon17/…) with identical digests, multi-arch, SBOM attested and cosign-signed keyless; the repository README carries the cosign verify line.

That is the whole install: no dependencies, nothing to pre-create. The chart mints a CA and a ten-year serving certificate at install time, embeds the caBundle in the webhook configuration, and reuses the Secret on upgrades so the trust does not silently rotate. The webhook terminates that TLS in-process — the API server speaks HTTPS to admission webhooks and nothing else.

The cert-manager mode

helm install dynamic-config deploy/helm \
  --set webhook.certManager.enabled=true \
  --set webhook.certManager.issuerRef.name=<your-issuer>

With cert-manager the certificate is renewed — cainjector maintains the caBundle, and the webhook picks up the renewed pair from disk without a restart (it polls the mounted files for a changed modification time). The trade between the two:

The selfRotate mode

helm install dynamic-config deploy/helm \
  --set webhook.selfRotate.enabled=true

The third answer, the Vault-agent-injector shape: the webhook is its own certificate authority. It mints a CA and leaf in memory at rotation time, writes the pair to its own Secret — every replica serves it through the same file hot-reload cert-manager uses — and patches the webhook configuration's caBundle itself. A fresh pair every 24 hours, leader-elected over a Lease so replicas do not race, jittered so a fleet restarted together does not rotate together.

The price is stated in values.yaml beside the toggle: a service-account token and three narrow, name-scoped permissions — the zero-RBAC purity of the other two modes, knowingly traded for rotation without a dependency.

self-signed (default)cert-managerselfRotate
dependenciesnonecert-manager installednone
renewalnone — ten-year cert; rotate by deleting the Secret and upgradingautomatic, well before expiryautomatic, every 24h, leader-elected
caBundleembedded at installmaintained by cainjectorpatched by the webhook itself
RBACnonenoneone Secret, one MWC, leases — by name
fitsgetting started, edge clusters, air-gappedanywhere cert-manager already runsrotation wanted, cert-manager not

All three mount the same Secret shape at the same path; the webhook's serving loop cannot tell them apart, which is what makes switching later a values change.

A private mirror for the agent image

The injected agent is pulled in application namespaces, so a fleet that mirrors images into a private registry needs two values:

helm install dynamic-config deploy/helm \
  --set agent.image=registry.internal/dynamic-config-agent \
  --set agent.pullSecret=mirror-cred

agent.pullSecret is appended to each injected pod's own imagePullSecrets, never replacing them. Pull secrets are namespaced — the Secret must exist in every namespace that injects, which is the usual replication job (kubectl create secret docker-registry … -n <each>, or a replicator you already run). The webhook's and operator's own images use the chart-level imagePullSecrets list instead, because those pods live in the release namespace.

failurePolicy

Ignore by default, and the book owns the trade: Fail would make the webhook a single point of failure for every pod creation in selected namespaces, while Ignore means an annotated pod created during a webhook outage starts without its agent — loudly, because the file its application waits for never appears. Flip it with --set webhook.failurePolicy=Fail once the webhook has earned it in your cluster; the security page carries the full argument.

Namespace gating

Clusters that prefer opt-in injection (the Istio shape) set webhook.namespaceGating=true and label the namespaces that want it:

kubectl label namespace payments dynamic-config.rs/injection=enabled

Everything else is invisible to the webhook — the security page explains what that buys.

What the chart hardens for you

The security page is the complete inventory; the short list: both deployments run as non-root with the restricted-PSS container posture, the webhook's ServiceAccount mounts no API token, kube-system and the release namespace are excluded from injection, two replicas ride a PodDisruptionBudget, and tag: latest fails the render. An optional NetworkPolicy writes down that the webhook accepts the API server and calls nobody.

One namespace of its own

Install into a dedicated namespace, always:

helm install dynamic-config deploy/helm -n dynamic-config --create-namespace

The webhook configuration excludes its own namespace by name — the self-deadlock guard — so a release installed into default silently excludes every workload sharing default with it. The chart's NOTES print a warning when that happens; the e2e smoke installs the dedicated way for the same reason.

Fleet-wide defaults, validated at the door

What a pod does not say per annotation, the installation says once — and it can say it twice, because defaults come in tiers: annotation > per-store default > fleet default > built-in. Every knob the annotations know is defaultable; there is no second vocabulary:

agent:
  defaults:
    cpuRequest: 10m        # agent-cpu-request still wins per pod
    memoryRequest: 32Mi
    memoryLimit: 64Mi
    cpuLimit: ""           # empty on purpose
    fileMode: "0640"       # empty = the agent's 0644; file-mode wins per pod
    watchSeconds: "30"     # empty = 15; watch-seconds wins per pod
    mode: "both"           # empty = sidecar
    volumeMedium: ""       # empty = memory
    nativeSidecar: ""      # empty = false
    runAsUser: "1000"      # empty = 65532; 0 refused, same as the annotation
    runAsGroup: "1000"
    metricsPort: "9102"    # empty = no metrics; pods opt out with metrics-port "0"
    env: "HTTPS_PROXY=http://egress.infra.svc:3128"  # every agent; pod's agent-env wins per name
    source: "consul"       # pods may omit source entirely
    path: "/config/rendered.toml"
    overridable: ""        # "false" pins every value set here; "!"/"?" per value
    perStore:              # the tier between annotation and fleet
      vault: "endpoint=https://vault.vault.svc:8200!, auth=kubernetes, watch-seconds=10"
      s3: "agent-memory-limit=128Mi"
webhook:
  agentEnvAllow: "payments: HTTPS_PROXY, AWS_*; *: RUST_LOG"
  sourceAllow: ""          # empty = every store, everywhere
  sourceDeny: "sandbox: git"

perStore keys are spelled exactly as the annotations spell them (watch-seconds, not watchSeconds) — one grammar for the value, whether it arrives per pod, per store, or per fleet — and they cover EVERY store-shaped annotation, so a developer can deploy knowing nothing but inject: "true". The ! above PINS the vault address: a pod annotating a different endpoint is refused, not silently corrected. agent.defaults.env needs no allowlist: the installer owns both the values and the gate.

Helm's schema refuses a malformed value at render time; the webhook re-validates ALL of it at startup and refuses to serve on a typo — so an installation written any of the three ways gets the same refusal at the same door. The readable form for kustomize is base/installation.yaml, a ConfigMap of the same settings as YAML (Installation Defaults); the variables below are the other way, and still work:

# kustomization.yaml, an overlay patch
patches:
  - target: { kind: Deployment, name: dynamic-config-webhook }
    patch: |
      - op: add
        path: /spec/template/spec/containers/0/env/-
        value: { name: DYNAMIC_CONFIG_AGENT_FILE_MODE, value: "0640" }

agentEnvAllow and the source gates are security gates, not defaults. Installation Defaults and Gates is the full treatment: every knob with its validation, per-store examples for all nine stores with every field filled, the gates' semantics and threat model, and the kustomize equivalents.

Values, all of them

The chart README is the full values reference — naming overrides, common labels, per-component images and pull policy, service accounts, probe and rollout tuning, extraEnv/extraVolumes escape hatches, namespace gating, the operator's RBAC toggle. Everything the templates read is in that table.

Without helm: kustomize

deploy/kustomize/ carries the same resources as a base — including installation.yaml, the fleet defaults and gates written as YAML rather than as environment-variable grammar — plus TLS overlays for cert-manager, bring-your-own-PEMs via secretGenerator, and the self-rotating mode. Its README is the three-step walkthrough, including the one caBundle patch kustomize cannot express. The CRDs ship inside the base, drift-gated against the operator's --crds output like every other copy.

Without a registry: an air-gapped install

./scripts/airgap-bundle.sh 0.3.0        # on the connected side
# move dynamic-config-0.3.0-airgap.tar.gz across
tar xzf dynamic-config-0.3.0-airgap.tar.gz
cd dynamic-config-0.3.0-airgap
./load.sh registry.internal:5000
./verify.sh
kubectl apply --server-side -f crds/
helm install dynamic-config ./chart --values values-airgap.yaml

The bundle carries the chart, the CRDs, the three image indexes and their signatures and attestations — because the supply-chain work this project does is worth nothing offline if the signatures are left behind with the registry. verify.sh runs cosign --offline against the images as they landed rather than as they were published.

Two things the procedure insists on. load.sh pushes by digest, because the digests are what the signatures cover; and the CRDs are applied by hand, because Helm installs crds/ once and never upgrades it — the same step an upgrade needs on a connected cluster.

The smoke test

The e2e smoke (e2e/smoke.sh) is the install, end to end, against a kind cluster: the chart in its zero-dependency default, a live Consul, one annotated pod, the rendered file read back out of it, and the injected container's security posture asserted on the running pod. CERT_MANAGER=1 e2e/smoke.sh runs the same flow through the other TLS mode.

Installation Defaults and Gates

Two kinds of installation-time decisions live in the webhook's configuration, and they must not be confused: defaults, which a pod may always override, and gates, which a pod may never override. Both arrive as chart values, or as a ConfigMap of YAML that kustomize can hand over too, or as environment variables on the webhook Deployment — three spellings of one thing, and every one of them is validated when the webhook starts. A mistyped value stops the install, never the first admission.

defaults:  annotation  >  per-store default  >  fleet default  >  built-in
pins:      a value marked "!" (or under overridable: "false") refuses a
           DIFFERING annotation — the tiers still fill what pods omit
gates:     the installer's word is final

Writing them as YAML

An installation reaches the webhook as strings, because that is what an environment variable is — and several of those strings are little grammars:

DYNAMIC_CONFIG_AGENT_STORE_DEFAULTS="vault: overridable=false, endpoint=https://vault:8200; s3: file-mode=0640?"
DYNAMIC_CONFIG_WEBHOOK_SOURCE_ALLOW="payments: vault, s3; *: consul"

Fine to parse and unpleasant to write. So the same settings may be written as maps:

# values.yaml
agent:
  defaults:
    perStore:
      vault:
        overridable: false
        endpoint: https://vault.vault.svc:8200
        auth: kubernetes
        watch-seconds: 30
      s3:
        agent-memory-limit: 128Mi
webhook:
  sourceAllow:
    payments: [vault, s3]
    "*": [consul]
  agentEnvAllow:
    payments: [HTTPS_PROXY, "AWS_*"]

A map is defined as what it renders to. It travels to the pod as a mounted ConfigMap and is turned into the grammar there, by the same parser the string form goes through — so there is one set of rules, one set of messages, and two spellings that cannot mean different things. Pins, markers and tiers work the same either way: overridable: false in a map is the overridable=false in a string.

The webhook reads that document through its own engine: the YAML reader it gives applications is the one it reads its own configuration with.

Kustomize gets the same thing

This is why the document exists at all rather than a chart-side rendering. A kustomize base has no template engine, so a hand-written ConfigMap is the only structured form it can hand over — deploy/kustomize/base/installation.yaml, which ships empty and commented:

apiVersion: v1
kind: ConfigMap
metadata:
  name: dynamic-config-webhook-installation
data:
  installation.yaml: |
    storeDefaults:
      vault:
        overridable: false
        endpoint: https://vault.vault.svc:8200
    sourceAllow:
      payments: [vault, s3]
      "*": [consul]

Which wins

An environment variable set on the container beats the document: a document is the installation written down once, a variable is somebody standing in front of it for this deployment — the more specific of the two, which is the same rule the configuration layers themselves follow.

A setting the document does not know is refused at startup, with the list of the ones there are. A default silently ignored is a default that never applied, and the first sign of it would be a pod running without the posture somebody thought they had set.

The knob vocabulary

Every defaultable knob is spelled exactly as its annotation is spelled — there is no second vocabulary to learn, and every tier is validated with the same rules as the annotation it stands in for:

knobbuilt-invalidation
agent-cpu-request10ma Kubernetes quantity
agent-memory-request32Mia Kubernetes quantity
agent-cpu-limitnonea Kubernetes quantity
agent-memory-limit64Mia Kubernetes quantity
file-modethe agent's 0644octal, at most 0777, owner-readable
watch-seconds15whole seconds
modesidecarinit / sidecar / both
volume-mediummemorymemory / disk
native-sidecarfalse"true" / "false"
agent-run-as-user65532numeric, 0 refused
agent-run-as-group65532numeric, 0 refused
metrics-portnonea port; a pod opts out of a default with "0"
pathnoneabsolute; the rendered file's location
sourcenonefleet level only (agent.defaults.source): one of the nine stores

Pod-wide knobs (mode, volume, resources, identity, metrics) resolve against the DEFAULT render's store; per-render knobs (watch-seconds, file-mode) resolve against each render's own store.

And per store, EVERY store-shaped annotation is defaultable — the address, the document, the credentials' Secret names, the auth flags, the templates:

endpoint   endpoint-secret   key          token-secret   password-secret
ca-configmap   tls-secret    ssh-secret   aws-secret     section
auth   auth-mount   auth-role   auth-username   auth-token-path
namespace   ref   api-url   template   template-configmap

Each is validated as its annotation would be (*-secret values need the <secret-name>/<key> slash), and the either-or pairs (endpoint/endpoint-secret, template/template-configmap) resolve as a LEVEL: any pod-side answer mutes both installation halves, so a pod that chose endpoint-secret never inherits a stray plain endpoint.

With source and path at the fleet tier and endpoint and key in a store's tier, the minimal pod is a single line of intent:

metadata:
  annotations:
    dynamic-config.rs/inject: "true"

Defaulting key deserves a sentence of caution: two pods leaning on the same store default read the SAME document. For a shared, cluster-wide configuration that is exactly right; for per-app documents, leave key to the pods.

What stays the pod's alone: env-inject and env-restart (they name a container only the pod knows), agent-env (gated per pod), and inject itself — a default that opts workloads in silently is not a default, it is a surprise.

Pinning: the override mode

Every installation value carries an override rule, and the closest word wins:

  1. a trailing ! on the value — pinned — or a trailing ? — overridable;
  2. no marker: the STORE's own flag, when its perStore group carries overridable=true|false;
  3. neither: whatever agent.defaults.overridable says ("true" unless set otherwise).

A pinned value REFUSES a differing annotation at admission — never silently outvotes it, because a value the author wrote and did not get is a debugging session. The same value restated passes, which keeps migrations painless. Knobs the installation never set are untouched by all of this: overridable: "false" pins what you SET, not the whole contract.

agent:
  defaults:
    overridable: "false"        # everything set below is pinned…
    fileMode: "0640"            # …like this
    watchSeconds: "30?"         # …except this one, explicitly opened
    perStore:
      # A store pins its own group: everything vault-shaped is the
      # platform team's word, except the watch interval.
      vault: "overridable=false, endpoint=https://vault.vault.svc:8200, auth=kubernetes, watch-seconds=30?"
      # And a store can OPEN its group under a strict fleet flag:
      # consul values stay the pods' even with the "false" above.
      consul: "overridable=true, endpoint=http://consul.infra.svc:8500"

All three rungs compose per store and per value: a ! inside an overridable=true group pins just that value; a ? inside an overridable=false group opens just that one. The rule always comes from the TIER that supplied the value — a store's flag never pins a fleet default.

The pin follows the pair rule too: a pinned endpoint also refuses a pod that answers with endpoint-secret — the address is one decision, and it is not the pod's to make. A pinned fleet source (source: "consul!") refuses pods that name any other store, which is the enforcement twin of the sourceAllow gate below.

Per-store defaults, all nine stores

agent.defaults.perStore is the tier between the annotation and the fleet. One realistic installation, every store present — each line pairs with the pod that uses it below:

The string form below is one of the two spellings; the same installation as maps reads the same.

# values.yaml — or, for kustomize, the same settings in
# base/installation.yaml, or joined with "; " into
# DYNAMIC_CONFIG_AGENT_STORE_DEFAULTS
agent:
  defaults:
    watchSeconds: "30"          # the fleet's floor
    fileMode: "0640"
    perStore:
      # Local, cheap to poll: tighten the interval — and the address
      # is the platform's, PINNED, so pods need not carry it and
      # cannot point elsewhere.
      consul: "endpoint=http://consul.infra.svc:8500!, watch-seconds=10"
      # Secrets: the whole group is the platform team's word
      # (address, auth, CA), and the file lands owner-only, owned by
      # the app's uid.
      vault: "overridable=false, endpoint=https://vault.vault.svc:8200, auth=kubernetes, ca-configmap=vault-ca, file-mode=0400, agent-run-as-user=1000, agent-run-as-group=1000"
      # A JVM config server answers slowly; give the render room.
      config-server: "watch-seconds=20, agent-memory-limit=96Mi"
      # Billed per read: poll gently.
      firestore: "watch-seconds=120"
      # A clone per render costs memory and remote quota.
      git: "watch-seconds=120, agent-memory-limit=128Mi"
      # In-memory store, near-free reads.
      redis: "watch-seconds=10"
      # Watches are pushed by etcd itself; the interval is a backstop.
      etcd: "watch-seconds=60"
      # JetStream KV is push-cheap too.
      nats: "watch-seconds=10"
      # A GET per poll is a line on a bill; and S3 documents are often
      # the big ones.
      s3: "watch-seconds=60, agent-memory-limit=128Mi"

The pods, every field filled. None of them repeats a knob the installation already set — the tier exists so they never have to:

# consul — plain HTTP, a KV key
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "consul"
    dynamic-config.rs/endpoint: "http://consul.infra.svc:8500"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/path: "/config/rendered.toml"
    # watch-seconds arrives from perStore.consul: 10
# vault — kubernetes auth through the pod's own ServiceAccount,
# a private CA, one section of the secret
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "vault"
    dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
    dynamic-config.rs/key: "secret/myapp"
    dynamic-config.rs/section: "db"
    dynamic-config.rs/auth: "kubernetes"
    dynamic-config.rs/auth-role: "myapp"
    dynamic-config.rs/ca-configmap: "vault-ca"
    dynamic-config.rs/path: "/config/rendered.yaml"
    # file-mode 0400 and uid/gid 1000 arrive from perStore.vault
# config-server — the Spring-style application/profile pair as the key
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "config-server"
    dynamic-config.rs/endpoint: "http://config-server.infra.svc:8888"
    dynamic-config.rs/key: "billing/prod"
    dynamic-config.rs/path: "/config/rendered.json"
    # watch-seconds 20 and the 96Mi limit arrive from perStore
# firestore — the endpoint is the GCP project, the key a document path
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "firestore"
    dynamic-config.rs/endpoint: "acme-prod"
    dynamic-config.rs/key: "config/billing"
    dynamic-config.rs/path: "/config/rendered.json"
    # watch-seconds 120 arrives from perStore.firestore
# git — a repository over ssh, a ref, a file inside the tree
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "git"
    dynamic-config.rs/endpoint: "git@github.com:acme/config.git"
    dynamic-config.rs/ref: "main"
    dynamic-config.rs/key: "billing/prod.yaml"
    dynamic-config.rs/ssh-secret: "config-deploy-key"
    dynamic-config.rs/path: "/config/rendered.yaml"
    # watch-seconds 120 and the 128Mi limit arrive from perStore.git
# redis — the password rides in the URL, so the WHOLE endpoint is a
# Secret instead of an annotation
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "redis"
    dynamic-config.rs/endpoint-secret: "redis-cred/url"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/path: "/config/rendered.toml"
    # watch-seconds 10 arrives from perStore.redis
# etcd — a REQUIRED client certificate and a private CA (the
# password-auth twin swaps tls-secret for password-secret)
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "etcd"
    dynamic-config.rs/endpoint: "https://etcd.infra.svc:2379"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/tls-secret: "etcd-client-tls"
    dynamic-config.rs/ca-configmap: "etcd-ca"
    dynamic-config.rs/path: "/config/rendered.toml"
    # watch-seconds 60 arrives from perStore.etcd
# nats — a JetStream KV bucket and the key inside it
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "nats"
    dynamic-config.rs/endpoint: "nats://nats.infra.svc:4222"
    dynamic-config.rs/key: "config/db.json"
    dynamic-config.rs/path: "/config/rendered.toml"
    # watch-seconds 10 arrives from perStore.nats
# s3 — the endpoint IS the bucket; api-url points at MinIO/Ceph/R2,
# and aws-secret carries static credentials those servers need
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "s3"
    dynamic-config.rs/endpoint: "myapp-config"
    dynamic-config.rs/key: "prod/db.json"
    dynamic-config.rs/api-url: "http://minio.infra.svc:9000"
    dynamic-config.rs/aws-secret: "minio-cred"
    dynamic-config.rs/path: "/config/rendered.toml"
    # watch-seconds 60 and the 128Mi limit arrive from perStore.s3

Any pod that DOES set watch-seconds (or any other knob) still wins — the tiers only answer for what the pod left unsaid.

The gates, in depth

Three gates, one authority model: they live in the webhook's own configuration, owned by whoever installs it. None of them is a namespace annotation — whoever edits a namespace is usually the tenant being gated, and a gate its subject can open is not a gate. None of them reads the namespace object either: the webhook holds no RBAC and asks the API server for nothing; the pod's namespace arrives inside the AdmissionReview, and the ruling comes from configuration alone.

All three share one grammar:

spec   = group *( ";" group )
group  = [ namespace ":" ] names     ; no head, or "*" = every namespace
names  = name *( "," name )

agent-env allow — closed until opened

webhook:
  agentEnvAllow: "payments: HTTPS_PROXY, AWS_*; *: RUST_LOG"

agent-env puts environment on the container that holds store credentials, and environment steers SDKs. Concretely, on this agent: HTTPS_PROXY reroutes every vault/S3/consul request through a proxy of the pod author's choosing — with the bearer tokens and signatures inside; AWS_CA_BUNDLE and SSL_CERT_FILE swap the trust roots those connections verify against; AWS_EC2_METADATA_DISABLED, AWS_PROFILE, NO_PROXY all change where credentials come from or where traffic goes. That is why this gate defaults to closed: an empty allowlist refuses the annotation everywhere, and every name a pod wants must be opened by the installer, optionally per namespace.

Name rules: UPPER_SNAKE, exact match, or a trailing * as a prefix glob (AWS_*). A bare * opens everything — legitimate on a single-team cluster, a finding on a shared one. The refusal a pod sees names the variable, the namespace, what IS allowed there, and the chart value that opens the gate; it does not enumerate other namespaces' rules — one tenant's refusal must not describe another's setup.

What the gate does NOT govern: agent-env values (only names), the app container's environment (the agent's only), and the fleet's own agent.defaults.env (next section).

The fleet environment — no gate, on purpose

agent:
  defaults:
    env: "HTTPS_PROXY=http://egress.infra.svc:3128, RUST_LOG=info"

agent.defaults.env is environment EVERY injected agent gets — the cluster-wide egress proxy, a fleet log level. It passes no allowlist because the installer sets both the values and the allowlist; a gate you hold both sides of checks nothing. Merging rules, exactly:

  • A pod's own agent-env overrides a fleet name — the pod said it more specifically. (The pod's name still needs the allowlist: the OVERRIDE is a pod-author action even when the name is fleet-known.)
  • When a pod uses aws-secret, the fleet's AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY step aside — one credential, one place, and the pod named its Secret.
  • Everything else rides along verbatim, on every render's agent.

Source allow and deny — open until narrowed

webhook:
  sourceAllow: "payments: vault, s3; *: consul"   # empty = every store
  sourceDeny: "sandbox: git"                       # subtractive, wins

The source gates decide which STORES a namespace may render from. Two lists, because the two postures are different jobs:

  • sourceAllow — empty means *: every store, everywhere, so an upgrade changes nothing until the installer says so. Non-empty flips the posture: ONLY the listed sources pass, per namespace. Use it when the store set is policy: payments reads vault and s3, everyone may read consul, and anything unlisted is refused with a message naming what IS allowed there.
  • sourceDeny — always subtractive, and it outranks the allowlist: a source both listed and denied is denied. Use it for the surgical cut that does not flip the posture: git is off in the sandbox namespace, everything else stays open.

Both gates are judged against EVERY render on the pod — the default one and each named suffix: a denied store cannot ride in as source.cache. Entries are validated against the real store list at webhook startup, so sourceDeny: "sandbox: got" fails the install instead of silently gating nothing — in a security control, a typo that fails open is the worst of the four outcomes.

Deciding between them:

you wantuse
nothing changes on upgradeleave both empty
this namespace uses exactly these storessourceAllow
this store is banned here, rest stays opensourceDeny
allow broadly, carve exceptionsboth — deny wins on overlap

Kustomize, same doors

Every value above is either a key in base/installation.yaml — the readable form, and usually what you want — or one environment variable on the webhook Deployment. The chart is convenience, not capability:

# kustomization.yaml, an overlay patch
patches:
  - target: { kind: Deployment, name: dynamic-config-webhook }
    patch: |
      - op: add
        path: /spec/template/spec/containers/0/env/-
        value:
          name: DYNAMIC_CONFIG_AGENT_STORE_DEFAULTS
          value: "vault: file-mode=0400; git: watch-seconds=120"
      - op: add
        path: /spec/template/spec/containers/0/env/-
        value:
          name: DYNAMIC_CONFIG_WEBHOOK_SOURCE_DENY
          value: "sandbox: git"

The chart's schema cannot see a kustomize patch — which is exactly why the webhook re-validates the complete installation at startup and refuses to serve on any error. Helm users get two doors; kustomize users get the one that matters.

chart valueenvironment variable
agent.defaults.cpuRequest … .memoryLimitDYNAMIC_CONFIG_AGENT_CPU_REQUEST … _MEMORY_LIMIT
agent.defaults.fileModeDYNAMIC_CONFIG_AGENT_FILE_MODE
agent.defaults.watchSecondsDYNAMIC_CONFIG_AGENT_WATCH_SECONDS
agent.defaults.modeDYNAMIC_CONFIG_AGENT_MODE
agent.defaults.volumeMediumDYNAMIC_CONFIG_AGENT_VOLUME_MEDIUM
agent.defaults.nativeSidecarDYNAMIC_CONFIG_AGENT_NATIVE_SIDECAR
agent.defaults.runAsUser / .runAsGroupDYNAMIC_CONFIG_AGENT_RUN_AS_USER / _GROUP
agent.defaults.metricsPortDYNAMIC_CONFIG_AGENT_METRICS_PORT
agent.defaults.envDYNAMIC_CONFIG_AGENT_ENV
agent.defaults.sourceDYNAMIC_CONFIG_AGENT_SOURCE
agent.defaults.pathDYNAMIC_CONFIG_AGENT_PATH
agent.defaults.overridableDYNAMIC_CONFIG_AGENT_DEFAULTS_OVERRIDABLE
agent.defaults.perStoreDYNAMIC_CONFIG_AGENT_STORE_DEFAULTS
webhook.agentEnvAllowDYNAMIC_CONFIG_WEBHOOK_AGENT_ENV_ALLOW
webhook.sourceAllowDYNAMIC_CONFIG_WEBHOOK_SOURCE_ALLOW
webhook.sourceDenyDYNAMIC_CONFIG_WEBHOOK_SOURCE_DENY

The Annotation Contract

v1, and it is the API: a change here is a breaking change of this integration, whatever the binaries think. Two golden files in the webhook's test suite byte-compare the full admission response, so the contract cannot move without a reviewed diff saying it moved.

The core seven

annotationrequiredmeaning
dynamic-config.rs/injectyes"true" asks; anything else but "false" fails the admission
dynamic-config.rs/sourceyesconsul, vault, config-server, firestore, git, redis, etcd, nats, s3
dynamic-config.rs/endpointone of the twothe store's address — a url; <project>[/<database>] for firestore
dynamic-config.rs/endpoint-secretone of the two<secret>/<key> holding the address, when the address carries a password (a redis url)
dynamic-config.rs/keyyesthe document's key; mount/path for vault, application/profile for config-server, the file's path for git
dynamic-config.rs/pathyeswhere the rendered file lands; the extension picks the format
dynamic-config.rs/modenoinit, sidecar (default), both

Naming a store once

    dynamic-config.rs/class: "team-vault"
    dynamic-config.rs/key: "billing/config.json"
    dynamic-config.rs/path: "/config/app.yaml"

A DynamicConfigClass — the object the operator has read since 0.1.1 — holds the source, the endpoint and the credential's Secret. A pod that names one says only what is its own.

The class supplies defaults: a pod that names both a class and its own endpoint keeps its own, which is the rule the installation document already follows. Its own namespace is looked in first, then the cluster scope, so a namespace can override a platform default without asking the platform.

Off unless an administrator turns it on (webhook.classes.enabled), and the reason is worth reading. The admission path calls the API server nowhere — that is what keeps a busy API server from failing every pod creation in the cluster — and enabling this does not change it: the classes are listed on a background timer into a map in memory, and admission reads the map. What it does change is that the webhook holds a credential and a cluster-wide read of two custom resources, which is worth having only if pods actually name classes.

The cost is a synchronisation delay a mounted ConfigMap already has. A class created seconds ago may not have been polled yet, and a pod naming it is refused with that sentence rather than admitted without the store the class was to supply.

Two refusals are about who may use what:

  • a ClusterDynamicConfigClass whose namespaces list does not include the pod's — the list is what keeps cluster-scoped from meaning anyone
  • a cluster class whose credential Secret lives in another namespace. A pod mounts Secrets from its own namespace and nowhere else, so that class is usable by a DynamicConfigRender, which reads the Secret itself, and not by an injected agent

The namespace gate, before any pod is read

Every key below is per-pod. One optional guard sits a level above: with webhook.namespaceGating=true the webhook selects only namespaces labeled dynamic-config.rs/injection: enabled — a label, not an annotation, because the gate lives in the webhook configuration's namespaceSelector and Kubernetes selectors cannot see annotations. Inside a gated namespace the per-pod dynamic-config.rs/inject: "true" is still required; the security page owns the trade (blast radius, per-namespace failurePolicy: Fail), and examples/namespace-gating.yaml is the ready-to-apply shape.

What the webhook writes back

annotationvaluemeaning
dynamic-config.rs/statusinjectedthis pod has been through admission and carries the agent

Written by the webhook, never by a pod. A mutating webhook is not called once — reinvocationPolicy: IfNeeded asks the API server to call it again whenever a later webhook changes the pod, and some controllers resubmit a spec that has already been admitted. A marked pod is passed through untouched; without the mark, the second pass would add the agent again, and two containers with one name is a pod the API server refuses.

A pod that sets this annotation to anything else is refused, with a message saying so. Setting it to injected by hand is a way of saying "do not inject me", which inject: "false" already says more clearly.

Behaviour

annotationdefaultmeaning
dynamic-config.rs/timeoutthe store's own, 10sthe deadline for one fetch attempt, where the store's client has a door for it. Ten seconds is right for a store on this network and wrong for a Git remote across a WAN or a bucket in another region — until this, a workload in that position had no way to say so and its only recourse was a fetch that kept timing out
dynamic-config.rs/agent-imagethe installation'sthe image for this pod's agent. How an agent upgrade stops being all-or-nothing: one Deployment tries it first. Refused unless the installation lists a prefix it starts with — an image on the injected container runs chosen code beside the application, holding the store's credential
dynamic-config.rs/watch-seconds15how often the sidecar asks a store that must be asked, and how often it re-reads one that pushes; whole seconds. See Watching
dynamic-config.rs/sectionwhole documentthe section key the document nests under
dynamic-config.rs/native-sidecar"false""true" injects the watcher as an init container with restartPolicy: Always (Kubernetes 1.29+); Jobs finish
dynamic-config.rs/volume-mediummemorywhere the rendered file lives: memory (tmpfs, off the node's disk) or disk
dynamic-config.rs/init-first"false""true" puts the injected init container ahead of the pod's own, for a pod whose init container reads the rendered file. Appending stays the default: another injector's init container that must run first keeps running first. Refused with mode: sidecar, where there is no init container
dynamic-config.rs/agent-run-as-same-user"false"take the agent's UID from the application container instead of naming one, so the rendered file's owner matches without two numbers to keep in step. The application's own runAsUser first, then the pod's; absent is refused rather than guessed — inheriting whatever the image runs as is a UID that moves when the image does. Root is refused, as always
dynamic-config.rs/extra-secretnonea Secret mounted read-only under /etc/dynamic-config/extra, into the agent alone. For the file a store's own client wants that this contract does not model. The path is fixed rather than chosen: an annotation that took a mount path could be aimed at the rendered volume or the service-account token
dynamic-config.rs/inject-containersevery containera comma-separated list of the pod's own containers that receive the rendered volume. The default matches the reference implementation's; naming a subset is for the pod that runs a log shipper or a mesh proxy beside its application — neither has any business holding a rendered credential, and file-mode cannot draw that line because a sidecar usually runs as the same UID. A name the pod does not have is refused, and leaving out the container that env-inject wraps is refused too
dynamic-config.rs/agent-cpu-request10mthe injected container's CPU request
dynamic-config.rs/agent-memory-request32Miits memory request
dynamic-config.rs/agent-cpu-limitnoneits CPU limit — none by default, on purpose
dynamic-config.rs/agent-ephemeral-requestnoneephemeral-storage request. Matters for exactly one configuration and is unbounded without it: on volume-medium: disk the rendered volume is node storage rather than the pod's memory, history keeps copies beside it, and nothing else here declares that resource — a pod could fill a node's disk without exceeding a limit it had declared
dynamic-config.rs/agent-ephemeral-limitnoneits ephemeral-storage limit. A request larger than it is refused, because the scheduler's own refusal would name the pod rather than these
dynamic-config.rs/agent-memory-limit64Miits memory limit
dynamic-config.rs/file-modeumask's answer (0644)the rendered file's octal permissions, e.g. "0640" — set on the scratch file before the atomic rename, so a reader never sees the final path in a mode it will not keep
dynamic-config.rs/agent-run-as-user65532the injected container's UID, so the rendered file's owner matches what the app runs as; 0 is refused — the agent stays nonroot in every configuration
dynamic-config.rs/agent-run-as-group65532its GID, same rule
dynamic-config.rs/aws-secretnones3 only: a Secret whose AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY keys become exactly those variables on the agent — static credentials for S3-compatibles that are not AWS (MinIO, Ceph, R2); on AWS, IRSA needs none of it
dynamic-config.rs/metrics-portnonethe agent serves Prometheus text on this port (the series); only meaningful with a watching agent, and "0" opts out of an installation-wide default
dynamic-config.rs/agent-envnonecomma-separated NAME=value pairs, set as environment on every injected agent — SDK knobs like RUST_LOG, AWS_CA_BUNDLE, proxy variables. Gated: only names the installation's allowlist permits in this namespace pass admission (the gate)
dynamic-config.rs/env-injectnonea container name: its command is wrapped in set -a; . <path>; set +a; exec …, so the rendered dotenv is the process's REAL environment. Needs mode: init or both (env freezes at container start — Kubernetes' rule) and an explicit command (an ENTRYPOINT is invisible to the webhook); both refusals name the fix
dynamic-config.rs/env-restart"false"with env-inject and mode: both: when the sidecar re-renders the dotenv, the kubelet restarts JUST the app container (a liveness probe compares the file against the fingerprint the wrapper exported at start) and the wrapper re-sources the new file — the closest thing to a live env update the kernel permits: seconds, no pod recreation, no new IP. Refused if the container already has its own livenessProbe
dynamic-config.rs/templatenonean inline minijinja template; it owns the output bytes
dynamic-config.rs/template-configmapnone<name> or <name>/<key> (default key template): the template from a ConfigMap, mounted read-only and re-read every render

Freshness, and what happens when a store is not there

The store is not always reachable, and the document does not always still exist. Five annotations say what the agent does about it; every default is the one that keeps a running application running.

annotationdefaultmeaning
dynamic-config.rs/startup-policyallow-cachedwhat a first fetch failure means. allow-cached serves the file already on the volume if it parses — the rendered volume survives a container restart, so this is a real cache rather than a hopeful one — and goes on watching. require-fresh refuses to start without a fresh document, for a pod that must never come up on stale credentials. best-effort starts regardless
dynamic-config.rs/max-stalenessnonehow old the document may get before /readyz reports 503 — "6h", "90s". Last-known-good answers is there a document; this answers is it still worth trusting. A credential may be worthless after five minutes and a feature flag fine after a day, which is why there is no default
dynamic-config.rs/on-deleteretainwhat a document disappearing from the store means. retain keeps serving the last render, remove truncates the file so a consumer reads nothing rather than a revoked secret, fail ends the agent so the pod's restart policy takes over. Whichever is chosen it is reported: dynamic_config_agent_absent moves and the log says whether the store answered gone or did not answer at all
dynamic-config.rs/require-ack"false"readiness waits for the application to POST the fingerprint it is running, not merely for a document to exist. What it is for
dynamic-config.rs/historynonekeep the last N replaced generations beside the render. What it is for
dynamic-config.rs/readiness"true"whether the webhook attaches a readiness probe to the injected container. Pod readiness is AND-ed across containers, so a Service sends no traffic to a pod whose configuration has not arrived. "false" opts out, for a deployment that would rather start
dynamic-config.rs/max-document-bytes8388608the largest document the agent will accept, checked before it is parsed. The injected container's memory limit is 64Mi by default and a document is held more than once while it is resolved

Dynamic secrets

For a Vault path that mints a credential rather than storing one — database/creds/…, pki/issue/…, aws/creds/….

annotationdefaultmeaning
dynamic-config.rs/dynamic"false"read the path as a dynamic engine: no KV data/ nesting, and keep the lease Vault answered with. A renewable lease is renewed at 65% of its TTL, spread; one Vault marked renewable: false — every pki/issue — is never sent a renewal at all and is re-issued at 90% instead. A renewal keeps the file; a re-issue is new credentials, so it re-renders. For a certificate, whichever expires first binds: the lease, or the certificate's own notAfter. Vault has the whole of it
dynamic-config.rs/revoke-grace5show long the agent may spend handing the lease back. Refused past the pod's terminationGracePeriodSeconds — beyond that the kubelet sends SIGKILL, the revocation is cut off mid-request, and the lease stays out anyway, which is what the annotation was set to avoid
dynamic-config.rs/revoke-on-shutdown"true"give the lease back on SIGTERM. Best-effort with a short deadline: a pod that cannot reach Vault while terminating still terminates. Setting it without dynamic is refused — there is no lease to revoke

Vault revokes an expired lease on its own eventually; revoking on the way out is what makes eventually into now. See Vault.

Reaching the store over TLS

Two settings beside ca-configmap and tls-secret, and they are not the same kind of thing: one moves which name is checked, the other stops checking.

annotationdefaultmeaning
dynamic-config.rs/tls-server-namethe endpoint's hostthe name the store's certificate must carry, for an endpoint written as an address it does not name — a Service's cluster IP, a load balancer, a NodePort. The server stays authenticated: the certificate still chains to a trusted authority and still has to carry this name
dynamic-config.rs/tls-skip-verify"false"connect without authenticating the store at all
dynamic-config.rs/tls-reload"true"rebuild the store's client when ca-configmap, tls-secret or ssh-secret changes on disk — without restarting the pod. "false" keeps the old client until something restarts it

Not every store can express either. The clients differ, and a store that cannot say something refuses the whole configuration by name rather than ignoring it:

storetls-server-nametls-skip-verify
vault, consul, firestorenoyes
config-serveryesyes
gitnoyes
etcdyesno
redis, nats, s3nono

The two columns are almost disjoint, and that is the clients rather than a choice: ureq can turn verification off and cannot override a name, tonic can override a name and cannot turn verification off.

What tls-skip-verify costs

It is not a weaker TLS. It is TLS without the part that makes it mean anything: any party on the network path can present any certificate, read what is sent, and rewrite what comes back — and what comes back from a store is the configuration and the credentials this pod is about to run on.

So it is the one annotation in this contract a workload cannot reach on its own. Four things gate it:

  • an administrator must set webhook.allowTlsSkipVerify: true, or the annotation is refused with that value named
  • it is refused alongside ca-configmap: naming an authority and then not checking it are two answers to one question
  • every pod using it earns an admission warning, which kubectl prints
  • the agent logs it at start and reports dynamic_config_agent_tls_verification_skipped 1, so one alert finds every pod doing it

Rotation, without a restart

The kubelet rewrites a mounted ConfigMap or Secret in place: the file the container sees changes and nothing tells the process. Every store client in this family reads its trust material once, when it builds — so before 0.3.0 a rotated CA meant a pod restart, and a rotation nobody restarted for meant a store that stopped answering at a moment unrelated to the rotation.

The agent now watches the files it was given and rebuilds the client when they move. A rebuild is not a restart: the process stays up, the rendered file never leaves the volume, the last-known-good stays, and every counter keeps counting — only the client is new. dynamic_config_agent_tls_reloads_total counts them, and a fleet where that stays at zero through a CA rotation is a fleet that did not notice.

A rotation writes more than one file, so a change is confirmed by reading twice a quarter-second apart: rebuilding a TLS client from a half-written certificate and key is a failure that reads as a bad certificate.

Tokens needed none of this and never did — the projected service-account token and the config server's bearer file are re-read on every use, so the rotation that actually happens hourly was always picked up.

Try ca-configmap first. The two situations people usually reach for skip-verify in — a development server with a self-signed certificate, an enterprise private CA — are both one more certificate to trust, which is one annotation and keeps the server authenticated.

Telling the application what it is running

annotationdefaultmeaning
dynamic-config.rs/meta"false"write a sibling .<name>.meta beside the render — /config/app.yaml gets /config/.app.yaml.meta — holding the digest of the bytes, the store's own revision, and when it landed. It describes the render and never contains it: no values, ever
dynamic-config.rs/schema-configmapnone<name> or <name>/<key> (default key schema.json): a JSON Schema the resolved document must satisfy before it is published. A document that fails is refused and the last good one keeps serving, which is the behaviour a consumer that is not Rust, Python or Node cannot get any other way

Telling the application the file moved

    dynamic-config.rs/notify-http: "http://127.0.0.1:8080/-/reload"

The rename is atomic, so a consumer never sees half a document — but nobody tells it the document changed. nginx, Prometheus and most legacy daemons reload on a request and on nothing else.

After the rename, never before: the whole promise of the notification is that the document is already there when it arrives. One attempt, a two second deadline, and never fatal — the file is already correct, so a notification that did not land has undone nothing. notifications_total and notification_failures_total say how it went.

Localhost only, by construction. http://127.0.0.1, http://localhost or http://[::1], and nothing else — an agent that will POST to an arbitrary URL is an SSRF primitive holding this pod's store credential. The address is checked at admission and in the agent, because the two run in different places. Refused with mode: init, where the container writes once and exits before the application it would notify has started.

There is no signal form. Signalling a sibling container needs shareProcessNamespace: true — a pod-wide change to the process boundary between containers, which the webhook will not make on a workload's behalf.

Failures where somebody is already looking

    dynamic-config.rs/events: "true"

kubectl describe pod is where an operator looks first, and a render failure was not there — it was in the sidecar's log, one kubectl logs -c dynamic-config-agent away from the question. With this the agent writes a Warning Event on its own pod for a render that failed (RenderFailed) and for a document that vanished from the store (DocumentAbsent).

Off, and twice opt-in. The sidecar carries no API credential in any other configuration — that is a property the rest of this design leans on — so writing Events means mounting the pod's service-account token beside the application. An administrator has to create the Role (agent.events.enabled with the namespaces, which grants create on events and nothing else) and set webhook.allowEvents, and only then may a pod ask. Asking without the installation offering it is refused, with the chart value named, rather than admitted into a 403 at the first failure.

An Event is commentary on work that already happened, so it is written off the loop: a render never waits on the API server, and an Event that could not be written is logged once and dropped.

Some of the fleet before all of it

    dynamic-config.rs/canary-configmap: "rollout"    # or "rollout/percent"
kubectl create configmap rollout --from-literal=percent=5
# look at the five per cent
kubectl patch configmap rollout --type merge -p '{"data":{"percent":"100"}}'

A change published to every pod at once either works or is an incident. This buys the third outcome: a few pods take it, somebody looks, and the rest follow or do not.

Which pods is the pod's own name hashed into a bucket from 0 to 99, so the cohort is deterministic and stable — a cohort that reshuffled as it widened would put every pod through the new document eventually and prove nothing about any of them. No coordination and no leader: five thousand agents each answer for themselves.

Who widens it is whoever edits the ConfigMap. That is why the percentage is a mounted file and not an annotation: the kubelet rewrites a mount in place, so the cohort grows with no pod restart — and a restart would discard the very state a canary exists to watch. A pod outside the cohort holds what it fetched and publishes it the moment the number passes its bucket, so nothing has to be re-fetched and the store need not say anything again.

0 holds everybody and 100 holds nobody; both ends read as what they mean. A file that is missing or does not parse is no canary at all rather than zero — a typo must not freeze the fleet on its current document.

Two series say what is happening: canary_holding is 1 while this pod is outside the cohort, and canary_percent is the number it last read. Together with applied from below they are the question somebody has to answer before widening: did the pods that took it actually run it?

Refused with mode: init, where the agent publishes once and exits.

One interaction to know. A held pod's document is by definition not the newest, so max-staleness will eventually call it stale. That is correct — it is stale — but a long-running canary under a short staleness ceiling will make held pods unready, and the two numbers should be chosen together.

Knowing the application actually applied it

    dynamic-config.rs/meta: "true"          # so the app can read the digest
    dynamic-config.rs/require-ack: "true"   # and readiness waits for it

renders_total says a document reached disk. It says nothing about whether the application read it, and an application still running the previous one while every dashboard reports success is the outage this closes.

The application POSTs the fingerprint it is running to /applied on the metrics port:

curl -fsS -XPOST --data-binary @- http://127.0.0.1:9110/applied   <<< "$(jq -r .fingerprint /config/.app.yaml.meta)"

200 means that is the current document, 409 means a different one is published and names it — a restarted application acknowledges what it read before the restart, and a slow one acknowledges a generation the store has already replaced. Neither is an error the application caused, and neither is convergence.

A Rust, Python or Node application does not need the meta file: fingerprint() is the same string, from the same digest.

Three series come out of it, and the third is the one to alert on:

serieswhat it says
acks_totalacknowledgements received
ack_mismatches_totalacknowledgements naming a document this agent never published. Climbing steadily means acknowledgements and renders are talking past each other
applied1 when the application is running what was published
unapplied_secondshow long the published document has gone unacknowledged

require-ack makes that readiness. The pod stays 0/2 until the application says it applied the document, so a Service sends no traffic to a pod running configuration nobody confirmed. Off by default and refused without a readiness probe or on mode: init: it needs the application's cooperation, and one that never acknowledges would never become ready.

An application that never acknowledges is not penalised. It leaves applied at zero and nothing else changes.

What the file was before

    dynamic-config.rs/history: "3"

The rename that publishes a new document is the same rename that destroys the old one, so what was the file before? — the question an incident starts with — has no answer by the time anybody asks it. This keeps the replaced generation beside the render, under /config/.app.yaml.history/<when>-<digest>.yaml, newest kept and oldest pruned past the count.

At most ten, and off unless asked: the rendered volume is the pod's own memory by default, so every kept generation is charged to a limit that is 64Mi by default. Each copy takes the mode the render had, so a history entry cannot be read by anything the document itself could not be.

Refused with volume-medium: disk, where a replaced secret would sit on node-backed storage, outlive the pod that held it, and survive a reboot.

This is not a rollback. Putting an old document back needs the application to say the new one is bad, and nothing here can hear that — half a feature that implies the other half is worse than neither. This keeps files; a person reads them.

When something else writes to the rendered file

    dynamic-config.rs/on-drift: "repair"     # or warn (default), fail

The agent owns the file; the volume is shared. A debug session, an init container with an opinion, an application that rewrites its own configuration — any of them can change it, and until the next change arrives from the store, nothing notices.

Checked on the same cadence the store is read. warn is the default because the agent cannot know whether the write was a mistake or the point; what it can do is stop the difference being invisible. repair writes the rendered document back, for a file whose contents are the store's and nobody else's. fail ends the agent, so the pod restarts and renders again.

Either way dynamic_config_agent_drift goes to 1 and drift_total counts it. An unreadable file counts as drifted too: one that has been deleted, or has become a directory, is not the file this agent wrote either.

Several files, one fetch

also.<name> cuts more files from the same fetched document, and publishes them all or none of them:

dynamic-config.rs/source: "vault"
dynamic-config.rs/key: "secret/myapp"
dynamic-config.rs/path: "/config/app.yaml"
# the same document, a section of it, as a second file:
dynamic-config.rs/also.db: "/config/db.env"
dynamic-config.rs/also-section.db: "database"

One fetch, one generation, one refusal if any of them cannot be written — so a failure in the third does not leave the first two published. Each write is still its own atomic rename: a reader can catch the microseconds between two of them, which is a rename apart rather than a fetch apart.

This is within one document. Several stores cannot share a generation — two stores have no common instant and no protocol between them can say "these two reads are the same" — so a second store stays a named render, below.

Several documents, one pod

Every store-shaped key accepts a .<name> suffix, and each name is one more injected agent writing one more file into the same directory:

# the default render:
dynamic-config.rs/source: "vault"
dynamic-config.rs/key: "secret/myapp"
dynamic-config.rs/path: "/config/app.yaml"
dynamic-config.rs/auth: "kubernetes"
# a second, named `cache`:
dynamic-config.rs/source.cache: "redis"
dynamic-config.rs/endpoint-secret.cache: "redis-url/url"
dynamic-config.rs/key.cache: "myapp/cache.json"
dynamic-config.rs/path.cache: "/config/cache.toml"

Per-name: everything a store needs — source, endpoint(+secret), key, path, section, watch cadence, every auth key, CA/TLS/ssh material (mounted under suffixed paths), templates, file-mode. Pod-wide, on purpose: mode, volume medium, resources, run-as identity, env-inject (the default render's file is the one that can become the environment). Two rules, refused with the fix named: every path lives in the default path's directory (one shared volume), and each name's source/key/path are required exactly like the default's. Container names follow the suffix — dynamic-config-agent-cache in kubectl get pod, so a broken render says which one it is.

Authentication

Each store takes its own methods; the store's page spells every one of them out with full manifests. The webhook forwards these to the agent verbatim, and the agent refuses a wrong combination at startup — in the pod's events, not as a store error twenty minutes later.

annotationmeaning
dynamic-config.rs/auththe method: consul token | kubernetes | jwt; vault token | kubernetes | approle | jwt | userpass | ldap | cert; firestore metadata-server | access-token | emulator; git anonymous | token | ssh-key
dynamic-config.rs/auth-mountvault: the auth method's mount path when not the default; consul: the auth method's name (required for kubernetes and jwt)
dynamic-config.rs/auth-rolevault kubernetes: the role to assume (required); vault approle: the role id; vault jwt/cert: optional
dynamic-config.rs/auth-usernamevault userpass/ldap: the user; git: the basic-auth user when the host wants one
dynamic-config.rs/auth-token-pathwhere the service-account token is mounted, when a projected volume moved it
dynamic-config.rs/namespacethe Vault namespace (Vault Enterprise)
dynamic-config.rs/refgit: main, branch:main, tag:v1.4, or commit:<sha>
dynamic-config.rs/api-urlfirestore: the API endpoint when it is not Google's — the emulator

Secrets and certificates

Secret material never rides an annotation — kubectl describe pod prints annotations and arguments to anyone with pod read access. These four name Kubernetes objects instead; the webhook mounts them and the agent reads them. The geography is fixed.

annotationformbecomes
dynamic-config.rs/token-secret<secret>/<key>env DYNAMIC_CONFIG_AGENT_TOKEN
dynamic-config.rs/password-secret<secret>/<key>env DYNAMIC_CONFIG_AGENT_PASSWORD — the approle secret id, the userpass/ldap password
dynamic-config.rs/endpoint-secret<secret>/<key>env DYNAMIC_CONFIG_AGENT_ENDPOINT
dynamic-config.rs/ca-configmap<name> or <name>/<key> (default key ca.crt)a read-only mount under /etc/dynamic-config/ca and the agent's --ca
dynamic-config.rs/tls-secret<name> — a kubernetes.io/tls Secreta read-only mount under /etc/dynamic-config/tls and --tls-cert/--tls-key (that Secret type fixed its two keys as tls.crt/tls.key)
dynamic-config.rs/ssh-secret<name> or <name>/<key> (default key ssh-privatekey, the kubernetes.io/ssh-auth convention)a 0400 mount under /etc/dynamic-config/ssh, --ssh-key, and auth: ssh-key implied when no auth was named

Value forms, source by source

What endpoint, key and auth take for each value of source — the store pages carry the full manifests, this table is the lookup:

sourceendpointkeyauth values
consulhttp(s)://host:8500KV path with extension: myapp/config.json(none), token, kubernetes, jwt
vaulthttp(s)://host:8200<mount>/<path>: secret/myapp(none = token), token, kubernetes, approle, jwt, userpass, ldap, cert
config-serverhttp(s)://host:8888<application>/<profile>: billing/prod(none — bearer via token-secret only)
firestore<project> or <project>/<database>: acme-prodcollection/document: config/billing(none = metadata-server), metadata-server, access-token, emulator
gitany clone url: https://…, git@host:org/repo.gitfile path in the repository: billing/prod.yaml(none = token if set, else anonymous), anonymous, token, ssh, ssh-key
redisredis:// / rediss:// url — via endpoint-secret when it carries a passwordkey with extension: myapp/config.json(none — credentials live in the url)
etcd(no auth key)tls-secret client certificates, or auth-username + password-secret — etcd's own two methods, both first-class--key is the etcd key
nats(no auth key)a .creds file via auth-token-path, or token-secret; anonymous otherwise--key is <bucket>/<key>
s3(no auth key)the ambient AWS chain — IRSA on EKS, the workload's own identity--endpoint is the bucket; api-url overrides for MinIO/Ceph/R2

The prefix is claimed territory

Every dynamic-config.rs/* annotation must be a key this page lists — an unknown one fails the admission. The rule exists for the typo: tokne-secret silently ignored would be a pod running without the authentication it declared, and nobody would know until the audit. Annotations outside the prefix are none of this webhook's business and pass untouched.

Retiring a key

Nothing here is deprecated. When something is, the contract has somewhere to say so: every key is one row of a registry inside the webhook, carrying the release that retired it and what to write instead. A pod that sets a retired key is admitted with a warning naming its replacement, rather than refused — a contract that breaks a working pod to make a point is a contract people pin an old version of.

Two properties come from the key list being one table rather than several. Whether a key may take a .name suffix is read off the same row, so the two can no longer disagree; and the documentation is checked against that table by a test, so a key cannot be accepted and left undocumented.

The one liberty the strictness buys back: because every key is validated, template and template-configmap could ship later without a migration — pods that used them early were refused, not silently ignored. They shipped; the Rendering page owns them.

What fails the admission

A wrong ask fails the admission. A pod that says inject: "true" and misspells the rest is refused with the reason, not started without its configuration — silence there is how an outage begins. The refusals, verbatim from the tests:

  • inject set to anything but "true"/"false"
  • a missing required annotation, named in the message
  • mode outside init | sidecar | both
  • watch-seconds that does not parse as whole seconds
  • a *-secret value without the <secret-name>/<key> slash
  • endpoint and endpoint-secret both set — one address, one place
  • ssh-secret alongside an auth other than ssh-key
  • volume-medium outside memory | disk; native-sidecar outside true | false
  • a resource annotation that is not a Kubernetes quantity
  • file-mode outside octal 0400–0777 — setuid bits answer no question, and an owner-unreadable file is write-only noise
  • agent-run-as-user/-group of 0 — the agent stays nonroot in every configuration
  • inject-containers naming a container the pod does not have, or set to nothing at all — a render nothing can read is a render nobody asked for
  • init-first with mode: sidecar, where there is no init container
  • agent-run-as-same-user with no runAsUser to inherit, alongside agent-run-as-user, or on an application that runs as root
  • revoke-grace past the pod's terminationGracePeriodSeconds, at "0", or without dynamic
  • notify-http at anything but a localhost address, or with mode: init
  • on-drift outside warn | repair | fail
  • history at "0", past ten, or alongside volume-medium: disk
  • agent-ephemeral-request larger than agent-ephemeral-limit
  • timeout at "0", which is no deadline rather than the store's own
  • agent-image naming an image no webhook.agentImageAllow prefix admits
  • require-ack without a readiness probe, or with mode: init
  • canary-configmap with mode: init
  • events where the installation does not offer it — the refusal names the chart value an administrator sets
  • class naming one that is not visible, one whose namespaces list excludes this pod, or one whose credential is in another namespace
  • tls-skip-verify where the installation does not offer it, or alongside ca-configmap
  • env-inject naming a container that inject-containers leaves out: the wrapper sources the rendered file, so it has to be able to read it
  • env-inject with mode: sidecar, a container the pod does not have, or a container with no explicit command
  • env-restart without env-inject, without mode: both, or on a container that already owns a livenessProbe
  • agent-env entries that are not NAME=value, names that are not UPPER_SNAKE, a name set twice, or a name that shadows what aws-secret already sets
  • an agent-env name outside the installation's allowlist for the pod's namespace — the refusal names the chart value that opens it
  • a source sourceDeny turns off in the pod's namespace, or one missing from a non-empty sourceAllow — checked on every render, named suffixes included
  • an annotation that differs from a PINNED installation value (!-marked, or any set value under overridable: "false") — the refusal names both values; restating the pinned value passes
  • any dynamic-config.rs/* key the contract does not list
  • template and template-configmap both set — one template, one place

What passes with a warning

Not every misconfiguration earns a refusal. Kubernetes lets an admission response carry warnings, which kubectl prints and a controller records, and four configurations earn one:

the pod saidthe warning
volume-medium: diskthe rendered document is on node-backed storage, where it outlives the pod and is readable by anything that can read the node's disk
a world-readable file-modeevery container in the pod can read it, including ones added later
watch-seconds below 5every replica polls the store at that rate, and a store with a native watch delivers changes without one
dynamic with revoke-on-shutdown: "false"the credential stays valid after the pod is gone, until its lease expires

The bar is high on purpose: a warning on every admission is a warning nobody reads. Each of these is a configuration that works and is probably not what was meant.

Checking a manifest before the cluster does

The webhook is a pure function of a pod and an installation, so the same decision is available without a cluster:

$ dynamic-config-webhook validate pod.yaml deployment.yaml
pod.yaml: allowed, and an agent would be injected
deployment.yaml: refused (InvalidAnnotation)
  dynamic-config.rs/auth-role is required for vault kubernetes auth

JSON or YAML, one exit code for the lot, installation defaults from the process's own environment — so a CI job that runs the webhook's image with the chart's environment gets the answer the cluster would give. The point of catching it here is that the alternative is catching it from a rollout that will not start.

Whatever passes the webhook is validated again by the agent, which knows the store-by-store rules (vault kubernetes needs auth-role, consul kubernetes needs auth-mount, a certificate needs its key). Those refusals land in the injected container's log and the pod's events.

DynamicConfigClass shrinks all of this to a class reference — the operator page carries the shape.

The agent-env gate

agent-env puts variables on the container that holds store credentials, and environment steers SDKs — HTTPS_PROXY reroutes the agent's traffic, AWS_CA_BUNDLE and SSL_CERT_FILE swap its trust roots. So the names that may pass are the installer's decision, declared once:

# values.yaml
webhook:
  agentEnvAllow: "payments: HTTPS_PROXY, AWS_*; *: RUST_LOG"

Semicolons separate groups; each group takes an optional namespace: head (absent or * means every namespace); a trailing * on a name is a prefix glob. Empty — the default — refuses the annotation everywhere. Kustomize installs set the same grammar in the DYNAMIC_CONFIG_WEBHOOK_AGENT_ENV_ALLOW variable, and either way the webhook validates it at startup and refuses to serve on a typo.

The gate is NOT a namespace annotation, by design: whoever edits a namespace is usually the tenant being gated, and a gate its subject can open is not a gate. It is also not read from the namespace object at admission time — the webhook holds no RBAC and asks the API server for nothing; the pod's namespace arrives inside the AdmissionReview, and the ruling comes from the webhook's own configuration.

The source gates

The same authority model gates which STORES a pod may use:

webhook:
  sourceAllow: "payments: vault, s3; *: consul"   # empty = every store
  sourceDeny: "sandbox: git"                       # subtractive, wins

sourceAllow empty means every store everywhere — the safe default for an upgrade; non-empty means ONLY the listed, per namespace. sourceDeny turns stores off outright and outranks the allowlist. Both cover every render on the pod — a denied store cannot ride in on a named suffix — and both refusals name the namespace and the value that opens the gate. Kustomize sets DYNAMIC_CONFIG_WEBHOOK_SOURCE_ALLOW / _SOURCE_DENY; a name that is not a real store fails webhook startup, so a typo cannot silently gate nothing.

Defaults come in tiers

Every default in the table above is only the LAST tier of four:

annotation  >  per-store default  >  fleet default  >  built-in

The middle tiers are the installation's (the full list): fleet defaults cover every knob — resources, file-mode, watch-seconds, mode, volume-medium, native-sidecar, agent-run-as-user/-group, metrics-port, path, source, and a fleet-wide agent environment — and agent.defaults.perStore sets, one store at a time, the same knobs PLUS every store-shaped annotation: endpoint, key, auth and its friends, the credential Secrets, the templates. A pod can deploy carrying nothing but inject: "true". Pod-wide knobs (mode, volume, resources, identity) take the DEFAULT render's store tier; per-render knobs resolve against each render's own store. Every tier is validated with the SAME rules as the annotation it stands in for, at webhook startup — and any value can be PINNED (!, or overridable: "false"), refusing a differing annotation instead of being overridden by it.

Installation Defaults and Gates carries the full knob table, a per-store example for all nine stores, and the gates in depth.

Secrets, Certificates, and What Shows Where

Before wiring any store's authentication, one map of where things are visible. Kubernetes shows different fields to different eyes:

where a value sitswho sees it
an annotationanyone with get pod — and every system that logs admission objects
a container argumentanyone with get pod; kubectl describe pod prints args in full
an environment variable from a Secretthe pod spec shows only the Secret's name; the value needs get secret rights
a mounted Secret/ConfigMapsame — the spec names the object, the bytes stay behind RBAC

That table decides the whole contract:

  • Names travel in annotations — a role name, an auth method's mount, a username. Reading them tells an attacker nothing they could not guess.
  • Secrets travel as environment variables drawn from Secrets — the agent reads three, and there is no flag for the second one on purpose:
variableannotation that fills itcarries
DYNAMIC_CONFIG_AGENT_TOKENtoken-secret: <secret>/<key>the bearer/access token
DYNAMIC_CONFIG_AGENT_PASSWORDpassword-secret: <secret>/<key>the second secret, where a method has one: approle's secret id, userpass/ldap's password
DYNAMIC_CONFIG_AGENT_ENDPOINTendpoint-secret: <secret>/<key>the address, when the address embeds a password — a redis url
  • Key material travels as mounts, read-only, into the agent container alone — the application containers never see them:
annotationobjectlands at
ca-configmap: <name>[/<key>]ConfigMap, default key ca.crt/etc/dynamic-config/ca/
tls-secret: <name>kubernetes.io/tls Secret/etc/dynamic-config/tls/tls.crt + tls.key
ssh-secret: <name>[/<key>]kubernetes.io/ssh-auth Secret, default key ssh-privatekey, mounted 0400/etc/dynamic-config/ssh/

A private CA, end to end

Most internal Vaults, Consuls and git hosts serve TLS from an internal PKI. The chain is one ConfigMap and one annotation:

kubectl create configmap vault-ca --from-file=ca.crt=./internal-ca.pem
metadata:
  annotations:
    dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
    dynamic-config.rs/ca-configmap: "vault-ca"

The agent gets --ca /etc/dynamic-config/ca/ca.crt, and the store crate under it adds the CA to its trust roots — the same TlsConfig every store crate takes, so the spelling is identical for all six stores.

There is no way to turn verification off. The store crates refuse that setting by design, and the agent adds no flag for it: a configuration channel that skips TLS verification is a configuration channel anyone on the path can write to.

A client certificate

Some stores authenticate with the certificate (vault's cert method), some merely allow mTLS in front. Either way it is one kubernetes.io/tls Secret:

kubectl create secret tls vault-client \
  --cert=./client.pem --key=./client-key.pem
    dynamic-config.rs/tls-secret: "vault-client"

The certificate and key must come together; the agent refuses one without the other before any byte leaves the pod.

Why the pod's own identity beats all of this

Three of the six stores can authenticate a pod with no distributed secret at all — the pod's service-account token or the node's cloud identity:

Where one of those is available, prefer it: nothing to rotate, nothing to leak, and revocation is the platform's own. The token-shaped methods on every page exist for the stores and shops where it is not.

The Security Posture

Everything this integration does to a cluster, listed where an auditor can find it. The one-line summary: injection never relaxes a pod's posture, secrets never appear where kubectl describe reaches, and the webhook holds no credentials at all.

The injected agent complies with restricted PSS

Every injected container — init, sidecar, or native sidecar — carries the full restricted Pod Security Standard posture, so injection works in namespaces that enforce pod-security.kubernetes.io/enforce: restricted and passes the audit in ones that only warn:

securityContext:
  runAsNonRoot: true
  runAsUser: 65532          # distroless nonroot
  runAsGroup: 65532
  allowPrivilegeEscalation: false
  capabilities: { drop: ["ALL"] }
  readOnlyRootFilesystem: true
  seccompProfile: { type: RuntimeDefault }

The root filesystem is read-only because the agent writes exactly one place: the shared volume. Which is —

The rendered file lives in memory

The shared emptyDir is medium: Memory by default: rendered configuration regularly carries credentials, and tmpfs keeps them off the node's disk and out of its backups. The file is gone when the pod is. A pod that prefers disk (a giant document, a memory-tight node) says so:

    dynamic-config.rs/volume-medium: "disk"

Note the accounting: tmpfs pages count against the pod's memory. Configuration documents are small; the default agent memory limit below leaves room.

The agent's resource ask

Injected with requests and limits so it can never be the reason the node evicts the app:

resources:
  requests: { cpu: 10m, memory: 32Mi }
  limits: { memory: 64Mi }        # no CPU limit: throttling a config
                                  # agent buys nothing, delays reloads

Four annotations move them per pod: agent-cpu-request, agent-memory-request, agent-cpu-limit, agent-memory-limit — each a Kubernetes quantity, refused at admission when it is not.

Native sidecars

On Kubernetes 1.29+, ask for the sidecar as the platform now spells it:

    dynamic-config.rs/native-sidecar: "true"

The watching agent becomes an init container with restartPolicy: Always — started before the app containers, stopped after them, and a Job with one finishes, where a classic sidecar would hold it in Running forever. With mode: "both" the one-shot init still lands first, so the file-exists-before-the-app guarantee survives the move.

Secrets: where each thing is allowed to appear

The contract's rule: names in annotations, secrets in Secret-backed environment variables, key material in read-only mounts into the agent container alone. The webhook enforces it — there is no annotation that accepts a token value, and the password slot has no flag on the agent at all.

Which containers see the rendered document

Every container in the pod, unless the pod says otherwise:

    dynamic-config.rs/inject-containers: "app"

The default is the reference implementation's, and it is the right default — a pod whose containers all serve the same application should not have to enumerate them. But a pod that runs a log shipper, a mesh proxy or a debug sidecar beside its application is a pod where one container needs the credential and the others do not, and file-mode cannot express that: a sidecar in the same pod usually runs as the same UID, and reads a 0600 file exactly as well as the application does.

Naming a subset is the only lever that draws the line, so it exists. A name the pod does not have is refused rather than ignored — a typo would leave the application without its configuration and every other container without it too.

Policies Kubernetes enforces better than this webhook would

policies/ carries four ValidatingAdmissionPolicy samples in CEL: volume-medium: disk refused, tls-skip-verify refused whatever the installation allows, a world-readable file-mode refused, and the injected agent required to be pinned by digest.

They are examples rather than a feature. Kubernetes has a policy engine and this project should not grow a second one; what was missing was worked starting points. Each ships as Warn so it can be installed and read before it is enforced, and the first three select on a namespace label rather than the whole cluster — a policy that fires everywhere on day one is a policy somebody removes on day two.

They match on the annotations rather than on the injected container, because admission policies run before this webhook does: what they see is what the pod's author wrote.

What falls with what

THREAT_MODEL.md names the assets and the trust boundaries, and walks what an attacker reaches from each thing they might compromise — an application container, the agent, the webhook, the operator, the network path.

The one it is worth reading this page for: a cluster-scoped DynamicConfigClass holds one credential and admits many namespaces, and a tenant in an admitted namespace chooses the key. Scope that credential to what its tenants may read, and prefer namespaced Classes where the tenancy allows — a key-prefix policy on the Class would make it structural, and that is not built.

Transport credentials are wiped when they go

The Vault token, the AppRole secret-id, the password, the AWS secret key and the SSH key are held in a type that zeroes its own memory on drop and redacts its own Debug. What that buys is narrow, and narrow in a specific way: a core dump, a /proc/<pid>/mem read or a swapped page taken after the agent is finished with a credential does not contain it.

The resolved document is not covered, and cannot be. It is plaintext by necessity — it is about to be written to a file the application reads — so the protection that matters for it is the volume being tmpfs and the file mode being what the pod asked for, not memory hygiene.

The webhook holds nothing

  • Its ServiceAccount sets automountServiceAccountToken: false, and the deployment repeats it. The webhook reads the AdmissionReview it is handed and answers; it never calls the API server, so it carries no credential to steal.
  • It terminates TLS in-process with the certificate the chart issued; the private key never leaves its mount. Renewals are picked up from disk without a restart.
  • The optional NetworkPolicy (networkPolicy.enabled=true) writes both facts down for the CNI: ingress only on 8443, egress empty.

The webhook cannot select itself

The webhook configuration excludes kube-system, kube-node-lease and the release's own namespace by name — a mutating webhook that can select its own pods can deadlock its own rollout, and one that can mutate the control plane is a cluster risk with no matching reward. Add more with webhook.excludeNamespaces.

failurePolicy: the whole trade

Ignore (default): an unreachable webhook lets pods through un-injected. The failure is visible where it matters — the annotated pod's application waits for a file that never comes — and invisible where it does not: un-annotated pods, which are most pods, never notice.

Fail: no annotated pod can start un-injected, and no pod at all can start in selected namespaces while the webhook is down. Two replicas, a PodDisruptionBudget and topology spread are the chart's mitigations; they shrink the window, they do not close it.

Start on Ignore, alert on the webhook's availability, and flip to Fail when the alert has been quiet long enough to trust.

Every admission leaves a line

The webhook logs one structured line per decision that matters — namespace, pod name, source, and patched or refused — and no annotation values: endpoints and role names belong in the cluster, not in every log aggregator downstream. Pods that never asked are counted but not logged.

GET /metrics on the serving port exposes the counters (Observability is the full map) in Prometheus text format:

dynamic_config_admissions_total{outcome="skipped"} 1042
dynamic_config_admissions_total{outcome="patched"} 63
dynamic_config_admissions_total{outcome="refused"} 2

A rising refused is somebody fighting the contract; alert on it.

Typos cannot pass

An unknown dynamic-config.rs/* annotation fails the admission — the reference explains the rule. The enterprise version of the argument: a misspelled token-secret that is silently ignored produces a pod that runs, connects anonymously, and reads whatever the store's anonymous policy allows. Refusing at admission turns a quiet posture downgrade into a loud create-time error.

Templates are code, and scoped like data

A template renders only the resolved document — the same value the application reads. There is no file access, no environment access, no network in the template language; a hostile template can misrender the config file it owns and nothing else. Undefined keys are strict errors, so a template cannot silently swallow a value either. Keeping templates in ConfigMaps puts them through the same review as the code they effectively are.

Namespace gating

webhook.namespaceGating=true flips injection to Istio-style opt-in: only namespaces labeled dynamic-config.rs/injection: enabled are selected at all. Two things follow:

  • the blast radius of the webhook is exactly the namespaces that asked;
  • failurePolicy: Fail becomes a per-namespace promise — a platform team can fail closed for its opted-in tenants without coupling every pod CREATE in the cluster to this webhook.
kubectl label namespace team-a dynamic-config.rs/injection=enabled

A label, and it could not be an annotation: the gate lives in the webhook configuration's namespaceSelector, and Kubernetes selectors match labels only — annotations are invisible to them. The alternative (the webhook reading each pod's Namespace object to check an annotation) would hand the webhook API access it pointedly does not have; the zero-RBAC posture outranks the spelling preference. Either way the per-POD dynamic-config.rs/inject: "true" annotation is still required — the namespace gate is an outer guard, never an implicit opt-in.

Fleet-wide agent defaults

The injected container's resource defaults come from the chart (agent.defaults.*), not from a constant in a binary — platform teams set the fleet's floor once, and the per-pod annotations still override it. The same is true of the rendered file's permissions (agent.defaults.fileMode) and the watch interval (agent.defaults.watchSeconds); the same values file pins the agent image the webhook injects. Every fleet default is validated when the webhook STARTS — a mistyped octal stops the process at install, never at the first admission — which also covers kustomize installs, where no chart schema stands in front of the env vars.

The source gates are the installer's too

webhook.sourceAllow / webhook.sourceDeny decide which stores may be rendered from, per namespace — an allowlist that admits only what it names, and a subtractive deny that outranks it. Empty allow means every store, so an upgrade changes nothing until the installer says so. The check covers every render on the pod, named suffixes included, and a gate entry that is not a real store name fails webhook startup instead of silently gating nothing.

The agent-env gate is the installer's

agent-env lets a pod put environment on its injected agent, and agent environment steers SDKs (proxies, trust roots). Which names may pass, and in which namespaces, is declared in webhook.agentEnvAllow — owned by whoever installs the webhook, not by the pod author and not by a namespace annotation the tenant could edit for themselves. The default is empty: everything refused.

Supply chain

  • Images are distroless, run as nonroot, and the chart refuses tag: latest at render time; a digest value pins harder than a tag can.
  • The release workflow signs images with cosign and attaches SBOMs; ghcr.io and Docker Hub carry the same digests.
  • The agent binary embeds the engine and the store crates from crates.io — the same audited path every other binding uses; there is no k8s-only fork of anything.

What the injected agent never does

No hostPath, no privileged, no capabilities added, no writes outside the shared volume, no API server calls from the webhook, and no credentials in flags or annotations.

Two of those used to be said of the whole integration, and 0.3.0 made that untrue in two places. Both are off by default and both are named here rather than left for a reader to find:

  • The node agent mounts hostPath and runs as root. It is a CSI node plugin, so the kubelet's plugin socket and pod directories are where its work is, and the kubelet creates those directories owned by root. It adds no capabilities, escalates no privileges, and — since it makes no mounts — asks for HostToContainer propagation rather than the Bidirectional that would require a privileged container. nodeAgent.enabled is false. See the node agent.
  • tls-skip-verify exists. A pod can ask to reach its store without authenticating it, and the answer is no unless the installer set webhook.allowTlsSkipVerify, which defaults to false. It is the one annotation that trades away the guarantee the rest of this page is arranged around, which is why it is the installer's to grant and not the pod author's to take. server-name is the answer to the problem that usually leads people here — a certificate whose name does not match the address — and it gives up nothing.

Rendering

The agent does one resolution, through the same engine every binding uses, and writes the resolved document — so the file on disk is what an in-process consumer would have computed, not a second dialect.

The output format follows --out's extension: .json, .toml, .yaml, .ini, .properties.

The flat formats are legal here and refused by the engine's save — both on purpose. save's contract is a typed round trip, which a string-widening format cannot keep. A rendered file for a consumer is a different contract, and the agent owns it, stated:

  • Nested tables become dotted keys (properties) or sections (INI).
  • A string that would widen on the way back in — "1.10", "true" — is double-quoted in INI, so the round trip through the engine's own parser answers the same document. There is a test that holds exactly this.
  • Arrays are refused, by path. Neither format has them; inventing an encoding would be a dialect of one. Render to json/toml/yaml when the document has lists.

Writes are write-then-rename, so a watching application sees whole files — the same courtesy an atomic-save editor pays, and the reason the engine's own watcher tolerates a 25ms grace.

On a fetch failure the sidecar keeps the last rendered file and says so in its log — keep-last-good, the organisation's standing behaviour. An init run with nothing yet rendered fails instead, which fails the pod, which is what an init container is for.

Templates

Without one, the agent renders the resolved document verbatim — same keys, same shapes, only the format changes with the extension. That is the right default and it stays the default.

A template takes over when verbatim cannot serve: an application that wants DATABASE_URL=postgres://… assembled from three keys, a framework with its own nesting, a file with a header. The template owns the output bytes, which also frees the extension — .env and .conf become legal exactly there.

apiVersion: v1
kind: ConfigMap
metadata:
  name: billing-template
data:
  template: |
    DATABASE_URL=postgres://{{ db.user }}@{{ db.host }}:{{ db.port }}/billing
    BETA={{ flags.beta }}
    dynamic-config.rs/path: "/config/app.env"
    dynamic-config.rs/template-configmap: "billing-template"
    # or, for one-liners:
    dynamic-config.rs/template: "db={{ db.host }}:{{ db.port }}"

The syntax is minijinja's — Jinja2: {{ value }}, {% for %}, {% if %}, filters. The template's context is the resolved document, the same value every binding reads, so a template cannot see anything the application could not.

The semantics that matter in production:

  • Undefined is strict. {{ db.hots }} is a render error, not an empty string — a typo that silently renders nothing would ship a broken file with a clean exit code. At startup the error is fatal and lands in the pod's events; during a watch, the running pod keeps its last good file, like any fetch failure.
  • Booleans render as true/false, not Python's True — a template writes config files, and every format this agent speaks spells them lowercase.
  • The trailing newline survives. Env files want one; what the template author wrote is what lands.
  • The ConfigMap is re-read at every render, so editing the template takes effect on the next tick — no rollout. It is also why the template belongs in a ConfigMap: it is code, and it gets reviewed and versioned like code.
  • template and template-configmap together are refused at admission: one template, one place.

The filters this agent adds

minijinja's built-ins, plus six that a configuration template needs and minijinja either does not ship or ships behind a feature this build does not carry:

filterfor
b64encodea Kubernetes Secret's data is base64, so a template writing one has to encode
b64decodethe other direction; a value that is not base64, or not UTF-8 once decoded, is a render error rather than mojibake
jsontojson is behind a disabled feature, and a template that cannot emit JSON is missing the format half its consumers read
yamlthe same, for the other half
quotea password with a # in it ends a line in half the formats here
requiredstrict undefined already refuses a missing key; this refuses one that is present and empty, which is the shape a missing secret usually arrives in. {{ db.password | required("no password in the vault path") }} names the field instead of writing a blank one

What is not here, and will not be: a filter that reaches the network, the filesystem or the environment. The pipeline is fetch → resolve → validate → template, and a template that could fetch would make the same input render differently on two pods.

Checking the document before it is published

    dynamic-config.rs/schema-configmap: "billing-schema"   # or "billing-schema/other-key.json"

The resolved document is validated against a JSON Schema before anything is written. A document that fails is refused, the last good file keeps serving, and the failure is a log line and a counter — the same shape as any other render failure.

The bindings already validate: a Rust, Python or Node application gets a typed refusal from the engine. This is the door for everyone else — the Java service reading a .properties file, the daemon reading YAML — for whom a port: "abc" would otherwise be discovered at startup, one restart after the bad document was published.

The schema is re-read every render, like a template, so tightening it does not need a rollout.

Several files, one fetch

    dynamic-config.rs/path: "/config/app.yaml"
    dynamic-config.rs/also.db: "/config/db.env"
    dynamic-config.rs/also-section.db: "database"

More files cut from the same fetched document, published all or none: one fetch, one generation, and a failure in the third does not leave the first two on disk. Each file is still its own atomic rename — a reader can catch the gap between two of them, which is a rename apart rather than a fetch apart.

Several stores cannot share a generation. Two stores have no common instant, and no protocol either of them speaks can say "these two reads are the same version", so a second store is a named render with a generation of its own.

What the application is running

    dynamic-config.rs/meta: "true"

Writes a sibling file — /config/app.yaml gets /config/.app.yaml.meta — holding the SHA-256 of the rendered bytes, the store's own revision, and when the render landed. Same atomic rename, same mode.

It answers a question an application cannot otherwise ask about itself: which configuration am I running? Two pods holding the same file is a claim nobody can check from inside either of them; two pods printing the same digest is one anybody can. It describes the render and never contains it — no values reach it, ever.

The Full Stack, One Deployment

Vault to sidecar to template to file to a live process — every hop on one page, as the manifests you would actually apply. Each piece has its own chapter; production is all of them at once, and the order they come up in.

Vault (secret/myapp)                        the values that move
  → injected agent, kubernetes auth        no distributed secret
  → minijinja template (ConfigMap)         the store's shape → the app's
  → /config/rendered.yaml (tmpfs)          off the node's disk
  → the app's own watcher                  live without a restart

0. The one-time pieces

$ helm install dynamic-config oci://ghcr.io/dynamic-config-rs/charts/dynamic-config \
    --namespace dynamic-config --create-namespace

$ vault auth enable kubernetes
$ vault write auth/kubernetes/config kubernetes_host=https://kubernetes.default.svc
$ vault write auth/kubernetes/role/billing \
    bound_service_account_names=billing \
    bound_service_account_namespaces=shop \
    policies=billing-read ttl=1h
$ kubectl -n shop create configmap vault-ca --from-file=ca.crt=./internal-ca.pem

1. The template — the store's shape becomes the app's

apiVersion: v1
kind: ConfigMap
metadata:
  name: billing-template
  namespace: shop
data:
  template: |
    db:
      host: {{ db.host }}
      pool_size: {{ db.pool_size | default(8) }}
    features:
      cache: {{ features.cache | default(false) }}

Re-read every render, so editing the ConfigMap is itself a live change — no rollout. The Rendering chapter owns the semantics; the one rule worth restating is that a template failure leaves the previous rendered file in place, which is last-known-good at the file layer.

2. The workload — annotations are the whole integration

apiVersion: apps/v1
kind: Deployment
metadata:
  name: billing
  namespace: shop
spec:
  replicas: 3
  selector: { matchLabels: { app: billing } }
  template:
    metadata:
      labels: { app: billing }
      annotations:
        dynamic-config.rs/inject: "true"
        dynamic-config.rs/source: "vault"
        dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
        dynamic-config.rs/key: "secret/myapp"
        dynamic-config.rs/auth: "kubernetes"
        dynamic-config.rs/auth-role: "billing"
        dynamic-config.rs/ca-configmap: "vault-ca"
        dynamic-config.rs/template-configmap: "billing-template"
        dynamic-config.rs/path: "/config/rendered.yaml"
        dynamic-config.rs/native-sidecar: "true"
        dynamic-config.rs/watch-seconds: "30"
    spec:
      serviceAccountName: billing
      containers:
        - name: app
          image: myapp:1
          # /config arrives injected: tmpfs, shared with the agent.
          readinessProbe:
            httpGet: { path: /readyz, port: 8080 }

No secret in the manifest, no volume stanza to write, no sidecar to maintain: the webhook injects the agent, the agent logs in with the pod's own service account, and the token never exists as a Kubernetes Secret. native-sidecar: "true" makes the agent an init container with restartPolicy: Always, so a Job with these annotations still finishes.

3. The app — the last hop is an ordinary file

The agent's output is a file, so the app's side is the engine's ordinary file story — any of the three languages, here Rust:

#[dynamic_config]
#[derive(Deserialize)]
struct Db {
    host: String,
    pool_size: u32,
}

Db::builder("db").file("/config/rendered.yaml").init()?;
Db::builder("db")
    .file("/config/rendered.yaml")
    .watch(Duration::from_millis(500))?
    .detach();

The write is atomic (rename), so the watcher never reads half a render — the same guarantee the kubelet's ..data swap gives a mounted ConfigMap, held at both layers.

What moves without a rollout, and what does not

ChangeTakes effect
the secret in Vaultnext watch-seconds tick → render → app's watcher
the template ConfigMapnext render (kubelet sync + render tick)
an annotationrollout — injection happens at pod creation
the chart's own valueshelm upgrade, webhook restart, no app rollout

The annotation row is the one that surprises: annotations are read by the webhook when the pod is admitted, so changing them is a Deployment edit and rolls pods — which is also why it is safe, because every replica converges through the same admission path.

Watching it work

$ kubectl -n shop logs deploy/billing -c dynamic-config-agent --tail=5
$ kubectl -n shop exec deploy/billing -c app -- cat /config/rendered.yaml
$ vault kv put secret/myapp db='{"host": "db-2.internal", "pool_size": 16}'
# … within watch-seconds, the same two commands show the new values,
# and the app's /readyz generation has moved. No pod restarted.

The Four Deliveries: File, Env, Secret, Volume

One engine, four ways a document reaches a workload — because real software disagrees about how it wants to be configured. Grafana re-reads files; Airflow reads environment variables at boot and nothing else; a Strimzi-shaped operator watches Kubernetes Secrets. Picking the delivery is a one-line decision here; what follows is the map.

File (webhook + agent)Env (env-inject)Secret (operator target)Volume (CSI, node agent)
The consumerreads/watches a filereads environ at startreads/watches a k8s Secret, or envFromreads/watches a file
Freshnesslive — atomic rename, watcher cadencefrozen at container start (Kubernetes' rule); env-restart opts into a kubelet container-restart on change — seconds, no pod recreationlive object; watchers react, envFrom at next startlive, same renames from a shared watch
Touches etcd?never — tmpfs emptyDirnever — same tmpfs file, sourcedyes — a Secret lives in etcd, stated out loudnever
Restart to update?noyes (next pod start)no for Secret-watchers; yes for envFromno
Containers per renderoneonenone in the workloadnone — one per node, shared
Set up bypod annotationspod annotations (+ explicit command)DynamicConfigRender CRa csi: volume, no annotations at all
Real exampleGrafanaAirflowKafka clientthe node agent's page

The decision procedure, in order: can the consumer read a file? File delivery — it is the only fully live one and the only one that never touches etcd. Does something else own the consumer (an operator that only reads Secrets)? The Secret target. Env-only software? env-inject, with the start-freeze stated rather than papered over.

And the fourth is a scale answer rather than a shape answer. The volume delivers the same bytes as the file does, from one process per node instead of one beside every render — 10,000 pods at 2.5 renders each is 25,000 sidecar containers, and that number is the only reason to reach for it. It costs the isolation the other three have: one process holds the store credentials of every pod on its node. Choose it from a measurement, not a preference, and read its page before you do.

Against the Vault Agent Injector

The closest relative of the webhook+agent half — same architecture (mutating webhook, injected init/sidecar, shared memory volume, pod service-account auth, file permissions and run-as knobs), different center of gravity:

Vault Agent Injectordynamic-config-k8s
BackendsVaultnine stores — Vault among them, plus Consul, etcd, git, Redis, NATS, S3, Firestore, config-server
Pod authKubernetes auth (SA token)the same, wherever the store speaks it (Vault, Consul); IRSA/Workload Identity for S3/Firestore; secret-based stays first-class where a store never will
What is deliveredrendered secrets, Consul-Template languagethe resolved configuration document — precedence, validation, provenance — in json/toml/yaml/ini/properties or a minijinja template
File perms / ownershipannotationsannotations (file-mode, agent-run-as-user/group), root refused at admission
Env variablesyou rewrite the command by hand to source the fileenv-inject writes the wrap for you, refuses the impossible cases by name
k8s Secret objectsnothe operator's secret: target, when the consumer requires one
Lease renewalrenews renewable leases, re-fetches non-renewable onesthe same, since 0.3.0: dynamic: "true" reads a dynamic engine, renews a renewable lease at 65% of its TTL, never sends a renewal to one the store marked non-renewable, re-issues at 90% instead, and hands the lease back on SIGTERM
PKI certificatestracks the certificate's own lifetimethe same: whichever expires first binds — the lease, or the certificate's notAfter
Failure semanticslast-known-good, retrieslast-known-good with a startup policy, a deletion policy, a staleness ceiling wired to readiness, and drift detection on the rendered file
Knowing it arrivedthe file existsthe file exists and the application said it applied it — require-ack makes that readiness
Rolling a changeall pods at oncecanary-configmap: a deterministic cohort takes it first, widened by editing a ConfigMap with no restart
Delivery shapesfile (and env by hand)file, env, k8s Secret, and a CSI volume from a per-node agent
Scopesecrets deliveryconfiguration delivery that treats secrets as first-class fields

What it still does that this does not, and what this declines on purpose, are written down rather than left to a table's silence: seven items in VAULT-PARITY-GAPS.md and twelve in VAULT-PARITY-REFUSED.md, both at the root of this repository. The short version is that the remaining gaps are ergonomics — a render-spec ConfigMap, two exit-on-failure policies, log format — and the refusals are the proxy, the cache, arbitrary commands, and anything that would let a pod rewrite the container this webhook injects.

The injector's template idiom, translated

The Vault injector spells "render me a connection string" as a per-secret annotation pair; here the same result is one source and one template, because the whole pod has one resolved document:

# Vault Agent Injector:
#   vault.hashicorp.com/agent-inject-secret-db-creds: "secret/data/db-app"
#   vault.hashicorp.com/agent-inject-template-db-creds: |
#     {{- with secret "secret/data/db-app" -}}
#     postgres://{{ .Data.data.username }}:{{ .Data.data.password }}@postgres:5432/appdb
#     {{- end }}

# dynamic-config-k8s, the same string from the same KV secret:
dynamic-config.rs/source: "vault"
dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
dynamic-config.rs/key: "secret/db-app"
dynamic-config.rs/auth: "kubernetes"
dynamic-config.rs/auth-role: "db-app"
dynamic-config.rs/path: "/config/db.env"
dynamic-config.rs/template: |
  DATABASE_URL=postgres://{{ username }}:{{ password }}@postgres:5432/appdb

The template owns the bytes (minijinja, strict-undefined: a typo is an error, not an empty string), so any shape works — a URL, an .env, a whole config file. What does NOT translate is the database/creds/… path in the injector's example: that is Vault's dynamic secrets engine, credentials minted per-request with leases — the boundary the next paragraph prices. This agent reads KV; for minted-with-TTL credentials, run the Vault Agent beside it.

One more idiom, matched: the injector's several -secret-<name> pairs per pod are this webhook's named renders — source.db, key.db, path.db beside the default, one agent and one file per name, all in one shared directory. When the pod wants them MERGED into a single document instead, the config server composes sections and the pod reads one endpoint. What has no counterpart is agent-inject-command (a post-render hook): env-restart covers the restart case, and anything richer belongs to the app.

Honest edge the other way: for dynamic Vault secrets with leases (database credentials minted per-pod, TTL renewal mid-life), the Vault Agent is the purpose-built tool and this is not — this agent re-fetches documents; it does not manage leases.

Against External Secrets Operator

The closest relative of the operator half — same split between a namespaced store and a platform-owned cluster store, different product:

External Secrets Operatordynamic-config-k8s
Store definitionSecretStore / ClusterSecretStoreDynamicConfigClass / ClusterDynamicConfigClass — the same two scopes, allowlist included
Outputa Kubernetes Secret, alwaysa ConfigMap, a Secret, or a file no etcd ever sees
The etcd tradeevery delivered secret lives in etcdonly the Secret target does, and choosing it is explicit — the file path exists precisely to avoid it
Data modelkey-by-key secret mappingwhole configuration documents: precedence, validation, formats, provenance
Env deliveryenvFrom the Secretthe same via envEntries — or env-inject, which needs no Secret at all
Backendsvery many secret managersnine configuration stores
TemplatingSecret templatesminijinja over the resolved document

Honest edge the other way: as a secret-synchronisation fleet tool across dozens of managers (AWS/GCP/Azure SM, Doppler, 1Password…), ESO has breadth this project does not chase — the store list here grows by demand, not by roadmap.

Use cases, mapped

  • Airflow / env-only software → env-inject over a rendered dotenv (example); add env-restart: "true" and a changed document restarts just that container in seconds — otherwise changes wait for the next pod start, and that limit is stated instead of hidden.
  • Grafana / anything that re-reads files → the sidecar; live updates, zero etcd, tmpfs only (example).
  • Strimzi-shaped operators / JVM client.properties → the Secret target, file or envEntries shape (example).
  • Multi-tenant platforms → ClusterDynamicConfigClass with a namespaces allowlist; tenants never see a credential (example).
  • A chart's existingSecret / an operator's secretName: → the Secret target with shape: entries — leaf keys verbatim, so the names some other chart already chose are met exactly (example).
  • All of it at once → the four-component shop stack: three secrets injected three ways (a chart's existingSecret, a secretKeyRef env, a mounted-and-live file), the API on env-inject + env-restart, the worker on live files — credentials existing only in the platform namespace.
  • Vault dynamic database credentials with TTLs → the Vault Agent Injector, genuinely. Use both: it owns the lease, this owns the configuration.

The Nine Stores

One page per store the agent speaks, each with the full pod YAML for every authentication method the store takes — copy, adjust names, apply. Everything is runnable; the consul flow is what the e2e harness runs on every pull request.

storespeaksauth methodsits page
ConsulKV over HTTPanonymous, token, kubernetes (login), jwtConsul
VaultKV v2token, kubernetes, approle, jwt, userpass, ldap, certVault
Config Serverthis project's own serverbearer tokenConfig Server
FirestoreGoogle Cloudmetadata-server (Workload Identity), access-token, emulatorFirestore
Gitany git hostanonymous, token, ssh keyGit
RedisRESPin the url (requirepass, ACL users)Redis

Common to all six:

  • The store's address rides dynamic-config.rs/endpoint — or endpoint-secret when the address itself carries a password.
  • The document's key rides dynamic-config.rs/key; the per-store syntax (mount/path, application/profile, a file path) is on the store's page.
  • Secret material rides Secrets, never annotations; the geography page has the one diagram.
  • A private CA is the same one annotation everywhere: dynamic-config.rs/ca-configmap.

Every pairing on these pages also exists as a ready-to-apply manifest in the repository's examples/ directory — twenty-three manifests plus six real-software walkthroughs, each self-contained with its Secret placeholders.

etcd, NATS, S3 — the async three

Since 0.1.1 the agent drives both of the engine's source traits: the blocking six run under a blocking task, and etcd, NATS and S3 — whose clients are async — are driven directly by the agent's own runtime. The 0.1 refusal-by-name retired with this.

The config server indirection remains the answer to a different question: a fleet of pods that should not each hold store credentials — the server holds them once.

A tenth store — GCP Secret Manager, Azure App Configuration — is a compile-time addition with a well-worn path: Adding a Store walks it end to end, worked example included, plus the two no-code compositions that cover the meantime.

Consul

A KV path over HTTP. Four ways in, ordered from development to production.

The key is a KV path with an extension — myapp/config.json — and the extension names the stored format; the rendered format is the path annotation's extension, and the two need not agree.

Anonymous

Correct for a Consul with ACLs disabled, and for a default policy that allows reads — both ordinary in development. No auth annotation at all:

apiVersion: v1
kind: Pod
metadata:
  name: billing
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "consul"
    dynamic-config.rs/endpoint: "http://consul.infra.svc:8500"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/path: "/config/rendered.toml"
spec:
  containers:
    - name: app
      image: myapp:1

This is the exact flow the e2e harness runs on every pull request.

An ACL token

The CONSUL_HTTP_TOKEN you already have, moved into a Secret:

kubectl create secret generic consul-token --from-literal=token=b3a7…
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "consul"
    dynamic-config.rs/endpoint: "http://consul.infra.svc:8500"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/path: "/config/rendered.toml"
    dynamic-config.rs/token-secret: "consul-token/token"

The webhook wires the Secret into the agent as DYNAMIC_CONFIG_AGENT_TOKEN; nothing token-shaped appears in the pod spec. A static token is also the one method that cannot recover on its own — when it expires or is revoked, the fetch fails until the Secret is updated. The two methods below fix that.

Kubernetes: login with the pod's identity

Consul's auth methods issue a token in exchange for a bearer the method trusts — for the kubernetes type, the pod's own service-account JWT. Nothing is distributed; the token is minted per login.

Consul side, once (the Consul docs on auth methods carry the full story):

consul acl auth-method create -type kubernetes -name k8s-pods \
  -kubernetes-host https://kubernetes.default.svc \
  -kubernetes-ca-cert @/path/to/cluster-ca.crt \
  -kubernetes-service-account-jwt "$REVIEWER_JWT"

consul acl binding-rule create -method k8s-pods \
  -bind-type role -bind-name 'config-reader' \
  -selector 'serviceaccount.namespace==default'

Pod side, two annotations — auth-mount carries the auth method's name, because that is the coordinate Consul's login endpoint wants:

metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "consul"
    dynamic-config.rs/endpoint: "http://consul.infra.svc:8500"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/path: "/config/rendered.toml"
    dynamic-config.rs/auth: "kubernetes"
    dynamic-config.rs/auth-mount: "k8s-pods"

The JWT is read from disk at every login rather than once: the kubelet rotates projected service-account tokens, and a copy taken at startup expires with the pod still running. If a projected volume moved the token off the conventional path, say where:

    dynamic-config.rs/auth-token-path: "/var/run/secrets/tokens/consul"

JWT: any bearer the method trusts

The same login endpoint, with the bearer supplied instead of read from the service-account mount — an OIDC id token, a JWT signed by something Consul trusts. The bearer is a secret, so it rides a Secret:

    dynamic-config.rs/auth: "jwt"
    dynamic-config.rs/auth-mount: "oidc-ci"
    dynamic-config.rs/token-secret: "ci-idtoken/jwt"

TLS

An internal Consul serving from a private PKI needs its CA trusted; one ConfigMap, one annotation:

    dynamic-config.rs/endpoint: "https://consul.infra.svc:8501"
    dynamic-config.rs/ca-configmap: "consul-ca"

A cluster fronting Consul with mTLS adds tls-secret.

When it fails

symptomlook atusual cause
agent log: 403consul acl token read -self with the same tokenthe token lacks key_prefix read on the path
agent log: login refusedconsul acl auth-method read -name k8s-podsbinding rule selector does not match the pod's namespace/SA
agent starts, file never updatesconsul kv get myapp/config.jsonthe key moved, or the watch interval is long — check watch-seconds

The agent keeps the last good render on any fetch failure — the rendering page spells out that guarantee.

Beyond the agent's flags, the consul crate can also scope reads to a datacenter and tune blocking-query waits; those knobs live on the store crate itself for applications embedding the engine directly.

Vault

KV v2, and the widest auth surface of the six — all seven of the vault crate's methods work through annotations. Ordered from development to production; if the pod runs on Kubernetes (it does — you are reading the k8s book), start with kubernetes.

The key is <mount>/<path>, the way vault CLI users write it: secret/myapp is the myapp path on the secret KV mount. What the agent renders is the secret's data — the fields under data.data in vault's own JSON.

A token

The method every tutorial starts with, and the only one that cannot recover on its own: there are no credentials behind it to log in again with. A renewable token is still renewed; a revoked one is the 3 a.m. page.

vault policy write myapp-read - <<'HCL'
path "secret/data/myapp" { capabilities = ["read"] }
HCL
vault token create -policy=myapp-read -ttl=768h -format=json \
  | jq -r .auth.client_token \
  | xargs -I{} kubectl create secret generic vault-token --from-literal=token={}
apiVersion: v1
kind: Pod
metadata:
  name: billing
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "vault"
    dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
    dynamic-config.rs/key: "secret/myapp"
    dynamic-config.rs/path: "/config/rendered.yaml"
    dynamic-config.rs/token-secret: "vault-token/token"
    dynamic-config.rs/ca-configmap: "vault-ca"
spec:
  containers:
    - name: app
      image: myapp:1

Kubernetes: the pod's own identity

No secret distributed anywhere: the agent presents the pod's service-account JWT to vault's kubernetes auth method and gets a token scoped to a role. This is the pod YAML the webhook's second golden file locks:

metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "vault"
    dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
    dynamic-config.rs/key: "secret/myapp"
    dynamic-config.rs/path: "/config/rendered.yaml"
    dynamic-config.rs/auth: "kubernetes"
    dynamic-config.rs/auth-role: "myapp"
    dynamic-config.rs/ca-configmap: "vault-ca"

Vault side, once:

vault auth enable kubernetes

vault write auth/kubernetes/config \
  kubernetes_host=https://kubernetes.default.svc

vault write auth/kubernetes/role/myapp \
  bound_service_account_names=billing \
  bound_service_account_namespaces=default \
  policies=myapp-read ttl=1h

Three details that bite:

  • The JWT is re-read from disk at every login, not cached: the kubelet rotates projected tokens, and a copy taken at startup expires with the pod still running. Nothing to configure — stated here so a security review can check the box.

  • A projected token with a custom audience lives off the conventional path; point at it:

        dynamic-config.rs/auth-token-path: "/var/run/secrets/tokens/vault"
    
  • A second kubernetes mount (multi-cluster Vaults mount one per cluster) is auth-mount:

        dynamic-config.rs/auth-mount: "kubernetes-prod-eu"
    

AppRole

The usual choice for a service outside Kubernetes, and for shops that want an identity Vault owns rather than the cluster. Two halves: the role id is public and rides an annotation; the secret id is a secret and rides a Secret.

vault auth enable approle
vault write auth/approle/role/myapp policies=myapp-read \
  secret_id_ttl=90d token_ttl=1h

vault read -field=role_id auth/approle/role/myapp/role-id
vault write -f -field=secret_id auth/approle/role/myapp/secret-id \
  | xargs -I{} kubectl create secret generic vault-approle --from-literal=secret-id={}
    dynamic-config.rs/auth: "approle"
    dynamic-config.rs/auth-role: "<the role id>"
    dynamic-config.rs/password-secret: "vault-approle/secret-id"

The secret id travels as DYNAMIC_CONFIG_AGENT_PASSWORD — the second secret slot, which has no flag equivalent on purpose.

JWT / OIDC

Any JWT a jwt mount trusts — a CI job's id token, a workload identity from another platform. The JWT itself is the credential, so it rides the token Secret:

    dynamic-config.rs/auth: "jwt"
    dynamic-config.rs/token-secret: "workload-jwt/jwt"
    dynamic-config.rs/auth-role: "myapp"      # when the mount has no default
    dynamic-config.rs/auth-mount: "jwt-ci"    # when not mounted at "jwt"

Userpass and LDAP

Same shape, different directory. The username is a name and rides an annotation; the password rides a Secret:

kubectl create secret generic vault-ldap --from-literal=password=…
    dynamic-config.rs/auth: "ldap"            # or "userpass"
    dynamic-config.rs/auth-username: "svc-billing"
    dynamic-config.rs/password-secret: "vault-ldap/password"

Cert: a TLS client certificate

The certificate IS the credential: vault's cert method authenticates the TLS handshake itself. One kubernetes.io/tls Secret carries the pair:

kubectl create secret tls vault-client --cert=client.pem --key=client-key.pem
    dynamic-config.rs/auth: "cert"
    dynamic-config.rs/tls-secret: "vault-client"
    dynamic-config.rs/auth-role: "myapp"      # when the mount does not pick by subject
    dynamic-config.rs/ca-configmap: "vault-ca"

Vault side:

vault auth enable cert
vault write auth/cert/certs/myapp certificate=@client-ca.pem \
  allowed_common_names=billing.default policies=myapp-read

Vault Enterprise namespaces

One annotation, passed through to the X-Vault-Namespace header:

    dynamic-config.rs/namespace: "team-payments"

How the token lifecycle behaves

Worth knowing before the first incident, whatever the method:

  • Close to expiry, the token is renewed — or, if it cannot be, replaced by a fresh login. This is the path that normally fires.
  • After a request, a 403 is treated as the token stopped working and triggers exactly one fresh login and retry. Clocks skew, Vault revokes, a lease is shorter than it said.
  • If a fresh token also gets 403, the problem is the policy rather than the lease, and the error says so instead of hanging in a retry loop.

Dynamic secrets

A path under database/, pki/ or aws/ does not hold a secret; it mints one, with a lease. One annotation says to read it that way:

    dynamic-config.rs/source: "vault"
    dynamic-config.rs/dynamic: "true"
    dynamic-config.rs/key: "database/creds/billing"
    dynamic-config.rs/path: "/config/db.env"
    dynamic-config.rs/auth: "kubernetes"
    dynamic-config.rs/auth-role: "billing"

The pod now holds a database credential nobody else has, for as long as it needs it. What changes:

  • The path is read as it is. KV v2's data/ nesting is not inserted and data.data is not unwrapped, because a dynamic engine has neither.
  • The lease is kept. lease_id, lease_duration and renewable come back with the document.
  • A lease that says it cannot be renewed is never sent a renewal. Vault answers renewable: false for every pki/issue, and for a database credential that has reached its role's maximum. Such a lease is re-issued at 90% of its life instead. Asking it to renew first would be a request that can only be refused — a round trip per cycle per pod, and a lease_renewal_failures_total that climbs steadily on a fleet where nothing is wrong.
  • A renewable lease is renewed at 65% of its TTL, spread, and what Vault grants is what the next renewal is scheduled from — a role's max_ttl is a ceiling the pod cannot see, so asking for an hour and being given ten minutes has to be believed rather than assumed.
  • The two fractions are not the same number by accident. A renewal is cheap and reversible: it fails, and a third of the lease is still there to get a new credential in. A re-issue is the new credential, so doing it early only shortens the one in use and wakes every application watching the file for it.
  • A renewal does not re-render. It extends the same credential; the file is already correct. A re-issue mints new credentials and does re-render, and that is what happens both when a renewal stops working and when the lease was never renewable.
  • A backoff never sleeps past the lease. This is the one place the usual ceiling is wrong: waiting five minutes to retry a credential that expires in twenty seconds is a pod that comes back to an expired secret.
  • SIGTERM revokes. Best-effort, with a short deadline — revoke-on-shutdown: "false" opts out, for a lease something else is still using. Vault expires the lease on its own eventually; revoking on the way out is what turns eventually into now.

Certificates are on their own clock

pki/issue is the one dynamic engine where the lease is not the whole truth. A PKI role can hand back a lease longer than the certificate it issued — the lease is Vault's accounting record, notAfter is what a TLS peer enforces — so the agent takes whichever runs out first.

Vault reports the certificate's notAfter as data.expiration, seconds since the epoch, which is why this needs no X.509 parsing: the number is already in the response.

One guard comes with it. That timestamp is Vault's wall clock compared against the pod's, and the two are not the same clock. A certificate that looks already expired from inside the pod is far more likely to be skew than an expired certificate — Vault would not have issued one — so the lease's own number wins there, rather than the agent re-issuing in a tight loop against a server that thinks everything is fine.

The watch capability drops to Interval: a dynamic engine has no version to poll, and asking whether it changed is not separable from asking for a new credential.

One path per render. Every read mints its own lease, so merging several paths into one document would leave every lease but one unrenewed and unrevoked. A second dynamic path is a second named render.

The four lease_* series in the agent's metrics are what to watch: lease_ttl_seconds is what Vault last granted, and a rising lease_renewal_failures_total is the signal that a re-fetch is coming.

When it fails

symptomlook atusual cause
403 on every readvault token capabilities <token> secret/data/myappthe policy grants secret/myapp, not secret/data/myapp — KV v2 inserts data/
login refused (kubernetes)vault write auth/kubernetes/login role=myapp jwt=@/tmp/jwt by handrole's bound SA/namespace does not match the pod's
x509 errorsopenssl s_client -connect vault:8200the CA in ca-configmap is not the one vault serves
works, then 403 after daysvault audit logorphan token hit its max TTL; move to kubernetes/approle auth

Config Server

This project's own server — the aggregation answer. The server side speaks all nine store crates (including the async three the agent cannot drive yet) and merges files, profiles and remote stores into one document per <application>/<profile>; the agent then speaks the server. The remote book owns the server's full story; what follows is the k8s-side wiring.

Two reasons to put it between the pods and the stores:

  • Credential concentration. A fleet of pods each holding a vault token is a fleet of tokens to rotate. The server holds the store credentials once; the pods hold short client tokens that grant exactly one application's sections.
  • Nine stores behind one address. The server speaks every store, so a pod's annotations name one endpoint whatever moves behind it.

Kubernetes auth: the pod's own identity, end to end

Since 0.1.1 the server indirection is zero-secret too: auth: "kubernetes" makes the agent present the pod's projected service-account token as the bearer (re-read every fetch — it rotates), and the server's [kubernetes] TokenReview grants map namespace:serviceaccount to applications:

    dynamic-config.rs/source: "config-server"
    dynamic-config.rs/endpoint: "https://config.infra.svc:8443"
    dynamic-config.rs/key: "shop/prod"
    dynamic-config.rs/auth: "kubernetes"

No client token minted, distributed, mounted or rotated — the remote book's server chapter carries the server-side TOML.

The wiring

The key is <application>/<profile>:

apiVersion: v1
kind: Pod
metadata:
  name: billing
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "config-server"
    dynamic-config.rs/endpoint: "http://config-server.infra.svc:8888"
    dynamic-config.rs/key: "billing/prod"
    dynamic-config.rs/path: "/config/rendered.json"
spec:
  containers:
    - name: app
      image: myapp:1

The bearer token

The server's [[server.clients]] blocks name each client and the applications it may read. The token is a secret; it rides a Secret:

# server.toml, server side
[[server.clients]]
name = "billing-pods"
token = "…at least 32 bytes…"
applications = ["billing"]
kubectl create secret generic config-server-token --from-literal=token=…
    dynamic-config.rs/token-secret: "config-server-token/token"

There is no other method on this store — one bearer, scoped server-side, is the whole model. Asking for auth: anything is refused by the agent with a sentence saying exactly that.

TLS

The server behind TLS from an internal PKI is the same one annotation as everywhere:

    dynamic-config.rs/endpoint: "https://config-server.infra.svc:8443"
    dynamic-config.rs/ca-configmap: "internal-ca"

A server requiring client certificates takes tls-secret alongside.

When it fails

symptomlook atusual cause
401server logthe token is not in any [[server.clients]] block
403server log, applications = […]the client's list does not include this application
404curl $SERVER/billing/prod with the tokenno [[server.sections]] matches the pair
stale valuesthe server's own watch configthe server polls its stores on its own cadence — two intervals stack

Firestore

A Google Cloud document read as configuration. The endpoint is not a url: it is <project> or <project>/<database>, and the key is collection/document (nest deeper as environments/prod/config/db).

On GKE the right method is the first one, and it involves no secret at all.

Metadata-server (Workload Identity)

The workload's own identity, from the metadata server — reachable from GKE, Cloud Run, GCE, and nowhere else, which is the security property that makes it the default. The agent asks for a token, gets a short-lived one, and renews it as it approaches expiry.

auth can be omitted entirely — metadata-server is what the agent does for firestore when nothing else is asked:

apiVersion: v1
kind: Pod
metadata:
  name: billing
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "firestore"
    dynamic-config.rs/endpoint: "acme-prod"
    dynamic-config.rs/key: "config/billing"
    dynamic-config.rs/path: "/config/rendered.json"
spec:
  serviceAccountName: billing
  containers:
    - name: app
      image: myapp:1

GKE side, the Workload Identity pairing (the GKE docs own the full ceremony):

gcloud iam service-accounts create billing-reader
gcloud projects add-iam-policy-binding acme-prod \
  --member serviceAccount:billing-reader@acme-prod.iam.gserviceaccount.com \
  --role roles/datastore.viewer

gcloud iam service-accounts add-iam-policy-binding \
  billing-reader@acme-prod.iam.gserviceaccount.com \
  --role roles/iam.workloadIdentityUser \
  --member "serviceAccount:acme-prod.svc.id.goog[default/billing]"

kubectl annotate serviceaccount billing \
  iam.gke.io/gcp-service-account=billing-reader@acme-prod.iam.gserviceaccount.com

Access token

A token somebody already obtained — gcloud auth print-access-token produces one. It expires within the hour and the agent cannot renew it, so this is a debugging method, not a deployment method; the honest use is a one-shot init container in a test cluster:

    dynamic-config.rs/mode: "init"
    dynamic-config.rs/auth: "access-token"
    dynamic-config.rs/token-secret: "gcp-token/token"

Emulator

The Firestore emulator wants no credential and a different endpoint; api-url points the API somewhere other than Google's:

    dynamic-config.rs/auth: "emulator"
    dynamic-config.rs/api-url: "http://firestore-emulator.test.svc:8080"
    dynamic-config.rs/endpoint: "demo-project"

The named database

The second database in a project is the endpoint's second segment:

    dynamic-config.rs/endpoint: "acme-prod/eu-config"

When it fails

symptomlook atusual cause
403 PERMISSION_DENIEDgcloud projects get-iam-policy acme-prodthe GSA lacks roles/datastore.viewer, or the WI binding names the wrong namespace/KSA pair
metadata server unreachablepod events, GKE node poolWorkload Identity not enabled on the pool
404gcloud firestore databases listthe document path or the named database is wrong
works locally, fails in-cluster—local gcloud credentials are not the pod's; the pod has only the metadata server

Git

Configuration that lives where its reviews live. Git has nothing to push, so the agent asks — but it asks the cheap question: the ref's advertisement is one handshake per tick, and a transfer happens only when the ref actually moved, so a 15-second watch does not hammer the host.

The endpoint is anything git understands (https://…, ssh://…, git@host:org/repo.git); the key is the file's path inside the repository; the ref defaults to branch main.

Anonymous

A public repository over HTTPS — no auth annotation:

apiVersion: v1
kind: Pod
metadata:
  name: billing
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "git"
    dynamic-config.rs/endpoint: "https://github.com/acme/config.git"
    dynamic-config.rs/key: "billing/prod.yaml"
    dynamic-config.rs/path: "/config/rendered.yaml"
    dynamic-config.rs/ref: "main"
spec:
  containers:
    - name: app
      image: myapp:1

A token over HTTPS

How every host takes a token: HTTP basic auth with the token in the password half. GitHub PATs and App installation tokens, GitLab deploy and project tokens, Azure DevOps PATs — all the same shape. The username half is filler that these hosts ignore; the agent sends x-access-token, the value GitHub documents:

kubectl create secret generic config-repo-token --from-literal=token=ghp_…
    dynamic-config.rs/auth: "token"
    dynamic-config.rs/token-secret: "config-repo-token/token"

For the rare host that does read the username, name it:

    dynamic-config.rs/auth-username: "deploy"

A GitLab deploy token is the least-privilege pick on that platform: scope read_repository, one repository, its own expiry.

An SSH deploy key

One kubernetes.io/ssh-auth Secret; its conventional key name is ssh-privatekey, and the webhook mounts it 0400 because ssh refuses group-readable keys:

ssh-keygen -t ed25519 -f deploy_key -N ""
# register deploy_key.pub as a read-only deploy key on the host
kubectl create secret generic config-deploy-key \
  --type=kubernetes.io/ssh-auth --from-file=ssh-privatekey=deploy_key
    dynamic-config.rs/endpoint: "git@github.com:acme/config.git"
    dynamic-config.rs/ssh-secret: "config-deploy-key"

auth: ssh-key is implied by ssh-secret when no auth is named. The key is offered with IdentitiesOnly=yes, so an agent holding other keys cannot exhaust the server's auth tries before the right one.

And the permission caveat, out loud: the kubelet writes secret files as root, 0400 means owner-read only, and the agent runs nonroot — so the mounted key is readable only when the pod sets a securityContext.fsGroup (the kubelet then group-owns the files) and the custom image's ssh client tolerates a group-readable key, or when the pod runs the agent's uid. This combination is honest-but-untested: the e2e suite covers HTTPS and token auth; ssh in-pod is documented, not gated.

The stock image caveat, out loud: git-over-SSH is carried by the ssh program, exactly as git itself does it — and the distroless agent image does not contain one. HTTPS works from the stock image; SSH needs an image with an ssh client:

FROM ghcr.io/dynamic-config-rs/dynamic-config-agent:0.1.0 AS agent
FROM alpine:3.20
RUN apk add --no-cache openssh-client ca-certificates \
 && printf 'github.com ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIOMqqnkVzrm0SdG6UOoqKLsabgH5C9okWi0dh2l9GKJl\n' \
    >> /etc/ssh/ssh_known_hosts
COPY --from=agent /dynamic-config-agent /dynamic-config-agent
ENTRYPOINT ["/dynamic-config-agent"]

…and the chart's agent.image value points at it. The known_hosts line is the second half ssh insists on; pin your own host's key, not a copy of this one.

Branch, tag, commit

ref takes four spellings:

    dynamic-config.rs/ref: "main"           # a branch, plainly
    dynamic-config.rs/ref: "branch:release" # the same, spelled out
    dynamic-config.rs/ref: "tag:v1.4"       # a tag
    dynamic-config.rs/ref: "commit:8f3a…"   # one exact tree, forever

A tag or commit still ticks the watch loop, and still never transfers — useful with mode: init for a pinned, reproducible render.

A self-hosted host with a private CA

The same annotation as every other store:

    dynamic-config.rs/endpoint: "https://git.internal.acme/config.git"
    dynamic-config.rs/ca-configmap: "internal-ca"

When it fails

symptomlook atusual cause
auth failed over HTTPStry the token in a git ls-remote by handtoken expired, or lacks read scope on the repo
Host key verification failedthe image's /etc/ssh/ssh_known_hoststhe custom image pinned no host key for this host
ssh: command not found—the stock distroless image; see the caveat above
file not foundgit ls-tree <ref> -- <path>the path is spelled from the repository root, and the ref matters

Redis

A key read as a document, watched over keyspace notifications — a change arrives as it happens rather than up to an interval later. The key carries its format in its extension (myapp/config.json); a key without one needs the document format to be guessable, so give it one.

Redis is the one store whose credentials travel in the url — requirepass and ACL users have no other place. That shapes the whole page: the moment the url grows a password, it stops being an annotation and becomes a Secret.

An open Redis

Development, a sidecar cache, a cluster-internal instance behind a NetworkPolicy:

apiVersion: v1
kind: Pod
metadata:
  name: billing
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "redis"
    dynamic-config.rs/endpoint: "redis://redis.infra.svc:6379/0"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/path: "/config/rendered.toml"
spec:
  containers:
    - name: app
      image: myapp:1

The trailing /0 is the database index; omit it for 0.

requirepass

The password goes into the url, and the url goes into a Secret — endpoint-secret replaces endpoint entirely, and the agent reads the address from DYNAMIC_CONFIG_AGENT_ENDPOINT:

kubectl create secret generic redis-url \
  --from-literal=url='redis://:s3cr3t@redis.infra.svc:6379/0'
metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "redis"
    dynamic-config.rs/endpoint-secret: "redis-url/url"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/path: "/config/rendered.toml"

Setting both endpoint and endpoint-secret fails the admission — one address, one place. Error messages redact the password even when the url cannot be parsed, because a parse error is the error most likely to be pasted somewhere.

An ACL user

Redis 6+ ACLs put a username before the password. Server side:

ACL SETUSER config-reader on >s3cr3t ~myapp/* +get resetchannels

The url names the user, read-only on exactly the config prefix:

kubectl create secret generic redis-url \
  --from-literal=url='redis://config-reader:s3cr3t@redis.infra.svc:6379/0'

TLS

rediss:// (two esses) plus the CA:

    dynamic-config.rs/endpoint: "rediss://redis.infra.svc:6380/0"
    dynamic-config.rs/ca-configmap: "redis-ca"

A redis:// url with TLS material is refused — a deployment that believes it is encrypted and is not — and there is no way to turn verification off. A client certificate is tls-secret, as everywhere.

With a password too, the whole rediss://user:pass@… url rides endpoint-secret and the CA annotation stays as it is.

When it fails

symptomlook atusual cause
NOAUTH / WRONGPASSredis-cli -u <the url> get myapp/config.jsonthe url in the Secret lost its password half, or the ACL user is off
NOPERMACL GETUSER config-readerthe key pattern does not cover the config key
refused: url is not rediss—TLS material with a redis:// url; add the second s
empty renderredis-cli … type myapp/config.jsonthe key holds a hash, not a string document

etcd

A key (or several) from etcd v3. etcd is the honest hard case for authentication: it has no Kubernetes auth method to log into — nothing that takes a projected service-account token the way Vault's kubernetes mount does. So this store lands with the two methods etcd itself speaks, both first-class: TLS client certificates, and username/password. Identity-first is the org's policy where a store can authenticate the pod; where it cannot, the secret-based method is supported without apology — a contract that punishes such stores punishes their users.

There is no auth annotation for etcd: the flags present ARE the method.

TLS client certificates

The method etcd operators already deploy — the certificate is the credential, delivered as a Secret the same way every store's TLS material is (the geography page):

apiVersion: v1
kind: Pod
metadata:
  name: billing
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "etcd"
    dynamic-config.rs/endpoint: "https://etcd.infra.svc:2379"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/path: "/config/rendered.toml"
    dynamic-config.rs/tls-secret: "etcd-client-tls"
    dynamic-config.rs/ca-configmap: "etcd-ca"
spec:
  containers:
    - name: app
      image: myapp:1
kubectl create secret tls etcd-client-tls \
  --cert=./billing.crt --key=./billing.key
kubectl create configmap etcd-ca --from-file=ca.crt=./etcd-ca.pem

Username and password

etcd's other method. The user rides an annotation; the password rides a Secret, never an annotation — kubectl describe pod prints annotations to anyone with pod read access:

    dynamic-config.rs/source: "etcd"
    dynamic-config.rs/endpoint: "https://etcd.infra.svc:2379"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/auth-username: "myapp"
    dynamic-config.rs/password-secret: "etcd-password/password"
    dynamic-config.rs/ca-configmap: "etcd-ca"

Both together is also legal, and is etcd's own combination: the certificate authenticates the channel, the user authenticates the principal.

Several keys

key takes a comma-free single key; several documents merging into one is the config server's job or a template's. The stored format is the key's extension, exactly as with Consul.

Ready-to-apply manifests: examples/etcd-tls.yaml, examples/etcd-password.yaml.

NATS

A JetStream key-value bucket. The key annotation is <bucket>/<key> — the bucket must already exist; a configuration reader that provisions storage would hide a misconfigured deployment behind an empty one.

A credentials file

The way a NATS account authenticates — a .creds file, delivered as a Secret and named by path:

apiVersion: v1
kind: Pod
metadata:
  name: billing
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "nats"
    dynamic-config.rs/endpoint: "nats://nats.infra.svc:4222"
    dynamic-config.rs/key: "config/db.json"
    dynamic-config.rs/path: "/config/rendered.toml"
    dynamic-config.rs/auth-token-path: "/etc/nats/nats.creds"
spec:
  containers:
    - name: app
      image: myapp:1
      volumeMounts:
        - { name: nats-creds, mountPath: /etc/nats, readOnly: true }
  volumes:
    - name: nats-creds
      secret: { secretName: nats-creds }

A token

token-secret, the same one annotation every token-shaped credential rides:

    dynamic-config.rs/source: "nats"
    dynamic-config.rs/endpoint: "nats://nats.infra.svc:4222"
    dynamic-config.rs/key: "config/db.json"
    dynamic-config.rs/token-secret: "nats-token/token"

Anonymous — no auth annotation at all — is correct for a NATS without auth, ordinary in development.

Ready-to-apply manifest: examples/nats-creds.yaml.

S3

An object from a bucket — AWS's S3, and everything that speaks its API: MinIO, Ceph, R2, B2. endpoint is the bucket; key is the object key; the stored format is the key's extension.

On EKS: IRSA, and nothing else

The default is the ambient AWS credential chain, which on EKS is IAM Roles for Service Accounts — the workload's own identity, no secret distributed, the token renewed by the platform. There is nothing to configure on this side beyond the service account:

apiVersion: v1
kind: Pod
metadata:
  name: billing
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/source: "s3"
    dynamic-config.rs/endpoint: "myapp-config"
    dynamic-config.rs/key: "prod/db.json"
    dynamic-config.rs/path: "/config/rendered.toml"
spec:
  serviceAccountName: billing   # annotated with the IAM role, IRSA's side
  containers:
    - name: app
      image: myapp:1

The injected agent inherits the pod's identity the same way the app does — that is the entire point of workload identity.

MinIO, Ceph, R2

api-url overrides the endpoint; a region is not required alongside it (the agent supplies a placeholder the server ignores — set AWS_REGION on the pod if yours cares); path-style addressing is always on, because virtual-hosted buckets need DNS entries only AWS has:

    dynamic-config.rs/source: "s3"
    dynamic-config.rs/endpoint: "myapp-config"
    dynamic-config.rs/key: "prod/db.json"
    dynamic-config.rs/api-url: "http://minio.infra.svc:9000"

Static credentials for a non-AWS store ride one annotation: dynamic-config.rs/aws-secret: "<secret>" — the Secret's AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY keys become exactly those variables on the injected agent, and nothing lands in the pod spec.

Ready-to-apply manifest: examples/s3-irsa.yaml.

Adding a Store

Can the agent learn a tenth store — Google Cloud Secret Manager, say? Yes, and what follows is the map. What it is NOT is a plugin system: stores are compiled in on purpose (a config agent that loads plugins is its own supply-chain problem), and the organisation's standing rule is stores land by demand, not by roadmap — the ask starts as an issue, and the walkthrough below is what the work looks like once agreed.

What a store is

One Rust type implementing one of the engine's two traits:

impl AsyncRemoteSource for GcpSecretManager {
    fn fetch(&self) -> BoxFuture<'_, Result<Fetched, Error>>;  // text + format
    fn describe(&self) -> String;                              // REDACTED — errors quote it
}

fetch answers the document's text and format; describe names the store in errors and logs and must never carry a credential. Blocking clients implement RemoteSource instead; the agent drives both — the blocking six run under a blocking task, the async three (etcd, NATS, S3) on the runtime directly. A new store picks whichever its client dialect is.

The walkthrough, with GCP Secret Manager as the worked example

The 0.1.1 change that landed etcd, NATS and S3 is the living template — git log --grep "async three" in the two repositories shows every touchpoint below with real diffs.

1. The store crate (dynamic-config-remote repository, dynamic-config-gcpsm/). The Firestore crate is the closest twin: Google auth, same project-scoped shape. The heart is small:

pub struct GcpSecretManager { project: String, secret: String, auth: Auth, /* … */ }

impl AsyncRemoteSource for GcpSecretManager {
    fn fetch(&self) -> BoxFuture<'_, Result<Fetched, Error>> {
        Box::pin(async move {
            // GET https://secretmanager.googleapis.com/v1/projects/{p}
            //     /secrets/{s}/versions/latest:access
            // Authorization: Bearer <metadata-server token>
            let text = self.access_latest().await?;

            Ok(Fetched { text, format: self.format })
        })
    }

    fn describe(&self) -> String {
        format!("gcp-sm {}/{}", self.project, self.secret) // no token, ever
    }
}

Identity-first applies: the default auth is the metadata server (Workload Identity — the pod's own identity, no distributed secret), exactly as the Firestore crate already does; an access-token arm stays a supported peer. House rules the crate must keep: describe redacted, TLS through the shared TlsConfig vocabulary, errors mapped to the engine's kinds (auth for 401/403, remote for the rest), #![forbid(unsafe_code)].

2. The agent (dynamic-config-agent/src/):

  • sources.rs — one async constructor arm beside etcd/nats/s3, returning Built::Async(Arc::new(store)).
  • spec.rs — the source joins the validated() list, its auth rules join validated_auth() (refusals name the fix; that is the house voice), and USAGE learns the flags. Tests beside the etcd ones.

3. The webhook (dynamic-config-webhook/src/annotations.rs) — the source name joins the admission allowlist, and its auth annotations join the validation match so a typo is refused at admission, in the pod's events, not twenty minutes later in a crashloop. A golden test pins the happy path.

4. The paper trail — a store page in this book (auth methods, full manifests, the identity-first default stated), a ready-to-apply file in examples/, a row in the annotation reference's source table, and — for a store with a live server in CI — an e2e leg shaped like e2e/stores-smoke.sh — one cluster, every live store, a stage per source.

5. The gates — just check in both repositories, the webhook goldens, the kind smoke. Nothing about the release train changes: the store crate ships from dynamic-config-remote, the agent picks it up on the next k8s release.

A store in another language

Three lanes, all of them already open:

  • In-process, Python: subclass the binding's RemoteSource ABC — fetch() and describe() — and hand the instance to DynamicConfig.remote(...). The Python book owns the chapter.
  • In-process, Node: any JS object with fetch() and describe() through useStore(config, store) — the Node remote package's contract.
  • Out-of-process, ANY language — the one the agent can use: the agent cannot load Go or Java, but it speaks the config server's small HTTP contract. Implement GET /{application}/{profile} (bearer-token auth, the resolved document as the body) in your language, and every consumer here — the agent's --source config-server, the engine, both bindings — reads it like any other store. A Go service fronting your in-house store is an afternoon: one endpoint, one auth header, and TLS. The remote book's server chapter tables the full route surface, /stream included for push.

Today, without code

Two honest compositions cover most "we need store X now" cases:

  • External Secrets Operator in front: ESO syncs GCP SM (or any of its many backends) into a Kubernetes Secret, and everything on The Four Deliveries consumes that Secret — mounted as a live file, envFrom, or secretKeyRef. You trade the document model for breadth, which is exactly the trade the comparison table prices.
  • A mirror job: a CronJob copies the document from the unsupported store into one the agent speaks (S3, git, Consul). Crude, visible, and often all a migration window needs.

GitOps: ArgoCD and Flux

Running the chart under a GitOps controller works, with two things worth knowing in advance: certificates, and the difference between the two Gits in play.

The two Gits

A GitOps setup around dynamic-config has two repositories doing two jobs, and conflating them causes most confusion:

  • ArgoCD's Git delivers manifests — the chart, its values, your Deployments with their dynamic-config.rs/* annotations. Changing an annotation is a Deployment change: ArgoCD applies it, the rollout creates new pods, the webhook injects the new agent config. Cadence: deploys.
  • The agent's Git store delivers configuration documents — the agent polls the repository directly, inside the pod, no sync loop involved. Cadence: seconds-to-minutes, no rollout, no ArgoCD.

Config-in-Git does not require routing config through ArgoCD. The agent's Git store gives you Git-audited configuration that updates without a pod restart; keep ArgoCD for the manifests.

Certificates under a sync loop

The chart's default TLS mode mints a CA and certificate at template time (helm template runs genCA). Under ArgoCD that means every sync renders a new certificate — a permanent diff, and a webhook whose caBundle churns on every sync.

Under GitOps, pick one of:

  • cert-manager mode (recommended where cert-manager runs): webhook.certManager.enabled: true. cert-manager issues and renews; cainjector maintains the caBundle; the rendered manifests are stable.
  • selfRotate mode: rotation without the dependency — the webhook patches its own caBundle at runtime, so tell ArgoCD that field is runtime-owned (the ignoreDifferences block below, caBundle entry).
  • Kustomize-native Flux/Argo setups get the same choices as overlays: cert-manager, own-cert, selfrotate, plus with-operator for the Render reconciler — composable in one Kustomization.
  • Provided-cert mode: mount your own Secret (Secrets & TLS) and set the caBundle in values — everything is declarative, nothing is generated.
  • If you must keep the self-signed default, tell ArgoCD to look away:
ignoreDifferences:
  - kind: Secret
    name: dynamic-config-webhook-tls
    jsonPointers: [/data]
  - group: admissionregistration.k8s.io
    kind: MutatingWebhookConfiguration
    jqPathExpressions: [".webhooks[].clientConfig.caBundle"]
syncPolicy:
  syncOptions: [ApplyOutOfSyncOnly=true]

CRDs

Helm installs the crds/ directory on install and never upgrades it — standard Helm behaviour, and under ArgoCD it depends on how the chart is rendered. Two reliable shapes:

  • ArgoCD renders charts with helm template --include-crds; the CRDs become ordinary tracked resources and upgrades apply. Add ServerSideApply=true to syncOptions — the rendering CRDs carry large schemas and can exceed the client-side annotation limit.
  • Or manage deploy/crds/ as its own Application (or Flux Kustomization) in an earlier sync wave, which is also the shape the operator's future CRD changes will assume.

Order of arrival

The webhook must be serving before annotated workloads are admitted. In one Application that is a race; in practice it is self-healing (failurePolicy: Ignore is the default — an early pod is admitted un-injected and the next rollout catches it), but a strict setup puts the chart in an earlier sync wave than the workloads:

metadata:
  annotations:
    argocd.argoproj.io/sync-wave: "-1"   # the chart's Application

With failurePolicy: Fail the wave separation stops being optional — see The Security Posture for that trade.

The Operator

Shipped as of 0.1.1: a DynamicConfigRender becomes a ConfigMap, reconciled through the SAME source construction and rendering the sidecar agent uses — one implementation, no drift between the two paths a document can take into a pod. The CRDs stay generated from the Rust types (deploy/crds.json, drift-gated in CI), and the e2e suite drives the full loop: apply, render, propagate, delete, garbage-collect.

DynamicConfigClass — name the store once

The class bundles source, endpoint and the token Secret, so pods stop repeating them:

apiVersion: dynamic-config.rs/v1alpha1
kind: DynamicConfigClass
metadata:
  name: infra-consul
  namespace: billing
spec:
  source: consul
  endpoint: http://consul.infra.svc:8500
  tokenSecret: consul-agent-token     # optional; its `token` key travels

A pod then says only what is its own:

metadata:
  annotations:
    dynamic-config.rs/inject: "true"
    dynamic-config.rs/class: "infra-consul"
    dynamic-config.rs/key: "myapp/config.json"
    dynamic-config.rs/path: "/config/rendered.toml"

The long-form annotations keep working forever; the class is sugar over them, not a replacement.

DynamicConfigRender — a ConfigMap instead of a sidecar

For workloads that cannot take an injected container — third-party charts, jobs, anything sidecar-averse — the operator renders into a ConfigMap on a cadence, and the workload mounts it like any other:

apiVersion: dynamic-config.rs/v1alpha1
kind: DynamicConfigRender
metadata:
  name: billing-config
  namespace: billing
spec:
  class: infra-consul
  key: myapp/config.json
  target:
    configMap: billing-rendered
    file: config.properties        # the extension picks the format here too
  intervalSeconds: 30
status:                             # written by the operator
  renderedAt: "2026-08-18T09:00:00Z"
  lastError: null                   # kind and path only, never a value

Classes: namespaced, or cluster-scoped

Two class kinds, the same split External Secrets Operator drew between SecretStore and ClusterSecretStore:

  • DynamicConfigClass (namespaced) — a team's own store, its tokenSecret read from the same namespace. Self-service, blast radius one namespace.
  • ClusterDynamicConfigClass (cluster-scoped) — the platform team's store, defined once. Its credential names the namespace it lives in explicitly, so tenants reference the class without ever being able to read the credential:
apiVersion: dynamic-config.rs/v1alpha1
kind: ClusterDynamicConfigClass
metadata:
  name: platform-consul
spec:
  source: consul
  endpoint: http://consul.infra.svc:8500
  tokenSecret:
    name: consul-token
    namespace: platform        # RBAC keeps tenants out of here
    key: token                 # `token` by default
  namespaces: [team-a, team-b] # the allowlist; absent = every namespace

A tenant's Render opts in by kind:

spec:
  class: platform-consul
  classKind: ClusterDynamicConfigClass

Two rules, out loud. The allowlist is enforced at reconcile: a Render in a namespace the class does not list fails with the class named, in status.lastError — the platform team's boundary, not a convention. And the credential read happens with the operator's identity, which is why the operator's RBAC has cluster get on Secrets: the tenant's own service account never touches the platform namespace. Editing a cluster class re-renders every Render referencing it, in every namespace — the same live wiring the namespaced class has.

Targets: a ConfigMap, or a Secret

target names exactly one destination:

  target:
    configMap: billing-rendered   # workloads that mount files
    file: config.properties
  target:
    secret: myapp-env             # workloads that read SECRETS natively
    shape: envEntries             # …or take environment through envFrom

The Secret target is for the consumers the file path cannot reach: an operator that watches Kubernetes Secrets reacts to every reconcile with no pod restart — the Secret is a live object — and an envFrom block turns shape: envEntries (every leaf of the resolved document, dotted paths upper-snaked: db.pool_size → DB_POOL_SIZE) into environment variables at the next container start. Environment freezes at start; that is Kubernetes' rule, and this page will not pretend otherwise. A vault class is allowed into the Secret target — that is the container a secret store's document belongs in — while the ConfigMap target keeps refusing it.

Feeding a name someone else already chose

The commonest enterprise shape: a helm chart or an operator demands a Secret by name — auth.existingSecret in half the chart ecosystem, a secretName: field in an operator's CRD — and refuses env vars or files. That named Secret is exactly what a Render produces:

spec:
  class: platform-vault
  classKind: ClusterDynamicConfigClass
  key: secret/postgres
  target:
    secret: pg-credentials
    shape: entries          # leaf keys VERBATIM: postgres-password, …
helm install db bitnami/postgresql --set auth.existingSecret=pg-credentials

Three shapes, three contracts: file when the consumer wants one document under one key; envEntries when it reads through envFrom (keys upper-snaked to the env dialect); entries when the key names are someone else's contract — every leaf verbatim, postgres-password staying postgres-password, because any mangling breaks a name you do not own. The password never appears in values, in git, or in a developer's hands; the store document is the single source, and the Secret follows it on the Render's interval.

Deleting the Render deletes its target unless it says otherwise:

spec:
  target:
    secret: billing-credentials
    deletionPolicy: Retain          # or Delete, the default

Delete owns the target via ownerReferences, so the cleanup is Kubernetes' own garbage collector — no finalizer to get wrong. Retain writes no owner reference at all, which is the right answer for a Secret something else still needs: deleting the Render that produced a credential should not be the same act as revoking it everywhere.

status.renderedAt says when the last render landed and status.lastError carries kind and shape only, never a value. status.observedGeneration is the metadata.generation this status was written for, so a controller or a kubectl wait can tell "this succeeded" from "this has not been looked at since it changed".

Why a render is not ready

The Ready condition carries a machine-readable reason, and the three are kept apart because they belong to different people:

reasonwhat it meanswho fixes it
ClassNotFoundthe class this render names does not existwhoever wrote the Render
ClassNotAllowedit exists, and does not admit this namespacethe platform team that owns the class
RenderFailedthe store could not be read, or the document could not be rendereddepends on the message

Before 0.3.0 all three were RenderFailed, so every alert grepped a message. These strings are API on the same terms as the annotation contract: a dashboard aggregates on them, and a rename is a breaking change.

Two honesty notes, stated before the reconciler shipped and still true:

  • ConfigMap propagation is slow — the kubelet syncs mounted ConfigMaps on its own cadence (up to a minute-plus; the engine book's Kubernetes Files page walks the mechanism). The sidecar's emptyDir is the low-latency path; DynamicConfigRender trades latency for no-sidecar.
  • A ConfigMap is not a Secret. The reconciler refuses vault classes into ConfigMaps with exactly that sentence; the secret: target is the lift, and the refusal message points at it.

The class annotation — still ahead

dynamic-config.rs/class on a pod (the webhook resolving a class so annotations shrink) is contract-only: it needs the webhook to read CRs, which trades away its zero-RBAC posture the way selfRotate does, and that trade is taken per-feature, not by default.

What the operator will not do

Own lifecycles inside pods, restart workloads on render, or template documents. It renders and it reports; reacting is the workload's business, and the whole engine exists so reacting is cheap.

One Agent per Node

Every other shape here puts an agent beside the application: one container per render, one fetch per pod. That is the right default — the credential is scoped to the workload that needs it, and a compromised agent reaches one application's configuration.

It is also 25,000 containers at 10,000 pods and 2.5 renders each, and that number is what this exists for.

volumes:
  - name: config
    csi:
      driver: config.dynamic-config.rs
      readOnly: true
      volumeAttributes:
        source: consul
        endpoint: http://consul.default.svc:8500
        key: myapp/config.json
        path: rendered.toml        # a name inside the volume
containers:
  - name: app
    volumeMounts:
      - name: config
        mountPath: /config

No annotations, no injected container, no webhook involved at all. The kubelet asks the node's agent for the volume, and the agent renders into it.

Why this is one component and not two

"A node-level agent" and "a CSI driver" were two entries on the same list, and building them separately would have been building a thing and its only delivery mechanism as though they were unrelated.

A DaemonSet that fetches for a whole node has to get bytes into a pod, and there are two ways: a hostPath the pod also mounts — which restricted Pod Security forbids, for the reason it forbids it — or a CSI volume, which is the interface Kubernetes added for exactly this and which the kubelet already knows how to mount, unmount and clean up after an eviction.

So this is a CSI node plugin whose backing store is the same engine the sidecar runs. Not a second implementation of fetch-resolve-render-watch: the sidecar's own crate, called as a library.

What it shares

Two pods on one node that want the same document from the same store under the same credential share one fetch and one watch. A node running a hundred pods that read one Consul key opens one connection to Consul.

The credential is part of that identity, not metadata beside it. Two pods reading one key under different tokens are two reads, and sharing them would hand one namespace's document to another under a credential it was never granted.

What they do not share is the rendered file: each pod gets its own bytes at its own path, in its own format, with its own mode. A .properties reader and a YAML reader on one node share the fetch and share nothing else.

Two series say whether any of it is working:

dynamic_config_node_agent_documents 12
dynamic_config_node_agent_readers   96

On a node where those two numbers are equal, nothing is being shared and a sidecar would have cost the same.

One property a sidecar cannot offer

The first render happens before the pod starts. The kubelet does not start a pod's containers until every volume is published, and publishing is what does the first fetch — so an application cannot observe a missing file, and there is no init container here because there is nothing for one to do.

What it costs

Said plainly, because it is the reason this is off by default and the sidecar is not.

  • Credentials for many workloads in one process. A compromised node agent reaches every store credential every pod on that node uses. The sidecar's isolation is exactly what is traded away.
  • It runs as root. The kubelet creates a pod's volume directories owned by root, and a plugin that cannot write into them cannot publish. Every other binary here runs nonroot; this one cannot.
  • It mounts host paths. The kubelet's plugin and pod directories, which is what a CSI driver is.

Not a better shape. A different trade, for a scale that makes the first one untenable — and a decision to make with a measurement rather than a preference.

Installing it

helm upgrade dynamic-config … --set nodeAgent.enabled=true

nodeAgent.kubeletPath is /var/lib/kubelet everywhere except a few distributions — k0s and some managed offerings move it, and a wrong value is a driver the kubelet registers and never calls.

The DaemonSet carries upstream's node-driver-registrar beside the agent. Its whole job is telling the kubelet this driver exists; writing that here would be reimplementing a thing Kubernetes ships.

The attributes a volume takes

The same vocabulary as the annotations, for the same reason the agent's flags are: one contract, learned once, whichever shape delivers it.

attribute
sourcerequired — consul, vault, etcd, …
keyrequired — the document's key, path or object
pathrequired — a name inside the volume; the extension picks the format
endpointthe store's address
section, auth, auth-mount, auth-role, auth-username, namespace, ref, api-url, file-mode, watch-secondsas the annotations of those names

path is checked rather than trusted: absolute paths and .. are refused, because the kubelet owns the directory and a volume that could write outside it could write into every other pod's volume on the node.

There is no mode: a CSI volume is published before the containers start and stays. There is no env-inject: there is no command here to wrap. And there is no metrics-port: the metrics are the node agent's, one endpoint for the whole node.

Observability

Three components, three Prometheus endpoints, and OTLP traces beside them when a collector is configured. The metrics are hand-rolled exposition text — a handful of counters do not earn a metrics crate — and are scrapeable by anything that speaks the format, the OTel Collector included.

What each component exposes

componentwherehow it turns on
webhookits own plain-HTTP port, and GET /metrics on the serving (TLS) port toowebhook.metrics.enabled, on by default
agentits own port, plain HTTPmetrics-port annotation, or the agent.defaults.metricsPort fleet default
operator0.0.0.0:9090, plain HTTPon by default; DYNAMIC_CONFIG_OPERATOR_METRICS_ADDR="" turns it off

The webhook's metrics

The webhook sits in the path of every pod creation in the cluster, so the first question about it is never "did it work" but "how long did it take" — the API server's own ten-second timeout turns a slow admission into a refused one, and only a histogram answers what the tail is doing.

# TYPE dynamic_config_admissions_total counter
dynamic_config_admissions_total{outcome="skipped"} 41
dynamic_config_admissions_total{outcome="patched"} 12
dynamic_config_admissions_total{outcome="refused"} 3
# TYPE dynamic_config_admission_refusals_total counter
dynamic_config_admission_refusals_total{reason="policy"} 2
dynamic_config_admission_refusals_total{reason="pinned"} 1
dynamic_config_admission_refusals_total{reason="conflict"} 0
dynamic_config_admission_refusals_total{reason="malformed"} 0
dynamic_config_admission_refusals_total{reason="other"} 0
# TYPE dynamic_config_admission_duration_seconds histogram
dynamic_config_admission_duration_seconds_bucket{le="0.0001"} 39
dynamic_config_admission_duration_seconds_bucket{le="0.001"} 55
dynamic_config_admission_duration_seconds_bucket{le="+Inf"} 56
dynamic_config_admission_duration_seconds_sum 0.031
dynamic_config_admission_duration_seconds_count 56
# TYPE dynamic_config_admission_patch_bytes_total counter
dynamic_config_admission_patch_bytes_total 24576

skipped is a pod that did not ask; patched asked and got the agent; refused asked wrongly. The refusals are labelled by kind because they are three different pages:

reasonwhat happenedwho fixes it
policya store, or an agent-env name, that this installation does not allow herewhoever wrote the pod, or whoever set the gate
pinneda value the installation fixed, overridden by a podwhoever is working around the installation
malformedan annotation that is not the shape it has to bewhoever wrote the pod — a typo
conflictthe pod already has a container by a name the injection needswhoever wrote the pod — rename it

The kind is a status.reason the refusal carries, not something a scrape works out by reading the message, so rewording an error does not silently re-label a metric.

In selfRotate mode two more say whether rotation is still happening:

# TYPE dynamic_config_certificate_rotations_total counter
dynamic_config_certificate_rotations_total 7
# TYPE dynamic_config_certificate_expires_at_seconds gauge
dynamic_config_certificate_expires_at_seconds 1755698600

A counter that should climb on a schedule, and a wall-clock second that should always be in the future. Alert on the gauge: a webhook whose certificate expires takes every pod creation in the cluster with it.

- alert: DynamicConfigWebhookCertificateExpiring
  expr: dynamic_config_certificate_expires_at_seconds - time() < 3600
- alert: DynamicConfigAdmissionSlow
  expr: histogram_quantile(0.99, rate(dynamic_config_admission_duration_seconds_bucket[5m])) > 1

Where to scrape it

A port of its own, in plain HTTP — webhook.metrics.port, 9091 by default. The admission port is mutual TLS against a CA the webhook mints for itself, so scraping that meant handing Prometheus a client certificate from that CA, and a deployment that did not simply had no metrics. The admission port still answers /metrics, so a scrape already configured against it keeps working.

With the Prometheus Operator, one toggle:

webhook:
  metrics:
    serviceMonitor:
      enabled: true      # needs the ServiceMonitor CRD
      interval: 30s

Without it, the Service carries a named metrics port:

- job_name: dynamic-config-webhook
  kubernetes_sd_configs:
    - role: endpoints
      namespaces: { names: [dynamic-config] }
  relabel_configs:
    - source_labels: [__meta_kubernetes_endpoint_port_name]
      action: keep
      regex: metrics

Watching

The sidecar watches; it does not poll. Each store says how it learns that its document changed, and the agent uses it — so a change in etcd, Consul, NATS, Redis or a config server arrives as it happens, and a store that must be asked is asked the cheapest question it offers (an S3 object is a HEAD, not a download).

watch-seconds keeps its spelling and means two things depending on the store:

the storewhat a watch iswhat watch-seconds means
etcd, Consul, NATS, Redis, config-serverthe store's own pusha resync — how often to re-read anyway
Vault, S3, git, Firestorea cheap "has it changed?"how often to ask

examples/watch-driven.yaml is a pod on the push path with both numbers set: a five-minute resync under etcd's stream, and the metrics port to scrape it from.

The resync is not belt and braces. The failure mode of a stream is silence: a subscription the broker forgot, a connection that dropped without an error, a blocking query answering an index that will never move again. All three look exactly like a store where nothing has changed, and the only way to tell them apart is to go and ask.

The agent's metrics

A watching agent with a metrics-port serves twenty-seven series; the port also lands on the container spec as a named port (metrics), so selector-based discovery finds it without configuration. Since 0.3.0 the chart gives every installation a default port (9110), so this is on unless somebody turns it off with metrics-port: "0":

# TYPE dynamic_config_agent_renders_total counter
dynamic_config_agent_renders_total 128
# TYPE dynamic_config_agent_render_failures_total counter
dynamic_config_agent_render_failures_total 2
# TYPE dynamic_config_agent_last_render_timestamp_seconds gauge
dynamic_config_agent_last_render_timestamp_seconds 1755612200
# TYPE dynamic_config_agent_deliveries_total counter
dynamic_config_agent_deliveries_total 12
# TYPE dynamic_config_agent_resyncs_total counter
dynamic_config_agent_resyncs_total 480
# TYPE dynamic_config_agent_watch_connected gauge
dynamic_config_agent_watch_connected 1
# TYPE dynamic_config_agent_watch_reconnects_total counter
dynamic_config_agent_watch_reconnects_total 3
# TYPE dynamic_config_agent_staleness_seconds gauge
dynamic_config_agent_staleness_seconds 4
# TYPE dynamic_config_agent_generation gauge
dynamic_config_agent_generation 41
# TYPE dynamic_config_agent_lease_renewals_total counter
dynamic_config_agent_lease_renewals_total 6
# TYPE dynamic_config_agent_lease_renewal_failures_total counter
dynamic_config_agent_lease_renewal_failures_total 0
# TYPE dynamic_config_agent_lease_revocations_total counter
dynamic_config_agent_lease_revocations_total 0
# TYPE dynamic_config_agent_lease_ttl_seconds gauge
dynamic_config_agent_lease_ttl_seconds 3600
# TYPE dynamic_config_agent_absent gauge
dynamic_config_agent_absent 0
# TYPE dynamic_config_agent_absent_total counter
dynamic_config_agent_absent_total 0
# TYPE dynamic_config_agent_notifications_total counter
dynamic_config_agent_notifications_total 12
# TYPE dynamic_config_agent_notification_failures_total counter
dynamic_config_agent_notification_failures_total 0
# TYPE dynamic_config_agent_drift gauge
dynamic_config_agent_drift 0
# TYPE dynamic_config_agent_drift_total counter
dynamic_config_agent_drift_total 0
# TYPE dynamic_config_agent_canary_holding gauge
dynamic_config_agent_canary_holding 0
# TYPE dynamic_config_agent_canary_percent gauge
dynamic_config_agent_canary_percent 100
# TYPE dynamic_config_agent_acks_total counter
dynamic_config_agent_acks_total 41
# TYPE dynamic_config_agent_ack_mismatches_total counter
dynamic_config_agent_ack_mismatches_total 0
# TYPE dynamic_config_agent_applied gauge
dynamic_config_agent_applied 1
# TYPE dynamic_config_agent_unapplied_seconds gauge
dynamic_config_agent_unapplied_seconds 0
# TYPE dynamic_config_agent_tls_reloads_total counter
dynamic_config_agent_tls_reloads_total 0
# TYPE dynamic_config_agent_tls_verification_skipped gauge
dynamic_config_agent_tls_verification_skipped 0

generation is the store's own revision of what was last rendered, where the store counts them — a Vault KV version, a Consul index, an etcd revision. Zero for a store whose revision is opaque, because an ETag has no number to report and inventing one would make a dashboard compare things that do not compare.

The four lease_* series are zero unless the source is a dynamic-secret engine. lease_ttl_seconds is what the store last granted, not what was asked for.

lease_renewal_failures_total is worth alerting on precisely because it should stay at zero. A lease the store marked renewable: false — every pki/issue, and a database credential past its role's maximum — is never sent a renewal in the first place; it is re-issued instead, and lease_renewals_total stays at zero while the credential keeps arriving. So a rising failure count is always a lease that was expected to renew and did not, which is a real thing to page on rather than the background noise it would be if every non-renewable lease were asked anyway.

tls_verification_skipped is 1 when the agent runs with tls-skip-verify. A gauge rather than a log line repeated per fetch, because the question it answers is a fleet question — which of these five thousand pods are doing this? — and one alert over a gauge answers it where five thousand log streams do not. It should stay at zero.

canary_holding is 1 while this pod is outside a canary cohort and is holding a document it fetched — without it, a held document looks exactly like a store with nothing new to say. canary_percent is the number it last read, and 100 when no canary is configured.

The only end-to-end answer

renders_total says a document reached disk. applied says the application is running it — the four ack series are the difference between "we published" and "it converged", and an application still running the previous document while every other series reports success is the outage they close.

They stay at zero unless the application acknowledges, which is not a failure: an application that never acknowledges is not penalised, it simply leaves applied at zero. unapplied_seconds is the one to alert on — ack_mismatches_total climbing steadily means acknowledgements and renders are talking past each other, which is a different problem from being behind.

The annotations page has the two-line contract, and require-ack turns it into readiness.

tls_reloads_total counts the times the store's client was rebuilt for rotated trust material — a counter rather than a gauge, because the question worth asking is whether a rotation was picked up at all.

drift is 1 while a rendered file is something other than what this agent wrote. notification_failures_total counts post-render calls that were refused or unanswered; neither is fatal, because the file is already published by the time either can happen.

absent is 1 while the store says the document is not there, and absent_total counts how many times it has gone missing. The distinction they encode is the one a stale file cannot make on its own: a store that does not answer is an outage, which waiting cures, and a store that answers gone is a deletion, which waiting does not. Before 0.3.0 a deleted Vault secret moved nothing at all and the last render went on being served with every health check reporting fine. What the agent does about it is on-delete.

/readyz means there is a document

The agent's readiness is stricter than the operator's, and since 0.3.0 the webhook attaches a probe to the injected container: /readyz answers 503 until a document has been rendered, and pod readiness is already AND-ed across containers — so a Service sends no traffic to a pod whose configuration does not exist yet.

The trade, said plainly: with the store unreachable at start and nothing cached, the pod never becomes ready. That is correct — the application has no configuration — but it turns a store outage into a visibly stalled rollout rather than pods that come up and misbehave. Two ways out, and they answer different questions:

  • dynamic-config.rs/startup-policy: allow-cached (the default) serves the file already on disk when the first fetch fails. The volume survives a container restart, so an agent coming back after a crash usually finds its own last render there.
  • dynamic-config.rs/readiness: "false" detaches the probe entirely, for a deployment that would rather start than wait.

And dynamic-config.rs/max-staleness: "6h" puts a ceiling on the first of those: last-known-good answers is there a document and leaves is it too old to trust open. A credential may be worthless after five minutes and a feature flag fine after a week, so there is no default — the ceiling is off unless somebody sets one.

staleness_seconds is the one to alert on. It is seconds since the store was last read successfully, delivered or resynced — the number a pager asks for. Everything else says what happened; this says how long ago it stopped happening. Zero until the first success, so a pod that has never reached its store does not look fresh.

# The document went stale: nothing read from the store for 10 minutes.
- alert: DynamicConfigStoreStale
  expr: dynamic_config_agent_staleness_seconds > 600
- alert: DynamicConfigRenderFailing
  expr: increase(dynamic_config_agent_render_failures_total[10m]) > 0
# A watch that keeps reopening is a store or a network that is not well.
- alert: DynamicConfigWatchFlapping
  expr: increase(dynamic_config_agent_watch_reconnects_total[15m]) > 5

Deliveries and resyncs are told apart because that pair is a diagnosis. Deliveries flat while resyncs climb is the stalled stream the resync exists to cover: the store is being read, changes are landing, and the push half is doing nothing. Nothing else here would show it.

watch_connected is 0 between a watch ending and the next one opening; on a store that is polled rather than pushed it stays 1 for as long as the loop runs.

A one-shot init agent serves nothing — it renders once and exits, and a scrape target that lives for two seconds is noise. Only the watching half carries the port, which is also why metrics-port pairs with mode: sidecar or both.

With a PodMonitor (Prometheus Operator), the named port makes discovery one selector:

apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
  name: dynamic-config-agents
spec:
  selector:
    matchExpressions:
      - { key: app, operator: Exists }   # your workload labels
  podMetricsEndpoints:
    - port: metrics

Fleet-wide, set agent.defaults.metricsPort: "9102" once and every injected watching agent serves on 9102; a pod that must not (a port collision, say) opts out with metrics-port: "0".

The operator's metrics

The agent's render series, dynamic_config_operator_-prefixed — _renders_total counts reconciles that produced or refreshed a target, _render_failures_total the ones that ended in a RenderFailed event, and the gauge the last success — plus a reconcile histogram:

# TYPE dynamic_config_operator_reconcile_duration_seconds histogram
dynamic_config_operator_reconcile_duration_seconds_bucket{le="0.5"} 88
dynamic_config_operator_reconcile_duration_seconds_bucket{le="+Inf"} 91
dynamic_config_operator_reconcile_duration_seconds_sum 12.4
dynamic_config_operator_reconcile_duration_seconds_count 91
# TYPE dynamic_config_operator_reconcile_failures_total counter
dynamic_config_operator_reconcile_failures_total 3

A reconcile is a store fetch and an API write, so a tail that moves is usually the store rather than the operator. On by default at 0.0.0.0:9090; point DYNAMIC_CONFIG_OPERATOR_METRICS_ADDR elsewhere or set it empty to turn it off. The same port answers /healthz and /readyz, which is what the chart's probes use.

The operator's /readyz means the process is up, and that is not an oversight: an operator with no DynamicConfigRender resources has rendered nothing and is working perfectly. The agent means something stricter by the same path — see below.

Only one replica reconciles. The operator elects a leader over a Lease, so operator.replicas: 2 is spare capacity rather than twice the work — and the metrics of a follower stay flat by design. A leader that dies is replaced within the lease's fifteen-second term.

Beside the metrics, the operator reports per-object: a Ready condition on every DynamicConfigRender's status and Kubernetes Events (Rendered / RenderFailed) on the object — kubectl describe dynamicconfigrender is the first debugging stop, before any dashboard.

Logs

Every component writes JSON to stdout through tracing, one line per event, RUST_LOG grammar via the environment. The webhook logs one audit line per non-skipped admission — namespace, pod, source, outcome, and NO annotation values, because store addresses and role names belong in the cluster, not in every log aggregator. The agent logs each render with byte counts, and failures with the store's error.

Turning up an agent's verbosity is the canonical agent-env use:

# the installation allows it:   webhook.agentEnvAllow: "*: RUST_LOG"
dynamic-config.rs/agent-env: "RUST_LOG=debug"

— or fleet-wide for a day with agent.defaults.env: "RUST_LOG=debug", no allowlist needed, the installer owns both.

The engine's own metrics

The series above measure the DELIVERY machinery: admissions, renders, staleness. The configuration engine inside your application has its own richer contract — dynamic_config_reload_total, generation and snapshot-age gauges, per-source health — defined once in the metrics contract of the engine book and exported by the language bindings. When an app consumes the rendered file through the engine's own watcher, alert on the engine's series (closest to the truth the app sees) and keep the agent's gauge as the delivery-side backstop.

A dashboard and the rules under it

deploy/observability/ carries both, over the series named above:

kubectl apply -f deploy/observability/rules.yaml    # recording rules, then alerts
# Grafana → Dashboards → Import → dashboard.json

The rules come first — four of the dashboard's fleet counters read recording rules, and without them those panels stay empty.

Two things worth knowing before tuning the thresholds. Staleness is the gauge to page on, not a failure counter: an agent that fails one fetch and recovers is doing what it was built to do, while one whose store has not answered in half an hour is serving a document nobody can vouch for. And lease_renewal_failures_total should sit at zero, which is what makes a low threshold on it reasonable — a lease the store marked renewable: false is never sent a renewal at all, so anything counted there was expected to renew and did not.

The recording rules aggregate pod away. A thousand agents cost four series through them and four thousand without; the alerts that do keep pod are the ones whose whole purpose is naming which pod.

OpenTelemetry

Since 0.3.0 the webhook, the agent and the operator export OTLP traces — and only when a collector is configured. Setting OTEL_EXPORTER_OTLP_ENDPOINT is the whole opt-in; leave it unset and nothing is exported, nothing is allocated, and the process behaves exactly as it did. The variable is OpenTelemetry's own, so a collector that already injects it into pods needs no annotation from this chart.

Resource attributes come from the downward API — service.name, service.version, k8s.pod.name, k8s.namespace.name, k8s.node.name — and are absent rather than invented when nobody wires them: a made-up pod name is worse than none.

Where the pipeline is built

The engine offers an otel feature that records into a meter somebody else configured, and refuses to build the pipeline. That is the right split: a pipeline owns a runtime, a batch processor and a shutdown, and a library that takes those over is one that has to be fought. These three are programs, and programs own main — so the exporter, and the flush before exit, are theirs.

If a collector cannot be reached, the process logs a warning and carries on without traces. A configuration agent that refused to start because its telemetry was down would be a worse outage than the one it was reporting.

The Prometheus text is untouched

Nothing above replaces the scrape. Many deployments have one already, and the collector reads it directly — neither way out is the price of the other:

  • Metrics: the OTel Collector's prometheus receiver scrapes all three endpoints and exports OTLP wherever you aggregate:

    receivers:
      prometheus:
        config:
          scrape_configs:
            - job_name: dynamic-config-agents
              kubernetes_sd_configs: [{ role: pod }]
              relabel_configs:
                - source_labels: [__meta_kubernetes_pod_container_port_name]
                  regex: metrics
                  action: keep
    
  • Logs: the JSON lines on stdout are structured input for the filelog receiver, no parsing regexes required.

Traces are spans over the work — an admission, a fetch, a render — and the scrape is the state of it. A deployment that wants only one of the two is not missing anything by taking only one.

Troubleshooting

Each symptom, its first command, and the usual cause. The agent's logs are JSON; kubectl logs <pod> -c dynamic-config-agent | jq . is the readable form.

The file never appears

kubectl get pod <name> -o jsonpath='{.spec.containers[*].name}'

No dynamic-config-agent in the list → the webhook never mutated. In order of likelihood:

  1. The pod predates the webhook install — admission only sees CREATE.
  2. failurePolicy: Ignore swallowed a webhook outage: kubectl get events -n <ns> | grep dynamic-config and the webhook deployment's own logs say which.
  3. The annotation said inject: "false" or misspelled the prefix — dynamic-config.rs/, with the dot and the slash.

Agent present, file absent → the agent is failing. Its log carries the store error verbatim minus values:

kubectl logs <pod> -c dynamic-config-agent | jq -r '.fields.error // .fields.message'

The admission was denied

That is the contract working: inject: "true" with a missing or malformed companion annotation fails the pod's creation, and the reason names the annotation:

Error creating: admission webhook "inject.dynamic-config.rs" denied the
request: dynamic-config.rs/inject is true, so dynamic-config.rs/path is required

Silently starting without configuration is the failure mode this refusal exists to prevent; add the named annotation.

"container name is duplicated" from the API server

Error creating: Pod "app" is invalid: spec.containers[2].name:
Duplicate value: "dynamic-config-agent"

Your pod already has a container by a name the injection needs. The webhook refuses that now, with the name to rename:

admission webhook "inject.dynamic-config.rs" denied the request: this pod
already has a container called "dynamic-config-agent", and the injection
needs that name — rename yours, or set dynamic-config.rs/inject to "false"

Seeing the API server's version of it instead means a webhook older than 0.2.0 patched the pod. Upgrade, or rename the container.

The injected names are dynamic-config-agent and dynamic-config-init, plus -1, -2… for each named render.

The same pod was injected twice

Two dynamic-config-agent containers in a pod nobody wrote twice is admission running twice. It happens with webhook.reinvocationPolicy: IfNeeded, which asks the API server to call this webhook again whenever a later webhook changes the pod, and with controllers that resubmit an already-admitted spec.

Since 0.2.0 the patch marks the pod — dynamic-config.rs/status: injected — and a marked pod is passed through untouched. If you are seeing this, check the webhook image is 0.2.0 or later:

kubectl -n dynamic-config get deploy dynamic-config-webhook \
  -o jsonpath='{.spec.template.spec.containers[0].image}'

401/403 from the store

The token travels in DYNAMIC_CONFIG_AGENT_TOKEN, not in annotations.

kubectl exec <pod> -c dynamic-config-agent -- env | grep -c DYNAMIC_CONFIG

0 means the Secret was never mounted onto the agent container. Config-server 401s specifically: the bearer must belong to a [[server.clients]] block whose applications list names the application in your --key — the server's audit log line for the refusal names the client it matched.

The rendered file is stale

The sidecar keeps the last good render on fetch failure, on purpose — staleness with a warning beats an empty file. The log says so at warn level. --watch-seconds too high is the boring cause; a store ACL that started refusing is the interesting one, and the 401 section above applies.

mode: init never refreshes by design: rotation there is a pod restart, which is the trade the Vault page states.

The pod never becomes ready

Since 0.3.0 the webhook attaches a readiness probe to the injected container, so a pod stays 0/2 until its first render lands. kubectl logs <pod> -c dynamic-config-agent says why the fetch is not succeeding.

Three ways out, and they answer different questions:

  • the store is genuinely down and a stale document is acceptable → startup-policy: allow-cached (the default) already serves the file on the volume if there is one; there is none on a first start
  • the document is fine but too old → dynamic-config.rs/max-staleness is what flipped readiness; the staleness_seconds gauge says by how much
  • the pod should start regardless → dynamic-config.rs/readiness: "false", and the application handles a configuration that is not there yet

The document was deleted and nothing happened

Check dynamic_config_agent_absent. If it is 1, the store is answering gone and the agent is doing what on-delete says — retain by default, which keeps serving the last render. remove truncates the file and fail ends the agent so the pod restarts.

If it is 0 while the key is definitely gone, the store is not reporting the deletion: a Conditional store polled on an interval takes up to one watch-seconds to notice.

The render is refused and the old file keeps serving

Two checks can refuse a document after it has been fetched, and both keep the last good file rather than publishing a bad one:

  • the schema (schema-configmap) — the log line names the failing path and the constraint, never the value
  • the size ceiling (max-document-bytes, 8 MiB) — the refusal names the limit and the store; a document that grew past it is usually a key that now points at something else

--out refused at startup

The extension picks the format, and only five are legal: .json .toml .yaml .ini .properties. The error lists them; .conf and .cfg are nobody's format and stay refused.

Arrays in the document, flat output requested

`hosts` is an array, and neither flat format has one; render to json,
toml or yaml instead

Exactly what it says: pick a structured output, or reshape the document. The Rendering page owns the reasoning.

Reading what the webhook actually did

The golden test's fixture is the contract, and a live pod can be compared against it:

kubectl get pod <name> -o json | jq '.spec.containers[].volumeMounts'
kubectl get pod <name> -o json | jq '.spec.volumes[] | select(.name == "dynamic-config")'

The pod was created, nothing was injected

In order of likelihood:

  1. The namespace is excluded. kube-system, kube-node-lease and the chart's own namespace never get injection, plus anything in webhook.excludeNamespaces:

    kubectl get mutatingwebhookconfiguration dynamic-config \
      -o jsonpath='{.webhooks[0].namespaceSelector}'
    
  2. The webhook was down and failurePolicy: Ignore let the pod through. The API server records exactly that:

    kubectl get events --field-selector reason=FailedAdmissionWebhook -A
    kubectl -n <chart-namespace> get pods -l app.kubernetes.io/component=webhook
    
  3. TLS trust is broken — the caBundle does not match what the webhook serves. With the self-signed default this happens when the Secret was deleted but the webhook configuration was not re-rendered; helm upgrade heals both sides. The API server's opinion:

    kubectl logs -n <chart-namespace> deploy/dynamic-config-webhook | tail
    

The webhook pod refuses to start: an installation setting

dynamic-config-webhook: /etc/dynamic-config/installation.yaml:
"watchSecnods" is not an installation setting; the ones there are:
cpuRequest, memoryRequest, …, sourceDeny

A typo in the mounted installation document, refused at startup rather than ignored — a default that silently never applied is a fleet running without the posture somebody thought they had set. Fix the key in agent.defaults / webhook.* (chart) or in base/installation.yaml (kustomize).

The same check covers shapes: a store's settings are a map, a gate's namespaces map to lists, and a setting is a word, a number or a boolean.

The webhook pod refuses to start: "no TLS material"

The exact message names the two paths it looked at. The Secret dynamic-config-webhook-tls is missing or empty — with cert-manager enabled, check the Certificate:

kubectl describe certificate dynamic-config-webhook

An injected pod is rejected by Pod Security admission

It should not be: the injected container carries the full restricted posture. If a namespace enforces something stricter than restricted (an OPA/Kyverno policy), read the denial message — the security page lists every field the injection sets, so the diff is one screen.

The template refuses to render

Two shapes, two places:

  • At startup (a parse error, an undefined key on the first render): the agent exits and the reason is in the injected container's log — strict on purpose, so a typo'd {{ db.hots }} cannot ship an empty string with a clean exit code.

  • During a watch (the ConfigMap was edited into an error): the pod keeps its last good file and the agent logs the render error every tick until the template is fixed. Check with:

    kubectl logs <pod> -c dynamic-config-agent | tail
    kubectl get configmap billing-template -o jsonpath='{.data.template}'
    

Stability & Versioning

Experimental, stated plainly: this is the youngest repository in the organisation and the annotation contract is v1 — and since 0.3.0 that contract is a registry rather than a list somebody keeps in step: every key is one row carrying whether it takes a .name suffix and whether it has been retired, a test checks this book against it, and a retired key will be admitted with a warning naming its replacement rather than refused. Nothing is retired yet. The operator's reconcilers shipped in 0.1.1 and it stays 0.x until they have soak history, whatever the rest of the family does.

  • The annotation contract is the API; a breaking change to it bumps the minor and regenerates the golden file in the same commit. It grew additively in 0.2.0: dynamic-config.rs/status, which the webhook writes on a pod it has patched so that a second admission does not patch it again.
  • The installation is one contract with three spellings — chart values, a mounted YAML document, or environment variables. A map is rendered to the same grammar the variables carry and goes through the same parser, so adding a spelling is not adding a semantics.
  • The three components version together; images are the artefacts.
  • The agent's store list grew additively (etcd, nats and s3 landed in 0.1.1) and is complete at nine; a tenth would follow the same rule.
  • The engine dependency is a caret: an engine patch reaches the images on rebuild.

The repository's ROADMAP carries the ladder in full — the async stores and etcd's two-methods-forever answer, the self-rotating webhook TLS mode and the one narrow RBAC it will cost, the operator's reconcilers.