dynamic-config on Kubernetes
Annotate a pod; an agent appears in it that renders configuration from a remote store to a file the application watches. The agent-injector shape, for configuration.
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "consul"
dynamic-config.rs/endpoint: "http://consul:8500"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/path: "/config/rendered.toml"
The application needs no store client, no store credential, and no code
change: it reads a file, and if it reads it with any dynamic-config
binding it also reloads on every re-render — the agent writes
atomically, exactly the whole-file event a watcher wants.
When not to use this. The engine runs in-process everywhere; a Rust,
Python or Node service that can hold a store credential should usually
use its own binding's remote support and skip the sidecar entirely. This
integration exists for the pods that want files rendered for them:
Java services reading .properties, anything that must not carry store
credentials in-process, and fleets standardising one injection pattern.
How a document REACHES a workload — a live file, real environment variables, or a native Kubernetes Secret — is its own decision, with a map and two honest comparison tables (Vault Agent Injector, External Secrets Operator) on The Four Deliveries.
The four pieces, staged
| piece | ships in | today |
|---|---|---|
| agent | 0.3.0 | all nine stores — the blocking six since 0.1.0, etcd/nats/s3 on the async path since 0.1.1 — watched rather than polled since 0.2.0, and since 0.3.0: last-known-good with a startup policy, dynamic secrets with lease renewal, a readiness probe that means there is a document, drift, history, acknowledgement and canary cohorts |
| webhook | 0.3.0 | golden-tested; the annotation contract is v1 and is now a registry a test checks this book against; three TLS modes incl. selfRotate; admission warnings and a validate CLI since 0.3.0 |
| operator | 0.3.0 | Render → ConfigMap reconciler shipped, Class watch wired, e2e-gated, leader-elected since 0.2.0; deletionPolicy, observedGeneration and three refusal reasons since 0.3.0 |
| node agent | 0.3.0 | one process per node instead of one beside every render, delivering as a CSI volume; pods sharing a document share a watch. Off by default — it holds the credentials of every pod on its node, so the sidecar stays the shape to reach for |
Install
# From the OCI registry (ArtifactHub lists the same chart):
helm install dynamic-config \
oci://ghcr.io/dynamic-config-rs/charts/dynamic-config --version 0.3.0
# Or from a checkout:
helm install dynamic-config deploy/helm
Every image the chart deploys is on ghcr and mirrored to Docker Hub
(docker.io/ctolon17/…) with identical digests, multi-arch, SBOM
attested and cosign-signed keyless; the repository README carries the
cosign verify line.
That is the whole install: no dependencies, nothing to pre-create. The
chart mints a CA and a ten-year serving certificate at install time,
embeds the caBundle in the webhook configuration, and reuses the
Secret on upgrades so the trust does not silently rotate. The webhook
terminates that TLS in-process — the API server speaks HTTPS to
admission webhooks and nothing else.
The cert-manager mode
helm install dynamic-config deploy/helm \
--set webhook.certManager.enabled=true \
--set webhook.certManager.issuerRef.name=<your-issuer>
With cert-manager the certificate is renewed — cainjector maintains
the caBundle, and the webhook picks up the renewed pair from disk
without a restart (it polls the mounted files for a changed
modification time). The trade between the two:
The selfRotate mode
helm install dynamic-config deploy/helm \
--set webhook.selfRotate.enabled=true
The third answer, the Vault-agent-injector shape: the webhook is its
own certificate authority. It mints a CA and leaf in memory at
rotation time, writes the pair to its own Secret — every replica
serves it through the same file hot-reload cert-manager uses — and
patches the webhook configuration's caBundle itself. A fresh pair
every 24 hours, leader-elected over a Lease so replicas do not race,
jittered so a fleet restarted together does not rotate together.
The price is stated in values.yaml beside the toggle: a
service-account token and three narrow, name-scoped permissions — the
zero-RBAC purity of the other two modes, knowingly traded for rotation
without a dependency.
| self-signed (default) | cert-manager | selfRotate | |
|---|---|---|---|
| dependencies | none | cert-manager installed | none |
| renewal | none — ten-year cert; rotate by deleting the Secret and upgrading | automatic, well before expiry | automatic, every 24h, leader-elected |
| caBundle | embedded at install | maintained by cainjector | patched by the webhook itself |
| RBAC | none | none | one Secret, one MWC, leases — by name |
| fits | getting started, edge clusters, air-gapped | anywhere cert-manager already runs | rotation wanted, cert-manager not |
All three mount the same Secret shape at the same path; the webhook's serving loop cannot tell them apart, which is what makes switching later a values change.
A private mirror for the agent image
The injected agent is pulled in application namespaces, so a fleet that mirrors images into a private registry needs two values:
helm install dynamic-config deploy/helm \
--set agent.image=registry.internal/dynamic-config-agent \
--set agent.pullSecret=mirror-cred
agent.pullSecret is appended to each injected pod's own
imagePullSecrets, never replacing them. Pull secrets are namespaced —
the Secret must exist in every namespace that injects, which is the
usual replication job (kubectl create secret docker-registry … -n <each>, or a replicator you already run). The webhook's and operator's
own images use the chart-level imagePullSecrets list instead, because
those pods live in the release namespace.
failurePolicy
Ignore by default, and the book owns the trade: Fail would make the
webhook a single point of failure for every pod creation in selected
namespaces, while Ignore means an annotated pod created during a
webhook outage starts without its agent — loudly, because the file
its application waits for never appears. Flip it with
--set webhook.failurePolicy=Fail once the webhook has earned it in
your cluster; the security page
carries the full argument.
Namespace gating
Clusters that prefer opt-in injection (the Istio shape) set
webhook.namespaceGating=true and label the namespaces that want it:
kubectl label namespace payments dynamic-config.rs/injection=enabled
Everything else is invisible to the webhook — the security page explains what that buys.
What the chart hardens for you
The security page is the complete inventory; the short
list: both deployments run as non-root with the restricted-PSS
container posture, the webhook's ServiceAccount mounts no API token,
kube-system and the release namespace are excluded from injection, two
replicas ride a PodDisruptionBudget, and tag: latest fails the
render. An optional NetworkPolicy writes down that the webhook accepts
the API server and calls nobody.
One namespace of its own
Install into a dedicated namespace, always:
helm install dynamic-config deploy/helm -n dynamic-config --create-namespace
The webhook configuration excludes its own namespace by name — the
self-deadlock guard — so a release installed into default silently
excludes every workload sharing default with it. The chart's NOTES
print a warning when that happens; the e2e smoke installs the dedicated
way for the same reason.
Fleet-wide defaults, validated at the door
What a pod does not say per annotation, the installation says once — and it can say it twice, because defaults come in tiers: annotation > per-store default > fleet default > built-in. Every knob the annotations know is defaultable; there is no second vocabulary:
agent:
defaults:
cpuRequest: 10m # agent-cpu-request still wins per pod
memoryRequest: 32Mi
memoryLimit: 64Mi
cpuLimit: "" # empty on purpose
fileMode: "0640" # empty = the agent's 0644; file-mode wins per pod
watchSeconds: "30" # empty = 15; watch-seconds wins per pod
mode: "both" # empty = sidecar
volumeMedium: "" # empty = memory
nativeSidecar: "" # empty = false
runAsUser: "1000" # empty = 65532; 0 refused, same as the annotation
runAsGroup: "1000"
metricsPort: "9102" # empty = no metrics; pods opt out with metrics-port "0"
env: "HTTPS_PROXY=http://egress.infra.svc:3128" # every agent; pod's agent-env wins per name
source: "consul" # pods may omit source entirely
path: "/config/rendered.toml"
overridable: "" # "false" pins every value set here; "!"/"?" per value
perStore: # the tier between annotation and fleet
vault: "endpoint=https://vault.vault.svc:8200!, auth=kubernetes, watch-seconds=10"
s3: "agent-memory-limit=128Mi"
webhook:
agentEnvAllow: "payments: HTTPS_PROXY, AWS_*; *: RUST_LOG"
sourceAllow: "" # empty = every store, everywhere
sourceDeny: "sandbox: git"
perStore keys are spelled exactly as the annotations spell them
(watch-seconds, not watchSeconds) — one grammar for the value,
whether it arrives per pod, per store, or per fleet — and they cover
EVERY store-shaped annotation, so a developer can deploy knowing
nothing but inject: "true". The ! above PINS the vault address: a
pod annotating a different endpoint is refused, not silently
corrected. agent.defaults.env needs no allowlist: the installer owns
both the values and the gate.
Helm's schema refuses a malformed value at render time; the webhook
re-validates ALL of it at startup and refuses to serve on a typo — so
an installation written any of the three ways gets the same refusal at
the same door. The readable form for kustomize is
base/installation.yaml, a ConfigMap of the same settings as YAML
(Installation Defaults);
the variables below are the other way, and still work:
# kustomization.yaml, an overlay patch
patches:
- target: { kind: Deployment, name: dynamic-config-webhook }
patch: |
- op: add
path: /spec/template/spec/containers/0/env/-
value: { name: DYNAMIC_CONFIG_AGENT_FILE_MODE, value: "0640" }
agentEnvAllow and the
source gates are security gates,
not defaults. Installation Defaults and Gates
is the full treatment: every knob with its validation, per-store
examples for all nine stores with every field filled, the gates'
semantics and threat model, and the kustomize equivalents.
Values, all of them
The chart README
is the full values reference — naming overrides, common labels,
per-component images and pull policy, service accounts, probe and
rollout tuning, extraEnv/extraVolumes escape hatches, namespace
gating, the operator's RBAC toggle. Everything the templates read is in
that table.
Without helm: kustomize
deploy/kustomize/ carries the same resources as a base — including
installation.yaml, the fleet defaults and gates written as YAML rather
than as environment-variable grammar — plus TLS overlays for
cert-manager, bring-your-own-PEMs via secretGenerator, and the
self-rotating mode.
Its README
is the three-step walkthrough, including the one caBundle patch
kustomize cannot express. The CRDs ship inside the base, drift-gated
against the operator's --crds output like every other copy.
Without a registry: an air-gapped install
./scripts/airgap-bundle.sh 0.3.0 # on the connected side
# move dynamic-config-0.3.0-airgap.tar.gz across
tar xzf dynamic-config-0.3.0-airgap.tar.gz
cd dynamic-config-0.3.0-airgap
./load.sh registry.internal:5000
./verify.sh
kubectl apply --server-side -f crds/
helm install dynamic-config ./chart --values values-airgap.yaml
The bundle carries the chart, the CRDs, the three image indexes and their
signatures and attestations — because the supply-chain work this project
does is worth nothing offline if the signatures are left behind with the
registry. verify.sh runs cosign --offline against the images as they
landed rather than as they were published.
Two things the procedure insists on. load.sh pushes by digest, because
the digests are what the signatures cover; and the CRDs are applied by hand,
because Helm installs crds/ once and never upgrades it — the same step an
upgrade needs on a connected cluster.
The smoke test
The e2e smoke (e2e/smoke.sh) is the install, end to end, against a
kind cluster: the chart in its zero-dependency default, a live Consul,
one annotated pod, the rendered file read back out of it, and the
injected container's security posture asserted on the running pod.
CERT_MANAGER=1 e2e/smoke.sh runs the same flow through the other TLS
mode.
Installation Defaults and Gates
Two kinds of installation-time decisions live in the webhook's configuration, and they must not be confused: defaults, which a pod may always override, and gates, which a pod may never override. Both arrive as chart values, or as a ConfigMap of YAML that kustomize can hand over too, or as environment variables on the webhook Deployment — three spellings of one thing, and every one of them is validated when the webhook starts. A mistyped value stops the install, never the first admission.
defaults: annotation > per-store default > fleet default > built-in
pins: a value marked "!" (or under overridable: "false") refuses a
DIFFERING annotation — the tiers still fill what pods omit
gates: the installer's word is final
Writing them as YAML
An installation reaches the webhook as strings, because that is what an environment variable is — and several of those strings are little grammars:
DYNAMIC_CONFIG_AGENT_STORE_DEFAULTS="vault: overridable=false, endpoint=https://vault:8200; s3: file-mode=0640?"
DYNAMIC_CONFIG_WEBHOOK_SOURCE_ALLOW="payments: vault, s3; *: consul"
Fine to parse and unpleasant to write. So the same settings may be written as maps:
# values.yaml
agent:
defaults:
perStore:
vault:
overridable: false
endpoint: https://vault.vault.svc:8200
auth: kubernetes
watch-seconds: 30
s3:
agent-memory-limit: 128Mi
webhook:
sourceAllow:
payments: [vault, s3]
"*": [consul]
agentEnvAllow:
payments: [HTTPS_PROXY, "AWS_*"]
A map is defined as what it renders to. It travels to the pod as a
mounted ConfigMap and is turned into the grammar there, by the same
parser the string form goes through — so there is one set of rules, one
set of messages, and two spellings that cannot mean different things.
Pins, markers and tiers work the same either way: overridable: false
in a map is the overridable=false in a string.
The webhook reads that document through its own engine: the YAML reader it gives applications is the one it reads its own configuration with.
Kustomize gets the same thing
This is why the document exists at all rather than a chart-side
rendering. A kustomize base has no template engine, so a hand-written
ConfigMap is the only structured form it can hand over —
deploy/kustomize/base/installation.yaml, which ships empty and
commented:
apiVersion: v1
kind: ConfigMap
metadata:
name: dynamic-config-webhook-installation
data:
installation.yaml: |
storeDefaults:
vault:
overridable: false
endpoint: https://vault.vault.svc:8200
sourceAllow:
payments: [vault, s3]
"*": [consul]
Which wins
An environment variable set on the container beats the document: a document is the installation written down once, a variable is somebody standing in front of it for this deployment — the more specific of the two, which is the same rule the configuration layers themselves follow.
A setting the document does not know is refused at startup, with the list of the ones there are. A default silently ignored is a default that never applied, and the first sign of it would be a pod running without the posture somebody thought they had set.
The knob vocabulary
Every defaultable knob is spelled exactly as its annotation is spelled — there is no second vocabulary to learn, and every tier is validated with the same rules as the annotation it stands in for:
| knob | built-in | validation |
|---|---|---|
agent-cpu-request | 10m | a Kubernetes quantity |
agent-memory-request | 32Mi | a Kubernetes quantity |
agent-cpu-limit | none | a Kubernetes quantity |
agent-memory-limit | 64Mi | a Kubernetes quantity |
file-mode | the agent's 0644 | octal, at most 0777, owner-readable |
watch-seconds | 15 | whole seconds |
mode | sidecar | init / sidecar / both |
volume-medium | memory | memory / disk |
native-sidecar | false | "true" / "false" |
agent-run-as-user | 65532 | numeric, 0 refused |
agent-run-as-group | 65532 | numeric, 0 refused |
metrics-port | none | a port; a pod opts out of a default with "0" |
path | none | absolute; the rendered file's location |
source | none | fleet level only (agent.defaults.source): one of the nine stores |
Pod-wide knobs (mode, volume, resources, identity, metrics) resolve
against the DEFAULT render's store; per-render knobs (watch-seconds,
file-mode) resolve against each render's own
store.
And per store, EVERY store-shaped annotation is defaultable — the address, the document, the credentials' Secret names, the auth flags, the templates:
endpoint endpoint-secret key token-secret password-secret
ca-configmap tls-secret ssh-secret aws-secret section
auth auth-mount auth-role auth-username auth-token-path
namespace ref api-url template template-configmap
Each is validated as its annotation would be (*-secret values need
the <secret-name>/<key> slash), and the either-or pairs
(endpoint/endpoint-secret, template/template-configmap) resolve
as a LEVEL: any pod-side answer mutes both installation halves, so a
pod that chose endpoint-secret never inherits a stray plain
endpoint.
With source and path at the fleet tier and endpoint and key in
a store's tier, the minimal pod is a single line of intent:
metadata:
annotations:
dynamic-config.rs/inject: "true"
Defaulting key deserves a sentence of caution: two pods leaning on
the same store default read the SAME document. For a shared,
cluster-wide configuration that is exactly right; for per-app
documents, leave key to the pods.
What stays the pod's alone: env-inject and env-restart (they name
a container only the pod knows), agent-env (gated per pod), and
inject itself — a default that opts workloads in silently is not a
default, it is a surprise.
Pinning: the override mode
Every installation value carries an override rule, and the closest word wins:
- a trailing
!on the value — pinned — or a trailing?— overridable; - no marker: the STORE's own flag, when its
perStoregroup carriesoverridable=true|false; - neither: whatever
agent.defaults.overridablesays ("true"unless set otherwise).
A pinned value REFUSES a differing annotation at admission — never
silently outvotes it, because a value the author wrote and did not
get is a debugging session. The same value restated passes, which
keeps migrations painless. Knobs the installation never set are
untouched by all of this: overridable: "false" pins what you SET,
not the whole contract.
agent:
defaults:
overridable: "false" # everything set below is pinned…
fileMode: "0640" # …like this
watchSeconds: "30?" # …except this one, explicitly opened
perStore:
# A store pins its own group: everything vault-shaped is the
# platform team's word, except the watch interval.
vault: "overridable=false, endpoint=https://vault.vault.svc:8200, auth=kubernetes, watch-seconds=30?"
# And a store can OPEN its group under a strict fleet flag:
# consul values stay the pods' even with the "false" above.
consul: "overridable=true, endpoint=http://consul.infra.svc:8500"
All three rungs compose per store and per value: a ! inside an
overridable=true group pins just that value; a ? inside an
overridable=false group opens just that one. The rule always comes
from the TIER that supplied the value — a store's flag never pins a
fleet default.
The pin follows the pair rule too: a pinned endpoint also refuses a
pod that answers with endpoint-secret — the address is one decision,
and it is not the pod's to make. A pinned fleet source
(source: "consul!") refuses pods that name any other store, which is
the enforcement twin of the sourceAllow gate below.
Per-store defaults, all nine stores
agent.defaults.perStore is the tier between the annotation and the
fleet. One realistic installation, every store present — each line
pairs with the pod that uses it below:
The string form below is one of the two spellings; the same installation as maps reads the same.
# values.yaml — or, for kustomize, the same settings in
# base/installation.yaml, or joined with "; " into
# DYNAMIC_CONFIG_AGENT_STORE_DEFAULTS
agent:
defaults:
watchSeconds: "30" # the fleet's floor
fileMode: "0640"
perStore:
# Local, cheap to poll: tighten the interval — and the address
# is the platform's, PINNED, so pods need not carry it and
# cannot point elsewhere.
consul: "endpoint=http://consul.infra.svc:8500!, watch-seconds=10"
# Secrets: the whole group is the platform team's word
# (address, auth, CA), and the file lands owner-only, owned by
# the app's uid.
vault: "overridable=false, endpoint=https://vault.vault.svc:8200, auth=kubernetes, ca-configmap=vault-ca, file-mode=0400, agent-run-as-user=1000, agent-run-as-group=1000"
# A JVM config server answers slowly; give the render room.
config-server: "watch-seconds=20, agent-memory-limit=96Mi"
# Billed per read: poll gently.
firestore: "watch-seconds=120"
# A clone per render costs memory and remote quota.
git: "watch-seconds=120, agent-memory-limit=128Mi"
# In-memory store, near-free reads.
redis: "watch-seconds=10"
# Watches are pushed by etcd itself; the interval is a backstop.
etcd: "watch-seconds=60"
# JetStream KV is push-cheap too.
nats: "watch-seconds=10"
# A GET per poll is a line on a bill; and S3 documents are often
# the big ones.
s3: "watch-seconds=60, agent-memory-limit=128Mi"
The pods, every field filled. None of them repeats a knob the installation already set — the tier exists so they never have to:
# consul — plain HTTP, a KV key
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "consul"
dynamic-config.rs/endpoint: "http://consul.infra.svc:8500"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/path: "/config/rendered.toml"
# watch-seconds arrives from perStore.consul: 10
# vault — kubernetes auth through the pod's own ServiceAccount,
# a private CA, one section of the secret
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "vault"
dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
dynamic-config.rs/key: "secret/myapp"
dynamic-config.rs/section: "db"
dynamic-config.rs/auth: "kubernetes"
dynamic-config.rs/auth-role: "myapp"
dynamic-config.rs/ca-configmap: "vault-ca"
dynamic-config.rs/path: "/config/rendered.yaml"
# file-mode 0400 and uid/gid 1000 arrive from perStore.vault
# config-server — the Spring-style application/profile pair as the key
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "config-server"
dynamic-config.rs/endpoint: "http://config-server.infra.svc:8888"
dynamic-config.rs/key: "billing/prod"
dynamic-config.rs/path: "/config/rendered.json"
# watch-seconds 20 and the 96Mi limit arrive from perStore
# firestore — the endpoint is the GCP project, the key a document path
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "firestore"
dynamic-config.rs/endpoint: "acme-prod"
dynamic-config.rs/key: "config/billing"
dynamic-config.rs/path: "/config/rendered.json"
# watch-seconds 120 arrives from perStore.firestore
# git — a repository over ssh, a ref, a file inside the tree
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "git"
dynamic-config.rs/endpoint: "git@github.com:acme/config.git"
dynamic-config.rs/ref: "main"
dynamic-config.rs/key: "billing/prod.yaml"
dynamic-config.rs/ssh-secret: "config-deploy-key"
dynamic-config.rs/path: "/config/rendered.yaml"
# watch-seconds 120 and the 128Mi limit arrive from perStore.git
# redis — the password rides in the URL, so the WHOLE endpoint is a
# Secret instead of an annotation
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "redis"
dynamic-config.rs/endpoint-secret: "redis-cred/url"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/path: "/config/rendered.toml"
# watch-seconds 10 arrives from perStore.redis
# etcd — a REQUIRED client certificate and a private CA (the
# password-auth twin swaps tls-secret for password-secret)
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "etcd"
dynamic-config.rs/endpoint: "https://etcd.infra.svc:2379"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/tls-secret: "etcd-client-tls"
dynamic-config.rs/ca-configmap: "etcd-ca"
dynamic-config.rs/path: "/config/rendered.toml"
# watch-seconds 60 arrives from perStore.etcd
# nats — a JetStream KV bucket and the key inside it
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "nats"
dynamic-config.rs/endpoint: "nats://nats.infra.svc:4222"
dynamic-config.rs/key: "config/db.json"
dynamic-config.rs/path: "/config/rendered.toml"
# watch-seconds 10 arrives from perStore.nats
# s3 — the endpoint IS the bucket; api-url points at MinIO/Ceph/R2,
# and aws-secret carries static credentials those servers need
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "s3"
dynamic-config.rs/endpoint: "myapp-config"
dynamic-config.rs/key: "prod/db.json"
dynamic-config.rs/api-url: "http://minio.infra.svc:9000"
dynamic-config.rs/aws-secret: "minio-cred"
dynamic-config.rs/path: "/config/rendered.toml"
# watch-seconds 60 and the 128Mi limit arrive from perStore.s3
Any pod that DOES set watch-seconds (or any other knob) still wins —
the tiers only answer for what the pod left unsaid.
The gates, in depth
Three gates, one authority model: they live in the webhook's own configuration, owned by whoever installs it. None of them is a namespace annotation — whoever edits a namespace is usually the tenant being gated, and a gate its subject can open is not a gate. None of them reads the namespace object either: the webhook holds no RBAC and asks the API server for nothing; the pod's namespace arrives inside the AdmissionReview, and the ruling comes from configuration alone.
All three share one grammar:
spec = group *( ";" group )
group = [ namespace ":" ] names ; no head, or "*" = every namespace
names = name *( "," name )
agent-env allow — closed until opened
webhook:
agentEnvAllow: "payments: HTTPS_PROXY, AWS_*; *: RUST_LOG"
agent-env puts environment on the container that holds store
credentials, and environment steers SDKs. Concretely, on this agent:
HTTPS_PROXY reroutes every vault/S3/consul request through a proxy
of the pod author's choosing — with the bearer tokens and signatures
inside; AWS_CA_BUNDLE and SSL_CERT_FILE swap the trust roots those
connections verify against; AWS_EC2_METADATA_DISABLED,
AWS_PROFILE, NO_PROXY all change where credentials come from or
where traffic goes. That is why this gate defaults to closed: an
empty allowlist refuses the annotation everywhere, and every name a
pod wants must be opened by the installer, optionally per namespace.
Name rules: UPPER_SNAKE, exact match, or a trailing * as a prefix
glob (AWS_*). A bare * opens everything — legitimate on a
single-team cluster, a finding on a shared one. The refusal a pod
sees names the variable, the namespace, what IS allowed there, and
the chart value that opens the gate; it does not enumerate other
namespaces' rules — one tenant's refusal must not describe another's
setup.
What the gate does NOT govern: agent-env values (only names), the
app container's environment (the agent's only), and the fleet's own
agent.defaults.env (next section).
The fleet environment — no gate, on purpose
agent:
defaults:
env: "HTTPS_PROXY=http://egress.infra.svc:3128, RUST_LOG=info"
agent.defaults.env is environment EVERY injected agent gets — the
cluster-wide egress proxy, a fleet log level. It passes no allowlist
because the installer sets both the values and the allowlist; a gate
you hold both sides of checks nothing. Merging rules, exactly:
- A pod's own
agent-envoverrides a fleet name — the pod said it more specifically. (The pod's name still needs the allowlist: the OVERRIDE is a pod-author action even when the name is fleet-known.) - When a pod uses
aws-secret, the fleet'sAWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEYstep aside — one credential, one place, and the pod named its Secret. - Everything else rides along verbatim, on every render's agent.
Source allow and deny — open until narrowed
webhook:
sourceAllow: "payments: vault, s3; *: consul" # empty = every store
sourceDeny: "sandbox: git" # subtractive, wins
The source gates decide which STORES a namespace may render from. Two lists, because the two postures are different jobs:
sourceAllow— empty means*: every store, everywhere, so an upgrade changes nothing until the installer says so. Non-empty flips the posture: ONLY the listed sources pass, per namespace. Use it when the store set is policy: payments reads vault and s3, everyone may read consul, and anything unlisted is refused with a message naming what IS allowed there.sourceDeny— always subtractive, and it outranks the allowlist: a source both listed and denied is denied. Use it for the surgical cut that does not flip the posture: git is off in the sandbox namespace, everything else stays open.
Both gates are judged against EVERY render on the pod — the default
one and each named suffix:
a denied store cannot ride in as source.cache. Entries are
validated against the real store list at webhook startup, so
sourceDeny: "sandbox: got" fails the install instead of silently
gating nothing — in a security control, a typo that fails open is the
worst of the four outcomes.
Deciding between them:
| you want | use |
|---|---|
| nothing changes on upgrade | leave both empty |
| this namespace uses exactly these stores | sourceAllow |
| this store is banned here, rest stays open | sourceDeny |
| allow broadly, carve exceptions | both — deny wins on overlap |
Kustomize, same doors
Every value above is either a key in
base/installation.yaml — the
readable form, and usually what you want — or one environment variable
on the webhook Deployment. The chart is convenience, not capability:
# kustomization.yaml, an overlay patch
patches:
- target: { kind: Deployment, name: dynamic-config-webhook }
patch: |
- op: add
path: /spec/template/spec/containers/0/env/-
value:
name: DYNAMIC_CONFIG_AGENT_STORE_DEFAULTS
value: "vault: file-mode=0400; git: watch-seconds=120"
- op: add
path: /spec/template/spec/containers/0/env/-
value:
name: DYNAMIC_CONFIG_WEBHOOK_SOURCE_DENY
value: "sandbox: git"
The chart's schema cannot see a kustomize patch — which is exactly why the webhook re-validates the complete installation at startup and refuses to serve on any error. Helm users get two doors; kustomize users get the one that matters.
| chart value | environment variable |
|---|---|
agent.defaults.cpuRequest … .memoryLimit | DYNAMIC_CONFIG_AGENT_CPU_REQUEST … _MEMORY_LIMIT |
agent.defaults.fileMode | DYNAMIC_CONFIG_AGENT_FILE_MODE |
agent.defaults.watchSeconds | DYNAMIC_CONFIG_AGENT_WATCH_SECONDS |
agent.defaults.mode | DYNAMIC_CONFIG_AGENT_MODE |
agent.defaults.volumeMedium | DYNAMIC_CONFIG_AGENT_VOLUME_MEDIUM |
agent.defaults.nativeSidecar | DYNAMIC_CONFIG_AGENT_NATIVE_SIDECAR |
agent.defaults.runAsUser / .runAsGroup | DYNAMIC_CONFIG_AGENT_RUN_AS_USER / _GROUP |
agent.defaults.metricsPort | DYNAMIC_CONFIG_AGENT_METRICS_PORT |
agent.defaults.env | DYNAMIC_CONFIG_AGENT_ENV |
agent.defaults.source | DYNAMIC_CONFIG_AGENT_SOURCE |
agent.defaults.path | DYNAMIC_CONFIG_AGENT_PATH |
agent.defaults.overridable | DYNAMIC_CONFIG_AGENT_DEFAULTS_OVERRIDABLE |
agent.defaults.perStore | DYNAMIC_CONFIG_AGENT_STORE_DEFAULTS |
webhook.agentEnvAllow | DYNAMIC_CONFIG_WEBHOOK_AGENT_ENV_ALLOW |
webhook.sourceAllow | DYNAMIC_CONFIG_WEBHOOK_SOURCE_ALLOW |
webhook.sourceDeny | DYNAMIC_CONFIG_WEBHOOK_SOURCE_DENY |
The Annotation Contract
v1, and it is the API: a change here is a breaking change of this integration, whatever the binaries think. Two golden files in the webhook's test suite byte-compare the full admission response, so the contract cannot move without a reviewed diff saying it moved.
The core seven
| annotation | required | meaning |
|---|---|---|
dynamic-config.rs/inject | yes | "true" asks; anything else but "false" fails the admission |
dynamic-config.rs/source | yes | consul, vault, config-server, firestore, git, redis, etcd, nats, s3 |
dynamic-config.rs/endpoint | one of the two | the store's address — a url; <project>[/<database>] for firestore |
dynamic-config.rs/endpoint-secret | one of the two | <secret>/<key> holding the address, when the address carries a password (a redis url) |
dynamic-config.rs/key | yes | the document's key; mount/path for vault, application/profile for config-server, the file's path for git |
dynamic-config.rs/path | yes | where the rendered file lands; the extension picks the format |
dynamic-config.rs/mode | no | init, sidecar (default), both |
Naming a store once
dynamic-config.rs/class: "team-vault"
dynamic-config.rs/key: "billing/config.json"
dynamic-config.rs/path: "/config/app.yaml"
A DynamicConfigClass — the object the operator has read since 0.1.1 —
holds the source, the endpoint and the credential's Secret. A pod that names
one says only what is its own.
The class supplies defaults: a pod that names both a class and its own
endpoint keeps its own, which is the rule the installation document
already follows. Its own namespace is looked in first, then the cluster
scope, so a namespace can override a platform default without asking the
platform.
Off unless an administrator turns it on (webhook.classes.enabled), and
the reason is worth reading. The admission path calls the API server
nowhere — that is what keeps a busy API server from failing every pod
creation in the cluster — and enabling this does not change it: the classes
are listed on a background timer into a map in memory, and admission reads
the map. What it does change is that the webhook holds a credential and a
cluster-wide read of two custom resources, which is worth having only if
pods actually name classes.
The cost is a synchronisation delay a mounted ConfigMap already has. A class created seconds ago may not have been polled yet, and a pod naming it is refused with that sentence rather than admitted without the store the class was to supply.
Two refusals are about who may use what:
- a
ClusterDynamicConfigClasswhosenamespaceslist does not include the pod's — the list is what keeps cluster-scoped from meaning anyone - a cluster class whose credential Secret lives in another namespace. A pod
mounts Secrets from its own namespace and nowhere else, so that class is
usable by a
DynamicConfigRender, which reads the Secret itself, and not by an injected agent
The namespace gate, before any pod is read
Every key below is per-pod. One optional guard sits a level
above: with webhook.namespaceGating=true the webhook selects only
namespaces labeled dynamic-config.rs/injection: enabled — a
label, not an annotation, because the gate lives in the webhook
configuration's namespaceSelector and Kubernetes selectors cannot
see annotations. Inside a gated namespace the per-pod
dynamic-config.rs/inject: "true" is still required; the
security page owns the trade
(blast radius, per-namespace failurePolicy: Fail), and
examples/namespace-gating.yaml
is the ready-to-apply shape.
What the webhook writes back
| annotation | value | meaning |
|---|---|---|
dynamic-config.rs/status | injected | this pod has been through admission and carries the agent |
Written by the webhook, never by a pod. A mutating webhook is not
called once — reinvocationPolicy: IfNeeded asks the API server to call
it again whenever a later webhook changes the pod, and some controllers
resubmit a spec that has already been admitted. A marked pod is passed
through untouched; without the mark, the second pass would add the agent
again, and two containers with one name is a pod the API server refuses.
A pod that sets this annotation to anything else is refused, with a
message saying so. Setting it to injected by hand is a way of saying
"do not inject me", which inject: "false" already says more clearly.
Behaviour
| annotation | default | meaning |
|---|---|---|
dynamic-config.rs/timeout | the store's own, 10s | the deadline for one fetch attempt, where the store's client has a door for it. Ten seconds is right for a store on this network and wrong for a Git remote across a WAN or a bucket in another region — until this, a workload in that position had no way to say so and its only recourse was a fetch that kept timing out |
dynamic-config.rs/agent-image | the installation's | the image for this pod's agent. How an agent upgrade stops being all-or-nothing: one Deployment tries it first. Refused unless the installation lists a prefix it starts with — an image on the injected container runs chosen code beside the application, holding the store's credential |
dynamic-config.rs/watch-seconds | 15 | how often the sidecar asks a store that must be asked, and how often it re-reads one that pushes; whole seconds. See Watching |
dynamic-config.rs/section | whole document | the section key the document nests under |
dynamic-config.rs/native-sidecar | "false" | "true" injects the watcher as an init container with restartPolicy: Always (Kubernetes 1.29+); Jobs finish |
dynamic-config.rs/volume-medium | memory | where the rendered file lives: memory (tmpfs, off the node's disk) or disk |
dynamic-config.rs/init-first | "false" | "true" puts the injected init container ahead of the pod's own, for a pod whose init container reads the rendered file. Appending stays the default: another injector's init container that must run first keeps running first. Refused with mode: sidecar, where there is no init container |
dynamic-config.rs/agent-run-as-same-user | "false" | take the agent's UID from the application container instead of naming one, so the rendered file's owner matches without two numbers to keep in step. The application's own runAsUser first, then the pod's; absent is refused rather than guessed — inheriting whatever the image runs as is a UID that moves when the image does. Root is refused, as always |
dynamic-config.rs/extra-secret | none | a Secret mounted read-only under /etc/dynamic-config/extra, into the agent alone. For the file a store's own client wants that this contract does not model. The path is fixed rather than chosen: an annotation that took a mount path could be aimed at the rendered volume or the service-account token |
dynamic-config.rs/inject-containers | every container | a comma-separated list of the pod's own containers that receive the rendered volume. The default matches the reference implementation's; naming a subset is for the pod that runs a log shipper or a mesh proxy beside its application — neither has any business holding a rendered credential, and file-mode cannot draw that line because a sidecar usually runs as the same UID. A name the pod does not have is refused, and leaving out the container that env-inject wraps is refused too |
dynamic-config.rs/agent-cpu-request | 10m | the injected container's CPU request |
dynamic-config.rs/agent-memory-request | 32Mi | its memory request |
dynamic-config.rs/agent-cpu-limit | none | its CPU limit — none by default, on purpose |
dynamic-config.rs/agent-ephemeral-request | none | ephemeral-storage request. Matters for exactly one configuration and is unbounded without it: on volume-medium: disk the rendered volume is node storage rather than the pod's memory, history keeps copies beside it, and nothing else here declares that resource — a pod could fill a node's disk without exceeding a limit it had declared |
dynamic-config.rs/agent-ephemeral-limit | none | its ephemeral-storage limit. A request larger than it is refused, because the scheduler's own refusal would name the pod rather than these |
dynamic-config.rs/agent-memory-limit | 64Mi | its memory limit |
dynamic-config.rs/file-mode | umask's answer (0644) | the rendered file's octal permissions, e.g. "0640" — set on the scratch file before the atomic rename, so a reader never sees the final path in a mode it will not keep |
dynamic-config.rs/agent-run-as-user | 65532 | the injected container's UID, so the rendered file's owner matches what the app runs as; 0 is refused — the agent stays nonroot in every configuration |
dynamic-config.rs/agent-run-as-group | 65532 | its GID, same rule |
dynamic-config.rs/aws-secret | none | s3 only: a Secret whose AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY keys become exactly those variables on the agent — static credentials for S3-compatibles that are not AWS (MinIO, Ceph, R2); on AWS, IRSA needs none of it |
dynamic-config.rs/metrics-port | none | the agent serves Prometheus text on this port (the series); only meaningful with a watching agent, and "0" opts out of an installation-wide default |
dynamic-config.rs/agent-env | none | comma-separated NAME=value pairs, set as environment on every injected agent — SDK knobs like RUST_LOG, AWS_CA_BUNDLE, proxy variables. Gated: only names the installation's allowlist permits in this namespace pass admission (the gate) |
dynamic-config.rs/env-inject | none | a container name: its command is wrapped in set -a; . <path>; set +a; exec …, so the rendered dotenv is the process's REAL environment. Needs mode: init or both (env freezes at container start — Kubernetes' rule) and an explicit command (an ENTRYPOINT is invisible to the webhook); both refusals name the fix |
dynamic-config.rs/env-restart | "false" | with env-inject and mode: both: when the sidecar re-renders the dotenv, the kubelet restarts JUST the app container (a liveness probe compares the file against the fingerprint the wrapper exported at start) and the wrapper re-sources the new file — the closest thing to a live env update the kernel permits: seconds, no pod recreation, no new IP. Refused if the container already has its own livenessProbe |
dynamic-config.rs/template | none | an inline minijinja template; it owns the output bytes |
dynamic-config.rs/template-configmap | none | <name> or <name>/<key> (default key template): the template from a ConfigMap, mounted read-only and re-read every render |
Freshness, and what happens when a store is not there
The store is not always reachable, and the document does not always still exist. Five annotations say what the agent does about it; every default is the one that keeps a running application running.
| annotation | default | meaning |
|---|---|---|
dynamic-config.rs/startup-policy | allow-cached | what a first fetch failure means. allow-cached serves the file already on the volume if it parses — the rendered volume survives a container restart, so this is a real cache rather than a hopeful one — and goes on watching. require-fresh refuses to start without a fresh document, for a pod that must never come up on stale credentials. best-effort starts regardless |
dynamic-config.rs/max-staleness | none | how old the document may get before /readyz reports 503 — "6h", "90s". Last-known-good answers is there a document; this answers is it still worth trusting. A credential may be worthless after five minutes and a feature flag fine after a day, which is why there is no default |
dynamic-config.rs/on-delete | retain | what a document disappearing from the store means. retain keeps serving the last render, remove truncates the file so a consumer reads nothing rather than a revoked secret, fail ends the agent so the pod's restart policy takes over. Whichever is chosen it is reported: dynamic_config_agent_absent moves and the log says whether the store answered gone or did not answer at all |
dynamic-config.rs/require-ack | "false" | readiness waits for the application to POST the fingerprint it is running, not merely for a document to exist. What it is for |
dynamic-config.rs/history | none | keep the last N replaced generations beside the render. What it is for |
dynamic-config.rs/readiness | "true" | whether the webhook attaches a readiness probe to the injected container. Pod readiness is AND-ed across containers, so a Service sends no traffic to a pod whose configuration has not arrived. "false" opts out, for a deployment that would rather start |
dynamic-config.rs/max-document-bytes | 8388608 | the largest document the agent will accept, checked before it is parsed. The injected container's memory limit is 64Mi by default and a document is held more than once while it is resolved |
Dynamic secrets
For a Vault path that mints a credential rather than storing one —
database/creds/…, pki/issue/…, aws/creds/….
| annotation | default | meaning |
|---|---|---|
dynamic-config.rs/dynamic | "false" | read the path as a dynamic engine: no KV data/ nesting, and keep the lease Vault answered with. A renewable lease is renewed at 65% of its TTL, spread; one Vault marked renewable: false — every pki/issue — is never sent a renewal at all and is re-issued at 90% instead. A renewal keeps the file; a re-issue is new credentials, so it re-renders. For a certificate, whichever expires first binds: the lease, or the certificate's own notAfter. Vault has the whole of it |
dynamic-config.rs/revoke-grace | 5s | how long the agent may spend handing the lease back. Refused past the pod's terminationGracePeriodSeconds — beyond that the kubelet sends SIGKILL, the revocation is cut off mid-request, and the lease stays out anyway, which is what the annotation was set to avoid |
dynamic-config.rs/revoke-on-shutdown | "true" | give the lease back on SIGTERM. Best-effort with a short deadline: a pod that cannot reach Vault while terminating still terminates. Setting it without dynamic is refused — there is no lease to revoke |
Vault revokes an expired lease on its own eventually; revoking on the way out is what makes eventually into now. See Vault.
Reaching the store over TLS
Two settings beside ca-configmap and tls-secret, and they are not the
same kind of thing: one moves which name is checked, the other stops
checking.
| annotation | default | meaning |
|---|---|---|
dynamic-config.rs/tls-server-name | the endpoint's host | the name the store's certificate must carry, for an endpoint written as an address it does not name — a Service's cluster IP, a load balancer, a NodePort. The server stays authenticated: the certificate still chains to a trusted authority and still has to carry this name |
dynamic-config.rs/tls-skip-verify | "false" | connect without authenticating the store at all |
dynamic-config.rs/tls-reload | "true" | rebuild the store's client when ca-configmap, tls-secret or ssh-secret changes on disk — without restarting the pod. "false" keeps the old client until something restarts it |
Not every store can express either. The clients differ, and a store that cannot say something refuses the whole configuration by name rather than ignoring it:
| store | tls-server-name | tls-skip-verify |
|---|---|---|
| vault, consul, firestore | no | yes |
| config-server | yes | yes |
| git | no | yes |
| etcd | yes | no |
| redis, nats, s3 | no | no |
The two columns are almost disjoint, and that is the clients rather than a
choice: ureq can turn verification off and cannot override a name,
tonic can override a name and cannot turn verification off.
What tls-skip-verify costs
It is not a weaker TLS. It is TLS without the part that makes it mean anything: any party on the network path can present any certificate, read what is sent, and rewrite what comes back — and what comes back from a store is the configuration and the credentials this pod is about to run on.
So it is the one annotation in this contract a workload cannot reach on its own. Four things gate it:
- an administrator must set
webhook.allowTlsSkipVerify: true, or the annotation is refused with that value named - it is refused alongside
ca-configmap: naming an authority and then not checking it are two answers to one question - every pod using it earns an admission warning, which
kubectlprints - the agent logs it at start and reports
dynamic_config_agent_tls_verification_skipped 1, so one alert finds every pod doing it
Rotation, without a restart
The kubelet rewrites a mounted ConfigMap or Secret in place: the file the container sees changes and nothing tells the process. Every store client in this family reads its trust material once, when it builds — so before 0.3.0 a rotated CA meant a pod restart, and a rotation nobody restarted for meant a store that stopped answering at a moment unrelated to the rotation.
The agent now watches the files it was given and rebuilds the client when
they move. A rebuild is not a restart: the process stays up, the
rendered file never leaves the volume, the last-known-good stays, and every
counter keeps counting — only the client is new.
dynamic_config_agent_tls_reloads_total counts them, and a fleet where that
stays at zero through a CA rotation is a fleet that did not notice.
A rotation writes more than one file, so a change is confirmed by reading twice a quarter-second apart: rebuilding a TLS client from a half-written certificate and key is a failure that reads as a bad certificate.
Tokens needed none of this and never did — the projected service-account token and the config server's bearer file are re-read on every use, so the rotation that actually happens hourly was always picked up.
Try ca-configmap first. The two situations people usually reach for
skip-verify in — a development server with a self-signed certificate, an
enterprise private CA — are both one more certificate to trust, which is one
annotation and keeps the server authenticated.
Telling the application what it is running
| annotation | default | meaning |
|---|---|---|
dynamic-config.rs/meta | "false" | write a sibling .<name>.meta beside the render — /config/app.yaml gets /config/.app.yaml.meta — holding the digest of the bytes, the store's own revision, and when it landed. It describes the render and never contains it: no values, ever |
dynamic-config.rs/schema-configmap | none | <name> or <name>/<key> (default key schema.json): a JSON Schema the resolved document must satisfy before it is published. A document that fails is refused and the last good one keeps serving, which is the behaviour a consumer that is not Rust, Python or Node cannot get any other way |
Telling the application the file moved
dynamic-config.rs/notify-http: "http://127.0.0.1:8080/-/reload"
The rename is atomic, so a consumer never sees half a document — but nobody tells it the document changed. nginx, Prometheus and most legacy daemons reload on a request and on nothing else.
After the rename, never before: the whole promise of the notification is
that the document is already there when it arrives. One attempt, a two
second deadline, and never fatal — the file is already correct, so a
notification that did not land has undone nothing. notifications_total and
notification_failures_total say how it went.
Localhost only, by construction. http://127.0.0.1, http://localhost
or http://[::1], and nothing else — an agent that will POST to an
arbitrary URL is an SSRF primitive holding this pod's store credential. The
address is checked at admission and in the agent, because the two run in
different places. Refused with mode: init, where the container writes once
and exits before the application it would notify has started.
There is no signal form. Signalling a sibling container needs
shareProcessNamespace: true — a pod-wide change to the process boundary
between containers, which the webhook will not make on a workload's behalf.
Failures where somebody is already looking
dynamic-config.rs/events: "true"
kubectl describe pod is where an operator looks first, and a render
failure was not there — it was in the sidecar's log, one
kubectl logs -c dynamic-config-agent away from the question. With this the
agent writes a Warning Event on its own pod for a render that failed
(RenderFailed) and for a document that vanished from the store
(DocumentAbsent).
Off, and twice opt-in. The sidecar carries no API credential in any
other configuration — that is a property the rest of this design leans on —
so writing Events means mounting the pod's service-account token beside the
application. An administrator has to create the Role (agent.events.enabled
with the namespaces, which grants create on events and nothing else) and
set webhook.allowEvents, and only then may a pod ask. Asking without the
installation offering it is refused, with the chart value named, rather than
admitted into a 403 at the first failure.
An Event is commentary on work that already happened, so it is written off the loop: a render never waits on the API server, and an Event that could not be written is logged once and dropped.
Some of the fleet before all of it
dynamic-config.rs/canary-configmap: "rollout" # or "rollout/percent"
kubectl create configmap rollout --from-literal=percent=5
# look at the five per cent
kubectl patch configmap rollout --type merge -p '{"data":{"percent":"100"}}'
A change published to every pod at once either works or is an incident. This buys the third outcome: a few pods take it, somebody looks, and the rest follow or do not.
Which pods is the pod's own name hashed into a bucket from 0 to 99, so the cohort is deterministic and stable — a cohort that reshuffled as it widened would put every pod through the new document eventually and prove nothing about any of them. No coordination and no leader: five thousand agents each answer for themselves.
Who widens it is whoever edits the ConfigMap. That is why the percentage is a mounted file and not an annotation: the kubelet rewrites a mount in place, so the cohort grows with no pod restart — and a restart would discard the very state a canary exists to watch. A pod outside the cohort holds what it fetched and publishes it the moment the number passes its bucket, so nothing has to be re-fetched and the store need not say anything again.
0 holds everybody and 100 holds nobody; both ends read as what they
mean. A file that is missing or does not parse is no canary at all
rather than zero — a typo must not freeze the fleet on its current
document.
Two series say what is happening: canary_holding is 1 while this pod is
outside the cohort, and canary_percent is the number it last read.
Together with applied from below they are the question somebody has to
answer before widening: did the pods that took it actually run it?
Refused with mode: init, where the agent publishes once and exits.
One interaction to know. A held pod's document is by definition not the
newest, so max-staleness will eventually call it stale. That is correct —
it is stale — but a long-running canary under a short staleness ceiling
will make held pods unready, and the two numbers should be chosen together.
Knowing the application actually applied it
dynamic-config.rs/meta: "true" # so the app can read the digest
dynamic-config.rs/require-ack: "true" # and readiness waits for it
renders_total says a document reached disk. It says nothing about whether
the application read it, and an application still running the previous one
while every dashboard reports success is the outage this closes.
The application POSTs the fingerprint it is running to /applied on the
metrics port:
curl -fsS -XPOST --data-binary @- http://127.0.0.1:9110/applied <<< "$(jq -r .fingerprint /config/.app.yaml.meta)"
200 means that is the current document, 409 means a different one is
published and names it — a restarted application acknowledges what it read
before the restart, and a slow one acknowledges a generation the store has
already replaced. Neither is an error the application caused, and neither is
convergence.
A Rust, Python or Node application does not need the meta file: fingerprint()
is the same string, from the same digest.
Three series come out of it, and the third is the one to alert on:
| series | what it says |
|---|---|
acks_total | acknowledgements received |
ack_mismatches_total | acknowledgements naming a document this agent never published. Climbing steadily means acknowledgements and renders are talking past each other |
applied | 1 when the application is running what was published |
unapplied_seconds | how long the published document has gone unacknowledged |
require-ack makes that readiness. The pod stays 0/2 until the
application says it applied the document, so a Service sends no traffic to a
pod running configuration nobody confirmed. Off by default and refused
without a readiness probe or on mode: init: it needs the application's
cooperation, and one that never acknowledges would never become ready.
An application that never acknowledges is not penalised. It leaves applied
at zero and nothing else changes.
What the file was before
dynamic-config.rs/history: "3"
The rename that publishes a new document is the same rename that destroys
the old one, so what was the file before? — the question an incident
starts with — has no answer by the time anybody asks it. This keeps the
replaced generation beside the render, under
/config/.app.yaml.history/<when>-<digest>.yaml, newest kept and oldest
pruned past the count.
At most ten, and off unless asked: the rendered volume is the pod's own memory by default, so every kept generation is charged to a limit that is 64Mi by default. Each copy takes the mode the render had, so a history entry cannot be read by anything the document itself could not be.
Refused with volume-medium: disk, where a replaced secret would sit on
node-backed storage, outlive the pod that held it, and survive a reboot.
This is not a rollback. Putting an old document back needs the application to say the new one is bad, and nothing here can hear that — half a feature that implies the other half is worse than neither. This keeps files; a person reads them.
When something else writes to the rendered file
dynamic-config.rs/on-drift: "repair" # or warn (default), fail
The agent owns the file; the volume is shared. A debug session, an init container with an opinion, an application that rewrites its own configuration — any of them can change it, and until the next change arrives from the store, nothing notices.
Checked on the same cadence the store is read. warn is the default because
the agent cannot know whether the write was a mistake or the point; what it
can do is stop the difference being invisible. repair writes the rendered
document back, for a file whose contents are the store's and nobody else's.
fail ends the agent, so the pod restarts and renders again.
Either way dynamic_config_agent_drift goes to 1 and drift_total counts
it. An unreadable file counts as drifted too: one that has been deleted, or
has become a directory, is not the file this agent wrote either.
Several files, one fetch
also.<name> cuts more files from the same fetched document, and
publishes them all or none of them:
dynamic-config.rs/source: "vault"
dynamic-config.rs/key: "secret/myapp"
dynamic-config.rs/path: "/config/app.yaml"
# the same document, a section of it, as a second file:
dynamic-config.rs/also.db: "/config/db.env"
dynamic-config.rs/also-section.db: "database"
One fetch, one generation, one refusal if any of them cannot be written — so a failure in the third does not leave the first two published. Each write is still its own atomic rename: a reader can catch the microseconds between two of them, which is a rename apart rather than a fetch apart.
This is within one document. Several stores cannot share a generation — two stores have no common instant and no protocol between them can say "these two reads are the same" — so a second store stays a named render, below.
Several documents, one pod
Every store-shaped key accepts a .<name> suffix, and each name is
one more injected agent writing one more file into the same directory:
# the default render:
dynamic-config.rs/source: "vault"
dynamic-config.rs/key: "secret/myapp"
dynamic-config.rs/path: "/config/app.yaml"
dynamic-config.rs/auth: "kubernetes"
# a second, named `cache`:
dynamic-config.rs/source.cache: "redis"
dynamic-config.rs/endpoint-secret.cache: "redis-url/url"
dynamic-config.rs/key.cache: "myapp/cache.json"
dynamic-config.rs/path.cache: "/config/cache.toml"
Per-name: everything a store needs — source, endpoint(+secret), key,
path, section, watch cadence, every auth key, CA/TLS/ssh material
(mounted under suffixed paths), templates, file-mode. Pod-wide, on
purpose: mode, volume medium, resources, run-as identity,
env-inject (the default render's file is the one that can become the
environment). Two rules, refused with the fix named: every path lives
in the default path's directory (one shared volume), and each
name's source/key/path are required exactly like the default's.
Container names follow the suffix — dynamic-config-agent-cache in
kubectl get pod, so a broken render says which one it is.
Authentication
Each store takes its own methods; the store's page spells every one of them out with full manifests. The webhook forwards these to the agent verbatim, and the agent refuses a wrong combination at startup — in the pod's events, not as a store error twenty minutes later.
| annotation | meaning |
|---|---|
dynamic-config.rs/auth | the method: consul token | kubernetes | jwt; vault token | kubernetes | approle | jwt | userpass | ldap | cert; firestore metadata-server | access-token | emulator; git anonymous | token | ssh-key |
dynamic-config.rs/auth-mount | vault: the auth method's mount path when not the default; consul: the auth method's name (required for kubernetes and jwt) |
dynamic-config.rs/auth-role | vault kubernetes: the role to assume (required); vault approle: the role id; vault jwt/cert: optional |
dynamic-config.rs/auth-username | vault userpass/ldap: the user; git: the basic-auth user when the host wants one |
dynamic-config.rs/auth-token-path | where the service-account token is mounted, when a projected volume moved it |
dynamic-config.rs/namespace | the Vault namespace (Vault Enterprise) |
dynamic-config.rs/ref | git: main, branch:main, tag:v1.4, or commit:<sha> |
dynamic-config.rs/api-url | firestore: the API endpoint when it is not Google's — the emulator |
Secrets and certificates
Secret material never rides an annotation — kubectl describe pod
prints annotations and arguments to anyone with pod read access. These
four name Kubernetes objects instead; the webhook mounts them and the
agent reads them. The geography is fixed.
| annotation | form | becomes |
|---|---|---|
dynamic-config.rs/token-secret | <secret>/<key> | env DYNAMIC_CONFIG_AGENT_TOKEN |
dynamic-config.rs/password-secret | <secret>/<key> | env DYNAMIC_CONFIG_AGENT_PASSWORD — the approle secret id, the userpass/ldap password |
dynamic-config.rs/endpoint-secret | <secret>/<key> | env DYNAMIC_CONFIG_AGENT_ENDPOINT |
dynamic-config.rs/ca-configmap | <name> or <name>/<key> (default key ca.crt) | a read-only mount under /etc/dynamic-config/ca and the agent's --ca |
dynamic-config.rs/tls-secret | <name> — a kubernetes.io/tls Secret | a read-only mount under /etc/dynamic-config/tls and --tls-cert/--tls-key (that Secret type fixed its two keys as tls.crt/tls.key) |
dynamic-config.rs/ssh-secret | <name> or <name>/<key> (default key ssh-privatekey, the kubernetes.io/ssh-auth convention) | a 0400 mount under /etc/dynamic-config/ssh, --ssh-key, and auth: ssh-key implied when no auth was named |
Value forms, source by source
What endpoint, key and auth take for each value of source — the
store pages carry the full manifests, this table is the lookup:
source | endpoint | key | auth values |
|---|---|---|---|
consul | http(s)://host:8500 | KV path with extension: myapp/config.json | (none), token, kubernetes, jwt |
vault | http(s)://host:8200 | <mount>/<path>: secret/myapp | (none = token), token, kubernetes, approle, jwt, userpass, ldap, cert |
config-server | http(s)://host:8888 | <application>/<profile>: billing/prod | (none — bearer via token-secret only) |
firestore | <project> or <project>/<database>: acme-prod | collection/document: config/billing | (none = metadata-server), metadata-server, access-token, emulator |
git | any clone url: https://…, git@host:org/repo.git | file path in the repository: billing/prod.yaml | (none = token if set, else anonymous), anonymous, token, ssh, ssh-key |
redis | redis:// / rediss:// url — via endpoint-secret when it carries a password | key with extension: myapp/config.json | (none — credentials live in the url) |
etcd | (no auth key) | tls-secret client certificates, or auth-username + password-secret — etcd's own two methods, both first-class | --key is the etcd key |
nats | (no auth key) | a .creds file via auth-token-path, or token-secret; anonymous otherwise | --key is <bucket>/<key> |
s3 | (no auth key) | the ambient AWS chain — IRSA on EKS, the workload's own identity | --endpoint is the bucket; api-url overrides for MinIO/Ceph/R2 |
The prefix is claimed territory
Every dynamic-config.rs/* annotation must be a key this page lists —
an unknown one fails the admission. The rule exists for the typo:
tokne-secret silently ignored would be a pod running without the
authentication it declared, and nobody would know until the audit.
Annotations outside the prefix are none of this webhook's business and
pass untouched.
Retiring a key
Nothing here is deprecated. When something is, the contract has somewhere to say so: every key is one row of a registry inside the webhook, carrying the release that retired it and what to write instead. A pod that sets a retired key is admitted with a warning naming its replacement, rather than refused — a contract that breaks a working pod to make a point is a contract people pin an old version of.
Two properties come from the key list being one table rather than several.
Whether a key may take a .name suffix is read off the same row, so the two
can no longer disagree; and the documentation is checked against that table
by a test, so a key cannot be accepted and left undocumented.
The one liberty the strictness buys back: because every key is
validated, template and template-configmap could ship later without
a migration — pods that used them early were refused, not silently
ignored. They shipped; the Rendering page
owns them.
What fails the admission
A wrong ask fails the admission. A pod that says inject: "true"
and misspells the rest is refused with the reason, not started without
its configuration — silence there is how an outage begins. The refusals,
verbatim from the tests:
injectset to anything but"true"/"false"- a missing required annotation, named in the message
modeoutsideinit | sidecar | bothwatch-secondsthat does not parse as whole seconds- a
*-secretvalue without the<secret-name>/<key>slash endpointandendpoint-secretboth set — one address, one placessh-secretalongside anauthother thanssh-keyvolume-mediumoutsidememory | disk;native-sidecaroutsidetrue | false- a resource annotation that is not a Kubernetes quantity
file-modeoutside octal0400–0777— setuid bits answer no question, and an owner-unreadable file is write-only noiseagent-run-as-user/-groupof0— the agent stays nonroot in every configurationinject-containersnaming a container the pod does not have, or set to nothing at all — a render nothing can read is a render nobody asked forinit-firstwithmode: sidecar, where there is no init containeragent-run-as-same-userwith norunAsUserto inherit, alongsideagent-run-as-user, or on an application that runs as rootrevoke-gracepast the pod'sterminationGracePeriodSeconds, at"0", or withoutdynamicnotify-httpat anything but a localhost address, or withmode: initon-driftoutsidewarn | repair | failhistoryat"0", past ten, or alongsidevolume-medium: diskagent-ephemeral-requestlarger thanagent-ephemeral-limittimeoutat"0", which is no deadline rather than the store's ownagent-imagenaming an image nowebhook.agentImageAllowprefix admitsrequire-ackwithout a readiness probe, or withmode: initcanary-configmapwithmode: initeventswhere the installation does not offer it — the refusal names the chart value an administrator setsclassnaming one that is not visible, one whosenamespaceslist excludes this pod, or one whose credential is in another namespacetls-skip-verifywhere the installation does not offer it, or alongsideca-configmapenv-injectnaming a container thatinject-containersleaves out: the wrapper sources the rendered file, so it has to be able to read itenv-injectwithmode: sidecar, a container the pod does not have, or a container with no explicitcommandenv-restartwithoutenv-inject, withoutmode: both, or on a container that already owns alivenessProbeagent-enventries that are notNAME=value, names that are not UPPER_SNAKE, a name set twice, or a name that shadows whataws-secretalready sets- an
agent-envname outside the installation's allowlist for the pod's namespace — the refusal names the chart value that opens it - a source
sourceDenyturns off in the pod's namespace, or one missing from a non-emptysourceAllow— checked on every render, named suffixes included - an annotation that differs from a PINNED installation value
(
!-marked, or any set value underoverridable: "false") — the refusal names both values; restating the pinned value passes - any
dynamic-config.rs/*key the contract does not list templateandtemplate-configmapboth set — one template, one place
What passes with a warning
Not every misconfiguration earns a refusal. Kubernetes lets an
admission response carry warnings, which kubectl prints and a
controller records, and four configurations earn one:
| the pod said | the warning |
|---|---|
volume-medium: disk | the rendered document is on node-backed storage, where it outlives the pod and is readable by anything that can read the node's disk |
a world-readable file-mode | every container in the pod can read it, including ones added later |
watch-seconds below 5 | every replica polls the store at that rate, and a store with a native watch delivers changes without one |
dynamic with revoke-on-shutdown: "false" | the credential stays valid after the pod is gone, until its lease expires |
The bar is high on purpose: a warning on every admission is a warning nobody reads. Each of these is a configuration that works and is probably not what was meant.
Checking a manifest before the cluster does
The webhook is a pure function of a pod and an installation, so the same decision is available without a cluster:
$ dynamic-config-webhook validate pod.yaml deployment.yaml
pod.yaml: allowed, and an agent would be injected
deployment.yaml: refused (InvalidAnnotation)
dynamic-config.rs/auth-role is required for vault kubernetes auth
JSON or YAML, one exit code for the lot, installation defaults from the process's own environment — so a CI job that runs the webhook's image with the chart's environment gets the answer the cluster would give. The point of catching it here is that the alternative is catching it from a rollout that will not start.
Whatever passes the webhook is validated again by the agent, which
knows the store-by-store rules (vault kubernetes needs auth-role,
consul kubernetes needs auth-mount, a certificate needs its key).
Those refusals land in the injected container's log and the pod's
events.
DynamicConfigClass shrinks all of this to a class reference —
the operator page carries the shape.
The agent-env gate
agent-env puts variables on the container that holds store
credentials, and environment steers SDKs — HTTPS_PROXY reroutes the
agent's traffic, AWS_CA_BUNDLE and SSL_CERT_FILE swap its trust
roots. So the names that may pass are the installer's decision,
declared once:
# values.yaml
webhook:
agentEnvAllow: "payments: HTTPS_PROXY, AWS_*; *: RUST_LOG"
Semicolons separate groups; each group takes an optional
namespace: head (absent or * means every namespace); a trailing
* on a name is a prefix glob. Empty — the default — refuses the
annotation everywhere. Kustomize installs set the same grammar in the
DYNAMIC_CONFIG_WEBHOOK_AGENT_ENV_ALLOW variable, and either way the
webhook validates it at startup and refuses to serve on a typo.
The gate is NOT a namespace annotation, by design: whoever edits a namespace is usually the tenant being gated, and a gate its subject can open is not a gate. It is also not read from the namespace object at admission time — the webhook holds no RBAC and asks the API server for nothing; the pod's namespace arrives inside the AdmissionReview, and the ruling comes from the webhook's own configuration.
The source gates
The same authority model gates which STORES a pod may use:
webhook:
sourceAllow: "payments: vault, s3; *: consul" # empty = every store
sourceDeny: "sandbox: git" # subtractive, wins
sourceAllow empty means every store everywhere — the safe default
for an upgrade; non-empty means ONLY the listed, per namespace.
sourceDeny turns stores off outright and outranks the allowlist.
Both cover every render on the pod — a denied store cannot ride in on
a named suffix — and both refusals name
the namespace and the value that opens the gate. Kustomize sets
DYNAMIC_CONFIG_WEBHOOK_SOURCE_ALLOW / _SOURCE_DENY; a name that is
not a real store fails webhook startup, so a typo cannot silently
gate nothing.
Defaults come in tiers
Every default in the table above is only the LAST tier of four:
annotation > per-store default > fleet default > built-in
The middle tiers are the installation's (the full
list): fleet
defaults cover every knob — resources, file-mode, watch-seconds,
mode, volume-medium, native-sidecar, agent-run-as-user/-group,
metrics-port, path, source, and a fleet-wide agent environment —
and agent.defaults.perStore sets, one store at a time, the same
knobs PLUS every store-shaped annotation: endpoint, key, auth
and its friends, the credential Secrets, the templates. A pod can
deploy carrying nothing but inject: "true". Pod-wide knobs (mode,
volume, resources, identity) take the DEFAULT render's store tier;
per-render knobs resolve against each render's own store. Every tier
is validated with the SAME rules as the annotation it stands in for,
at webhook startup — and any value can be PINNED (!, or
overridable: "false"), refusing a differing annotation instead of
being overridden by it.
Installation Defaults and Gates carries the full knob table, a per-store example for all nine stores, and the gates in depth.
Secrets, Certificates, and What Shows Where
Before wiring any store's authentication, one map of where things are visible. Kubernetes shows different fields to different eyes:
| where a value sits | who sees it |
|---|---|
| an annotation | anyone with get pod — and every system that logs admission objects |
| a container argument | anyone with get pod; kubectl describe pod prints args in full |
| an environment variable from a Secret | the pod spec shows only the Secret's name; the value needs get secret rights |
| a mounted Secret/ConfigMap | same — the spec names the object, the bytes stay behind RBAC |
That table decides the whole contract:
- Names travel in annotations — a role name, an auth method's mount, a username. Reading them tells an attacker nothing they could not guess.
- Secrets travel as environment variables drawn from Secrets — the agent reads three, and there is no flag for the second one on purpose:
| variable | annotation that fills it | carries |
|---|---|---|
DYNAMIC_CONFIG_AGENT_TOKEN | token-secret: <secret>/<key> | the bearer/access token |
DYNAMIC_CONFIG_AGENT_PASSWORD | password-secret: <secret>/<key> | the second secret, where a method has one: approle's secret id, userpass/ldap's password |
DYNAMIC_CONFIG_AGENT_ENDPOINT | endpoint-secret: <secret>/<key> | the address, when the address embeds a password — a redis url |
- Key material travels as mounts, read-only, into the agent container alone — the application containers never see them:
| annotation | object | lands at |
|---|---|---|
ca-configmap: <name>[/<key>] | ConfigMap, default key ca.crt | /etc/dynamic-config/ca/ |
tls-secret: <name> | kubernetes.io/tls Secret | /etc/dynamic-config/tls/tls.crt + tls.key |
ssh-secret: <name>[/<key>] | kubernetes.io/ssh-auth Secret, default key ssh-privatekey, mounted 0400 | /etc/dynamic-config/ssh/ |
A private CA, end to end
Most internal Vaults, Consuls and git hosts serve TLS from an internal PKI. The chain is one ConfigMap and one annotation:
kubectl create configmap vault-ca --from-file=ca.crt=./internal-ca.pem
metadata:
annotations:
dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
dynamic-config.rs/ca-configmap: "vault-ca"
The agent gets --ca /etc/dynamic-config/ca/ca.crt, and the store
crate under it adds the CA to its trust roots — the same
TlsConfig every store crate takes, so the spelling is identical for
all six stores.
There is no way to turn verification off. The store crates refuse that setting by design, and the agent adds no flag for it: a configuration channel that skips TLS verification is a configuration channel anyone on the path can write to.
A client certificate
Some stores authenticate with the certificate (vault's cert
method), some merely allow
mTLS in front. Either way it is one kubernetes.io/tls Secret:
kubectl create secret tls vault-client \
--cert=./client.pem --key=./client-key.pem
dynamic-config.rs/tls-secret: "vault-client"
The certificate and key must come together; the agent refuses one without the other before any byte leaves the pod.
Why the pod's own identity beats all of this
Three of the six stores can authenticate a pod with no distributed secret at all — the pod's service-account token or the node's cloud identity:
- vault:
auth: kubernetes - consul:
auth: kubernetes - firestore:
auth: metadata-server
Where one of those is available, prefer it: nothing to rotate, nothing to leak, and revocation is the platform's own. The token-shaped methods on every page exist for the stores and shops where it is not.
The Security Posture
Everything this integration does to a cluster, listed where an auditor
can find it. The one-line summary: injection never relaxes a pod's
posture, secrets never appear where kubectl describe reaches, and the
webhook holds no credentials at all.
The injected agent complies with restricted PSS
Every injected container — init, sidecar, or native sidecar — carries
the full restricted Pod Security Standard
posture, so injection works in namespaces that enforce
pod-security.kubernetes.io/enforce: restricted and passes the audit
in ones that only warn:
securityContext:
runAsNonRoot: true
runAsUser: 65532 # distroless nonroot
runAsGroup: 65532
allowPrivilegeEscalation: false
capabilities: { drop: ["ALL"] }
readOnlyRootFilesystem: true
seccompProfile: { type: RuntimeDefault }
The root filesystem is read-only because the agent writes exactly one place: the shared volume. Which is —
The rendered file lives in memory
The shared emptyDir is medium: Memory by default: rendered
configuration regularly carries credentials, and tmpfs keeps them off
the node's disk and out of its backups. The file is gone when the pod
is. A pod that prefers disk (a giant document, a memory-tight node)
says so:
dynamic-config.rs/volume-medium: "disk"
Note the accounting: tmpfs pages count against the pod's memory. Configuration documents are small; the default agent memory limit below leaves room.
The agent's resource ask
Injected with requests and limits so it can never be the reason the node evicts the app:
resources:
requests: { cpu: 10m, memory: 32Mi }
limits: { memory: 64Mi } # no CPU limit: throttling a config
# agent buys nothing, delays reloads
Four annotations move them per pod: agent-cpu-request,
agent-memory-request, agent-cpu-limit, agent-memory-limit — each
a Kubernetes quantity, refused at admission when it is not.
Native sidecars
On Kubernetes 1.29+, ask for the sidecar as the platform now spells it:
dynamic-config.rs/native-sidecar: "true"
The watching agent becomes an init container with
restartPolicy: Always — started before the app containers, stopped
after them, and a Job with one finishes, where a classic sidecar
would hold it in Running forever. With mode: "both" the one-shot
init still lands first, so the file-exists-before-the-app guarantee
survives the move.
Secrets: where each thing is allowed to appear
The contract's rule: names in annotations, secrets in Secret-backed environment variables, key material in read-only mounts into the agent container alone. The webhook enforces it — there is no annotation that accepts a token value, and the password slot has no flag on the agent at all.
Which containers see the rendered document
Every container in the pod, unless the pod says otherwise:
dynamic-config.rs/inject-containers: "app"
The default is the reference implementation's, and it is the right default
— a pod whose containers all serve the same application should not have to
enumerate them. But a pod that runs a log shipper, a mesh proxy or a debug
sidecar beside its application is a pod where one container needs the
credential and the others do not, and file-mode cannot express that:
a sidecar in the same pod usually runs as the same UID, and reads a 0600
file exactly as well as the application does.
Naming a subset is the only lever that draws the line, so it exists. A name the pod does not have is refused rather than ignored — a typo would leave the application without its configuration and every other container without it too.
Policies Kubernetes enforces better than this webhook would
policies/ carries four ValidatingAdmissionPolicy samples in CEL:
volume-medium: disk refused, tls-skip-verify refused whatever the
installation allows, a world-readable file-mode refused, and the injected
agent required to be pinned by digest.
They are examples rather than a feature. Kubernetes has a policy engine and
this project should not grow a second one; what was missing was worked
starting points. Each ships as Warn so it can be installed and read before
it is enforced, and the first three select on a namespace label rather than
the whole cluster — a policy that fires everywhere on day one is a policy
somebody removes on day two.
They match on the annotations rather than on the injected container, because admission policies run before this webhook does: what they see is what the pod's author wrote.
What falls with what
THREAT_MODEL.md
names the assets and the trust boundaries, and walks what an attacker
reaches from each thing they might compromise — an application container,
the agent, the webhook, the operator, the network path.
The one it is worth reading this page for: a cluster-scoped
DynamicConfigClass holds one credential and admits many namespaces, and a
tenant in an admitted namespace chooses the key. Scope that credential to
what its tenants may read, and prefer namespaced Classes where the tenancy
allows — a key-prefix policy on the Class would make it structural, and that
is not built.
Transport credentials are wiped when they go
The Vault token, the AppRole secret-id, the password, the AWS secret key
and the SSH key are held in a type that zeroes its own memory on drop and
redacts its own Debug. What that buys is narrow, and narrow in a
specific way: a core dump, a /proc/<pid>/mem read or a swapped page taken
after the agent is finished with a credential does not contain it.
The resolved document is not covered, and cannot be. It is plaintext by necessity — it is about to be written to a file the application reads — so the protection that matters for it is the volume being tmpfs and the file mode being what the pod asked for, not memory hygiene.
The webhook holds nothing
- Its ServiceAccount sets
automountServiceAccountToken: false, and the deployment repeats it. The webhook reads the AdmissionReview it is handed and answers; it never calls the API server, so it carries no credential to steal. - It terminates TLS in-process with the certificate the chart issued; the private key never leaves its mount. Renewals are picked up from disk without a restart.
- The optional NetworkPolicy (
networkPolicy.enabled=true) writes both facts down for the CNI: ingress only on 8443, egress empty.
The webhook cannot select itself
The webhook configuration excludes kube-system, kube-node-lease
and the release's own namespace by name — a mutating webhook that can
select its own pods can deadlock its own rollout, and one that can
mutate the control plane is a cluster risk with no matching reward.
Add more with webhook.excludeNamespaces.
failurePolicy: the whole trade
Ignore (default): an unreachable webhook lets pods through
un-injected. The failure is visible where it matters — the annotated
pod's application waits for a file that never comes — and invisible
where it does not: un-annotated pods, which are most pods, never notice.
Fail: no annotated pod can start un-injected, and no pod at all can
start in selected namespaces while the webhook is down. Two replicas,
a PodDisruptionBudget and topology spread are the chart's mitigations;
they shrink the window, they do not close it.
Start on Ignore, alert on the webhook's availability, and flip to
Fail when the alert has been quiet long enough to trust.
Every admission leaves a line
The webhook logs one structured line per decision that matters —
namespace, pod name, source, and patched or refused — and no
annotation values: endpoints and role names belong in the cluster,
not in every log aggregator downstream. Pods that never asked are
counted but not logged.
GET /metrics on the serving port exposes the counters
(Observability is the full map) in Prometheus
text format:
dynamic_config_admissions_total{outcome="skipped"} 1042
dynamic_config_admissions_total{outcome="patched"} 63
dynamic_config_admissions_total{outcome="refused"} 2
A rising refused is somebody fighting the contract; alert on it.
Typos cannot pass
An unknown dynamic-config.rs/* annotation fails the admission — the
reference explains the rule.
The enterprise version of the argument: a misspelled token-secret that
is silently ignored produces a pod that runs, connects anonymously, and
reads whatever the store's anonymous policy allows. Refusing at
admission turns a quiet posture downgrade into a loud create-time error.
Templates are code, and scoped like data
A template renders only the resolved document — the same value the application reads. There is no file access, no environment access, no network in the template language; a hostile template can misrender the config file it owns and nothing else. Undefined keys are strict errors, so a template cannot silently swallow a value either. Keeping templates in ConfigMaps puts them through the same review as the code they effectively are.
Namespace gating
webhook.namespaceGating=true flips injection to Istio-style opt-in:
only namespaces labeled dynamic-config.rs/injection: enabled are
selected at all. Two things follow:
- the blast radius of the webhook is exactly the namespaces that asked;
failurePolicy: Failbecomes a per-namespace promise — a platform team can fail closed for its opted-in tenants without coupling every pod CREATE in the cluster to this webhook.
kubectl label namespace team-a dynamic-config.rs/injection=enabled
A label, and it could not be an annotation: the gate lives in the
webhook configuration's namespaceSelector, and Kubernetes selectors
match labels only — annotations are invisible to them. The alternative
(the webhook reading each pod's Namespace object to check an
annotation) would hand the webhook API access it pointedly does not
have; the zero-RBAC posture outranks the spelling preference. Either
way the per-POD dynamic-config.rs/inject: "true" annotation is still
required — the namespace gate is an outer guard, never an implicit
opt-in.
Fleet-wide agent defaults
The injected container's resource defaults come from the chart
(agent.defaults.*), not from a constant in a binary — platform teams
set the fleet's floor once, and the per-pod annotations still override
it. The same is true of the rendered file's permissions
(agent.defaults.fileMode) and the watch interval
(agent.defaults.watchSeconds); the same values file pins the agent
image the webhook injects. Every fleet default is validated when the
webhook STARTS — a mistyped octal stops the process at install, never
at the first admission — which also covers kustomize installs, where
no chart schema stands in front of the env vars.
The source gates are the installer's too
webhook.sourceAllow / webhook.sourceDeny decide which stores may
be rendered from, per namespace — an allowlist that admits only what
it names, and a subtractive deny that outranks it. Empty allow means
every store, so an upgrade changes nothing until the installer says
so. The check covers every render on the pod, named suffixes
included, and a gate entry that is not a real store name fails
webhook startup instead of silently gating nothing.
The agent-env gate is the installer's
agent-env lets a pod put
environment on its injected agent, and agent environment steers SDKs
(proxies, trust roots). Which names may pass, and in which namespaces,
is declared in webhook.agentEnvAllow — owned by whoever installs the
webhook, not by the pod author and not by a namespace annotation the
tenant could edit for themselves. The default is empty: everything
refused.
Supply chain
- Images are distroless, run as
nonroot, and the chart refusestag: latestat render time; adigestvalue pins harder than a tag can. - The release workflow signs images with cosign and attaches SBOMs; ghcr.io and Docker Hub carry the same digests.
- The agent binary embeds the engine and the store crates from crates.io — the same audited path every other binding uses; there is no k8s-only fork of anything.
What the injected agent never does
No hostPath, no privileged, no capabilities added, no writes
outside the shared volume, no API server calls from the webhook, and no
credentials in flags or annotations.
Two of those used to be said of the whole integration, and 0.3.0 made that untrue in two places. Both are off by default and both are named here rather than left for a reader to find:
- The node agent mounts
hostPathand runs as root. It is a CSI node plugin, so the kubelet's plugin socket and pod directories are where its work is, and the kubelet creates those directories owned by root. It adds no capabilities, escalates no privileges, and — since it makes no mounts — asks forHostToContainerpropagation rather than theBidirectionalthat would require a privileged container.nodeAgent.enabledisfalse. See the node agent. tls-skip-verifyexists. A pod can ask to reach its store without authenticating it, and the answer is no unless the installer setwebhook.allowTlsSkipVerify, which defaults to false. It is the one annotation that trades away the guarantee the rest of this page is arranged around, which is why it is the installer's to grant and not the pod author's to take.server-nameis the answer to the problem that usually leads people here — a certificate whose name does not match the address — and it gives up nothing.
Rendering
The agent does one resolution, through the same engine every binding uses, and writes the resolved document — so the file on disk is what an in-process consumer would have computed, not a second dialect.
The output format follows --out's extension: .json, .toml,
.yaml, .ini, .properties.
The flat formats are legal here and refused by the engine's save —
both on purpose. save's contract is a typed round trip, which a
string-widening format cannot keep. A rendered file for a consumer is a
different contract, and the agent owns it, stated:
- Nested tables become dotted keys (properties) or sections (INI).
- A string that would widen on the way back in —
"1.10","true"— is double-quoted in INI, so the round trip through the engine's own parser answers the same document. There is a test that holds exactly this. - Arrays are refused, by path. Neither format has them; inventing an encoding would be a dialect of one. Render to json/toml/yaml when the document has lists.
Writes are write-then-rename, so a watching application sees whole files — the same courtesy an atomic-save editor pays, and the reason the engine's own watcher tolerates a 25ms grace.
On a fetch failure the sidecar keeps the last rendered file and says so in its log — keep-last-good, the organisation's standing behaviour. An init run with nothing yet rendered fails instead, which fails the pod, which is what an init container is for.
Templates
Without one, the agent renders the resolved document verbatim — same keys, same shapes, only the format changes with the extension. That is the right default and it stays the default.
A template takes over when verbatim cannot serve: an application that
wants DATABASE_URL=postgres://… assembled from three keys, a
framework with its own nesting, a file with a header. The template owns
the output bytes, which also frees the extension — .env and .conf
become legal exactly there.
apiVersion: v1
kind: ConfigMap
metadata:
name: billing-template
data:
template: |
DATABASE_URL=postgres://{{ db.user }}@{{ db.host }}:{{ db.port }}/billing
BETA={{ flags.beta }}
dynamic-config.rs/path: "/config/app.env"
dynamic-config.rs/template-configmap: "billing-template"
# or, for one-liners:
dynamic-config.rs/template: "db={{ db.host }}:{{ db.port }}"
The syntax is minijinja's — Jinja2:
{{ value }}, {% for %}, {% if %}, filters. The template's context
is the resolved document, the same value every binding reads, so a
template cannot see anything the application could not.
The semantics that matter in production:
- Undefined is strict.
{{ db.hots }}is a render error, not an empty string — a typo that silently renders nothing would ship a broken file with a clean exit code. At startup the error is fatal and lands in the pod's events; during a watch, the running pod keeps its last good file, like any fetch failure. - Booleans render as
true/false, not Python'sTrue— a template writes config files, and every format this agent speaks spells them lowercase. - The trailing newline survives. Env files want one; what the template author wrote is what lands.
- The ConfigMap is re-read at every render, so editing the template takes effect on the next tick — no rollout. It is also why the template belongs in a ConfigMap: it is code, and it gets reviewed and versioned like code.
templateandtemplate-configmaptogether are refused at admission: one template, one place.
The filters this agent adds
minijinja's built-ins, plus six that a configuration template needs and minijinja either does not ship or ships behind a feature this build does not carry:
| filter | for |
|---|---|
b64encode | a Kubernetes Secret's data is base64, so a template writing one has to encode |
b64decode | the other direction; a value that is not base64, or not UTF-8 once decoded, is a render error rather than mojibake |
json | tojson is behind a disabled feature, and a template that cannot emit JSON is missing the format half its consumers read |
yaml | the same, for the other half |
quote | a password with a # in it ends a line in half the formats here |
required | strict undefined already refuses a missing key; this refuses one that is present and empty, which is the shape a missing secret usually arrives in. {{ db.password | required("no password in the vault path") }} names the field instead of writing a blank one |
What is not here, and will not be: a filter that reaches the network, the filesystem or the environment. The pipeline is fetch → resolve → validate → template, and a template that could fetch would make the same input render differently on two pods.
Checking the document before it is published
dynamic-config.rs/schema-configmap: "billing-schema" # or "billing-schema/other-key.json"
The resolved document is validated against a JSON Schema before anything is written. A document that fails is refused, the last good file keeps serving, and the failure is a log line and a counter — the same shape as any other render failure.
The bindings already validate: a Rust, Python or Node application gets a
typed refusal from the engine. This is the door for everyone else — the
Java service reading a .properties file, the daemon reading YAML — for
whom a port: "abc" would otherwise be discovered at startup, one
restart after the bad document was published.
The schema is re-read every render, like a template, so tightening it does not need a rollout.
Several files, one fetch
dynamic-config.rs/path: "/config/app.yaml"
dynamic-config.rs/also.db: "/config/db.env"
dynamic-config.rs/also-section.db: "database"
More files cut from the same fetched document, published all or none: one fetch, one generation, and a failure in the third does not leave the first two on disk. Each file is still its own atomic rename — a reader can catch the gap between two of them, which is a rename apart rather than a fetch apart.
Several stores cannot share a generation. Two stores have no common instant, and no protocol either of them speaks can say "these two reads are the same version", so a second store is a named render with a generation of its own.
What the application is running
dynamic-config.rs/meta: "true"
Writes a sibling file — /config/app.yaml gets /config/.app.yaml.meta —
holding the SHA-256 of the rendered bytes, the store's own revision, and
when the render landed. Same atomic rename, same mode.
It answers a question an application cannot otherwise ask about itself: which configuration am I running? Two pods holding the same file is a claim nobody can check from inside either of them; two pods printing the same digest is one anybody can. It describes the render and never contains it — no values reach it, ever.
The Full Stack, One Deployment
Vault to sidecar to template to file to a live process — every hop on one page, as the manifests you would actually apply. Each piece has its own chapter; production is all of them at once, and the order they come up in.
Vault (secret/myapp) the values that move
→ injected agent, kubernetes auth no distributed secret
→ minijinja template (ConfigMap) the store's shape → the app's
→ /config/rendered.yaml (tmpfs) off the node's disk
→ the app's own watcher live without a restart
0. The one-time pieces
$ helm install dynamic-config oci://ghcr.io/dynamic-config-rs/charts/dynamic-config \
--namespace dynamic-config --create-namespace
$ vault auth enable kubernetes
$ vault write auth/kubernetes/config kubernetes_host=https://kubernetes.default.svc
$ vault write auth/kubernetes/role/billing \
bound_service_account_names=billing \
bound_service_account_namespaces=shop \
policies=billing-read ttl=1h
$ kubectl -n shop create configmap vault-ca --from-file=ca.crt=./internal-ca.pem
1. The template — the store's shape becomes the app's
apiVersion: v1
kind: ConfigMap
metadata:
name: billing-template
namespace: shop
data:
template: |
db:
host: {{ db.host }}
pool_size: {{ db.pool_size | default(8) }}
features:
cache: {{ features.cache | default(false) }}
Re-read every render, so editing the ConfigMap is itself a live change — no rollout. The Rendering chapter owns the semantics; the one rule worth restating is that a template failure leaves the previous rendered file in place, which is last-known-good at the file layer.
2. The workload — annotations are the whole integration
apiVersion: apps/v1
kind: Deployment
metadata:
name: billing
namespace: shop
spec:
replicas: 3
selector: { matchLabels: { app: billing } }
template:
metadata:
labels: { app: billing }
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "vault"
dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
dynamic-config.rs/key: "secret/myapp"
dynamic-config.rs/auth: "kubernetes"
dynamic-config.rs/auth-role: "billing"
dynamic-config.rs/ca-configmap: "vault-ca"
dynamic-config.rs/template-configmap: "billing-template"
dynamic-config.rs/path: "/config/rendered.yaml"
dynamic-config.rs/native-sidecar: "true"
dynamic-config.rs/watch-seconds: "30"
spec:
serviceAccountName: billing
containers:
- name: app
image: myapp:1
# /config arrives injected: tmpfs, shared with the agent.
readinessProbe:
httpGet: { path: /readyz, port: 8080 }
No secret in the manifest, no volume stanza to write, no sidecar to
maintain: the webhook injects the agent, the agent logs in with the
pod's own service account, and the token never exists as a Kubernetes
Secret. native-sidecar: "true" makes the agent an init container with
restartPolicy: Always, so a Job with these annotations still
finishes.
3. The app — the last hop is an ordinary file
The agent's output is a file, so the app's side is the engine's ordinary file story — any of the three languages, here Rust:
#[dynamic_config]
#[derive(Deserialize)]
struct Db {
host: String,
pool_size: u32,
}
Db::builder("db").file("/config/rendered.yaml").init()?;
Db::builder("db")
.file("/config/rendered.yaml")
.watch(Duration::from_millis(500))?
.detach();
The write is atomic (rename), so the watcher never reads half a
render — the same guarantee the kubelet's ..data swap gives a mounted
ConfigMap, held at both layers.
What moves without a rollout, and what does not
| Change | Takes effect |
|---|---|
| the secret in Vault | next watch-seconds tick → render → app's watcher |
| the template ConfigMap | next render (kubelet sync + render tick) |
| an annotation | rollout — injection happens at pod creation |
| the chart's own values | helm upgrade, webhook restart, no app rollout |
The annotation row is the one that surprises: annotations are read by the webhook when the pod is admitted, so changing them is a Deployment edit and rolls pods — which is also why it is safe, because every replica converges through the same admission path.
Watching it work
$ kubectl -n shop logs deploy/billing -c dynamic-config-agent --tail=5
$ kubectl -n shop exec deploy/billing -c app -- cat /config/rendered.yaml
$ vault kv put secret/myapp db='{"host": "db-2.internal", "pool_size": 16}'
# … within watch-seconds, the same two commands show the new values,
# and the app's /readyz generation has moved. No pod restarted.
The Four Deliveries: File, Env, Secret, Volume
One engine, four ways a document reaches a workload — because real software disagrees about how it wants to be configured. Grafana re-reads files; Airflow reads environment variables at boot and nothing else; a Strimzi-shaped operator watches Kubernetes Secrets. Picking the delivery is a one-line decision here; what follows is the map.
| File (webhook + agent) | Env (env-inject) | Secret (operator target) | Volume (CSI, node agent) | |
|---|---|---|---|---|
| The consumer | reads/watches a file | reads environ at start | reads/watches a k8s Secret, or envFrom | reads/watches a file |
| Freshness | live — atomic rename, watcher cadence | frozen at container start (Kubernetes' rule); env-restart opts into a kubelet container-restart on change — seconds, no pod recreation | live object; watchers react, envFrom at next start | live, same renames from a shared watch |
| Touches etcd? | never — tmpfs emptyDir | never — same tmpfs file, sourced | yes — a Secret lives in etcd, stated out loud | never |
| Restart to update? | no | yes (next pod start) | no for Secret-watchers; yes for envFrom | no |
| Containers per render | one | one | none in the workload | none — one per node, shared |
| Set up by | pod annotations | pod annotations (+ explicit command) | DynamicConfigRender CR | a csi: volume, no annotations at all |
| Real example | Grafana | Airflow | Kafka client | the node agent's page |
The decision procedure, in order: can the consumer read a file?
File delivery — it is the only fully live one and the only one that
never touches etcd. Does something else own the consumer (an
operator that only reads Secrets)? The Secret target. Env-only
software? env-inject, with the start-freeze stated rather than
papered over.
And the fourth is a scale answer rather than a shape answer. The volume delivers the same bytes as the file does, from one process per node instead of one beside every render — 10,000 pods at 2.5 renders each is 25,000 sidecar containers, and that number is the only reason to reach for it. It costs the isolation the other three have: one process holds the store credentials of every pod on its node. Choose it from a measurement, not a preference, and read its page before you do.
Against the Vault Agent Injector
The closest relative of the webhook+agent half — same architecture (mutating webhook, injected init/sidecar, shared memory volume, pod service-account auth, file permissions and run-as knobs), different center of gravity:
| Vault Agent Injector | dynamic-config-k8s | |
|---|---|---|
| Backends | Vault | nine stores — Vault among them, plus Consul, etcd, git, Redis, NATS, S3, Firestore, config-server |
| Pod auth | Kubernetes auth (SA token) | the same, wherever the store speaks it (Vault, Consul); IRSA/Workload Identity for S3/Firestore; secret-based stays first-class where a store never will |
| What is delivered | rendered secrets, Consul-Template language | the resolved configuration document — precedence, validation, provenance — in json/toml/yaml/ini/properties or a minijinja template |
| File perms / ownership | annotations | annotations (file-mode, agent-run-as-user/group), root refused at admission |
| Env variables | you rewrite the command by hand to source the file | env-inject writes the wrap for you, refuses the impossible cases by name |
| k8s Secret objects | no | the operator's secret: target, when the consumer requires one |
| Lease renewal | renews renewable leases, re-fetches non-renewable ones | the same, since 0.3.0: dynamic: "true" reads a dynamic engine, renews a renewable lease at 65% of its TTL, never sends a renewal to one the store marked non-renewable, re-issues at 90% instead, and hands the lease back on SIGTERM |
| PKI certificates | tracks the certificate's own lifetime | the same: whichever expires first binds — the lease, or the certificate's notAfter |
| Failure semantics | last-known-good, retries | last-known-good with a startup policy, a deletion policy, a staleness ceiling wired to readiness, and drift detection on the rendered file |
| Knowing it arrived | the file exists | the file exists and the application said it applied it — require-ack makes that readiness |
| Rolling a change | all pods at once | canary-configmap: a deterministic cohort takes it first, widened by editing a ConfigMap with no restart |
| Delivery shapes | file (and env by hand) | file, env, k8s Secret, and a CSI volume from a per-node agent |
| Scope | secrets delivery | configuration delivery that treats secrets as first-class fields |
What it still does that this does not, and what this declines on
purpose, are written down rather than left to a table's silence: seven
items in VAULT-PARITY-GAPS.md and twelve in VAULT-PARITY-REFUSED.md,
both at the root of this repository. The short version is that the
remaining gaps are ergonomics — a render-spec ConfigMap, two
exit-on-failure policies, log format — and the refusals are the proxy, the
cache, arbitrary commands, and anything that would let a pod rewrite the
container this webhook injects.
The injector's template idiom, translated
The Vault injector spells "render me a connection string" as a per-secret annotation pair; here the same result is one source and one template, because the whole pod has one resolved document:
# Vault Agent Injector:
# vault.hashicorp.com/agent-inject-secret-db-creds: "secret/data/db-app"
# vault.hashicorp.com/agent-inject-template-db-creds: |
# {{- with secret "secret/data/db-app" -}}
# postgres://{{ .Data.data.username }}:{{ .Data.data.password }}@postgres:5432/appdb
# {{- end }}
# dynamic-config-k8s, the same string from the same KV secret:
dynamic-config.rs/source: "vault"
dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
dynamic-config.rs/key: "secret/db-app"
dynamic-config.rs/auth: "kubernetes"
dynamic-config.rs/auth-role: "db-app"
dynamic-config.rs/path: "/config/db.env"
dynamic-config.rs/template: |
DATABASE_URL=postgres://{{ username }}:{{ password }}@postgres:5432/appdb
The template owns the bytes (minijinja, strict-undefined: a typo is an
error, not an empty string), so any shape works — a URL, an .env, a
whole config file. What does NOT translate is the database/creds/…
path in the injector's example: that is Vault's dynamic secrets
engine, credentials minted per-request with leases — the boundary the
next paragraph prices. This agent reads KV; for minted-with-TTL
credentials, run the Vault Agent beside it.
One more idiom, matched: the injector's several -secret-<name>
pairs per pod are this webhook's named renders —
source.db, key.db, path.db beside the default, one agent and one
file per name, all in one shared directory. When the pod wants them
MERGED into a single document instead, the
config server composes sections and the pod reads
one endpoint. What has no counterpart is agent-inject-command (a
post-render hook): env-restart covers the restart
case, and anything richer belongs to the app.
Honest edge the other way: for dynamic Vault secrets with leases (database credentials minted per-pod, TTL renewal mid-life), the Vault Agent is the purpose-built tool and this is not — this agent re-fetches documents; it does not manage leases.
Against External Secrets Operator
The closest relative of the operator half — same split between a namespaced store and a platform-owned cluster store, different product:
| External Secrets Operator | dynamic-config-k8s | |
|---|---|---|
| Store definition | SecretStore / ClusterSecretStore | DynamicConfigClass / ClusterDynamicConfigClass — the same two scopes, allowlist included |
| Output | a Kubernetes Secret, always | a ConfigMap, a Secret, or a file no etcd ever sees |
| The etcd trade | every delivered secret lives in etcd | only the Secret target does, and choosing it is explicit — the file path exists precisely to avoid it |
| Data model | key-by-key secret mapping | whole configuration documents: precedence, validation, formats, provenance |
| Env delivery | envFrom the Secret | the same via envEntries — or env-inject, which needs no Secret at all |
| Backends | very many secret managers | nine configuration stores |
| Templating | Secret templates | minijinja over the resolved document |
Honest edge the other way: as a secret-synchronisation fleet tool across dozens of managers (AWS/GCP/Azure SM, Doppler, 1Password…), ESO has breadth this project does not chase — the store list here grows by demand, not by roadmap.
Use cases, mapped
- Airflow / env-only software →
env-injectover a rendered dotenv (example); addenv-restart: "true"and a changed document restarts just that container in seconds — otherwise changes wait for the next pod start, and that limit is stated instead of hidden. - Grafana / anything that re-reads files → the sidecar; live updates, zero etcd, tmpfs only (example).
- Strimzi-shaped operators / JVM
client.properties→ the Secret target,fileorenvEntriesshape (example). - Multi-tenant platforms →
ClusterDynamicConfigClasswith anamespacesallowlist; tenants never see a credential (example). - A chart's
existingSecret/ an operator'ssecretName:→ the Secret target withshape: entries— leaf keys verbatim, so the names some other chart already chose are met exactly (example). - All of it at once → the four-component
shop stack:
three secrets injected three ways (a chart's
existingSecret, asecretKeyRefenv, a mounted-and-live file), the API onenv-inject+env-restart, the worker on live files — credentials existing only in the platform namespace. - Vault dynamic database credentials with TTLs → the Vault Agent Injector, genuinely. Use both: it owns the lease, this owns the configuration.
The Nine Stores
One page per store the agent speaks, each with the full pod YAML for every authentication method the store takes — copy, adjust names, apply. Everything is runnable; the consul flow is what the e2e harness runs on every pull request.
| store | speaks | auth methods | its page |
|---|---|---|---|
| Consul | KV over HTTP | anonymous, token, kubernetes (login), jwt | Consul |
| Vault | KV v2 | token, kubernetes, approle, jwt, userpass, ldap, cert | Vault |
| Config Server | this project's own server | bearer token | Config Server |
| Firestore | Google Cloud | metadata-server (Workload Identity), access-token, emulator | Firestore |
| Git | any git host | anonymous, token, ssh key | Git |
| Redis | RESP | in the url (requirepass, ACL users) | Redis |
Common to all six:
- The store's address rides
dynamic-config.rs/endpoint— orendpoint-secretwhen the address itself carries a password. - The document's key rides
dynamic-config.rs/key; the per-store syntax (mount/path,application/profile, a file path) is on the store's page. - Secret material rides Secrets, never annotations; the geography page has the one diagram.
- A private CA is the same one annotation everywhere:
dynamic-config.rs/ca-configmap.
Every pairing on these pages also exists as a ready-to-apply manifest
in the repository's examples/ directory — twenty-three manifests plus six real-software walkthroughs, each
self-contained with its Secret placeholders.
etcd, NATS, S3 — the async three
Since 0.1.1 the agent drives both of the engine's source traits: the blocking six run under a blocking task, and etcd, NATS and S3 — whose clients are async — are driven directly by the agent's own runtime. The 0.1 refusal-by-name retired with this.
The config server indirection remains the answer to a different question: a fleet of pods that should not each hold store credentials — the server holds them once.
A tenth store — GCP Secret Manager, Azure App Configuration — is a compile-time addition with a well-worn path: Adding a Store walks it end to end, worked example included, plus the two no-code compositions that cover the meantime.
Consul
A KV path over HTTP. Four ways in, ordered from development to production.
The key is a KV path with an extension — myapp/config.json — and the
extension names the stored format; the rendered format is the path
annotation's extension, and the two need not agree.
Anonymous
Correct for a Consul with ACLs disabled, and for a default policy
that allows reads — both ordinary in development. No auth annotation at
all:
apiVersion: v1
kind: Pod
metadata:
name: billing
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "consul"
dynamic-config.rs/endpoint: "http://consul.infra.svc:8500"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/path: "/config/rendered.toml"
spec:
containers:
- name: app
image: myapp:1
This is the exact flow the e2e harness runs on every pull request.
An ACL token
The CONSUL_HTTP_TOKEN you already have, moved into a Secret:
kubectl create secret generic consul-token --from-literal=token=b3a7…
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "consul"
dynamic-config.rs/endpoint: "http://consul.infra.svc:8500"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/path: "/config/rendered.toml"
dynamic-config.rs/token-secret: "consul-token/token"
The webhook wires the Secret into the agent as
DYNAMIC_CONFIG_AGENT_TOKEN; nothing token-shaped appears in the pod
spec. A static token is also the one method that cannot recover on its
own — when it expires or is revoked, the fetch fails until the Secret
is updated. The two methods below fix that.
Kubernetes: login with the pod's identity
Consul's auth methods issue a token in exchange for a bearer the method
trusts — for the kubernetes type, the pod's own service-account JWT.
Nothing is distributed; the token is minted per login.
Consul side, once (the Consul docs on auth methods carry the full story):
consul acl auth-method create -type kubernetes -name k8s-pods \
-kubernetes-host https://kubernetes.default.svc \
-kubernetes-ca-cert @/path/to/cluster-ca.crt \
-kubernetes-service-account-jwt "$REVIEWER_JWT"
consul acl binding-rule create -method k8s-pods \
-bind-type role -bind-name 'config-reader' \
-selector 'serviceaccount.namespace==default'
Pod side, two annotations — auth-mount carries the auth method's
name, because that is the coordinate Consul's login endpoint wants:
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "consul"
dynamic-config.rs/endpoint: "http://consul.infra.svc:8500"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/path: "/config/rendered.toml"
dynamic-config.rs/auth: "kubernetes"
dynamic-config.rs/auth-mount: "k8s-pods"
The JWT is read from disk at every login rather than once: the kubelet rotates projected service-account tokens, and a copy taken at startup expires with the pod still running. If a projected volume moved the token off the conventional path, say where:
dynamic-config.rs/auth-token-path: "/var/run/secrets/tokens/consul"
JWT: any bearer the method trusts
The same login endpoint, with the bearer supplied instead of read from the service-account mount — an OIDC id token, a JWT signed by something Consul trusts. The bearer is a secret, so it rides a Secret:
dynamic-config.rs/auth: "jwt"
dynamic-config.rs/auth-mount: "oidc-ci"
dynamic-config.rs/token-secret: "ci-idtoken/jwt"
TLS
An internal Consul serving from a private PKI needs its CA trusted; one ConfigMap, one annotation:
dynamic-config.rs/endpoint: "https://consul.infra.svc:8501"
dynamic-config.rs/ca-configmap: "consul-ca"
A cluster fronting Consul with mTLS adds
tls-secret.
When it fails
| symptom | look at | usual cause |
|---|---|---|
agent log: 403 | consul acl token read -self with the same token | the token lacks key_prefix read on the path |
| agent log: login refused | consul acl auth-method read -name k8s-pods | binding rule selector does not match the pod's namespace/SA |
| agent starts, file never updates | consul kv get myapp/config.json | the key moved, or the watch interval is long — check watch-seconds |
The agent keeps the last good render on any fetch failure — the rendering page spells out that guarantee.
Beyond the agent's flags, the consul crate can also scope reads to a datacenter and tune blocking-query waits; those knobs live on the store crate itself for applications embedding the engine directly.
Vault
KV v2, and the widest auth surface of the six — all seven of the vault crate's methods work through annotations. Ordered from development to production; if the pod runs on Kubernetes (it does — you are reading the k8s book), start with kubernetes.
The key is <mount>/<path>, the way vault CLI users write it:
secret/myapp is the myapp path on the secret KV mount. What the
agent renders is the secret's data — the fields under data.data in
vault's own JSON.
A token
The method every tutorial starts with, and the only one that cannot recover on its own: there are no credentials behind it to log in again with. A renewable token is still renewed; a revoked one is the 3 a.m. page.
vault policy write myapp-read - <<'HCL'
path "secret/data/myapp" { capabilities = ["read"] }
HCL
vault token create -policy=myapp-read -ttl=768h -format=json \
| jq -r .auth.client_token \
| xargs -I{} kubectl create secret generic vault-token --from-literal=token={}
apiVersion: v1
kind: Pod
metadata:
name: billing
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "vault"
dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
dynamic-config.rs/key: "secret/myapp"
dynamic-config.rs/path: "/config/rendered.yaml"
dynamic-config.rs/token-secret: "vault-token/token"
dynamic-config.rs/ca-configmap: "vault-ca"
spec:
containers:
- name: app
image: myapp:1
Kubernetes: the pod's own identity
No secret distributed anywhere: the agent presents the pod's
service-account JWT to vault's kubernetes auth method and gets a
token scoped to a role. This is the pod YAML the webhook's second
golden file locks:
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "vault"
dynamic-config.rs/endpoint: "https://vault.vault.svc:8200"
dynamic-config.rs/key: "secret/myapp"
dynamic-config.rs/path: "/config/rendered.yaml"
dynamic-config.rs/auth: "kubernetes"
dynamic-config.rs/auth-role: "myapp"
dynamic-config.rs/ca-configmap: "vault-ca"
Vault side, once:
vault auth enable kubernetes
vault write auth/kubernetes/config \
kubernetes_host=https://kubernetes.default.svc
vault write auth/kubernetes/role/myapp \
bound_service_account_names=billing \
bound_service_account_namespaces=default \
policies=myapp-read ttl=1h
Three details that bite:
-
The JWT is re-read from disk at every login, not cached: the kubelet rotates projected tokens, and a copy taken at startup expires with the pod still running. Nothing to configure — stated here so a security review can check the box.
-
A projected token with a custom audience lives off the conventional path; point at it:
dynamic-config.rs/auth-token-path: "/var/run/secrets/tokens/vault" -
A second
kubernetesmount (multi-cluster Vaults mount one per cluster) isauth-mount:dynamic-config.rs/auth-mount: "kubernetes-prod-eu"
AppRole
The usual choice for a service outside Kubernetes, and for shops that want an identity Vault owns rather than the cluster. Two halves: the role id is public and rides an annotation; the secret id is a secret and rides a Secret.
vault auth enable approle
vault write auth/approle/role/myapp policies=myapp-read \
secret_id_ttl=90d token_ttl=1h
vault read -field=role_id auth/approle/role/myapp/role-id
vault write -f -field=secret_id auth/approle/role/myapp/secret-id \
| xargs -I{} kubectl create secret generic vault-approle --from-literal=secret-id={}
dynamic-config.rs/auth: "approle"
dynamic-config.rs/auth-role: "<the role id>"
dynamic-config.rs/password-secret: "vault-approle/secret-id"
The secret id travels as DYNAMIC_CONFIG_AGENT_PASSWORD — the second
secret slot, which has no flag equivalent on purpose.
JWT / OIDC
Any JWT a jwt mount trusts — a CI job's id token, a workload identity
from another platform. The JWT itself is the credential, so it rides
the token Secret:
dynamic-config.rs/auth: "jwt"
dynamic-config.rs/token-secret: "workload-jwt/jwt"
dynamic-config.rs/auth-role: "myapp" # when the mount has no default
dynamic-config.rs/auth-mount: "jwt-ci" # when not mounted at "jwt"
Userpass and LDAP
Same shape, different directory. The username is a name and rides an annotation; the password rides a Secret:
kubectl create secret generic vault-ldap --from-literal=password=…
dynamic-config.rs/auth: "ldap" # or "userpass"
dynamic-config.rs/auth-username: "svc-billing"
dynamic-config.rs/password-secret: "vault-ldap/password"
Cert: a TLS client certificate
The certificate IS the credential: vault's cert method authenticates
the TLS handshake itself. One kubernetes.io/tls Secret carries the
pair:
kubectl create secret tls vault-client --cert=client.pem --key=client-key.pem
dynamic-config.rs/auth: "cert"
dynamic-config.rs/tls-secret: "vault-client"
dynamic-config.rs/auth-role: "myapp" # when the mount does not pick by subject
dynamic-config.rs/ca-configmap: "vault-ca"
Vault side:
vault auth enable cert
vault write auth/cert/certs/myapp certificate=@client-ca.pem \
allowed_common_names=billing.default policies=myapp-read
Vault Enterprise namespaces
One annotation, passed through to the X-Vault-Namespace header:
dynamic-config.rs/namespace: "team-payments"
How the token lifecycle behaves
Worth knowing before the first incident, whatever the method:
- Close to expiry, the token is renewed — or, if it cannot be, replaced by a fresh login. This is the path that normally fires.
- After a request, a
403is treated as the token stopped working and triggers exactly one fresh login and retry. Clocks skew, Vault revokes, a lease is shorter than it said. - If a fresh token also gets
403, the problem is the policy rather than the lease, and the error says so instead of hanging in a retry loop.
Dynamic secrets
A path under database/, pki/ or aws/ does not hold a secret; it
mints one, with a lease. One annotation says to read it that way:
dynamic-config.rs/source: "vault"
dynamic-config.rs/dynamic: "true"
dynamic-config.rs/key: "database/creds/billing"
dynamic-config.rs/path: "/config/db.env"
dynamic-config.rs/auth: "kubernetes"
dynamic-config.rs/auth-role: "billing"
The pod now holds a database credential nobody else has, for as long as it needs it. What changes:
- The path is read as it is. KV v2's
data/nesting is not inserted anddata.datais not unwrapped, because a dynamic engine has neither. - The lease is kept.
lease_id,lease_durationandrenewablecome back with the document. - A lease that says it cannot be renewed is never sent a renewal.
Vault answers
renewable: falsefor everypki/issue, and for a database credential that has reached its role's maximum. Such a lease is re-issued at 90% of its life instead. Asking it to renew first would be a request that can only be refused — a round trip per cycle per pod, and alease_renewal_failures_totalthat climbs steadily on a fleet where nothing is wrong. - A renewable lease is renewed at 65% of its TTL, spread, and what
Vault grants is what the next renewal is scheduled from — a role's
max_ttlis a ceiling the pod cannot see, so asking for an hour and being given ten minutes has to be believed rather than assumed. - The two fractions are not the same number by accident. A renewal is cheap and reversible: it fails, and a third of the lease is still there to get a new credential in. A re-issue is the new credential, so doing it early only shortens the one in use and wakes every application watching the file for it.
- A renewal does not re-render. It extends the same credential; the file is already correct. A re-issue mints new credentials and does re-render, and that is what happens both when a renewal stops working and when the lease was never renewable.
- A backoff never sleeps past the lease. This is the one place the usual ceiling is wrong: waiting five minutes to retry a credential that expires in twenty seconds is a pod that comes back to an expired secret.
- SIGTERM revokes. Best-effort, with a short deadline —
revoke-on-shutdown: "false"opts out, for a lease something else is still using. Vault expires the lease on its own eventually; revoking on the way out is what turns eventually into now.
Certificates are on their own clock
pki/issue is the one dynamic engine where the lease is not the whole
truth. A PKI role can hand back a lease longer than the certificate it
issued — the lease is Vault's accounting record, notAfter is what a TLS
peer enforces — so the agent takes whichever runs out first.
Vault reports the certificate's notAfter as data.expiration, seconds
since the epoch, which is why this needs no X.509 parsing: the number is
already in the response.
One guard comes with it. That timestamp is Vault's wall clock compared against the pod's, and the two are not the same clock. A certificate that looks already expired from inside the pod is far more likely to be skew than an expired certificate — Vault would not have issued one — so the lease's own number wins there, rather than the agent re-issuing in a tight loop against a server that thinks everything is fine.
The watch capability drops to Interval: a dynamic engine has no version
to poll, and asking whether it changed is not separable from asking for a
new credential.
One path per render. Every read mints its own lease, so merging several paths into one document would leave every lease but one unrenewed and unrevoked. A second dynamic path is a second named render.
The four lease_* series in
the agent's metrics are what to
watch: lease_ttl_seconds is what Vault last granted, and a rising
lease_renewal_failures_total is the signal that a re-fetch is coming.
When it fails
| symptom | look at | usual cause |
|---|---|---|
403 on every read | vault token capabilities <token> secret/data/myapp | the policy grants secret/myapp, not secret/data/myapp — KV v2 inserts data/ |
| login refused (kubernetes) | vault write auth/kubernetes/login role=myapp jwt=@/tmp/jwt by hand | role's bound SA/namespace does not match the pod's |
| x509 errors | openssl s_client -connect vault:8200 | the CA in ca-configmap is not the one vault serves |
works, then 403 after days | vault audit log | orphan token hit its max TTL; move to kubernetes/approle auth |
Config Server
This project's own server — the aggregation answer. The server side
speaks all nine store crates (including the async three the agent
cannot drive yet) and merges files, profiles and remote stores into one
document per <application>/<profile>; the agent then speaks the
server. The remote book
owns the server's full story; what follows is the k8s-side wiring.
Two reasons to put it between the pods and the stores:
- Credential concentration. A fleet of pods each holding a vault token is a fleet of tokens to rotate. The server holds the store credentials once; the pods hold short client tokens that grant exactly one application's sections.
- Nine stores behind one address. The server speaks every store, so a pod's annotations name one endpoint whatever moves behind it.
Kubernetes auth: the pod's own identity, end to end
Since 0.1.1 the server indirection is zero-secret too:
auth: "kubernetes" makes the agent present the pod's projected
service-account token as the bearer (re-read every fetch — it
rotates), and the server's [kubernetes] TokenReview grants map
namespace:serviceaccount to applications:
dynamic-config.rs/source: "config-server"
dynamic-config.rs/endpoint: "https://config.infra.svc:8443"
dynamic-config.rs/key: "shop/prod"
dynamic-config.rs/auth: "kubernetes"
No client token minted, distributed, mounted or rotated — the remote book's server chapter carries the server-side TOML.
The wiring
The key is <application>/<profile>:
apiVersion: v1
kind: Pod
metadata:
name: billing
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "config-server"
dynamic-config.rs/endpoint: "http://config-server.infra.svc:8888"
dynamic-config.rs/key: "billing/prod"
dynamic-config.rs/path: "/config/rendered.json"
spec:
containers:
- name: app
image: myapp:1
The bearer token
The server's [[server.clients]] blocks name each client and the
applications it may read. The token is a secret; it rides a Secret:
# server.toml, server side
[[server.clients]]
name = "billing-pods"
token = "…at least 32 bytes…"
applications = ["billing"]
kubectl create secret generic config-server-token --from-literal=token=…
dynamic-config.rs/token-secret: "config-server-token/token"
There is no other method on this store — one bearer, scoped
server-side, is the whole model. Asking for auth: anything is refused
by the agent with a sentence saying exactly that.
TLS
The server behind TLS from an internal PKI is the same one annotation as everywhere:
dynamic-config.rs/endpoint: "https://config-server.infra.svc:8443"
dynamic-config.rs/ca-configmap: "internal-ca"
A server requiring client certificates takes
tls-secret alongside.
When it fails
| symptom | look at | usual cause |
|---|---|---|
401 | server log | the token is not in any [[server.clients]] block |
403 | server log, applications = […] | the client's list does not include this application |
404 | curl $SERVER/billing/prod with the token | no [[server.sections]] matches the pair |
| stale values | the server's own watch config | the server polls its stores on its own cadence — two intervals stack |
Firestore
A Google Cloud document read as configuration. The endpoint is not a
url: it is <project> or <project>/<database>, and the key is
collection/document (nest deeper as
environments/prod/config/db).
On GKE the right method is the first one, and it involves no secret at all.
Metadata-server (Workload Identity)
The workload's own identity, from the metadata server — reachable from GKE, Cloud Run, GCE, and nowhere else, which is the security property that makes it the default. The agent asks for a token, gets a short-lived one, and renews it as it approaches expiry.
auth can be omitted entirely — metadata-server is what the agent does
for firestore when nothing else is asked:
apiVersion: v1
kind: Pod
metadata:
name: billing
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "firestore"
dynamic-config.rs/endpoint: "acme-prod"
dynamic-config.rs/key: "config/billing"
dynamic-config.rs/path: "/config/rendered.json"
spec:
serviceAccountName: billing
containers:
- name: app
image: myapp:1
GKE side, the Workload Identity pairing (the GKE docs own the full ceremony):
gcloud iam service-accounts create billing-reader
gcloud projects add-iam-policy-binding acme-prod \
--member serviceAccount:billing-reader@acme-prod.iam.gserviceaccount.com \
--role roles/datastore.viewer
gcloud iam service-accounts add-iam-policy-binding \
billing-reader@acme-prod.iam.gserviceaccount.com \
--role roles/iam.workloadIdentityUser \
--member "serviceAccount:acme-prod.svc.id.goog[default/billing]"
kubectl annotate serviceaccount billing \
iam.gke.io/gcp-service-account=billing-reader@acme-prod.iam.gserviceaccount.com
Access token
A token somebody already obtained — gcloud auth print-access-token
produces one. It expires within the hour and the agent cannot renew it,
so this is a debugging method, not a deployment method; the honest use
is a one-shot init container in a test cluster:
dynamic-config.rs/mode: "init"
dynamic-config.rs/auth: "access-token"
dynamic-config.rs/token-secret: "gcp-token/token"
Emulator
The Firestore emulator wants no credential and a different endpoint;
api-url points the API somewhere other than Google's:
dynamic-config.rs/auth: "emulator"
dynamic-config.rs/api-url: "http://firestore-emulator.test.svc:8080"
dynamic-config.rs/endpoint: "demo-project"
The named database
The second database in a project is the endpoint's second segment:
dynamic-config.rs/endpoint: "acme-prod/eu-config"
When it fails
| symptom | look at | usual cause |
|---|---|---|
403 PERMISSION_DENIED | gcloud projects get-iam-policy acme-prod | the GSA lacks roles/datastore.viewer, or the WI binding names the wrong namespace/KSA pair |
| metadata server unreachable | pod events, GKE node pool | Workload Identity not enabled on the pool |
404 | gcloud firestore databases list | the document path or the named database is wrong |
| works locally, fails in-cluster | — | local gcloud credentials are not the pod's; the pod has only the metadata server |
Git
Configuration that lives where its reviews live. Git has nothing to push, so the agent asks — but it asks the cheap question: the ref's advertisement is one handshake per tick, and a transfer happens only when the ref actually moved, so a 15-second watch does not hammer the host.
The endpoint is anything git understands (https://…, ssh://…,
git@host:org/repo.git); the key is the file's path inside the
repository; the ref defaults to branch main.
Anonymous
A public repository over HTTPS — no auth annotation:
apiVersion: v1
kind: Pod
metadata:
name: billing
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "git"
dynamic-config.rs/endpoint: "https://github.com/acme/config.git"
dynamic-config.rs/key: "billing/prod.yaml"
dynamic-config.rs/path: "/config/rendered.yaml"
dynamic-config.rs/ref: "main"
spec:
containers:
- name: app
image: myapp:1
A token over HTTPS
How every host takes a token: HTTP basic auth with the token in the
password half. GitHub PATs and App installation tokens, GitLab deploy
and project tokens, Azure DevOps PATs — all the same shape. The
username half is filler that these hosts ignore; the agent sends
x-access-token, the value GitHub documents:
kubectl create secret generic config-repo-token --from-literal=token=ghp_…
dynamic-config.rs/auth: "token"
dynamic-config.rs/token-secret: "config-repo-token/token"
For the rare host that does read the username, name it:
dynamic-config.rs/auth-username: "deploy"
A GitLab deploy token is the least-privilege pick on that platform:
scope read_repository, one repository, its own expiry.
An SSH deploy key
One kubernetes.io/ssh-auth Secret; its conventional key name is
ssh-privatekey, and the webhook mounts it 0400 because ssh refuses
group-readable keys:
ssh-keygen -t ed25519 -f deploy_key -N ""
# register deploy_key.pub as a read-only deploy key on the host
kubectl create secret generic config-deploy-key \
--type=kubernetes.io/ssh-auth --from-file=ssh-privatekey=deploy_key
dynamic-config.rs/endpoint: "git@github.com:acme/config.git"
dynamic-config.rs/ssh-secret: "config-deploy-key"
auth: ssh-key is implied by ssh-secret when no auth is named. The
key is offered with IdentitiesOnly=yes, so an agent holding other
keys cannot exhaust the server's auth tries before the right one.
And the permission caveat, out loud: the kubelet writes secret
files as root, 0400 means owner-read only, and the agent runs
nonroot — so the mounted key is readable only when the pod sets a
securityContext.fsGroup (the kubelet then group-owns the files) and
the custom image's ssh client tolerates a group-readable key, or when
the pod runs the agent's uid. This combination is honest-but-untested:
the e2e suite covers HTTPS and token auth; ssh in-pod is documented,
not gated.
The stock image caveat, out loud: git-over-SSH is carried by the
ssh program, exactly as git itself does it — and the distroless
agent image does not contain one. HTTPS works from the stock image;
SSH needs an image with an ssh client:
FROM ghcr.io/dynamic-config-rs/dynamic-config-agent:0.1.0 AS agent
FROM alpine:3.20
RUN apk add --no-cache openssh-client ca-certificates \
&& printf 'github.com ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIOMqqnkVzrm0SdG6UOoqKLsabgH5C9okWi0dh2l9GKJl\n' \
>> /etc/ssh/ssh_known_hosts
COPY --from=agent /dynamic-config-agent /dynamic-config-agent
ENTRYPOINT ["/dynamic-config-agent"]
…and the chart's agent.image value points at it. The known_hosts
line is the second half ssh insists on; pin your own host's key, not a
copy of this one.
Branch, tag, commit
ref takes four spellings:
dynamic-config.rs/ref: "main" # a branch, plainly
dynamic-config.rs/ref: "branch:release" # the same, spelled out
dynamic-config.rs/ref: "tag:v1.4" # a tag
dynamic-config.rs/ref: "commit:8f3a…" # one exact tree, forever
A tag or commit still ticks the watch loop, and still never transfers —
useful with mode: init for a pinned, reproducible render.
A self-hosted host with a private CA
The same annotation as every other store:
dynamic-config.rs/endpoint: "https://git.internal.acme/config.git"
dynamic-config.rs/ca-configmap: "internal-ca"
When it fails
| symptom | look at | usual cause |
|---|---|---|
| auth failed over HTTPS | try the token in a git ls-remote by hand | token expired, or lacks read scope on the repo |
Host key verification failed | the image's /etc/ssh/ssh_known_hosts | the custom image pinned no host key for this host |
| ssh: command not found | — | the stock distroless image; see the caveat above |
| file not found | git ls-tree <ref> -- <path> | the path is spelled from the repository root, and the ref matters |
Redis
A key read as a document, watched over keyspace notifications — a
change arrives as it happens rather than up to an interval later. The key carries its
format in its extension (myapp/config.json); a key without one needs
the document format to be guessable, so give it one.
Redis is the one store whose credentials travel in the url —
requirepass and ACL users have no other place. That shapes the whole
page: the moment the url grows a password, it stops being an annotation
and becomes a Secret.
An open Redis
Development, a sidecar cache, a cluster-internal instance behind a NetworkPolicy:
apiVersion: v1
kind: Pod
metadata:
name: billing
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "redis"
dynamic-config.rs/endpoint: "redis://redis.infra.svc:6379/0"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/path: "/config/rendered.toml"
spec:
containers:
- name: app
image: myapp:1
The trailing /0 is the database index; omit it for 0.
requirepass
The password goes into the url, and the url goes into a Secret —
endpoint-secret replaces endpoint entirely, and the agent reads the
address from DYNAMIC_CONFIG_AGENT_ENDPOINT:
kubectl create secret generic redis-url \
--from-literal=url='redis://:s3cr3t@redis.infra.svc:6379/0'
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "redis"
dynamic-config.rs/endpoint-secret: "redis-url/url"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/path: "/config/rendered.toml"
Setting both endpoint and endpoint-secret fails the admission —
one address, one place. Error messages redact the password even when
the url cannot be parsed, because a parse error is the error most
likely to be pasted somewhere.
An ACL user
Redis 6+ ACLs put a username before the password. Server side:
ACL SETUSER config-reader on >s3cr3t ~myapp/* +get resetchannels
The url names the user, read-only on exactly the config prefix:
kubectl create secret generic redis-url \
--from-literal=url='redis://config-reader:s3cr3t@redis.infra.svc:6379/0'
TLS
rediss:// (two esses) plus the CA:
dynamic-config.rs/endpoint: "rediss://redis.infra.svc:6380/0"
dynamic-config.rs/ca-configmap: "redis-ca"
A redis:// url with TLS material is refused — a deployment that
believes it is encrypted and is not — and there is no way to turn
verification off. A client certificate is
tls-secret, as everywhere.
With a password too, the whole rediss://user:pass@… url rides
endpoint-secret and the CA annotation stays as it is.
When it fails
| symptom | look at | usual cause |
|---|---|---|
NOAUTH / WRONGPASS | redis-cli -u <the url> get myapp/config.json | the url in the Secret lost its password half, or the ACL user is off |
NOPERM | ACL GETUSER config-reader | the key pattern does not cover the config key |
| refused: url is not rediss | — | TLS material with a redis:// url; add the second s |
| empty render | redis-cli … type myapp/config.json | the key holds a hash, not a string document |
etcd
A key (or several) from etcd v3. etcd is the honest hard case for
authentication: it has no Kubernetes auth method to log into — nothing
that takes a projected service-account token the way Vault's
kubernetes mount does. So this store lands with the two methods etcd
itself speaks, both first-class: TLS client certificates, and
username/password. Identity-first is the org's policy where a store
can authenticate the pod; where it cannot, the secret-based method is
supported without apology — a contract that punishes such stores
punishes their users.
There is no auth annotation for etcd: the flags present ARE the
method.
TLS client certificates
The method etcd operators already deploy — the certificate is the credential, delivered as a Secret the same way every store's TLS material is (the geography page):
apiVersion: v1
kind: Pod
metadata:
name: billing
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "etcd"
dynamic-config.rs/endpoint: "https://etcd.infra.svc:2379"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/path: "/config/rendered.toml"
dynamic-config.rs/tls-secret: "etcd-client-tls"
dynamic-config.rs/ca-configmap: "etcd-ca"
spec:
containers:
- name: app
image: myapp:1
kubectl create secret tls etcd-client-tls \
--cert=./billing.crt --key=./billing.key
kubectl create configmap etcd-ca --from-file=ca.crt=./etcd-ca.pem
Username and password
etcd's other method. The user rides an annotation; the password rides
a Secret, never an annotation — kubectl describe pod prints
annotations to anyone with pod read access:
dynamic-config.rs/source: "etcd"
dynamic-config.rs/endpoint: "https://etcd.infra.svc:2379"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/auth-username: "myapp"
dynamic-config.rs/password-secret: "etcd-password/password"
dynamic-config.rs/ca-configmap: "etcd-ca"
Both together is also legal, and is etcd's own combination: the certificate authenticates the channel, the user authenticates the principal.
Several keys
key takes a comma-free single key; several documents merging into
one is the config server's job or a template's.
The stored format is the key's extension, exactly as with
Consul.
Ready-to-apply manifests: examples/etcd-tls.yaml,
examples/etcd-password.yaml.
NATS
A JetStream key-value bucket. The key annotation is
<bucket>/<key> — the bucket must already exist; a configuration
reader that provisions storage would hide a misconfigured deployment
behind an empty one.
A credentials file
The way a NATS account authenticates — a .creds file, delivered as a
Secret and named by path:
apiVersion: v1
kind: Pod
metadata:
name: billing
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "nats"
dynamic-config.rs/endpoint: "nats://nats.infra.svc:4222"
dynamic-config.rs/key: "config/db.json"
dynamic-config.rs/path: "/config/rendered.toml"
dynamic-config.rs/auth-token-path: "/etc/nats/nats.creds"
spec:
containers:
- name: app
image: myapp:1
volumeMounts:
- { name: nats-creds, mountPath: /etc/nats, readOnly: true }
volumes:
- name: nats-creds
secret: { secretName: nats-creds }
A token
token-secret, the same one annotation every token-shaped credential
rides:
dynamic-config.rs/source: "nats"
dynamic-config.rs/endpoint: "nats://nats.infra.svc:4222"
dynamic-config.rs/key: "config/db.json"
dynamic-config.rs/token-secret: "nats-token/token"
Anonymous — no auth annotation at all — is correct for a NATS without auth, ordinary in development.
Ready-to-apply manifest: examples/nats-creds.yaml.
S3
An object from a bucket — AWS's S3, and everything that speaks its API:
MinIO, Ceph, R2, B2. endpoint is the bucket; key is the object
key; the stored format is the key's extension.
On EKS: IRSA, and nothing else
The default is the ambient AWS credential chain, which on EKS is IAM Roles for Service Accounts — the workload's own identity, no secret distributed, the token renewed by the platform. There is nothing to configure on this side beyond the service account:
apiVersion: v1
kind: Pod
metadata:
name: billing
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/source: "s3"
dynamic-config.rs/endpoint: "myapp-config"
dynamic-config.rs/key: "prod/db.json"
dynamic-config.rs/path: "/config/rendered.toml"
spec:
serviceAccountName: billing # annotated with the IAM role, IRSA's side
containers:
- name: app
image: myapp:1
The injected agent inherits the pod's identity the same way the app does — that is the entire point of workload identity.
MinIO, Ceph, R2
api-url overrides the endpoint; a region is not required alongside
it (the agent supplies a placeholder the server ignores — set
AWS_REGION on the pod if yours cares); path-style addressing is
always on,
because virtual-hosted buckets need DNS entries only AWS has:
dynamic-config.rs/source: "s3"
dynamic-config.rs/endpoint: "myapp-config"
dynamic-config.rs/key: "prod/db.json"
dynamic-config.rs/api-url: "http://minio.infra.svc:9000"
Static credentials for a non-AWS store ride one annotation:
dynamic-config.rs/aws-secret: "<secret>" — the Secret's
AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY keys become exactly
those variables on the injected agent, and nothing lands in the pod
spec.
Ready-to-apply manifest: examples/s3-irsa.yaml.
Adding a Store
Can the agent learn a tenth store — Google Cloud Secret Manager, say? Yes, and what follows is the map. What it is NOT is a plugin system: stores are compiled in on purpose (a config agent that loads plugins is its own supply-chain problem), and the organisation's standing rule is stores land by demand, not by roadmap — the ask starts as an issue, and the walkthrough below is what the work looks like once agreed.
What a store is
One Rust type implementing one of the engine's two traits:
impl AsyncRemoteSource for GcpSecretManager {
fn fetch(&self) -> BoxFuture<'_, Result<Fetched, Error>>; // text + format
fn describe(&self) -> String; // REDACTED — errors quote it
}
fetch answers the document's text and format; describe names the
store in errors and logs and must never carry a credential. Blocking
clients implement RemoteSource instead; the agent drives both — the
blocking six run under a blocking task, the async three (etcd, NATS,
S3) on the runtime directly. A new store picks whichever its client
dialect is.
The walkthrough, with GCP Secret Manager as the worked example
The 0.1.1 change that landed etcd, NATS and S3 is the living template —
git log --grep "async three" in the two repositories shows every
touchpoint below with real diffs.
1. The store crate (dynamic-config-remote repository,
dynamic-config-gcpsm/). The Firestore crate is the closest twin:
Google auth, same project-scoped shape. The heart is small:
pub struct GcpSecretManager { project: String, secret: String, auth: Auth, /* … */ }
impl AsyncRemoteSource for GcpSecretManager {
fn fetch(&self) -> BoxFuture<'_, Result<Fetched, Error>> {
Box::pin(async move {
// GET https://secretmanager.googleapis.com/v1/projects/{p}
// /secrets/{s}/versions/latest:access
// Authorization: Bearer <metadata-server token>
let text = self.access_latest().await?;
Ok(Fetched { text, format: self.format })
})
}
fn describe(&self) -> String {
format!("gcp-sm {}/{}", self.project, self.secret) // no token, ever
}
}
Identity-first applies: the default auth is the metadata server
(Workload Identity — the pod's own identity, no distributed secret),
exactly as the Firestore crate already does; an access-token arm
stays a supported peer. House rules the crate must keep: describe
redacted, TLS through the shared TlsConfig vocabulary, errors mapped
to the engine's kinds (auth for 401/403, remote for the rest),
#![forbid(unsafe_code)].
2. The agent (dynamic-config-agent/src/):
sources.rs— one async constructor arm besideetcd/nats/s3, returningBuilt::Async(Arc::new(store)).spec.rs— the source joins thevalidated()list, its auth rules joinvalidated_auth()(refusals name the fix; that is the house voice), andUSAGElearns the flags. Tests beside the etcd ones.
3. The webhook (dynamic-config-webhook/src/annotations.rs) — the
source name joins the admission allowlist, and its auth annotations
join the validation match so a typo is refused at admission, in
the pod's events, not twenty minutes later in a crashloop. A golden
test pins the happy path.
4. The paper trail — a store page in this book (auth methods, full
manifests, the identity-first default stated), a ready-to-apply file
in examples/, a row in the annotation reference's source table, and
— for a store with a live server in CI — an e2e leg shaped like
e2e/stores-smoke.sh — one cluster, every live store, a stage per source.
5. The gates — just check in both repositories, the webhook
goldens, the kind smoke. Nothing about the release train changes: the
store crate ships from dynamic-config-remote, the agent picks it up
on the next k8s release.
A store in another language
Three lanes, all of them already open:
- In-process, Python: subclass the binding's
RemoteSourceABC —fetch()anddescribe()— and hand the instance toDynamicConfig.remote(...). The Python book owns the chapter. - In-process, Node: any JS object with
fetch()anddescribe()throughuseStore(config, store)— the Node remote package's contract. - Out-of-process, ANY language — the one the agent can use: the
agent cannot load Go or Java, but it speaks the
config server's small HTTP contract. Implement
GET /{application}/{profile}(bearer-token auth, the resolved document as the body) in your language, and every consumer here — the agent's--source config-server, the engine, both bindings — reads it like any other store. A Go service fronting your in-house store is an afternoon: one endpoint, one auth header, and TLS. The remote book's server chapter tables the full route surface,/streamincluded for push.
Today, without code
Two honest compositions cover most "we need store X now" cases:
- External Secrets Operator in front: ESO syncs GCP SM (or any of
its many backends) into a Kubernetes Secret, and everything on
The Four Deliveries consumes that Secret —
mounted as a live file,
envFrom, orsecretKeyRef. You trade the document model for breadth, which is exactly the trade the comparison table prices. - A mirror job: a CronJob copies the document from the unsupported store into one the agent speaks (S3, git, Consul). Crude, visible, and often all a migration window needs.
GitOps: ArgoCD and Flux
Running the chart under a GitOps controller works, with two things worth knowing in advance: certificates, and the difference between the two Gits in play.
The two Gits
A GitOps setup around dynamic-config has two repositories doing two jobs, and conflating them causes most confusion:
- ArgoCD's Git delivers manifests — the chart, its values, your
Deployments with their
dynamic-config.rs/*annotations. Changing an annotation is a Deployment change: ArgoCD applies it, the rollout creates new pods, the webhook injects the new agent config. Cadence: deploys. - The agent's Git store delivers configuration documents — the agent polls the repository directly, inside the pod, no sync loop involved. Cadence: seconds-to-minutes, no rollout, no ArgoCD.
Config-in-Git does not require routing config through ArgoCD. The agent's Git store gives you Git-audited configuration that updates without a pod restart; keep ArgoCD for the manifests.
Certificates under a sync loop
The chart's default TLS mode mints a CA and certificate at template
time (helm template runs genCA). Under ArgoCD that means every
sync renders a new certificate — a permanent diff, and a webhook
whose caBundle churns on every sync.
Under GitOps, pick one of:
- cert-manager mode (recommended where cert-manager runs):
webhook.certManager.enabled: true. cert-manager issues and renews; cainjector maintains the caBundle; the rendered manifests are stable. - selfRotate mode: rotation without the dependency — the webhook
patches its own caBundle at runtime, so tell ArgoCD that field is
runtime-owned (the
ignoreDifferencesblock below, caBundle entry). - Kustomize-native Flux/Argo setups get the same choices as
overlays:
cert-manager,own-cert,selfrotate, pluswith-operatorfor the Render reconciler — composable in one Kustomization. - Provided-cert mode: mount your own Secret (Secrets & TLS) and set the caBundle in values — everything is declarative, nothing is generated.
- If you must keep the self-signed default, tell ArgoCD to look away:
ignoreDifferences:
- kind: Secret
name: dynamic-config-webhook-tls
jsonPointers: [/data]
- group: admissionregistration.k8s.io
kind: MutatingWebhookConfiguration
jqPathExpressions: [".webhooks[].clientConfig.caBundle"]
syncPolicy:
syncOptions: [ApplyOutOfSyncOnly=true]
CRDs
Helm installs the crds/ directory on install and never upgrades
it — standard Helm behaviour, and under ArgoCD it depends on how the
chart is rendered. Two reliable shapes:
- ArgoCD renders charts with
helm template --include-crds; the CRDs become ordinary tracked resources and upgrades apply. AddServerSideApply=truetosyncOptions— the rendering CRDs carry large schemas and can exceed the client-side annotation limit. - Or manage
deploy/crds/as its own Application (or Flux Kustomization) in an earlier sync wave, which is also the shape the operator's future CRD changes will assume.
Order of arrival
The webhook must be serving before annotated workloads are admitted.
In one Application that is a race; in practice it is self-healing
(failurePolicy: Ignore is the default — an early pod is admitted
un-injected and the next rollout catches it), but a strict setup puts
the chart in an earlier sync wave than the workloads:
metadata:
annotations:
argocd.argoproj.io/sync-wave: "-1" # the chart's Application
With failurePolicy: Fail the wave separation stops being optional —
see The Security Posture for that trade.
The Operator
Shipped as of 0.1.1: a DynamicConfigRender becomes a ConfigMap,
reconciled through the SAME source construction and rendering the
sidecar agent uses — one implementation, no drift between the two
paths a document can take into a pod. The CRDs stay generated from the
Rust types (deploy/crds.json, drift-gated in CI), and the e2e suite
drives the full loop: apply, render, propagate, delete,
garbage-collect.
DynamicConfigClass — name the store once
The class bundles source, endpoint and the token Secret, so pods stop repeating them:
apiVersion: dynamic-config.rs/v1alpha1
kind: DynamicConfigClass
metadata:
name: infra-consul
namespace: billing
spec:
source: consul
endpoint: http://consul.infra.svc:8500
tokenSecret: consul-agent-token # optional; its `token` key travels
A pod then says only what is its own:
metadata:
annotations:
dynamic-config.rs/inject: "true"
dynamic-config.rs/class: "infra-consul"
dynamic-config.rs/key: "myapp/config.json"
dynamic-config.rs/path: "/config/rendered.toml"
The long-form annotations keep working forever; the class is sugar over them, not a replacement.
DynamicConfigRender — a ConfigMap instead of a sidecar
For workloads that cannot take an injected container — third-party charts, jobs, anything sidecar-averse — the operator renders into a ConfigMap on a cadence, and the workload mounts it like any other:
apiVersion: dynamic-config.rs/v1alpha1
kind: DynamicConfigRender
metadata:
name: billing-config
namespace: billing
spec:
class: infra-consul
key: myapp/config.json
target:
configMap: billing-rendered
file: config.properties # the extension picks the format here too
intervalSeconds: 30
status: # written by the operator
renderedAt: "2026-08-18T09:00:00Z"
lastError: null # kind and path only, never a value
Classes: namespaced, or cluster-scoped
Two class kinds, the same split External Secrets Operator drew between
SecretStore and ClusterSecretStore:
DynamicConfigClass(namespaced) — a team's own store, itstokenSecretread from the same namespace. Self-service, blast radius one namespace.ClusterDynamicConfigClass(cluster-scoped) — the platform team's store, defined once. Its credential names the namespace it lives in explicitly, so tenants reference the class without ever being able to read the credential:
apiVersion: dynamic-config.rs/v1alpha1
kind: ClusterDynamicConfigClass
metadata:
name: platform-consul
spec:
source: consul
endpoint: http://consul.infra.svc:8500
tokenSecret:
name: consul-token
namespace: platform # RBAC keeps tenants out of here
key: token # `token` by default
namespaces: [team-a, team-b] # the allowlist; absent = every namespace
A tenant's Render opts in by kind:
spec:
class: platform-consul
classKind: ClusterDynamicConfigClass
Two rules, out loud. The allowlist is enforced at
reconcile: a Render in a namespace the class does not list fails with
the class named, in status.lastError — the platform team's boundary,
not a convention. And the credential read happens with the
operator's identity, which is why the operator's RBAC has cluster
get on Secrets: the tenant's own service account never touches the
platform namespace. Editing a cluster class re-renders every Render
referencing it, in every namespace — the same live wiring the
namespaced class has.
Targets: a ConfigMap, or a Secret
target names exactly one destination:
target:
configMap: billing-rendered # workloads that mount files
file: config.properties
target:
secret: myapp-env # workloads that read SECRETS natively
shape: envEntries # …or take environment through envFrom
The Secret target is for the consumers the file path cannot reach: an
operator that watches Kubernetes Secrets reacts to
every reconcile with no pod restart — the Secret is a live object —
and an envFrom block turns shape: envEntries (every leaf of the
resolved document, dotted paths upper-snaked: db.pool_size →
DB_POOL_SIZE) into environment variables at the next container start.
Environment freezes at start; that is Kubernetes' rule, and this page
will not pretend otherwise. A vault class is allowed into the Secret
target — that is the container a secret store's document belongs in —
while the ConfigMap target keeps refusing it.
Feeding a name someone else already chose
The commonest enterprise shape: a helm chart or an operator demands a
Secret by name — auth.existingSecret in half the chart ecosystem, a
secretName: field in an operator's CRD — and refuses env vars or
files. That named Secret is exactly what a Render produces:
spec:
class: platform-vault
classKind: ClusterDynamicConfigClass
key: secret/postgres
target:
secret: pg-credentials
shape: entries # leaf keys VERBATIM: postgres-password, …
helm install db bitnami/postgresql --set auth.existingSecret=pg-credentials
Three shapes, three contracts: file when the consumer wants one
document under one key; envEntries when it reads through envFrom
(keys upper-snaked to the env dialect); entries when the key names
are someone else's contract — every leaf verbatim, postgres-password
staying postgres-password, because any mangling breaks a name you do
not own. The password never appears in values, in git, or in a
developer's hands; the store document is the single source, and the
Secret follows it on the Render's interval.
Deleting the Render deletes its target unless it says otherwise:
spec:
target:
secret: billing-credentials
deletionPolicy: Retain # or Delete, the default
Delete owns the target via ownerReferences, so the cleanup is
Kubernetes' own garbage collector — no finalizer to get wrong. Retain
writes no owner reference at all, which is the right answer for a Secret
something else still needs: deleting the Render that produced a
credential should not be the same act as revoking it everywhere.
status.renderedAt says when the last render landed and
status.lastError carries kind and shape only, never a value.
status.observedGeneration is the metadata.generation this status was
written for, so a controller or a kubectl wait can tell "this succeeded"
from "this has not been looked at since it changed".
Why a render is not ready
The Ready condition carries a machine-readable reason, and the three are
kept apart because they belong to different people:
| reason | what it means | who fixes it |
|---|---|---|
ClassNotFound | the class this render names does not exist | whoever wrote the Render |
ClassNotAllowed | it exists, and does not admit this namespace | the platform team that owns the class |
RenderFailed | the store could not be read, or the document could not be rendered | depends on the message |
Before 0.3.0 all three were RenderFailed, so every alert grepped a
message. These strings are API on the same terms as the annotation
contract: a dashboard aggregates on them, and a rename is a breaking
change.
Two honesty notes, stated before the reconciler shipped and still true:
- ConfigMap propagation is slow — the kubelet syncs mounted
ConfigMaps on its own cadence (up to a minute-plus; the
engine book's Kubernetes Files
page
walks the mechanism). The sidecar's emptyDir is the low-latency
path;
DynamicConfigRendertrades latency for no-sidecar. - A ConfigMap is not a Secret. The reconciler refuses vault
classes into ConfigMaps with exactly that sentence; the
secret:target is the lift, and the refusal message points at it.
The class annotation — still ahead
dynamic-config.rs/class on a pod (the webhook resolving a class so
annotations shrink) is contract-only: it needs the webhook to read
CRs, which trades away its zero-RBAC posture the way
selfRotate does, and that trade is taken
per-feature, not by default.
What the operator will not do
Own lifecycles inside pods, restart workloads on render, or template documents. It renders and it reports; reacting is the workload's business, and the whole engine exists so reacting is cheap.
One Agent per Node
Every other shape here puts an agent beside the application: one container per render, one fetch per pod. That is the right default — the credential is scoped to the workload that needs it, and a compromised agent reaches one application's configuration.
It is also 25,000 containers at 10,000 pods and 2.5 renders each, and that number is what this exists for.
volumes:
- name: config
csi:
driver: config.dynamic-config.rs
readOnly: true
volumeAttributes:
source: consul
endpoint: http://consul.default.svc:8500
key: myapp/config.json
path: rendered.toml # a name inside the volume
containers:
- name: app
volumeMounts:
- name: config
mountPath: /config
No annotations, no injected container, no webhook involved at all. The kubelet asks the node's agent for the volume, and the agent renders into it.
Why this is one component and not two
"A node-level agent" and "a CSI driver" were two entries on the same list, and building them separately would have been building a thing and its only delivery mechanism as though they were unrelated.
A DaemonSet that fetches for a whole node has to get bytes into a pod, and
there are two ways: a hostPath the pod also mounts — which restricted Pod
Security forbids, for the reason it forbids it — or a CSI volume, which
is the interface Kubernetes added for exactly this and which the kubelet
already knows how to mount, unmount and clean up after an eviction.
So this is a CSI node plugin whose backing store is the same engine the sidecar runs. Not a second implementation of fetch-resolve-render-watch: the sidecar's own crate, called as a library.
What it shares
Two pods on one node that want the same document from the same store under the same credential share one fetch and one watch. A node running a hundred pods that read one Consul key opens one connection to Consul.
The credential is part of that identity, not metadata beside it. Two pods reading one key under different tokens are two reads, and sharing them would hand one namespace's document to another under a credential it was never granted.
What they do not share is the rendered file: each pod gets its own bytes at
its own path, in its own format, with its own mode. A .properties reader
and a YAML reader on one node share the fetch and share nothing else.
Two series say whether any of it is working:
dynamic_config_node_agent_documents 12
dynamic_config_node_agent_readers 96
On a node where those two numbers are equal, nothing is being shared and a sidecar would have cost the same.
One property a sidecar cannot offer
The first render happens before the pod starts. The kubelet does not start a pod's containers until every volume is published, and publishing is what does the first fetch — so an application cannot observe a missing file, and there is no init container here because there is nothing for one to do.
What it costs
Said plainly, because it is the reason this is off by default and the sidecar is not.
- Credentials for many workloads in one process. A compromised node agent reaches every store credential every pod on that node uses. The sidecar's isolation is exactly what is traded away.
- It runs as root. The kubelet creates a pod's volume directories owned by root, and a plugin that cannot write into them cannot publish. Every other binary here runs nonroot; this one cannot.
- It mounts host paths. The kubelet's plugin and pod directories, which is what a CSI driver is.
Not a better shape. A different trade, for a scale that makes the first one untenable — and a decision to make with a measurement rather than a preference.
Installing it
helm upgrade dynamic-config … --set nodeAgent.enabled=true
nodeAgent.kubeletPath is /var/lib/kubelet everywhere except a few
distributions — k0s and some managed offerings move it, and a wrong value is
a driver the kubelet registers and never calls.
The DaemonSet carries upstream's node-driver-registrar beside the agent.
Its whole job is telling the kubelet this driver exists; writing that here
would be reimplementing a thing Kubernetes ships.
The attributes a volume takes
The same vocabulary as the annotations, for the same reason the agent's flags are: one contract, learned once, whichever shape delivers it.
| attribute | |
|---|---|
source | required — consul, vault, etcd, … |
key | required — the document's key, path or object |
path | required — a name inside the volume; the extension picks the format |
endpoint | the store's address |
section, auth, auth-mount, auth-role, auth-username, namespace, ref, api-url, file-mode, watch-seconds | as the annotations of those names |
path is checked rather than trusted: absolute paths and .. are refused,
because the kubelet owns the directory and a volume that could write outside
it could write into every other pod's volume on the node.
There is no mode: a CSI volume is published before the containers start
and stays. There is no env-inject: there is no command here to wrap. And
there is no metrics-port: the metrics are the node agent's, one endpoint
for the whole node.
Observability
Three components, three Prometheus endpoints, and OTLP traces beside them when a collector is configured. The metrics are hand-rolled exposition text — a handful of counters do not earn a metrics crate — and are scrapeable by anything that speaks the format, the OTel Collector included.
What each component exposes
| component | where | how it turns on |
|---|---|---|
| webhook | its own plain-HTTP port, and GET /metrics on the serving (TLS) port too | webhook.metrics.enabled, on by default |
| agent | its own port, plain HTTP | metrics-port annotation, or the agent.defaults.metricsPort fleet default |
| operator | 0.0.0.0:9090, plain HTTP | on by default; DYNAMIC_CONFIG_OPERATOR_METRICS_ADDR="" turns it off |
The webhook's metrics
The webhook sits in the path of every pod creation in the cluster, so the first question about it is never "did it work" but "how long did it take" — the API server's own ten-second timeout turns a slow admission into a refused one, and only a histogram answers what the tail is doing.
# TYPE dynamic_config_admissions_total counter
dynamic_config_admissions_total{outcome="skipped"} 41
dynamic_config_admissions_total{outcome="patched"} 12
dynamic_config_admissions_total{outcome="refused"} 3
# TYPE dynamic_config_admission_refusals_total counter
dynamic_config_admission_refusals_total{reason="policy"} 2
dynamic_config_admission_refusals_total{reason="pinned"} 1
dynamic_config_admission_refusals_total{reason="conflict"} 0
dynamic_config_admission_refusals_total{reason="malformed"} 0
dynamic_config_admission_refusals_total{reason="other"} 0
# TYPE dynamic_config_admission_duration_seconds histogram
dynamic_config_admission_duration_seconds_bucket{le="0.0001"} 39
dynamic_config_admission_duration_seconds_bucket{le="0.001"} 55
dynamic_config_admission_duration_seconds_bucket{le="+Inf"} 56
dynamic_config_admission_duration_seconds_sum 0.031
dynamic_config_admission_duration_seconds_count 56
# TYPE dynamic_config_admission_patch_bytes_total counter
dynamic_config_admission_patch_bytes_total 24576
skipped is a pod that did not ask; patched asked and got the
agent; refused asked wrongly. The refusals are labelled by kind
because they are three different pages:
| reason | what happened | who fixes it |
|---|---|---|
policy | a store, or an agent-env name, that this installation does not allow here | whoever wrote the pod, or whoever set the gate |
pinned | a value the installation fixed, overridden by a pod | whoever is working around the installation |
malformed | an annotation that is not the shape it has to be | whoever wrote the pod — a typo |
conflict | the pod already has a container by a name the injection needs | whoever wrote the pod — rename it |
The kind is a status.reason the refusal carries, not something a
scrape works out by reading the message, so rewording an error does not
silently re-label a metric.
In selfRotate mode two more say whether rotation is still happening:
# TYPE dynamic_config_certificate_rotations_total counter
dynamic_config_certificate_rotations_total 7
# TYPE dynamic_config_certificate_expires_at_seconds gauge
dynamic_config_certificate_expires_at_seconds 1755698600
A counter that should climb on a schedule, and a wall-clock second that should always be in the future. Alert on the gauge: a webhook whose certificate expires takes every pod creation in the cluster with it.
- alert: DynamicConfigWebhookCertificateExpiring
expr: dynamic_config_certificate_expires_at_seconds - time() < 3600
- alert: DynamicConfigAdmissionSlow
expr: histogram_quantile(0.99, rate(dynamic_config_admission_duration_seconds_bucket[5m])) > 1
Where to scrape it
A port of its own, in plain HTTP — webhook.metrics.port, 9091 by
default. The admission port is mutual TLS against a CA the webhook
mints for itself, so scraping that meant handing Prometheus a client
certificate from that CA, and a deployment that did not simply had no
metrics. The admission port still answers /metrics, so a scrape
already configured against it keeps working.
With the Prometheus Operator, one toggle:
webhook:
metrics:
serviceMonitor:
enabled: true # needs the ServiceMonitor CRD
interval: 30s
Without it, the Service carries a named metrics port:
- job_name: dynamic-config-webhook
kubernetes_sd_configs:
- role: endpoints
namespaces: { names: [dynamic-config] }
relabel_configs:
- source_labels: [__meta_kubernetes_endpoint_port_name]
action: keep
regex: metrics
Watching
The sidecar watches; it does not poll. Each store says how it learns
that its document changed, and the agent uses it — so a change in etcd,
Consul, NATS, Redis or a config server arrives as it happens, and a store
that must be asked is asked the cheapest question it offers (an S3 object
is a HEAD, not a download).
watch-seconds keeps its spelling and means two things depending on the
store:
| the store | what a watch is | what watch-seconds means |
|---|---|---|
| etcd, Consul, NATS, Redis, config-server | the store's own push | a resync — how often to re-read anyway |
| Vault, S3, git, Firestore | a cheap "has it changed?" | how often to ask |
examples/watch-driven.yaml
is a pod on the push path with both numbers set: a five-minute resync
under etcd's stream, and the metrics port to scrape it from.
The resync is not belt and braces. The failure mode of a stream is silence: a subscription the broker forgot, a connection that dropped without an error, a blocking query answering an index that will never move again. All three look exactly like a store where nothing has changed, and the only way to tell them apart is to go and ask.
The agent's metrics
A watching agent with a metrics-port serves twenty-seven series; the port
also lands on the container spec as a named port (metrics), so
selector-based discovery finds it without configuration. Since 0.3.0 the
chart gives every installation a default port (9110), so this is on unless
somebody turns it off with metrics-port: "0":
# TYPE dynamic_config_agent_renders_total counter
dynamic_config_agent_renders_total 128
# TYPE dynamic_config_agent_render_failures_total counter
dynamic_config_agent_render_failures_total 2
# TYPE dynamic_config_agent_last_render_timestamp_seconds gauge
dynamic_config_agent_last_render_timestamp_seconds 1755612200
# TYPE dynamic_config_agent_deliveries_total counter
dynamic_config_agent_deliveries_total 12
# TYPE dynamic_config_agent_resyncs_total counter
dynamic_config_agent_resyncs_total 480
# TYPE dynamic_config_agent_watch_connected gauge
dynamic_config_agent_watch_connected 1
# TYPE dynamic_config_agent_watch_reconnects_total counter
dynamic_config_agent_watch_reconnects_total 3
# TYPE dynamic_config_agent_staleness_seconds gauge
dynamic_config_agent_staleness_seconds 4
# TYPE dynamic_config_agent_generation gauge
dynamic_config_agent_generation 41
# TYPE dynamic_config_agent_lease_renewals_total counter
dynamic_config_agent_lease_renewals_total 6
# TYPE dynamic_config_agent_lease_renewal_failures_total counter
dynamic_config_agent_lease_renewal_failures_total 0
# TYPE dynamic_config_agent_lease_revocations_total counter
dynamic_config_agent_lease_revocations_total 0
# TYPE dynamic_config_agent_lease_ttl_seconds gauge
dynamic_config_agent_lease_ttl_seconds 3600
# TYPE dynamic_config_agent_absent gauge
dynamic_config_agent_absent 0
# TYPE dynamic_config_agent_absent_total counter
dynamic_config_agent_absent_total 0
# TYPE dynamic_config_agent_notifications_total counter
dynamic_config_agent_notifications_total 12
# TYPE dynamic_config_agent_notification_failures_total counter
dynamic_config_agent_notification_failures_total 0
# TYPE dynamic_config_agent_drift gauge
dynamic_config_agent_drift 0
# TYPE dynamic_config_agent_drift_total counter
dynamic_config_agent_drift_total 0
# TYPE dynamic_config_agent_canary_holding gauge
dynamic_config_agent_canary_holding 0
# TYPE dynamic_config_agent_canary_percent gauge
dynamic_config_agent_canary_percent 100
# TYPE dynamic_config_agent_acks_total counter
dynamic_config_agent_acks_total 41
# TYPE dynamic_config_agent_ack_mismatches_total counter
dynamic_config_agent_ack_mismatches_total 0
# TYPE dynamic_config_agent_applied gauge
dynamic_config_agent_applied 1
# TYPE dynamic_config_agent_unapplied_seconds gauge
dynamic_config_agent_unapplied_seconds 0
# TYPE dynamic_config_agent_tls_reloads_total counter
dynamic_config_agent_tls_reloads_total 0
# TYPE dynamic_config_agent_tls_verification_skipped gauge
dynamic_config_agent_tls_verification_skipped 0
generation is the store's own revision of what was last rendered, where
the store counts them — a Vault KV version, a Consul index, an etcd
revision. Zero for a store whose revision is opaque, because an ETag has no
number to report and inventing one would make a dashboard compare things
that do not compare.
The four lease_* series are zero unless the source is a dynamic-secret
engine. lease_ttl_seconds is what the store last granted, not what was
asked for.
lease_renewal_failures_total is worth alerting on precisely because it
should stay at zero. A lease the store marked renewable: false — every
pki/issue, and a database credential past its role's maximum — is never
sent a renewal in the first place; it is re-issued instead, and
lease_renewals_total stays at zero while the credential keeps arriving.
So a rising failure count is always a lease that was expected to renew
and did not, which is a real thing to page on rather than the background
noise it would be if every non-renewable lease were asked anyway.
tls_verification_skipped is 1 when the agent runs with
tls-skip-verify. A gauge
rather than a log line repeated per fetch, because the question it answers
is a fleet question — which of these five thousand pods are doing this? —
and one alert over a gauge answers it where five thousand log streams do
not. It should stay at zero.
canary_holding is 1 while this pod is outside a
canary cohort and is
holding a document it fetched — without it, a held document looks exactly
like a store with nothing new to say. canary_percent is the number it last
read, and 100 when no canary is configured.
The only end-to-end answer
renders_total says a document reached disk. applied says the
application is running it — the four ack series are the difference
between "we published" and "it converged", and an application still running
the previous document while every other series reports success is the
outage they close.
They stay at zero unless the application acknowledges, which is not a
failure: an application that never acknowledges is not penalised, it simply
leaves applied at zero. unapplied_seconds is the one to alert on —
ack_mismatches_total climbing steadily means acknowledgements and renders
are talking past each other, which is a different problem from being behind.
The annotations page
has the two-line contract, and require-ack turns it into readiness.
tls_reloads_total counts the times the store's client was rebuilt for
rotated trust material — a counter rather than a gauge, because the question
worth asking is whether a rotation was picked up at all.
drift is 1 while a rendered file is something other than what this agent
wrote. notification_failures_total counts post-render calls that were
refused or unanswered; neither is fatal, because the file is already
published by the time either can happen.
absent is 1 while the store says the document is not there, and
absent_total counts how many times it has gone missing. The distinction
they encode is the one a stale file cannot make on its own: a store that
does not answer is an outage, which waiting cures, and a store that answers
gone is a deletion, which waiting does not. Before 0.3.0 a deleted Vault
secret moved nothing at all and the last render went on being served with
every health check reporting fine. What the agent does about it is
on-delete.
/readyz means there is a document
The agent's readiness is stricter than the operator's, and since 0.3.0 the
webhook attaches a probe to the injected container: /readyz answers 503
until a document has been rendered, and pod readiness is already AND-ed
across containers — so a Service sends no traffic to a pod whose
configuration does not exist yet.
The trade, said plainly: with the store unreachable at start and nothing cached, the pod never becomes ready. That is correct — the application has no configuration — but it turns a store outage into a visibly stalled rollout rather than pods that come up and misbehave. Two ways out, and they answer different questions:
dynamic-config.rs/startup-policy: allow-cached(the default) serves the file already on disk when the first fetch fails. The volume survives a container restart, so an agent coming back after a crash usually finds its own last render there.dynamic-config.rs/readiness: "false"detaches the probe entirely, for a deployment that would rather start than wait.
And dynamic-config.rs/max-staleness: "6h" puts a ceiling on the first of
those: last-known-good answers is there a document and leaves is it too
old to trust open. A credential may be worthless after five minutes and a
feature flag fine after a week, so there is no default — the ceiling is off
unless somebody sets one.
staleness_seconds is the one to alert on. It is seconds since the
store was last read successfully, delivered or resynced — the number a
pager asks for. Everything else says what happened; this says how long
ago it stopped happening. Zero until the first success, so a pod that has
never reached its store does not look fresh.
# The document went stale: nothing read from the store for 10 minutes.
- alert: DynamicConfigStoreStale
expr: dynamic_config_agent_staleness_seconds > 600
- alert: DynamicConfigRenderFailing
expr: increase(dynamic_config_agent_render_failures_total[10m]) > 0
# A watch that keeps reopening is a store or a network that is not well.
- alert: DynamicConfigWatchFlapping
expr: increase(dynamic_config_agent_watch_reconnects_total[15m]) > 5
Deliveries and resyncs are told apart because that pair is a diagnosis. Deliveries flat while resyncs climb is the stalled stream the resync exists to cover: the store is being read, changes are landing, and the push half is doing nothing. Nothing else here would show it.
watch_connected is 0 between a watch ending and the next one opening;
on a store that is polled rather than pushed it stays 1 for as long as
the loop runs.
A one-shot init agent serves nothing — it renders once and exits, and
a scrape target that lives for two seconds is noise. Only the watching
half carries the port, which is also why metrics-port pairs with
mode: sidecar or both.
With a PodMonitor (Prometheus Operator), the named port makes discovery one selector:
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: dynamic-config-agents
spec:
selector:
matchExpressions:
- { key: app, operator: Exists } # your workload labels
podMetricsEndpoints:
- port: metrics
Fleet-wide, set agent.defaults.metricsPort: "9102" once and every
injected watching agent serves on 9102; a pod that must not (a
port collision, say) opts out with metrics-port: "0".
The operator's metrics
The agent's render series, dynamic_config_operator_-prefixed —
_renders_total counts reconciles that produced or refreshed a target,
_render_failures_total the ones that ended in a RenderFailed event,
and the gauge the last success — plus a reconcile histogram:
# TYPE dynamic_config_operator_reconcile_duration_seconds histogram
dynamic_config_operator_reconcile_duration_seconds_bucket{le="0.5"} 88
dynamic_config_operator_reconcile_duration_seconds_bucket{le="+Inf"} 91
dynamic_config_operator_reconcile_duration_seconds_sum 12.4
dynamic_config_operator_reconcile_duration_seconds_count 91
# TYPE dynamic_config_operator_reconcile_failures_total counter
dynamic_config_operator_reconcile_failures_total 3
A reconcile is a store fetch and an API write, so a tail that moves is
usually the store rather than the operator. On by default at
0.0.0.0:9090; point DYNAMIC_CONFIG_OPERATOR_METRICS_ADDR elsewhere
or set it empty to turn it off. The same port answers /healthz and
/readyz, which is what the chart's probes use.
The operator's /readyz means the process is up, and that is not an
oversight: an operator with no DynamicConfigRender resources has rendered
nothing and is working perfectly. The agent means something stricter by
the same path — see below.
Only one replica reconciles. The operator elects a leader over a
Lease, so operator.replicas: 2 is spare capacity rather than twice the
work — and the metrics of a follower stay flat by design. A leader that
dies is replaced within the lease's fifteen-second term.
Beside the metrics, the operator reports per-object: a Ready
condition on every DynamicConfigRender's status and Kubernetes
Events (Rendered / RenderFailed) on the object — kubectl describe dynamicconfigrender is the first debugging stop, before any
dashboard.
Logs
Every component writes JSON to stdout through tracing, one line per
event, RUST_LOG grammar via the environment. The webhook logs one
audit line per non-skipped admission — namespace, pod, source, outcome,
and NO annotation values, because store addresses and role names
belong in the cluster, not in every log aggregator. The agent logs
each render with byte counts, and failures with the store's error.
Turning up an agent's verbosity is the canonical agent-env use:
# the installation allows it: webhook.agentEnvAllow: "*: RUST_LOG"
dynamic-config.rs/agent-env: "RUST_LOG=debug"
— or fleet-wide for a day with agent.defaults.env: "RUST_LOG=debug",
no allowlist needed, the installer owns
both.
The engine's own metrics
The series above measure the DELIVERY machinery: admissions, renders,
staleness. The configuration engine inside your application has its
own richer contract — dynamic_config_reload_total, generation and
snapshot-age gauges, per-source health — defined once in
the metrics contract
of the engine book and exported by the language bindings. When an app
consumes the rendered file through the engine's own watcher, alert on
the engine's series (closest to the truth the app sees) and keep the
agent's gauge as the delivery-side backstop.
A dashboard and the rules under it
deploy/observability/ carries both, over the series named above:
kubectl apply -f deploy/observability/rules.yaml # recording rules, then alerts
# Grafana → Dashboards → Import → dashboard.json
The rules come first — four of the dashboard's fleet counters read recording rules, and without them those panels stay empty.
Two things worth knowing before tuning the thresholds. Staleness is the
gauge to page on, not a failure counter: an agent that fails one fetch and
recovers is doing what it was built to do, while one whose store has not
answered in half an hour is serving a document nobody can vouch for. And
lease_renewal_failures_total should sit at zero, which is what makes a
low threshold on it reasonable — a lease the store marked renewable: false
is never sent a renewal at all, so anything counted there was expected to
renew and did not.
The recording rules aggregate pod away. A thousand agents cost four series
through them and four thousand without; the alerts that do keep pod are
the ones whose whole purpose is naming which pod.
OpenTelemetry
Since 0.3.0 the webhook, the agent and the operator export OTLP traces — and only when a
collector is configured. Setting OTEL_EXPORTER_OTLP_ENDPOINT is the whole
opt-in; leave it unset and nothing is exported, nothing is allocated, and
the process behaves exactly as it did. The variable is OpenTelemetry's own,
so a collector that already injects it into pods needs no annotation from
this chart.
Resource attributes come from the downward API — service.name,
service.version, k8s.pod.name, k8s.namespace.name, k8s.node.name —
and are absent rather than invented when nobody wires them: a made-up
pod name is worse than none.
Where the pipeline is built
The engine offers an otel feature that records into a meter somebody else
configured, and refuses to build the pipeline. That is the right split: a
pipeline owns a runtime, a batch processor and a shutdown, and a library
that takes those over is one that has to be fought. These three are
programs, and programs own main — so the exporter, and the flush before
exit, are theirs.
If a collector cannot be reached, the process logs a warning and carries on without traces. A configuration agent that refused to start because its telemetry was down would be a worse outage than the one it was reporting.
The Prometheus text is untouched
Nothing above replaces the scrape. Many deployments have one already, and the collector reads it directly — neither way out is the price of the other:
-
Metrics: the OTel Collector's
prometheusreceiver scrapes all three endpoints and exports OTLP wherever you aggregate:receivers: prometheus: config: scrape_configs: - job_name: dynamic-config-agents kubernetes_sd_configs: [{ role: pod }] relabel_configs: - source_labels: [__meta_kubernetes_pod_container_port_name] regex: metrics action: keep -
Logs: the JSON lines on stdout are structured input for the
filelogreceiver, no parsing regexes required.
Traces are spans over the work — an admission, a fetch, a render — and the scrape is the state of it. A deployment that wants only one of the two is not missing anything by taking only one.
Troubleshooting
Each symptom, its first command, and the usual cause. The agent's logs
are JSON; kubectl logs <pod> -c dynamic-config-agent | jq . is the
readable form.
The file never appears
kubectl get pod <name> -o jsonpath='{.spec.containers[*].name}'
No dynamic-config-agent in the list → the webhook never mutated.
In order of likelihood:
- The pod predates the webhook install — admission only sees CREATE.
failurePolicy: Ignoreswallowed a webhook outage:kubectl get events -n <ns> | grep dynamic-configand the webhook deployment's own logs say which.- The annotation said
inject: "false"or misspelled the prefix —dynamic-config.rs/, with the dot and the slash.
Agent present, file absent → the agent is failing. Its log carries the store error verbatim minus values:
kubectl logs <pod> -c dynamic-config-agent | jq -r '.fields.error // .fields.message'
The admission was denied
That is the contract working: inject: "true" with a missing or
malformed companion annotation fails the pod's creation, and the
reason names the annotation:
Error creating: admission webhook "inject.dynamic-config.rs" denied the
request: dynamic-config.rs/inject is true, so dynamic-config.rs/path is required
Silently starting without configuration is the failure mode this refusal exists to prevent; add the named annotation.
"container name is duplicated" from the API server
Error creating: Pod "app" is invalid: spec.containers[2].name:
Duplicate value: "dynamic-config-agent"
Your pod already has a container by a name the injection needs. The webhook refuses that now, with the name to rename:
admission webhook "inject.dynamic-config.rs" denied the request: this pod
already has a container called "dynamic-config-agent", and the injection
needs that name — rename yours, or set dynamic-config.rs/inject to "false"
Seeing the API server's version of it instead means a webhook older than 0.2.0 patched the pod. Upgrade, or rename the container.
The injected names are dynamic-config-agent and
dynamic-config-init, plus -1, -2… for each
named render.
The same pod was injected twice
Two dynamic-config-agent containers in a pod nobody wrote twice is
admission running twice. It happens with
webhook.reinvocationPolicy: IfNeeded, which asks the API server to
call this webhook again whenever a later webhook changes the pod, and
with controllers that resubmit an already-admitted spec.
Since 0.2.0 the patch marks the pod — dynamic-config.rs/status: injected — and a marked pod is passed through untouched. If you are
seeing this, check the webhook image is 0.2.0 or later:
kubectl -n dynamic-config get deploy dynamic-config-webhook \
-o jsonpath='{.spec.template.spec.containers[0].image}'
401/403 from the store
The token travels in DYNAMIC_CONFIG_AGENT_TOKEN, not in annotations.
kubectl exec <pod> -c dynamic-config-agent -- env | grep -c DYNAMIC_CONFIG
0 means the Secret was never mounted onto the agent container.
Config-server 401s specifically: the bearer must belong to a
[[server.clients]] block whose applications list names the
application in your --key — the server's audit log line for the
refusal names the client it matched.
The rendered file is stale
The sidecar keeps the last good render on fetch failure, on purpose —
staleness with a warning beats an empty file. The log says so at
warn level. --watch-seconds too high is the boring cause; a store
ACL that started refusing is the interesting one, and the 401 section
above applies.
mode: init never refreshes by design: rotation there is a pod
restart, which is the trade the Vault page states.
The pod never becomes ready
Since 0.3.0 the webhook attaches a readiness probe to the injected
container, so a pod stays 0/2 until its first render lands. kubectl logs <pod> -c dynamic-config-agent says why the fetch is not succeeding.
Three ways out, and they answer different questions:
- the store is genuinely down and a stale document is acceptable →
startup-policy: allow-cached(the default) already serves the file on the volume if there is one; there is none on a first start - the document is fine but too old →
dynamic-config.rs/max-stalenessis what flipped readiness; thestaleness_secondsgauge says by how much - the pod should start regardless →
dynamic-config.rs/readiness: "false", and the application handles a configuration that is not there yet
The document was deleted and nothing happened
Check dynamic_config_agent_absent. If it is 1, the store is answering
gone and the agent is doing what on-delete says — retain by default,
which keeps serving the last render. remove truncates the file and
fail ends the agent so the pod restarts.
If it is 0 while the key is definitely gone, the store is not reporting
the deletion: a Conditional store polled on an interval takes up to one
watch-seconds to notice.
The render is refused and the old file keeps serving
Two checks can refuse a document after it has been fetched, and both keep the last good file rather than publishing a bad one:
- the schema (
schema-configmap) — the log line names the failing path and the constraint, never the value - the size ceiling (
max-document-bytes, 8 MiB) — the refusal names the limit and the store; a document that grew past it is usually a key that now points at something else
--out refused at startup
The extension picks the format, and only five are legal: .json
.toml .yaml .ini .properties. The error lists them; .conf and
.cfg are nobody's format and stay refused.
Arrays in the document, flat output requested
`hosts` is an array, and neither flat format has one; render to json,
toml or yaml instead
Exactly what it says: pick a structured output, or reshape the document. The Rendering page owns the reasoning.
Reading what the webhook actually did
The golden test's fixture is the contract, and a live pod can be compared against it:
kubectl get pod <name> -o json | jq '.spec.containers[].volumeMounts'
kubectl get pod <name> -o json | jq '.spec.volumes[] | select(.name == "dynamic-config")'
The pod was created, nothing was injected
In order of likelihood:
-
The namespace is excluded. kube-system, kube-node-lease and the chart's own namespace never get injection, plus anything in
webhook.excludeNamespaces:kubectl get mutatingwebhookconfiguration dynamic-config \ -o jsonpath='{.webhooks[0].namespaceSelector}' -
The webhook was down and
failurePolicy: Ignorelet the pod through. The API server records exactly that:kubectl get events --field-selector reason=FailedAdmissionWebhook -A kubectl -n <chart-namespace> get pods -l app.kubernetes.io/component=webhook -
TLS trust is broken — the caBundle does not match what the webhook serves. With the self-signed default this happens when the Secret was deleted but the webhook configuration was not re-rendered;
helm upgradeheals both sides. The API server's opinion:kubectl logs -n <chart-namespace> deploy/dynamic-config-webhook | tail
The webhook pod refuses to start: an installation setting
dynamic-config-webhook: /etc/dynamic-config/installation.yaml:
"watchSecnods" is not an installation setting; the ones there are:
cpuRequest, memoryRequest, …, sourceDeny
A typo in the mounted installation document, refused at startup rather
than ignored — a default that silently never applied is a fleet running
without the posture somebody thought they had set. Fix the key in
agent.defaults / webhook.* (chart) or in
base/installation.yaml (kustomize).
The same check covers shapes: a store's settings are a map, a gate's namespaces map to lists, and a setting is a word, a number or a boolean.
The webhook pod refuses to start: "no TLS material"
The exact message names the two paths it looked at. The Secret
dynamic-config-webhook-tls is missing or empty — with cert-manager
enabled, check the Certificate:
kubectl describe certificate dynamic-config-webhook
An injected pod is rejected by Pod Security admission
It should not be: the injected container carries the full restricted posture. If a namespace enforces something stricter than restricted (an OPA/Kyverno policy), read the denial message — the security page lists every field the injection sets, so the diff is one screen.
The template refuses to render
Two shapes, two places:
-
At startup (a parse error, an undefined key on the first render): the agent exits and the reason is in the injected container's log — strict on purpose, so a typo'd
{{ db.hots }}cannot ship an empty string with a clean exit code. -
During a watch (the ConfigMap was edited into an error): the pod keeps its last good file and the agent logs the render error every tick until the template is fixed. Check with:
kubectl logs <pod> -c dynamic-config-agent | tail kubectl get configmap billing-template -o jsonpath='{.data.template}'
Stability & Versioning
Experimental, stated plainly: this is the youngest repository in the
organisation and the annotation contract is v1 — and since 0.3.0 that
contract is a registry rather than a list somebody keeps in step: every
key is one row carrying whether it takes a .name suffix and whether it has
been retired, a test checks this book against it, and a retired key will be
admitted with a warning naming its replacement rather than refused. Nothing
is retired yet. The operator's
reconcilers shipped in 0.1.1 and it stays 0.x until they have soak
history, whatever the rest of the family does.
- The annotation contract is the API; a breaking change to it bumps
the minor and regenerates the golden file in the same commit. It grew
additively in 0.2.0:
dynamic-config.rs/status, which the webhook writes on a pod it has patched so that a second admission does not patch it again. - The installation is one contract with three spellings — chart values, a mounted YAML document, or environment variables. A map is rendered to the same grammar the variables carry and goes through the same parser, so adding a spelling is not adding a semantics.
- The three components version together; images are the artefacts.
- The agent's store list grew additively (etcd, nats and s3 landed in 0.1.1) and is complete at nine; a tenth would follow the same rule.
- The engine dependency is a caret: an engine patch reaches the images on rebuild.
The repository's ROADMAP carries the ladder in full — the async stores and etcd's two-methods-forever answer, the self-rotating webhook TLS mode and the one narrow RBAC it will cost, the operator's reconcilers.