Wire kube-apiserver and kubelet tracing to a Jaeger collector on docker-29, deploy Grafana Alloy in-cluster to ship logs/metrics, and stand up Mimir + Loki on docker-30 as their backing stores. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
6.8 KiB
Tracing enable/disable runbook for K8s components
Context
The pod-creation tracing plan (plans/2026-05-21 20:15 - k8s-pod-creation-tracing.md) covers the one-time setup of OTLP tracing across kube-apiserver and kubelet, plus the Jaeger backend on docker-29. What's missing is a short operational doc for toggling tracing on/off as quickly as possible on the running cluster — useful for short measurement experiments where we want overhead only when we're actively recording.
The Terraform/cloud-init changes already in flight (master.tf, files/manifests/kube-apiserver.yaml, files/tracing-config.yaml) cover rebuilds. This runbook covers the live cluster: kube-master-31 + kube-node-32 + kube-node-33, K8s 1.32.
Output of this plan: a single new file docs/kubernetes-tracing.md containing three runbooks (quick suppress, full enable, full disable) and a verification section. No code or manifest changes — pure documentation.
What "as quickly as possible" means
There are two speeds of toggling:
- Quick suppress — leave the tracing flags/config in place on the components; just stop Jaeger so span exports go nowhere. No K8s component restart, takes seconds, but components still pay the (small) cost of building spans and attempting to export.
- Full enable / full disable — edit configs and restart components. apiserver static-pod restart is automatic via kubelet manifest watch (~5–10s); kubelet itself needs
systemctl restart kubeleton each node (brief NotReady blip per node, no pod evictions). This is what we'd use to truly remove overhead between experiments.
The runbook will present both, with quick-suppress as the default for "I'm done measuring for now" and full-disable for "we don't need this on the cluster anymore".
Deliverable: docs/kubernetes-tracing.md
Structure
# Kubernetes component tracing
## Components and where the toggles live
- table: component | config file on node | restart trigger
- kube-apiserver | /etc/kubernetes/manifests/kube-apiserver.yaml (flag) +
/etc/kubernetes/tracing-config.yaml (config) | kubelet auto-restarts static pod
- kubelet | /var/lib/kubelet/config.yaml (`tracing:` block) +
kube-system/kubelet-config CM (cluster-wide source of truth) | systemctl restart kubelet
## Quick suppress (no component restart)
1. `ssh novakj@192.168.0.29 'cd /path/to/tracing && docker compose stop jaeger'`
2. Components continue running with tracing wired up but spans go nowhere.
3. To resume: `docker compose start jaeger`.
Caveat: components still build spans and attempt OTLP export — small CPU/network overhead. For a real "off", use full disable.
## Full enable (live cluster)
Prereqs: Jaeger is up on docker-29 (`docker compose up -d` in `docker-29/tracing/`).
### kube-apiserver (master only)
On 192.168.0.31:
```bash
# 1. Drop the TracingConfiguration file (same content as files/tracing-config.yaml)
sudo tee /etc/kubernetes/tracing-config.yaml > /dev/null <<'EOF'
apiVersion: apiserver.config.k8s.io/v1beta1
kind: TracingConfiguration
endpoint: 192.168.0.29:4317
samplingRatePerMillion: 1000000
EOF
# 2. Back up the apiserver manifest
sudo cp /etc/kubernetes/manifests/kube-apiserver.yaml /root/kube-apiserver.yaml.pre-tracing
# 3. Add the flag (insert after --tls-private-key-file line)
sudo sed -i '/--tls-private-key-file=/a\ - --tracing-config-file=/etc/kubernetes/tracing-config.yaml' /etc/kubernetes/manifests/kube-apiserver.yaml
# 4. Wait for kubelet to restart the static pod (~5-10s)
sudo crictl ps | grep apiserver # new container id == restarted
kubectl get --raw /livez # should return "ok"
kubelet (every node: master + workers)
Per-node edit:
# Append the tracing block under kubelet config (idempotent: check first)
grep -q '^tracing:' /var/lib/kubelet/config.yaml || sudo tee -a /var/lib/kubelet/config.yaml > /dev/null <<'EOF'
tracing:
endpoint: 192.168.0.29:4317
samplingRatePerMillion: 1000000
EOF
sudo systemctl restart kubelet
Stagger across nodes (one at a time) so we don't NotReady the whole cluster at once.
Then make it survive node rebuilds/joins by updating the cluster ConfigMap:
kubectl -n kube-system edit cm kubelet-config
# add the same `tracing:` block under the `kubelet:` key
Full disable (live cluster)
kube-apiserver
# Restore the pre-tracing manifest (preferred — exact revert)
sudo cp /root/kube-apiserver.yaml.pre-tracing /etc/kubernetes/manifests/kube-apiserver.yaml
# OR remove just the flag in place:
sudo sed -i '/--tracing-config-file=/d' /etc/kubernetes/manifests/kube-apiserver.yaml
# Optional: remove the now-unused config
sudo rm /etc/kubernetes/tracing-config.yaml
kubelet restarts the static pod automatically.
kubelet (every node)
sudo sed -i '/^tracing:/,/samplingRatePerMillion:/d' /var/lib/kubelet/config.yaml
sudo systemctl restart kubelet
Then strip the tracing: block from kube-system/kubelet-config CM so node rejoins don't reintroduce it.
Verification (after enable)
- Open
http://192.168.0.29:16686— Services dropdown showsapiserver(appears after the first traced request) andkubelet. - Generate end-to-end traffic:
kubectl run trace-probe --image=registry.k8s.io/pause:3.10 --restart=Never kubectl wait --for=condition=Ready pod/trace-probe --timeout=60s kubectl delete pod trace-probe - In Jaeger: filter service
apiserver, look for spans tagged with the probe pod name; filter servicekubelet, look forsyncPodspans. - To verify disable: after running disable steps, repeat (2). Within a minute, no new spans for
apiserver/kubeletshould appear in Jaeger.
Notes
- 100% sampling is fine for short experiments; for steady-state leave tracing off (quick suppress is enough between sessions).
- The flag/config are GA in K8s 1.32 — no
--feature-gatesneeded. - This runbook only covers apiserver + kubelet. etcd and kube-scheduler tracing are documented as future steps in the parent plan.
## Files to change
| File | Change |
|---|---|
| `docs/kubernetes-tracing.md` | **new** — the runbook described above |
No other files modified. The Terraform/cloud-init side stays as already drafted in the parent plan.
## Verification of this plan's output
After the runbook is written:
1. Skim `docs/kubernetes-tracing.md` end-to-end and check that someone unfamiliar with the setup can follow it without opening the parent plan.
2. Cross-check the sed/tee commands against the actual files on the live master via SSH — confirm line patterns match (e.g., the `--tls-private-key-file=` line is present so the sed insertion lands in the right spot).
3. Dry-run the quick-suppress flow first (it's the lowest-risk path) and confirm Jaeger UI stops receiving new spans within ~30 s.