Files
Jan Novak d43ffd488e observability: add k8s API/kubelet tracing, Alloy, Mimir and Loki
Wire kube-apiserver and kubelet tracing to a Jaeger collector on
docker-29, deploy Grafana Alloy in-cluster to ship logs/metrics, and
stand up Mimir + Loki on docker-30 as their backing stores.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-09 00:05:21 +02:00

143 lines
5.8 KiB
Markdown

# Alloy: deployment & operations
Three Grafana Alloy instances run in the `monitoring` namespace, all deployed via
the Grafana `alloy` Helm chart (`>=1.8.0 <2.0.0`, image `v1.17.1`) and managed by
Flux with `driftDetection: enabled` (manual `kubectl edit`s get reverted — always
change config via git). Logs and events ship to Loki; metrics ship to Mimir. Every
instance stamps `cluster="homelab"` as an external label.
---
## Instances
| Instance | Controller | Collects | Sink |
|---|---|---|---|
| `alloy-logs` | DaemonSet (`runAsUser: 0`, one pod per node) | pod logs + node journald | Loki `http://192.168.0.30:3100/loki/api/v1/push` |
| `alloy-events` | Deployment (1 replica) | Kubernetes events | Loki (same endpoint) |
| `alloy-metrics` | Deployment (1 replica) | all ServiceMonitors + PodMonitors cluster-wide | Mimir `http://192.168.0.30:9009/api/v1/push` |
`alloy-logs` and `alloy-events` are one-per-node / cluster-scoped respectively —
`alloy-metrics` is also a single Deployment, never a DaemonSet, since its
Prometheus-Operator discovery is already cluster-wide (a DaemonSet would scrape
every target once per node).
---
## Files & Flux wiring
All manifests live in `gitops/home-kubernetes/alloy/`:
- `helmrepository_grafana.yaml` — the shared `grafana` HelmRepository.
- `helmrelease_alloy-logs.yaml` — logs DaemonSet.
- `helmrelease_allow-events.yaml` — events Deployment (note the `allow` filename typo, kept as-is).
- `helmrelease_alloy-metrics.yaml` — metrics Deployment.
- `rbac_log-collector.yaml` — ClusterRole `alloy-log-reader`, bound to the `alloy-logs`/`alloy-events` ServiceAccounts.
- `rbac_metrics-collector.yaml` — ClusterRole `alloy-metrics-reader`, bound to the `alloy-metrics` ServiceAccount.
Each HelmRelease sets `rbac.create: false` — RBAC is supplied manually by the two
files above rather than by the chart, since the chart's default ClusterRole
doesn't cover everything each pipeline needs (journald/events reads, or
ServiceMonitor/PodMonitor/kubelet-metrics reads).
Reconciliation: a Flux `Kustomization` named `alloy` in
`gitops/home-kubernetes/flux-system/extra-kustomizations.yaml`, pointing at
`./gitops/home-kubernetes/alloy`, with `dependsOn: kube-prometheus` — it needs
the `monitoring` namespace and the Prometheus Operator CRDs (ServiceMonitor/
PodMonitor) to exist first.
---
## How metrics discovery works
`alloy-metrics` runs two Alloy components — `prometheus.operator.servicemonitors`
and `prometheus.operator.podmonitors` — which discover and scrape **every**
ServiceMonitor/PodMonitor in the cluster, then forward to
`prometheus.remote_write` targeting Mimir. There's no bespoke scrape config for
"standard Kubernetes metrics": kube-prometheus-stack already ships
ServiceMonitors for kubelet, cAdvisor, kube-state-metrics, node-exporter,
apiserver, and coredns, so those are picked up automatically. Any future
ServiceMonitor/PodMonitor added anywhere in the cluster is picked up the same way
with no Alloy config change.
This is **additive** — kube-prometheus-stack's bundled Prometheus keeps scraping
the same targets into its own local 60d TSDB. Alloy scraping the same targets a
second time and shipping to Mimir is expected double-collection, not a bug.
---
## Operating
Status:
```bash
flux get helmreleases -n monitoring
kubectl get pods -n monitoring -l app.kubernetes.io/name=alloy -o wide
```
Force reconcile:
```bash
flux reconcile kustomization alloy --with-source
flux reconcile helmrelease alloy-metrics -n monitoring
```
Logs:
```bash
kubectl -n monitoring logs deploy/alloy-metrics
kubectl -n monitoring logs ds/alloy-logs
kubectl -n monitoring logs deploy/alloy-events
```
Alloy UI (component graph, target health, remote_write queue status):
```bash
kubectl -n monitoring port-forward deploy/alloy-metrics 12345:12345
# open http://localhost:12345
```
Editing config: change the inline `alloy.configMap.content` block in the
relevant `helmrelease_*.yaml`, commit, push, then force-reconcile (above).
Do not `kubectl edit` the generated ConfigMap directly — `driftDetection` will
revert it on the next reconcile.
---
## Verification / smoke tests
**Metrics reaching Mimir:**
```bash
curl -s 'http://192.168.0.30:9009/prometheus/api/v1/query?query=up' | jq '.data.result | length'
# expect > 0
```
Spot-check a few standard series exist and carry the cluster label:
```bash
curl -s 'http://192.168.0.30:9009/prometheus/api/v1/query?query=kube_pod_info' | jq '.data.result[0].metric'
curl -s 'http://192.168.0.30:9009/prometheus/api/v1/query?query=node_cpu_seconds_total' | jq '.data.result | length'
curl -s 'http://192.168.0.30:9009/prometheus/api/v1/query?query=container_cpu_usage_seconds_total' | jq '.data.result | length'
# confirm "cluster": "homelab" is present in the returned label sets
```
**Logs reaching Loki:**
```bash
curl -s http://192.168.0.30:3100/ready
# then in Grafana / logcli, query {cluster="homelab"}
```
---
## Troubleshooting
| Symptom | Likely cause |
|---|---|
| `forbidden` errors in `alloy-metrics` logs | `alloy-metrics-reader` ClusterRole missing a resource/verb — check `rbac_metrics-collector.yaml` against what the component is trying to read |
| No targets shown in the Alloy UI | Prometheus Operator CRDs not installed, or `kube-prometheus` Flux Kustomization not Ready yet (metrics depends on it) |
| `prometheus.remote_write` shows failed sends / 4xx-5xx | Mimir down, or wrong endpoint — must be `http://192.168.0.30:9009/api/v1/push` exactly |
| Loki push failing | Loki down at `192.168.0.30:3100`, check with `curl http://192.168.0.30:3100/ready` |
| Same series appearing to be scraped twice | Expected — kube-prometheus-stack's bundled Prometheus and `alloy-metrics` both scrape the same ServiceMonitors by design (additive, not deduplicated) |
| Manual `kubectl edit` on the Alloy ConfigMap reverts itself | `driftDetection: enabled` — edit the HelmRelease in git instead |