A Green Dashboard Can Hide a Dead Data Pipeline
A panel can render old data successfully after the collector behind it has already stopped.
TOPIC
Lessons, explainers, experiments, and implementation notes.
A panel can render old data successfully after the collector behind it has already stopped.
Presence, motion and occupancy are conclusions; retain enough underlying evidence to debug those conclusions.
The useful baseline is the room's normal operating condition, not an imaginary perfectly quiet environment.
A topology UI can pass synthetic tests and still misread production traffic at the first adapter.
Label names and cardinality are query contracts for dashboards, recording rules and alerts.
A production-engineering deep dive into monitoring tax: why exporters need resource budgets, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into temporary i/o is a query-behavior signal, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into disk full and inodes full are two different outages, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into page cache is not a memory leak, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into one postgresql deadlock is worth recording, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into the voip dashboard must follow a call across layers, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into desired state, observed state and user-visible state are three different things, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into waiting sessions tell me more than cpu during database contention, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into disk throughput is not disk latency, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into buffer hits, disk reads and the shape of postgresql cache pressure, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into swap usage alone is a bad memory alert, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into thermals on a 2014 mac mini are a production signal, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into one machine, many failure domains: mapping the 2014 mac mini, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into the server cannot report its own death, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into monitoring tcp retransmission on a server that also runs voip, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into what production-grade means on decade-old hardware, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into how cadvisor became one of my largest workloads, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into grouping and inhibition stop one failure becoming twenty pages, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into why 99.9% uptime does not tell me when to wake someone up, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into linux psi changed how i think about saturation, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into fast burn and slow burn are different incidents, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into clock synchronization is an availability dependency, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into container running, healthy and useful are three different states, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into restart counts without context create noise, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into oom events are better evidence than “high memory” alone, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into memory limit utilization vs working set: which one should page me?, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into container block i/o and i/o pressure tell different stories, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into a docker storage cache needs its own freshness metric, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into prometheus is my metrics control plane, not just a scraper, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into why every target should not be scraped every 15 seconds, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into 27.5k active series on a small server: what that number actually means, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into cardinality is usually a schema problem, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into metric relabeling is resource control, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into recording rules are a cpu trade: compute once or query repeatedly, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into 30-day retention and a 15 gb cap were capacity decisions, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into wal, head block and why prometheus memory does not equal stored data, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into rule evaluation failures mean monitoring logic is broken, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into how much prometheus is too much for an 8 gb machine?, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into dashboards should answer questions, not display metrics, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into building a noc dashboard for a phone screen, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into overview vs drill-down: one dashboard cannot do both well, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into why green dashboards can still lie, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into alert panels and investigation panels serve different humans, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into information density without dashboard wall art, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into the observability dashboard that watches the observability stack, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into metrics tell me that something happened; logs tell me what happened, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into why loki labels need to stay boring, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into ip addresses belong in log content more often than labels, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into turning ssh failures into a production security signal, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into journald cursor persistence prevents duplicate or missing logs, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into the stale docker container problem in alloy, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into how logging itself can become an i/o workload, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into seven days of logs was an engineering choice, not a default, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into alert fatigue is an architecture bug, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into threshold alerts vs symptom alerts, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into what prometheus `for:` really buys you—and what it cannot, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into warning and critical are different response contracts, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into monitoring is a failure-modeling problem, not a dashboard problem, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into the observability tax on a 7.1 gib linux server, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into why every monitoring component had to justify its ram, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into why my monitoring system needed monitoring too, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into why “up” is one of the weakest signals in production, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into redis eviction means policy is already affecting data, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into redis fragmentation can look like a memory leak, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into mysql slow queries belong in infrastructure monitoring, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into four database engines on one small server: observability without exporter sprawl, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into why a 302 redirect can mean the service is healthy, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into internal probe and public probe answer different questions, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into dns is a production dependency, not plumbing, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into tls certificate expiration is an availability incident waiting to happen, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into building a dead-man monitor outside the machine it watches, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into github actions as five-minute external infrastructure, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into i deliberately broke the probe to test the monitoring system, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into dead-man monitoring should not need a credential from the dead system, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into monitoring sip is not monitoring rtp, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into a healthy freeswitch process can still mean broken calls, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into heartbeat freshness is more useful than “container running” for voice workers, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into no healthy worker is a user-facing signal, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into opensips 5xx spikes are more useful than total sip traffic, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into rtpengine timeout reasons tell different media stories, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into a sealed secret server can be secure and still be down, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into how i exposed openbao metrics without exposing openbao, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into config drift is a monitoring signal, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into monitoring secrets without monitoring secret values, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into ssh, sudo, authelia and docker errors as one security telemetry plane, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into backup success is not recoverability, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into a fresh backup can still be corrupt, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into restore verification is the most important backup metric i have, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into monitoring the system that is supposed to save the system, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into rpo, rto and evidence: when a home server becomes a production system, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into how i took cadvisor from ~428 mib to ~20–28 mib, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into multi-window burn-rate alerts on a tiny self-hosted stack, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into resolved notifications are part of the incident lifecycle, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into why i monitor memavailable instead of “free ram”, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into docker filesystem scanning was more expensive than the metric was worth, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into load average without cpu count is almost meaningless, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into metrics, logs, probes and state checks answer different questions, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into every alert should suggest the next investigation, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into a successful tcp connection does not mean the database is healthy, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into postgresql connection utilization needs a denominator, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A target that disappeared can leave its last sample available long enough for threshold expressions to evaluate against stale data.
The host looked busy enough that a simple percentage could easily become the whole diagnosis, but Linux memory reclaim makes that misleading.
A few megabytes of swap on an old Linux host did not automatically mean an incident, especially after long uptime.
A load average of four means something very different on a two-core system than on an eight-core system.
The host can report high CPU usage even when the useful question is whether time is going to user work, system work, steal, or I/O wait.
Some periods looked acceptable in average CPU and RAM graphs while interactive services still felt slow.
Logs, TLS checks, backup ages and distributed event ordering all become harder to trust when the host clock drifts.
A small Mac mini running many containers can hit thermal constraints before ordinary CPU graphs explain why performance changed.
CPU percentage alone did not reveal when the kernel was spending more work scheduling tasks or servicing device activity.
Container monitoring does not cover host services such as Docker, networking, tunnels, backup timers or other systemd-managed dependencies.
A disk can have plenty of free capacity while the underlying SSD is reporting temperature or device-health problems.
The Docker daemon can report a container as running even when the application inside it has stopped serving useful traffic.
A service can look healthy now and still have restarted repeatedly overnight, erasing the evidence from a simple current-state view.
An application disappearing under load can be either a container memory-limit event or a host-wide memory emergency.
A container using one full CPU may be expected on an unrestricted worker and catastrophic for a service capped at a fraction of a core.
Container memory graphs become noisy when cache and reclaimable pages are treated exactly like unreclaimable application working memory.
RX and TX graphs looked busy enough, but throughput by itself could not tell whether traffic was healthy.
Host disk latency can rise because one container is performing heavy reads or writes while every other service only sees the consequence.
A workload can perform modest disk throughput and still suffer because requests are waiting behind slow storage or competing I/O.
Privileged containers, Docker socket mounts, root users and published ports are configuration facts that can silently drift after deployment.
A probe may keep returning success while taking progressively longer, showing a dependency slowdown before Docker marks the container unhealthy.
A single disk threshold gives operators no distinction between early cleanup work and a filesystem that is close to stopping writes.
Free-space graphs can remain green while a workload creates huge numbers of tiny files and consumes the available inode table.
The host can show modest megabytes per second while applications still wait because each storage request takes too long.
A database doing many tiny synchronous operations and a backup streaming large files can show similar disk utilization with very different access patterns.
A device at high utilization can be handling work efficiently, while a lower-utilization device can still be returning slow requests.
Docker storage size is useful, but asking the daemon for deep filesystem inventory on every fast scrape consumed too much monitoring overhead.
A large Docker volume is not actionable if the dashboard cannot tell which service owns it or whether that growth is expected.
Docker images can quietly accumulate through repeated deployments even when application volumes and databases remain stable.
The production monitoring stack measured roughly 428 MiB of cAdvisor memory with filesystem disk collection enabled on a small host.
Caching Docker storage inventory reduced overhead, but a silent refresh failure could otherwise leave old values looking current indefinitely.
A service can be healthy on the LAN and unreachable through Cloudflare, TLS, DNS or authentication at the public edge.
TLS can work perfectly today and still have a known future outage date embedded in the certificate.
A slow public request is not always slow application code; DNS resolution can consume a meaningful part of the user-visible path.
A single total probe duration hides whether slowness came from name resolution, TCP connection, TLS negotiation or server response.
High interface traffic can be completely healthy, while a small but sustained drop rate can damage voice, APIs and tunnel reliability.
A few retransmissions during heavy traffic may be normal, while the same count during low traffic can represent a serious quality problem.
The host collector originally relied on `iw`, but driver and interface behavior did not always expose the active connection signal consistently.
An external endpoint can flap because the tunnel reconnects even when the local service and LAN probe remain healthy.
A rising established-connection count can represent normal load, a leak, slow clients or a downstream dependency holding sockets open.
A health probe that automatically follows redirects may end on an authentication page and report successful HTTP even though the original service route is wrong.
A listening PostgreSQL or Redis port proves that something accepted a TCP connection, not that queries, authentication or storage are working correctly.
A count of sixty active PostgreSQL connections is meaningless until it is compared with the configured maximum for that server.
A PostgreSQL deadlock can resolve automatically by aborting one transaction, leaving the service apparently healthy after the incident.
Database traffic volume looked normal even when applications were rolling back more transactions than usual.
A database normally grows, so alerting on size alone would create noise while ignoring the important question of growth rate and disk headroom.
Redis can remain responsive while silently evicting keys because the configured memory limit has been reached.
A Redis instance serving disposable cache and one serving authentication sessions can show the same memory growth with very different operational risk.
A MongoDB service can accumulate client connections without a matching rise in useful query work, which can point to pooling or application lifecycle problems.
High database container CPU or memory can be an important symptom, but it cannot identify whether the engine is busy with useful work, blocked transactions or internal maintenance.
Copying every application database password into the observability stack would have expanded the secret blast radius just to collect metrics.
Restarting a log collector can replay old journal entries or skip new ones if it does not persist its position correctly.
A logging stack can look healthy at the process level while silently discarding entries because of relabeling, backpressure or write failures.
A few failed SSH logins are normal on an administered server, while a rapid burst can indicate brute-force activity or a broken automation credential.
Repeated sudo failures may be operator error, expired credentials or suspicious privilege-escalation attempts, and they often happen outside application logs.
A process disappearing can look like an application crash unless the kernel journal is checked for OOM-killer activity.
Individual container logs can all look normal while the Docker daemon is failing image, network, storage or runtime operations underneath them.
A two-factor login system naturally records failed credentials, expired sessions and rejected access, so any single failure is not automatically an attack.
A container can remain healthy by its HTTP probe while its logs suddenly fill with exceptions, retries or failed dependency calls.
A public request can fail before reaching Authelia, inside the authentication flow, or after authentication while the upstream application is unavailable.
Security logs contain useful source addresses, but promoting every IP to a Prometheus or Loki index label would create unbounded cardinality on an Internet-facing service.
The SIP proxy can be running as a process while its control connection to drachtio is down, leaving signaling logic unable to operate correctly.
A multi-worker voice stack can keep serving calls after one FreeSWITCH node fails, so one global up/down flag hides degraded capacity.
A FreeSWITCH process may remain alive while its worker integration stops reporting useful state to the signaling layer.
The most important load-balancer failure is not that a backend probe failed; it is that an incoming SIP request could not be assigned to any healthy worker.
A single SIP proxy error may be harmless noise, but repeated failures over a short interval can indicate backend, routing or dependency trouble.
OpenSIPS can continue processing calls while the MI metrics collector fails, leaving the service healthy but observability blind.
Registration counts and dialog counts can look stable while transaction failures increase for a subset of calls.
Media sessions can close for rejection, timeout, silent timeout, final timeout or offer timeout, and those reasons point to different failure paths.
A healthy total session count can hide a load-balancing problem if one FreeSWITCH node carries nearly all calls while another remains idle.
A sudden drop in registered SIP endpoints can indicate network reachability, credential, expiry or registrar problems even when call processing components are up.
A recent backup timestamp can look reassuring even when the archive is incomplete, corrupt or impossible to restore.
A backup directory can exist with the expected filenames while one archive is truncated or modified after creation.
A restore drill that passed months ago does not prove that today's schema, credentials and backup format can still be recovered.
A systemd timer can fire on schedule while the backup service itself fails, times out or exits before producing a valid set.
A backup that suddenly becomes much smaller may have completed successfully while silently omitting a database, artifact directory or other expected state.
A red backup alert identifies the outcome but usually does not explain which command, mount or permission caused the failure.
A backup process can be perfectly configured and still fail when the destination filesystem no longer has enough capacity for the next archive.
Old Docker volumes can consume storage after services are removed, and their names alone do not always reveal whether anything still depends on them.
No recent restore verification and a recent restore verification that actively failed are both bad, but they communicate different operational urgency.
Least-privilege monitoring sometimes cannot read protected backup evidence, and treating that access failure as healthy would be dangerous.
A dashboard that only shows firing alerts hides whether conditions are approaching thresholds or whether the expected rule set is even loaded.
CPU, packet loss and endpoint probes can cross thresholds for a few seconds during harmless transitions or deployment activity.
An alert can fire correctly and still never reach the operator if Alertmanager or the external delivery path fails.
The alert sink originally targeted a public ntfy service, which made production notification depend on infrastructure outside the hserver control boundary.
Moving alert delivery to authenticated self-hosted ntfy introduced a publisher token that the alert-sink needs at runtime.
A monitoring system that detects failures but cannot notify anyone is partially failed even if every Prometheus target remains green.
If every alert is critical, operators lose the distinction between conditions that require immediate intervention and those that need scheduled review.
Delivery counters only change when an alert is sent, so a quiet system could leave a dead notification service unnoticed for hours.
A deployment may push a metric over threshold without immediately firing because the configured `for` period has not elapsed.
A broken rule evaluator or missing target can produce a beautifully quiet alert dashboard while the system is blind.
The hserver stack carried roughly twenty-seven thousand active Prometheus series, enough that label growth and exporter changes could materially change memory and storage cost.
CPU can change meaningfully in seconds while Docker image storage or Grafana process memory does not need the same fifteen-second collection cadence.
Grafana, Loki and Alloy expose many internal metrics that are useful for development but unnecessary on a small production Prometheus.
Observability was becoming one of the larger workloads on a small production server, which is dangerous when monitoring competes with the services it protects.
Keeping thirty days of Prometheus data is useful until series growth causes the TSDB to consume more disk than the host can safely spare.
Placeholder or obsolete probe targets can stay in configuration after architecture changes and permanently pollute availability dashboards.
A wall of attractive graphs is slow during an incident if related signals are scattered by exporter rather than by the question an operator is trying to answer.
The phone-sized operations view cannot carry hundreds of panels without turning urgent information into scrolling noise.
A monitoring stack can drift through dashboard edits, local files and runtime tuning until nobody knows whether Git can reproduce what is currently trusted in production.