The Disciplines Plate VI Monitoring & Observability
Monitoring &
Observability.
Knowing before they do.
Dashboards, alerts and telemetry so the first person to know something is wrong is you - not a customer, not a board, not a headline on a Monday morning. Prometheus, Grafana, Zabbix, Checkmk.
Return to the workshopNot every business can - or should - send monitoring data to the cloud. Some environments require air-gapped solutions for compliance or security. Others benefit from cloud-based convenience. Most need something in between.
I design monitoring that matches your requirements, not a one-size-fits-all platform. A handful of servers or hundreds of endpoints - I find the right balance of visibility, security, and cost. Where possible I lean on open-source tools (Prometheus, Grafana, Uptime Kuma, Checkmk) that provide enterprise-grade capability without enterprise licence fees. You pay for expertise, not software.
§ What I monitor
-
I.
Infrastructure health
CPU, memory, disk, network - the fundamentals that keep systems running. Threshold-based alerts before you hit capacity. Historical trending to plan upgrades before they become emergencies.
-
II.
Uptime & availability
Synthetic monitoring from multiple locations. SSL certificate expiry warnings, DNS monitoring, response-time tracking with SLA reporting. Know when the site goes down before customers do.
-
III.
Application & database
Query performance, connection pools, replication lag, slow queries. Application-level metrics that matter for your specific stack - custom instrumentation where off-the-shelf falls short.
-
IV.
Security events
Failed logins, firewall blocks, unusual network activity. Integration with SIEM tools. Event correlation across systems to spot patterns that indicate threats.
-
V.
Log aggregation
Centralised logging from every system. Search across months of logs in seconds. Structured logging that makes troubleshooting faster. Retention policies balanced between compliance and storage cost.
-
VI.
Smart alerting
Alerts that matter, not noise. Escalation policies, on-call schedules, integration with Slack, Teams, email, SMS. Thresholds tuned to reduce false positives - alert fatigue is real.
§ The process
-
Step 1
Discovery
Audit current infrastructure and understand what matters most. Critical systems? Tolerance for downtime? Compliance requirements? Cloud, on-prem, air-gapped, or hybrid?
-
Step 2
Design
Propose an architecture that fits your requirements and budget. Clear recommendations on tools, deployment model, and what metrics to actually track.
-
Step 3
Deploy
Install and configure agents, set up dashboards, configure alerting, integrate with existing tools. Full documentation of what's deployed and why.
-
Step 4
Tune
The first month is refinement. Adjust thresholds, reduce noise, add metrics I missed, drop ones that aren't useful. Monitoring improves with time.
-
Step 5
Support
Ongoing support to keep monitoring healthy. Adding new systems, adjusting for change, responding to incidents with you. Or a clean handover if you want to run it yourself.
Questions, answered
We already get alerts and ignore them. How is this different?
Alert fatigue is a design failure, not a discipline failure. The work is mostly deciding what genuinely warrants waking someone and deleting the rest, which usually means fewer alerts than you have now.
Do we need to buy a monitoring platform?
Generally not. The open-source tools named above are capable enough that for a business of this size they do everything the licensed products do, and the budget is better spent on someone setting them up properly.
What does a commission cost?
Work is quoted per commission, after a conversation. There is no day rate to quote at you and no package to fit into, because the shape of the job decides the price and nobody knows the shape yet. The first half hour is free, and it usually settles the question of whether this is a small piece of work or a large one before any money is discussed.
Who owns the code when it is finished?
You do. Outright, including the source. It is handed over with the runbooks needed to keep it running, and there is no licence to renew and no seat to pay for. Commissions, not subscriptions, is meant literally.
What happens if you are unavailable?
Everything is built to be picked up by somebody else: ordinary technologies, readable code, written runbooks and no proprietary lock-in. It is a fair question to ask a one-person workshop, and the answer has to be in the work rather than in a promise.
Do you work outside Suffolk and Cambridgeshire?
Yes. Local work gets the option of somebody in the room, which is worth more than it sounds for the first conversation and the handover. Everything after that is done remotely for most clients anyway.