The Business Cost of Skipping Production Monitoring

Production monitoring is often treated as a technical operating expense: another platform to configure, another set of dashboards to maintain and another stream of telemetry to pay for.

That makes it easy to underestimate what the organisation is actually buying.

The main business value of monitoring is time and visibility.

It shortens the gap between something going wrong and somebody knowing about it. It helps teams understand how much of the application is affected, gives them evidence to begin investigating and reduces the chance that customers become the first people to discover a serious problem.

Without that visibility, even relatively ordinary technical issues can remain active for longer than they should.

That is where the real cost of skipping production monitoring comes from.

‍

The cost starts with a longer detection window

Imagine checkout starts failing at 10:00.

With effective monitoring, the team might detect the change within minutes.

Without it, the first signal may be a support ticket at 10:40. By then, more customers have encountered the problem and the business has lost forty minutes simply discovering that an incident exists.

IBM describes mean time to detect (MTTD) as the average time required to identify that an application or part of an IT system has failed or moved outside acceptable performance. Monitoring and automated notifications are specifically intended to reduce that interval. IBM: What Is Application Monitoring?

That difference matters because detection time sits at the beginning of almost every incident.

The application cannot be investigated, mitigated or repaired until somebody knows there is something to investigate.

AWS makes the same point in its reliability guidance: appropriate monitoring reduces recovery time partly by reducing time to detection, and it explicitly identifies missing alarms or thresholds that provide too little time to react as operational anti-patterns. AWS Well-Architected: Monitor All Components to Detect Failures

So the first cost of weak monitoring is not necessarily the failure itself.

It is the extra time the failure remains invisible.

‍

Longer incidents create more opportunity for revenue loss

For revenue-generating applications, the relationship is straightforward.

If customers cannot complete a transaction, renew a subscription, submit an order or use the product they are paying for, every additional minute of unresolved degradation creates more commercial exposure.

The exact cost depends entirely on the application.

A payment failure during peak trading hours may have a direct revenue impact. A slower internal reporting tool may create mostly productivity costs. A customer-facing SaaS outage could affect renewals or contract commitments rather than immediate transactions.

This is why universal “cost per minute of downtime” figures are rarely useful.

Uptime Institute's 2024 outage analysis illustrates how wide the consequences can become: among respondents whose organisations experienced significant, serious or severe outages, 54% said their most recent outage cost more than $100,000, while 16% reported costs above $1 million. Uptime also cautions that outage-cost data varies considerably and should be interpreted carefully. Uptime Institute: Annual Outage Analysis 2024

The useful lesson is not that every unmonitored app risks a million-dollar incident.

It is that duration matters once business activity depends on the application.

Monitoring gives the organisation a chance to reduce that duration.

[Internal Link: “How to Detect App Problems Before Your Users Report Them”]

‍

Customers absorb the cost when the business does not detect the problem first

Weak monitoring effectively transfers quality assurance into production.

Customers discover that a workflow is broken, contact support and explain the problem back to the organisation.

That has several costs at once.

Support teams spend time collecting information engineering could have obtained automatically. Customers have to interrupt whatever they were trying to accomplish, and the organisation often starts the incident investigation with incomplete descriptions rather than technical evidence.

AWS explicitly identifies user-reported issues as a warning sign that proactive telemetry and business-KPI monitoring are inadequate. AWS Well-Architected: Implement Application Telemetry

For Levelworks, this is one reason Observable is an App Health Aspect rather than merely an infrastructure capability.

A healthy application should communicate important deterioration to the people responsible for it before customers have to become the detection mechanism.

‍

Engineering spends more time reconstructing what happened

Skipping monitoring can also make incidents more expensive after they have been detected.

Suppose support reports that users experienced intermittent failures sometime during the previous three hours.

Without useful production evidence, engineers may need to reproduce the issue manually, compare deployment times, inspect infrastructure retrospectively or add instrumentation and wait for the failure to happen again.

That investigation is engineering time that cannot be spent on planned work.

Good telemetry does not eliminate investigation, but it changes the starting point.

Instead of:

“Something apparently went wrong this afternoon.”

the team may know:

“Checkout errors began at 14:07, affect one payment route and increased immediately after version 6.2 was deployed.”

Google Cloud's reliability guidance recommends metrics, logs and traces partly because they can expose potential failures before outages and help teams diagnose incidents when they do occur. Google Cloud: Detect Potential Failures by Using Observability

The business cost here is less visible than lost sales, but no less real: unplanned investigation consumes expensive engineering capacity.

‍

Incidents create unplanned work that pushes planned work aside

A production problem rarely occupies only the person who fixes the code.

Depending on severity, it can involve engineering, product, customer support, operations, account managers and leadership. Roadmap work pauses while the incident becomes the highest priority.

DORA's current software-delivery model explicitly tracks deployment rework rate—the proportion of deployments that were unplanned and required because of user-facing bugs—as well as failed-deployment recovery time and change fail rate. DORA: Software Delivery Performance Metrics

The connection to monitoring is important.

Monitoring cannot stop every incident, but earlier detection can limit how long the problem develops before remediation begins. Better production evidence can also shorten diagnosis, which reduces the amount of unplanned work required to understand and contain the issue.

The opportunity cost is easy to miss because nobody sends an invoice for “features not built while the team was firefighting.”

It still affects the business.

‍

Poor monitoring can hide performance problems that quietly lose customers

Not every costly production problem is an outage.

An application may become gradually slower while remaining completely available.

If the business does not monitor real performance or important business outcomes, that deterioration can continue for months.

AWS recommends linking workload KPIs to business goals rather than monitoring only system-level measurements. Its guidance specifically notes that worsening page-load performance can directly damage user experience and contribute to customer loss. AWS Well-Architected: Establish KPIs to Measure Workload Health and Performance

The business cost is therefore broader than downtime.

Monitoring can reveal:

  • gradually rising latency;
  • increasing transaction failures;
  • abandoned critical workflows;
  • growing queue delays;
  • worsening application errors;
  • resource constraints approaching capacity.

These signals give the team a chance to intervene before degradation becomes severe enough to trigger obvious complaints.

[Internal Link: “Application Monitoring: What Should You Monitor in a Production App?”]

‍

SLA and customer-trust costs are harder to quantify, but still matter

For products with contractual availability or service commitments, extended incidents may create service credits, escalations or commercial conversations with important customers.

Even without formal SLAs, repeated surprises change how customers perceive the product.

One isolated incident that is detected quickly, communicated clearly and resolved promptly is very different from customers repeatedly informing a vendor that its own application is broken.

Monitoring therefore contributes indirectly to trust.

It gives teams an earlier opportunity to acknowledge a problem, investigate it and communicate from evidence rather than uncertainty.

That does not mean monitoring prevents reputational damage automatically.

It means the organisation is less likely to discover the state of its own service through frustrated customers.

‍

The cost of monitoring should be compared with the cost of blindness

Production monitoring is not free.

Telemetry platforms cost money. Logs and traces require storage. Teams spend time defining alerts, tuning thresholds and reviewing signals.

It is reasonable to manage those costs.

The wrong comparison, however, is:

monitoring cost vs zero cost.

The real comparison is:

monitoring cost vs the cost of operating without enough visibility.

A simple way to think about that exposure is:

Incident cost ≈ duration × business impact per unit of time + response and recovery cost + downstream consequences

Monitoring influences several parts of that equation by reducing detection time, providing diagnostic evidence and making gradual deterioration visible before a major incident occurs.

The exact investment should match the application.

A low-risk internal utility does not need the same monitoring depth as a payment platform. The important thing is that the monitoring capability reflects how much the business depends on the software.

‍

Production monitoring is business-risk control

Monitoring is often discussed as an engineering capability because engineers build and operate it.

Its consequences extend well beyond engineering.

When production is poorly monitored, customers discover problems first, incidents stay active longer, support carries more load and engineering spends more time reconstructing events after they have happened. Revenue-impacting failures can continue unnoticed, while gradual degradation remains invisible until it becomes severe.

That is why Levelworks treats Observable as part of application health.

The value is not in collecting more graphs.

It is in reducing how long the business operates without knowing that something important has gone wrong.

[Internal Link: “How to Conduct an Application Health Check”]

‍

Continue reading