
Maintenance problems rarely arrive as one obvious event.
More often, several small things start happening at roughly the same time. A screen that used to open instantly now takes a few seconds. Support has learned a workaround for an error customers occasionally encounter. Engineers postpone a framework upgrade because nobody is sure what it will break, while the last two production releases required unexpected fixes.
Individually, none of these may justify sounding an alarm. Together, they describe an application whose condition is beginning to change.
That is what makes maintenance difficult to time. Waiting for a serious outage means waiting until the need for maintenance has already become urgent, while reacting to every minor imperfection creates unnecessary work.
The useful signals sit between those extremes.
An app needs maintenance when there is credible evidence that its performance, reliability, security, supportability or ability to change is deteriorating—even if users can still access it normally.
Here are 10 warning signs worth paying attention to.
Warning sign
What it may be telling you
1. The app is getting slower
Performance or capacity is deteriorating
2. The same problems keep returning
Symptoms are being fixed without underlying causes
3. Users find problems before your team does
Monitoring and observability are weak
4. Workarounds are becoming normal
Defects are being absorbed rather than resolved
5. Dependencies and frameworks are falling behind
Security and supportability risk is accumulating
6. Releases feel increasingly risky
Testing or deployment health is weakening
7. Integrations need more frequent rescue work
External dependencies are drifting
8. Growth is pushing the system toward its limits
Scalability assumptions may no longer hold
9. Small changes take surprisingly long
Maintainability is deteriorating
10. Too much knowledge lives with too few people
Support and recovery depend on individuals
Performance deterioration is one of the easiest maintenance signals for users to notice and one of the easiest for teams to gradually accept.
A page that once loaded in two seconds takes three. A report that completed almost immediately now needs noticeably longer as the underlying dataset grows. Search becomes sluggish on large accounts even though smaller accounts remain unaffected.
The dangerous phrase is:
“It isn't that slow yet.”
Maintenance should not wait until performance becomes unusable.
Look at trends rather than isolated measurements. If response times are steadily increasing across releases, database queries are taking longer as records accumulate, or memory consumption has been creeping upward, the direction itself deserves investigation.
Google's Site Reliability Engineering guidance treats latency and saturation as core signals of service health and notes that latency can rise before a constrained resource reaches full utilisation. Google SRE: Monitoring Distributed Systems
For native applications, Apple similarly recommends analysing production metrics such as launch time, memory use and interface responsiveness across app versions so regressions can be identified rather than judged from a single snapshot. Apple: Analyzing the Performance of Your Shipping App
For Levelworks, this is primarily a warning around Performant, but sustained deterioration can also indicate that Scalable is weakening underneath it.
The key signal is not simply “the app is slow.”
It is “the app is becoming slower and we do not fully understand why.”
A recurring defect tells you something different from a new defect.
Perhaps a scheduled job fails every few weeks and somebody restarts it. A particular workflow repeatedly produces the same error after releases, while a memory problem occasionally returns even though the service has already been patched several times.
If a problem keeps coming back, maintenance work may be correcting the symptom without addressing the condition that produces it.
That does not always mean a major architectural flaw exists. Sometimes the permanent fix is modest: better validation, more appropriate retry logic, improved error handling or a regression test that stops the same defect being reintroduced.
The warning appears when the organisation begins regarding recurrence as normal.
NIST's Secure Software Development Framework makes the same principle explicit in the security context: addressing vulnerabilities should include examining root causes so similar problems are less likely to recur. NIST Secure Software Development Framework
The logic applies more broadly to application maintenance.
Repeated problems can indicate weakness across Error Free, Tested or Explainable & Debuggable, depending on why they keep returning.
A closed ticket is not necessarily a resolved maintenance problem.
Customers will occasionally discover unusual edge cases. That is unavoidable.
It becomes a maintenance warning when users are routinely the first people to discover significant production problems.
Perhaps checkout has been failing for an hour before the first complaint arrives. An important background process stopped yesterday, but the team only notices because somebody asks why their data has not updated. Response times have deteriorated for weeks and customer support has a clearer picture of the problem than engineering.
These are not only reliability issues. They suggest the application is not sufficiently Observable.
Good monitoring should reveal important changes in application behaviour early enough for the team to investigate them. Google SRE recommends paying particular attention to latency, traffic, errors and saturation precisely because these signals provide different views of what users are experiencing and what may be approaching failure.
This does not mean teams need an alert for everything.
In fact, an application producing thousands of notifications that everyone ignores may have nearly as serious an observability problem as one producing none.
The warning sign is simpler:
Something important can be wrong for a meaningful period of time without the people responsible for the application knowing about it.
[Internal Link: “How to Detect App Problems Before Your Users Report Them”]
Workarounds are useful during incidents.
They become concerning when they quietly turn into permanent business processes.
Perhaps support tells users to clear something and try again whenever a particular action fails. Operations manually reruns a scheduled process every few days, or employees export data to a spreadsheet because an application report occasionally produces unreliable results.
The underlying application may continue functioning, especially when experienced staff know exactly how to compensate for its weaknesses.
That can make the maintenance need surprisingly difficult to see.
A useful question is:
If everyone stopped applying the workarounds tomorrow, what would break?
The answer reveals how much hidden maintenance debt the organisation is carrying through human effort.
This frequently affects the Supported and Error Free dimensions of App Health. A genuinely supported application does not merely have people available to rescue users; recurring support effort should also provide information about where the application itself needs attention.
The strongest warning is when a workaround is so familiar that nobody remembers it was meant to be temporary.
Software can remain perfectly functional while the technology underneath it becomes increasingly difficult to support.
A framework may still run despite being several releases behind. A library continues doing its job even after its maintainers stop supporting the version you use, while a runtime can remain stable long after the surrounding ecosystem has moved on.
The maintenance signal appears when keeping the application current is repeatedly postponed without a deliberate plan.
This matters for more than access to new functionality.
OWASP's 2025 guidance on software supply-chain failures explicitly identifies unsupported, outdated and vulnerable components as risk conditions and recommends continuously tracking dependency versions and handling updates in a timely, risk-based way. OWASP Top 10:2025 — Software Supply Chain Failures
NIST similarly frames patching and updating as preventive maintenance because it helps reduce security compromises and operational disruption. NIST Guide to Enterprise Patch Management Planning
The important warning is not that you are one version behind.
It is that upgrading has started to feel increasingly difficult, nobody has a clear view of what remains supported, or necessary patches are being delayed because other parts of the stack must first be modernised.
That is where Current & Updated starts affecting Secure, Compatible and eventually Seamlessly Updatable.
[Internal Link: “Why Dependency Updates Are Essential to App Maintenance”]
A healthy application should be changeable.
If releasing a small update requires several people to watch production nervously, maintenance may be overdue even when the application itself is running well.
Warning signs include deployments that depend on undocumented manual steps, rollbacks that are theoretically possible but rarely tested, unreliable automated tests or a pattern of emergency fixes following otherwise routine releases.
DORA's current software-delivery metrics include change fail rate, deployment rework rate and failed deployment recovery time, all of which help expose instability around production changes. DORA Software Delivery Performance Metrics
The issue here is not deployment frequency. Some applications reasonably release many times a day, while others release much less often.
What matters is confidence.
Can engineers make a necessary security update or bug fix without treating production as fragile?
Within Levelworks' App Health framework, persistent release anxiety often points toward some combination of Tested, Version Controlled and Seamlessly Updatable.
An application that remains stable mainly because nobody wants to change it is giving you a maintenance warning.
An integration that was once almost invisible can gradually become one of the most troublesome parts of an application.
Authentication tokens begin expiring unexpectedly. API requests hit limits that were irrelevant when the product was smaller, while a provider announces that the version you use will be retired or response times become inconsistent enough to affect your own users.
These problems often arrive gradually because the systems on either side continue evolving independently.
Watch for integrations that require increasing numbers of retries, manual interventions or exceptions in your code. Also pay attention to providers announcing authentication changes, version migrations or deprecation dates.
The maintenance issue is not simply whether the integration works today.
It is whether your application remains able to work with that external system tomorrow.
This is primarily an Integrable warning sign, but poorly contained integration failures can quickly affect Performant and Resilient too. If one optional external service becoming slow can hold up an entire page, the integration has become a broader health problem.
Growth is normally good news for a product.
It can also expose assumptions that were perfectly reasonable when the application was smaller.
A database designed for one volume of data now contains many times more records. A background process takes longer each month because its workload continues growing, while queue depths during peak periods are consistently higher than they used to be.
Nothing has necessarily failed.
That is what makes capacity one of the more valuable early maintenance signals.
Google SRE describes saturation as how close a service is to the limit of its most constrained resource and recommends paying attention not only to current utilisation but also to impending saturation.
Maintenance is needed when normal business growth starts eroding the application's operating headroom.
The response does not automatically need to be “add more infrastructure.” Sometimes the right work is query optimisation, data archiving, caching, queue redesign or reducing unnecessary processing.
The warning sign is that demand is changing while the technical assumptions remain fixed.
That is the point where Scalable begins exerting pressure on Performant and Resilient.
Perhaps the clearest maintenance warning inside a development team is when apparently simple requests keep turning into unexpectedly large pieces of work.
A small UI adjustment requires changes across several unrelated modules. Engineers avoid touching one part of the application because nobody is certain what depends on it, while adding a new integration first requires upgrading a framework that was postponed several times.
The product team experiences this as slow delivery.
Engineering experiences it as increasing complexity and risk.
This is where technical debt becomes visible as a maintenance problem rather than an abstract code-quality concern.
The important signal is not that older code exists. Old code can be stable, understandable and inexpensive to maintain.
The warning is that the cost of ordinary change is rising.
If every estimate contains an increasing amount of investigation, regression checking and work required simply to make the code safe to modify, the application's maintainability is deteriorating.
This usually cuts across several App Health Aspects rather than belonging neatly to one: Tested, Current & Updated, Explainable & Debuggable, Well Documented and Seamlessly Updatable may all contribute.
[Internal Link: “How Technical Debt Affects the Health of Your Application”]
Ask what would happen if the two people who know the application best were unavailable for a month.
Could the remaining team deploy it confidently? Would they know how the major integrations work, where important configuration lives and what to do during a serious incident?
If the answer is uncertain, maintenance is needed even if production is currently stable.
Documentation often becomes outdated quietly because the application changes faster than the knowledge around it is recorded. At first this causes little trouble because the original team remembers what happened.
The consequences appear during handovers, incidents or unfamiliar maintenance work.
Recovery processes deserve particular attention here. AWS recommends periodically restoring backups and testing recovery because the existence of a backup does not prove that the data can actually be restored within the required conditions. AWS: Verify Backup Integrity Through Recovery
A similar principle applies to operational knowledge: something should not be considered well understood merely because one person knows how to do it.
This is where Well Documented, Supported, Resilient and Explainable & Debuggable can weaken together.
Not every slow endpoint or outdated package means an application is in poor health.
Maintenance decisions need context.
The more useful signal is often the relationship between several warning signs.
If releases are difficult and dependencies are badly outdated, necessary updates may become increasingly hard to make.
If performance is slowing and infrastructure is approaching capacity, the issue may be bigger than one inefficient query.
If users find problems first and only one engineer can diagnose them, an ordinary incident can take far longer to resolve than it should.
If recurring bugs are being handled through support workarounds while automated tests remain unreliable, the application may be compensating for weaknesses rather than correcting them.
These clusters are where Levelworks' App Health perspective becomes particularly useful. The problem is rarely that one dimension has failed completely. More often, several aspects—such as Observable, Performant, Tested or Current & Updated—begin deteriorating together.
You do not need to wait until all ten appear.
A formal application health check becomes useful when the team knows something is deteriorating but does not yet understand the full extent of the problem.
That might be because releases have become noticeably harder, incidents are increasing, a major technology upgrade has been postponed repeatedly or several different teams are reporting seemingly unrelated weaknesses.
A health check can establish which warning signs are isolated and which share the same underlying cause.
[Internal Link: “How to Conduct an Application Health Check”]
The objective should not be to create a giant list of everything that could possibly be improved.
It is to identify which weaknesses are real, what evidence supports them and which ones are likely to become materially more expensive if they are ignored.
The most expensive warning sign is usually the one you only recognise afterwards.
A production outage makes the need for maintenance obvious, but by then the organisation has already lost the opportunity to handle the problem under normal conditions.
The useful signs come earlier: slower response times, recurring defects, support workarounds, ageing dependencies, nervous releases, integration drift, diminishing capacity or an increasing dependence on individual knowledge.
None automatically means the application is failing.
They mean its condition is changing.
That is the point at which maintenance provides the most value—while the application still works, the problem can still be investigated deliberately, and the team still has choices about how to address it.