MNT.2 - Select and Reopen the Maintenance Policy
Type: Method Status: Stable
MNT.2:1 - Problem frame
Use this pattern when a recurring maintenance task, inspection interval or proposed technology needs a reason to be retained or changed. A calendar says “replace every year,” but the planner cannot explain which failure this prevents or why a different task would be worse.
Choose a reusable maintenance policy for the identified functioning and failure situation. Here a policy states which task applies and under what condition it is performed. The first result can be a justified decision to retain the current policy. A one-off response to today’s condition belongs in MNT.6; a fleet-wide support and work decision belongs in MNT.13.
MNT.2:2 - Problem
Task frequency and technical sophistication are poor substitutes for policy reasoning. An age-based replacement may remove a wear-related failure, do little for a largely age-independent fault or introduce a new installation fault. A sensor may observe degradation without leaving enough time to obtain a part and act.
A policy therefore needs a credible connection between failure behaviour, what the task can change, the consequence of waiting and the practical burden of intervention.
MNT.2:3 - Forces
Prevention can avoid an expensive failure but consume useful component life and introduce disturbance. Frequent inspection can shorten detection delay while consuming access and interpretation effort. Keeping stock shortens response time but ties up resources and can leave obsolete parts. Hidden protective functions create a different trade-off from failures that operation immediately notices. Compare these consequences within the maintained use; do not reduce every value to downtime alone.
MNT.2:4 - Solution
Start from the required function and a failure account sufficient to distinguish policy alternatives. Identify the loss of function, credible mechanisms and consequences. When one policy covers several mechanisms, check each mechanism that could justify a different task. Retain uncertainty where it matters rather than inventing a lifetime distribution from sparse events.
Construct the plausible task alternatives. The following distinctions help select a task; they are not a ladder of maturity.
| Task family | What can make it useful | What can defeat it |
|---|---|---|
| Corrective or deliberate run-to-failure | Failure is detectable and its consequence, recovery demand and collateral effects are acceptable for the use. | A hidden protection loss, unacceptable exposure or unavailable recovery capability makes waiting untenable. |
| Age- or usage-based action | Failure behaviour and restoration effectiveness support acting before a relevant age or usage condition. | Calendar age is a weak predictor, or replacement repeatedly introduces faults. |
| Condition-based action | An observable condition changes early enough to support a useful response. | The relevant failure is not detected, the alarm is unreliable, or response lead time exceeds the usable warning. |
| Failure-finding | A task reveals loss of a function that normal operation does not expose, such as a standby protective function. | The test misses the relevant failure or creates unacceptable exposure without adequate controls. |
| Opportunity-based combination | Shared access or downtime makes coordinated tasks worthwhile. | Bundling consumes useful life, creates interference or overcommits the outage. |
For each credible alternative, explain what performing the task changes. Include imperfect repair, task-induced defects and residual failure modes when material. If no acceptable maintenance action addresses a serious failure, return the specific redesign, redundancy or changed-use question to engineering or operation. Adding inspections that cannot reveal or alter the failure is not a substitute.
For condition-based work, reason across the complete response. The usable interval between detectable degradation and unacceptable functioning must accommodate detection delay, interpretation, obtaining support, access and the selected intervention, with uncertainty appropriate to the consequence. This is an applicability question, not a universal formula for a safe interval. A local detection threshold or interval needs the equipment and operating evidence that supports it.
Compare alternatives using the receiving use’s relevant values: loss of service, exposure, labour, material, disturbance, environmental effects and uncertainty. Use comparable operating horizons and state deliberately accepted trade-offs. Existing adequate evidence may already settle the choice. Select further investigation only when an attainable answer could change it enough to justify acquisition, interpretation, delay and displaced work.
State the chosen policy in actionable terms. Name the applicable population or configuration, task, trigger or interval basis, relevant support assumptions, response to an out-of-scope condition and the evidence or changed use that would reopen it. Where an operative requirement fixes a task, preserve its current force. A separate merits appraisal may support a request to the authorized rule holder; it does not authorize unilateral relaxation.
MNT.2:5 - Archetypal Grounding
For PS17, the constructed condition history supports a bearing-related deterioration concern. The team retains a condition-informed policy for this failure family because the interpreted signal can support a planned response in the stated use. The present recommendation still depends on actual support and access; an alarm alone does not establish that replacement can fit the next outage.
Consider instead a cheap, accessible indicator lamp whose failure is obvious and has no protection role. If its loss and replacement demand are acceptable, deliberate run-to-failure can be a sound policy. Scheduling repeated intrusive replacement needs an additional gain to justify its burden.
The same choice is unsuitable for an otherwise unobserved protective trip function. Normal production can continue while that function has failed. A suitable failure-finding task addresses that detection problem; an operator’s observation that “the line still runs” does not. The applicable specialist basis determines the test and interval.
In the 48-pump fleet example, the operating exposures differ, and policy selection is not randomized. Four failures versus six do not by themselves justify replacing the policy. Retain the current policy when its existing failure-and-response basis remains adequate; address the known spare-support deficiency on its own existing evidence. Failure to prove a better rival does not establish that the current policy is adequate.
MNT.2:6 - Bias-Annotation
Breakdowns are conspicuous; unnecessary preventive work and failures introduced by maintenance can disappear into ordinary cost codes. Include those consequences in a comparison. A vendor’s technology categories can also favour its own sensor or software offering over a simpler adequate task.
MNT.2:7 - Conformance Checklist
Does the policy address a stated functioning loss and credible failure behaviour? Can the selected task detect, prevent, mitigate or restore the relevant failure in time? Are consequence, support and task-induced effects considered where they distinguish the alternatives? Is the chosen applicability and trigger usable by the intended practitioner?
A retain decision can close the question. New trials, a complete failure model and numerical optimization are warranted only when their obtainable contribution changes this choice.
MNT.2:8 - Common Anti-Patterns and How to Avoid Them
“Predictive is better than preventive” ranks technology without the failure and response conditions. Compare what each arrangement can actually change. “We have always replaced annually” conceals the interval’s basis; recover that basis and reopen only the unsupported choice.
A protective function that has never been demanded can appear failure-free. Examine how its failed state would be detected before choosing run-to-failure.
MNT.2:9 - Consequences
The chosen task has an explicit maintenance contribution and a useful reopen condition. Some existing work can be retained; some unnecessary work can stop through the applicable decision. Where no task provides an acceptable answer, the Method makes the engineering or operating problem visible instead of disguising it as a maintenance backlog.
MNT.2:10 - Architectural Rationale
A reusable policy differs from a current intervention choice: one establishes when a task generally applies, while the other resolves a present case. Keeping them related but separate permits a valid policy to coexist with an exceptional current response.
The task alternatives remain plural because failure visibility, mechanism, consequence and support differ. A single escalating technology sequence would discard legitimate corrective and failure-finding policies.
MNT.2:11 - SoTA-Echoing
For the question “Which task is worth performing for this failure?”, this pattern adapts the applicable/effective task reasoning of NASA’s historical 2008 RCM guide and the mechanism-plus-decision-model line discussed by Arts and colleagues in 2025. It rejects a technology ladder as the policy rule: the comparison table and complete-response reasoning preserve corrective, hidden-function and support-limited cases that the ladder obscures. More demanding models remain useful when their decision gain warrants their effort. Neither source supplies a universal interval or permission rule. Reopen when failure behaviour, warning time, restoration effectiveness or the operating consequence changes. See the RCM task discussion and maintenance optimization account.
MNT.2:12 - Relations
MNT.3 supplies failure evidence and MNT.4 supplies condition interpretation. MNT.5 and MNT.7 qualify response feasibility; MNT.6 selects the present intervention. MNT.13 uses policy consequences across a fleet and MNT.14 compares Method changes. FPF C.11 supports comparison, and C.11.DUA governs a disputed evidence demand or requirement appraisal.