In brief
Condition monitoring starts with the effect of a failure on the IT load and the required time to respond. A failure mode and effects analysis informs the instrumentation, alarms, tests and maintenance regime for pumps, valves and other critical equipment.
A data center cooling system contains a great deal that does not fail in practice. Pipework, heat exchangers and valve bodies last for decades. What fails are the parts with moving components: compressors, pumps, fans and valve actuators.
The usual response is to fit vibration measurement to the pumps and consider the matter handled. But that is an answer to a question nobody has asked yet. The right order is to establish first what can go wrong in the plant and what follows from it, and only then decide what to monitor and how. Condition monitoring is a conclusion, not a starting point. The wider picture is in our complete guide to data center cooling.
What a buyer should provide
Share the cooling P&ID, equipment list, resilience objective, control narrative and existing alarms. Finkova can discuss the critical-machine monitoring and alarm responsibilities relevant to its scope.
Data centers are too often designed as buildings
In process industry, a plant’s failure scenarios are analysed systematically before it is built. The methods are old and well established, and their use is taken for granted in a power station, a chemical plant or a paper mill.
In a data center the same analysis is often never done. The reason is that a data center is conceived as a building with building services attached, rather than as a process plant whose only job is to manage the heat its equipment produces. Building services design practice answers the question of whether the system meets its design values under normal conditions. It does not answer the question of what happens when something breaks.
The distinction is not academic. A liquid-cooled hall running tens of kilowatts per rack is a heat transfer process whose interruption damages equipment within minutes. It deserves the same treatment as any other critical process.
An example of what the analysis would have found
In August 2026 Proton published an unusually open report on the total loss of cooling at its Frankfurt data center. The case is instructive, because it contains exactly the things a failure analysis looks for.
The root cause was not an equipment failure. It was a maintenance action: the data center operator replaced the air filters on the compressors powering the cooling system, on both redundant units, in the middle of the night and without prior notice. The redundancy existed, but a single routine task removed it entirely.
The hall temperature rose from about 21.8 °C to nearly 52 °C in under half an hour, with some probes reading 60 °C. According to the report, the time available before conditions became critical had previously been three to four hours, but growth in equipment density had cut it to roughly twenty minutes. The time budget had changed, but the preparations built on it had not been updated.
The knock-on effects were ones nobody had anticipated. Network cards reached 105 °C, which triggered their thermal protection and left them disabled until a cold system reset. A deliberately restricted out-of-band management path meant additional staff had to be woken during the night to recover the site. Some servers were lost to heat, and the remaining life of the survivors was left an open question.
None of this is exotic or improbable in hindsight. These are ordinary chains of events that a systematic analysis finds in advance.
Failure mode and effects analysis
Failure mode and effects analysis, FMEA, is a method in which every component, its possible failure modes and their consequences are worked through systematically. When an assessment of criticality is included, the method is called FMECA. It is standardised, and there is nothing new or experimental about it.
The analysis answers five questions that otherwise go unanswered:
What can break. The analysis surfaces failure scenarios nobody had thought possible. Simultaneous maintenance on two redundant units is a good example.
What follows from it. Knock-on effects are usually the part that surprises. A single fault rarely takes a site down, but a chain does.
Whether the system survives the fault. The N-1 principle familiar from transmission grids means the system has to keep working when one component fails. The analysis shows whether this actually holds and how many simultaneous faults are tolerated.
How much time there is to act. In data centers this has changed decisively as density has grown. The analysis establishes the time budget and what has to be available within it: spare parts, tools, competence and an on-call arrangement.
What is worth monitoring. The analysis is the basis of the maintenance strategy, and it drives the choice of condition monitoring and machinery protection directly.
A documented analysis is also evidence that the risks were identified and managed. That carries weight with insurers and in questions of liability if damage occurs regardless.
What can go wrong in a cooling system
The value of the analysis comes from making the list exhaustively rather than covering only the obvious failures. A typical starting list for a liquid-cooled site looks like this:
Pumps
- the pump does not pump
- the pump leaks but keeps pumping, so the system loses fluid slowly
- some pumps run and others do not, so flow distributes wrongly
- whether the system is designed so that all pumps must run, or whether fluid is routed to the wrong place when one stops
Heat rejection
- a dry cooler leaks
- a dry cooler does not leak but its fan does not run
- some fans run, so capacity falls without an alarm being raised
Pipework and valves
- a water leak in the pipework
- a valve stuck open, closed or part-way
- an actuator failure, or loss of control air or power
- no isolation valve where one is needed
Rack and cabinet level
- a fault or leak in an individual cabinet’s cooling
- complete loss of cooling to a cabinet
- a quick disconnect leaking after maintenance
Control
- a control fault at cabinet level
- a control fault at system level
- a sensor fault that reports wrong data without alarming
- a shared point of failure between automation and mechanics, such as the same supply board or the same controller
External and human causes
- a maintenance action that removes redundancy
- a power supply transfer during a load change
- unannounced work on site
The second item on the pump list deserves particular attention. A pump that leaks but keeps running is a failure mode neither conventional safeguard detects: machinery protection does not trip, because the machine is running normally, and vibration measurement shows nothing unusual. It is detected from make-up water consumption, which is the measurement most often left out. Metering points are covered in the article on flow and energy metering in data center cooling loops.
Criticality classification drives the maintenance strategy
Once failure modes and their consequences have been worked through, equipment is classified by criticality according to what a failure would cause for production, safety and cost. Common practice is to carry out a detailed failure mode analysis only for the highest class.
The classification produces a maintenance strategy, and from that follows the level of monitoring. Machinery protection for the most critical machines, continuous condition monitoring for the next level, wireless monitoring below that, and periodic route-based measurement at the lowest level. This is the point at which condition monitoring is decided, and it is the outcome of the analysis rather than a substitute for it.
Machinery protection and condition monitoring
The two are often conflated, although they solve different problems.
Machinery protection stops a machine before it is damaged. It is fast, automatic and tied to fixed limits. Its job is not to tell you anything in advance but to prevent damage when something happens suddenly.
Condition monitoring tells you about a machine’s health before anything happens. It follows the development of vibration and identifies a developing fault, so that maintenance can be scheduled rather than forced.
The criteria for evaluating vibration measurement are set out in ISO 20816-1, which replaced the earlier ISO 10816-1. The central idea is that evaluation rests on both the level of vibration and its change, and in practice the change matters more. Two identical pumps can sit at different vibration levels with nothing wrong with either, because installation, foundation and pipe support all affect the level. A rise in the same machine’s level from its own baseline, on the other hand, always means something.
It follows that the baseline has to be measured at commissioning. It can only be done once, and that moment is at handover. Handover baselines are covered in the article on data center cooling: commissioning and handover.
What vibration tells you
Imbalance appears at rotational frequency and is typical of fans that accumulate dirt.
Misalignment appears at multiples of rotational frequency and is a common consequence in pumps of installation or foundation settlement.
Bearing faults appear at high frequencies and develop gradually. This is the fault mode condition monitoring exists for, because the warning time is long.
Cavitation appears as broadband noise and indicates a problem on the pump suction side, such as insufficient pressure or a blocked filter. Filtration and loop condition are covered in the article on liquid cooling water chemistry, filtration and material compatibility.
Looseness and resonance point to the mounting or the structure rather than to the machine itself.
Alarm limits and responsibilities
Condition monitoring creates value only when someone sees the result and knows what to do with it. Alarm limits are set from the baseline, not from manufacturer defaults, which tend to be either so loose that the alarm arrives late or so tight that alarms are learned to be ignored. The data goes into the same system the site is otherwise monitored from. Responsibility is assigned: who receives the alarm, and who decides when maintenance happens.
One consequence of the analysis is not technical at all. The notification and approval practice for maintenance work is part of managing failure scenarios. In the case described above, the root cause was precisely that redundancy-removing maintenance was carried out at night without notice.
Common mistakes
Failure scenarios are never analysed. The plant is designed for its design values and for normal operation.
Redundancy is assumed to mean fault tolerance. A duplicated machine does not help if both draw power from the same board, take control from the same controller, or get maintained at the same time. Redundancy is covered in the article on TCS and FWS: data center cooling loop architecture and scope splits.
The time budget is not updated as density grows. Preparations built on old assumptions are at the wrong scale.
The wrong machines are monitored. Monitoring goes where it is easy to install rather than where a failure would be most expensive.
The baseline is never measured. Everything measured later then becomes guesswork.
Nobody looks at the results. The most common reason condition monitoring fails to produce savings.
Finkova’s scope and next step
Finkova installs and commissions data center cooling in the Nordics, with instrumentation and lifecycle support. The work scope, OEM requirements and acceptance criteria are agreed for each project.
Contact Finkova with the project inputs above to discuss the relevant installation, testing and handover interfaces.
Frequently asked questions
What is failure mode and effects analysis?
Failure mode and effects analysis, FMEA, is a standardised method in which every component, its possible failure modes and their consequences are worked through systematically. When an assessment of criticality is included, the method is called FMECA.
Why should a data center have a failure mode and effects analysis?
Because it surfaces failure scenarios and knock-on effects nobody had thought possible. It shows whether the system survives a fault, how much time there is to act and what is worth monitoring. A documented analysis is also evidence with insurers and in questions of liability.
What does N-1 mean?
N-1 is a principle familiar from transmission grids: the system must keep working when one component fails. In data center cooling it means no single pump, dry cooler or CDU may stop the site. A failure mode and effects analysis shows whether the principle actually holds.
What is criticality classification?
Criticality classification sorts equipment into classes according to what a failure would cause for production, safety and cost. In Finland the PSK 6800 standard is most commonly used, with three classes. The classification drives the maintenance strategy and the level of monitoring.
What is the difference between machinery protection and condition monitoring?
Machinery protection stops a machine before it is damaged. It is fast, automatic and tied to fixed limits. Condition monitoring tells you about a machine’s health before anything happens, so maintenance can be scheduled. A data center needs both, but on different machines.
Why can redundant cooling still fail?
Because redundancy is not the same as fault tolerance. A duplicated machine does not help if both units draw power from the same board, take control from the same controller or get maintained at the same time. In Proton’s August 2026 outage, the filters on both redundant compressors were changed at once.
How fast does a data center heat up if cooling fails?
Faster than it used to. According to Proton’s report, time to a critical state had previously been three to four hours, but growth in equipment density had cut it to around twenty minutes. The hall temperature rose to nearly 52 °C in under half an hour.
What faults does condition monitoring miss?
A pump that leaks but keeps running, for example. Machinery protection does not trip because the machine runs normally, and vibration measurement shows nothing unusual. The fault is detected from make-up water consumption. This is why monitoring decisions should follow a failure mode analysis.
Which standard covers vibration measurement?
The evaluation criteria for vibration measurement are set out in ISO 20816-1, which replaced the earlier ISO 10816-1. The standard bases evaluation on both the level of vibration and its change.
Why must a vibration baseline be taken at commissioning?
Because a change in a machine’s own level says more than its absolute value. Two identical pumps can sit at different levels with nothing wrong, because installation and foundation affect the level. Without a baseline, later measurements give a value with nothing to compare against.