Skip to content
Guide · Manufacturing

How to measure machine downtime

Downtime is the biggest loss on any line — and the least measured. Very few plants can say how long, for what reason and on which shift. Here's the whole method: what to record, how to classify it, the two metrics that matter (MTBF and MTTR) and the four steps between a paper sheet and automatic PLC data.

Machine list from a monitoring system: each piece of equipment with model, code, connection status and last recorded output
Machine registry from a monitoring system we delivered, with client details removed. Interface shown in Portuguese.

What to record for every stop

Four fields. Fewer than this and the data is useless; more than this and the operator won't fill it in.

Start and end, by the clock

Not an estimated duration written later. The actual time, as it happens. Fifteen stops of “about ten minutes” turn into two and a half hours nobody can account for.

Machine, shift and operator

Without these the data can't be cross-referenced with anything. This is what lets you discover the breakdown is always on the same machine, or always on the same shift.

The reason, from a closed list

A free-text field turns into “problem” and “other”. A list of 10 to 15 standardised reason codes is what makes the Pareto work later.

Planned or unplanned

Changeovers, meal breaks and preventive maintenance are planned stops — they come out of available time. Breakdowns and material shortages are losses. Mixing the two ruins the calculation.

The two numbers that come out of it

MTBF = run time ÷ number of failures

Mean Time Between Failures. The higher, the more reliable the machine. It drops when preventive maintenance is overdue or the machine is running at its limit.

MTTR = total failure downtime ÷ number of failures

Mean Time To Repair. The lower, the faster the response. It rises when the spare part isn't in stock, when the technician is far away or when nobody was notified in time.

Together they tell you what the problem is. Low MTBF calls for preventive maintenance and root-cause analysis. High MTTR calls for spare parts inventory, immediate alerts and a defined response path — and no investment in a new machine fixes that.

A worked example

One machine, one week of five 8-hour shifts, each with a planned 30-minute meal break. During the week there were 6 breakdowns adding up to 3 hours of downtime.

Planned time
5 × (480 − 30) = 2,250 min
Breakdown downtime
6 failures, 180 min in total
Run time
2,250 − 180 = 2,070 min
MTBF
2,070 ÷ 6 = 345 min (5 h 45)
MTTR
180 ÷ 6 = 30 min

Availability = 2,070 ÷ 2,250 = 92%

An MTTR of 30 minutes looks acceptable until you read the detail: 4 of the 6 failures lasted 10 minutes and the other 2 lasted 70 — both waiting on a spare part that wasn't in stock. The average hides that; the reason list shows it. That's why you record every stop, not just the shift total.

The four steps of downtime tracking

Nobody jumps from paper to sensors. Each step has a gain and a limit, and the right move is to climb when the current limit starts costing money.

  1. 01

    Paper at the machine

    One sheet per shift with start, end and reason. Costs nothing and already changes the conversation. Limit: short stops never make it in, and someone has to type it all up later.

  2. 02

    Spreadsheet

    The same log, typed in. You gain the Pareto and a history. Limit: the data arrives hours later and depends on someone's typing discipline.

  3. 03

    Shop floor data collection

    A tablet or terminal next to the machine; the operator logs the stop as it happens, from a closed list. You gain accuracy and real time. Limit: micro-stops still depend on someone remembering.

  4. 04

    PLC data + operator reason

    The PLC or a sensor reports that the machine stopped and for how long; the operator only classifies it. No stop escapes, and the alert fires while there's still time to act.

The mistakes that make the log useless

Almost every plant that “tried tracking downtime and it didn't work” tripped over one of these.

  • Logging at the end of the shift, from memory — short stops vanish and long ones shrink.
  • Leaving “other” on the reason list. Within three months it's the plant's biggest reason and says nothing.
  • Counting changeovers as breakdowns, or breakdowns as changeovers. One calls for process engineering, the other for maintenance.
  • Measuring time stopped and never looking at frequency. Twenty 3-minute stops cost more than one 1-hour stop — and have a different cause.
  • Keeping the log and never running the Pareto. Data that doesn't turn into a decision is just paperwork for the operator.

Frequently asked questions

What are MTBF and MTTR?
MTBF (Mean Time Between Failures) is run time divided by the number of failures. It tells you how reliable the machine is. MTTR (Mean Time To Repair) is total failure downtime divided by the number of failures. It tells you how fast maintenance responds. A machine can have a high MTBF and a terrible MTTR — it rarely breaks, but when it does it's down for a day — and the fix is completely different from the opposite case.
How many downtime reason codes should I have?
Between 10 and 15 is usually the sweet spot. Fewer than that and everything falls into two or three generic buckets; more and the operator can't find the right one and picks any. Start with the reasons your team already names off the top of their heads, run it for a month and adjust: codes that never get used come out, and whatever keeps showing up under “other” gets its own line.
Can I track downtime without sensors?
Yes, and that's how most plants start: a sheet or a spreadsheet per machine, with the operator logging start, end and reason. The limit of that method is short stops — 30 seconds to 2 minutes — which add up fast and nobody writes down. When the manual Pareto no longer explains the gap between what the machine should produce and what it does, it's time to read the running signal straight from the PLC or a sensor.
What is a downtime Pareto?
Ranking the reasons from most to least accumulated downtime. In almost every plant, two or three reasons account for 70% to 80% of time lost. The Pareto is what stops maintenance from spending the week on the most annoying problem instead of the most expensive one — and it only exists if the reason was logged from a closed list.
How often should I look at downtime data?
The Pareto weekly — that's the rhythm at which you can attack one cause and see the effect. The log itself, at the moment of the stop. And the alert for a machine stopped longer than normal, in real time: the difference between knowing at 9:05 and knowing at the end of the shift is the difference between 5 minutes and 3 hours of loss.
Where does Volvi come in?
We build the fourth step: the machine reports it has stopped (via PLC or sensor), the operator picks the reason from a list your plant defines, and the software calculates MTBF, MTTR and the Pareto per machine and per shift — with a configurable long-stop alert. It's already running in a wire drawing and straightening plant. The team has 25 years of manufacturing experience and works remotely, available for plants in the UK, the US and beyond. If your plant is still on paper, we help design the sheet and the reason list first; the software comes once the method exists.

Want to know how much your machines stop, and why?

Tell us how many machines you have and how downtime is logged today. We reply within one business day with a path forward — starting with the machine that hurts most.

Get in touch