HelpFab

How to track machine downtime

Most small plants already track downtime. It just lives on a clipboard by the press, in a supervisor's notebook, or in a shared spreadsheet that one person updates on Friday afternoon. You know the line stopped. You cannot say for how long, how often, or why.

This guide walks through how to get from that state to downtime data you can act on, without buying an MES.

Start with a definition, not a tool

Before you log anything, write down what counts as downtime at your plant. Two people on the same line will otherwise record the same event differently.

A workable definition for most shops: downtime is any period the machine was scheduled to run but did not produce good parts. That single sentence forces three decisions:

Write the definition on a laminated card and hang it at the machine. Consistency matters more than precision here.

  • Unscheduled shifts do not count. If you run one shift, the other sixteen hours are not downtime.
  • Breaks and planned changeovers count, but as planned downtime, tracked separately from breakdowns.
  • Running while producing scrap is not downtime — it is a quality loss. Keep those separate or your Pareto will mix two different problems.

Set a threshold so operators log the right things

If you ask operators to log every stop, they will log nothing. Micro-stops — a jam cleared in twenty seconds — are real losses but they are not what your log is for at the start.

Set a minimum duration. Five minutes is a common starting point for job shops; two minutes for high-volume lines. Anything shorter gets captured later as a rate loss, once you have the basics working.

Build a short reason code list

This is where most downtime programs fail. Someone builds a 60-code taxonomy copied from a textbook, operators cannot find the right code, and everything gets logged as "Other."

Start with eight to twelve codes. Group them into planned and unplanned:

Planned

Unplanned

The "Other" code needs a mandatory comment field. Review those comments monthly — that is how you discover the codes you are missing. When "Other" comments cluster around one cause three months running, promote it to its own code.

Keep the list under fifteen codes for the first year. You can always split a code later; you can never un-confuse an operator who faced 60 options at 2am.

  • Changeover / setup
  • Preventive maintenance
  • Break / lunch
  • No schedule
  • Mechanical failure
  • Electrical / controls fault
  • Tooling failure or die issue
  • Material shortage
  • No operator
  • Quality hold
  • Waiting on inspection
  • Other (with a required note)

Decide who logs and when

Three workable models:

Most small manufacturers should start with model 1 on their two worst machines, not model 3 across the plant. Prove the process works before you wire anything up.

For every logged event you need five fields: machine, start time, end time, reason code, and a free-text note. That is it. Resist adding fields — every extra field cuts logging compliance.

  • Operator logs at the machine. Best data quality, requires a tablet or phone at each station. Log the stop when it happens, close it when the machine restarts.
  • Supervisor logs at end of shift. Lower friction, but durations become estimates and short stops disappear. Acceptable to start with.
  • Automatic capture from the PLC or HMI. The machine reports state changes; a human only assigns the reason code. Best of both, but it needs integration work and only pays off on lines with consistent equipment.

A worked example: 4 weeks on a stamping line

Here is what a month of data looks like on a 300-ton press running two shifts, five days a week. Scheduled run time was 320 hours.

That is 48.7 hours of downtime against 320 scheduled hours — availability of 84.8%.

Two things jump out that you cannot see from a clipboard:

Die changes are the single biggest bucket, but they are planned. 41 changeovers averaging 30 minutes each. This is not a breakdown problem, it is a SMED problem. Cutting the average changeover from 30 to 20 minutes returns 6.8 hours a month — more than eliminating every mechanical failure would.

Tooling failure is only 12 events but 720 minutes. Average 60 minutes per event, five times worse per occurrence than anything else. That is where a maintenance conversation belongs. Frequency and duration tell different stories, which is why you track both.

Material shortage is 18 events. That is not a machine problem at all — it is scheduling or inventory. Downtime tracking regularly surfaces problems that live in another department, which is exactly why the log needs to be visible outside maintenance.

The three reports that matter

Once you have four weeks of data, you need exactly three views. More than that and nobody reads any of them.

1. Pareto by total downtime minutes

Bars sorted descending, with a cumulative percentage line. This answers "where does the time go" and is the report you take to the weekly production meeting. In the example above, two codes account for 67% of all downtime.

2. Pareto by event count

Same chart, different measure. This answers "what goes wrong most often" and surfaces the chronic irritants that a minutes-based Pareto hides. Frequent short stops erode operator confidence and often precede a big failure.

3. Trend over time

Downtime minutes per week, split planned versus unplanned. This is the only report that tells you whether your improvements are working. A single month of data tells you where the problems are; twelve weeks tells you whether you fixed them.

Plot the trend by week, not by day — daily data on a single machine is too noisy to read.

Common mistakes

  • Tracking everything at once. Two machines for two months beats twenty machines for two weeks.
  • Letting the log become a blame tool. The moment operators believe downtime data is used to evaluate them, the data becomes fiction. Say out loud that you are measuring the process, not the person, and then act like it.
  • Never closing the loop. If you collect data for three months and change nothing, logging compliance collapses. Fix one thing from the first Pareto and tell the floor that the data is what found it.
  • Comparing availability between dissimilar machines. A press and a CNC cell have different natural availability. Compare each machine to its own history.
  • Rounding everything to 15 minutes. Estimated durations cluster on round numbers and destroy your averages. Log actual start and stop times.

Where HelpFab fits

HelpFab's Issues and downtime tracking module is built for this workflow: operators log an issue against a machine from a phone or tablet, choose from your own configurable issue types, and the Issue Reports page builds Pareto charts and trend views over any date range. Because it sits alongside maintenance schedules and production tracking, the downtime you log connects to the work orders and PM plans it affects.

If you are still on a clipboard, though, the tool is the last decision, not the first. Write the definition, build the code list, pick two machines, and run it for a month. The data will tell you what to automate next.

Explore HelpFab

  • Home
  • Learn more
  • Pricing & plans
  • Issues tracking
  • Project tracking & job costing
  • Inventory software
  • Production tracking
  • Shop floor tracking
  • Excel alternatives
  • Factory production tracking
  • Blog