2
Types: Planned & Unplanned
6
Big Losses
Code
Every Minute of Downtime
Pareto
Focus on the Top 3

Why Downtime Analysis Matters

Downtime is the most visible and measurable form of lost capacity. Every minute your bottleneck is down is a unit you will never make. Yet most plants have only a vague idea of why their equipment stops — "it was down for maintenance" or "we had some issues." Without precise categorization, you cannot improve.

Structured downtime analysis transforms vague complaints into specific, actionable improvement targets. It answers: what stops the equipment most often, for how long, and what can we do about it?

Planned vs. Unplanned Downtime

TypeDefinitionExamplesGoal
PlannedScheduled stops that are part of the production planChangeovers, PM, breaks, meetings, cleaning, material replenishmentMinimize through SMED, efficient PM, structured handoffs
UnplannedUnexpected stops that disrupt the scheduleBreakdowns, quality holds, material shortages, operator absence, utility failuresEliminate through TPM, poka-yoke, supplier development

The 6 Big Losses

The TPM framework categorizes all equipment losses into 6 types, grouped by the OEE factor they affect:

OEE FactorLossDescriptionCountermeasure
Availability1. Equipment FailureBreakdowns that stop productionPM/PdM, AM
2. Setup & AdjustmentChangeovers, startups, adjustmentsSMED, standardized setup
Performance3. Idling & Minor StopsBrief stops <5 min (jams, sensor trips, feeding errors)Root cause analysis, error-proofing
4. Reduced SpeedRunning below rated speedRestore design conditions, operator training
Quality5. Process DefectsScrap and rework during steady-state productionCapability improvement, SPC
6. Startup RejectsDefects during warmup, changeover, startupStandardized startup procedures, standard work

Building a Downtime Coding System

A coding system is essential for consistent categorization. Without it, the same event gets logged differently by each operator ("breakdown" vs "machine down" vs "maintenance issue") making analysis impossible.

Define 10-15 codes (no more)Too few codes = no detail. Too many = inconsistent use. Start with broad categories that match the 6 big losses, then add 2-3 specific codes for your plant's most common issues.
Make codes mutually exclusiveEvery downtime event should fit in exactly one code. If operators cannot decide between two codes, the definitions are unclear — revise them.
Post the code list at every machineLaminated reference card at the operator station. If they have to remember codes from memory, they will guess or skip it.
Record start time, end time, and codeMinimum data: when it started, when it ended, and the code. Duration is calculated. If possible, add a brief description for context.

Example Coding System

CodeCategoryDescription
BDBreakdownUnplanned equipment failure requiring maintenance
COChangeoverPlanned product change (setup + adjustment)
PMPlanned MaintenanceScheduled PM activity
MSMaterial ShortageLine waiting for raw material or components
QHQuality HoldStopped for quality investigation or containment
MNMinor StopBrief stop <5 min (jam, sensor, feeding)
SPSpeed LossRunning below rated speed
NPNo PlanNo production scheduled (no demand)
STStartup/ShutdownWarmup, cooldown, end-of-shift shutdown
OTOtherDoes not fit above categories (review periodically — if >10%, add a code)
Frequency is not minutes Minor Stops is first of six by number of stops and last of six by minutes lost; Operator Absence is last by stops and fourth by minutes. Sort the weekly Pareto by frequency and it points at a different code from the one that is actually costing the time.

The prose on this page carries no dataset, so every number below comes from the one worked example the page does contain: the six default rows of the Interactive Demo further down, each of which the prose also names, and which the reader can perturb to see this figure's argument bend. Mechanical Failure 8 × 45 = 360 min; Changeover 12 × 25 = 300; Material Shortage 5 × 60 = 300; Operator Absence 3 × 90 = 270; Quality Hold 6 × 20 = 120; Minor Stops 20 × 5 = 100 — 54 stops and 1,450 minutes in the month. Minor Stops' five-minute average is the page's own definition of the code ("brief stop under 5 min"), and Changeover is the only planned code of the six, which is why it alone is drawn as a hollow square. Everything else is arithmetic on those twelve numbers: Minor Stops is 20 of 54 stops (37%) but 100 of 1,450 minutes (7%); Operator Absence is 3 of 54 stops (6%) but 270 minutes (19%). The top two by count are 32 of 54 stops and only 400 minutes (28%); the top two by minutes are 660 minutes (46%) from just 20 stops. Changeover and Material Shortage tie at 300 minutes and are shown in the demo's own row order. The countermeasure phrases in the corners are the page's own, from the 6 Big Losses table.

Analyzing Downtime

Weekly Pareto by codeSort downtime minutes by code. The top 2-3 codes are your improvement targets. Post the Pareto on the visual board.
Trend over timeTrack total downtime and unplanned downtime as a % of available time, weekly. Is it improving? If the trend is flat, your countermeasures are not working.
Deep dive on top causeTake the #1 Pareto category and break it down further: which machine? Which component? Which shift? Use RCCA on the specific failure mode.
Calculate the costConvert downtime minutes to lost production value using the downtime cost calculator. This makes the business case for improvement undeniable.
✅ Good Downtime Tracking
  • Every stop coded in real time by the operator
  • Simple, clear coding system posted at every station
  • Weekly Pareto review at T2 meeting
  • Top causes get formal RCCA projects
  • Trend tracked monthly — downtime % is declining
❌ Downtime Guessing
  • "We were down for a while" — no duration, no code
  • Data entered at end of shift from memory
  • 50 codes that no one remembers
  • "Other" is the #1 category
  • Data collected but never analyzed or acted on

🎯 Key Takeaway

You cannot reduce downtime you do not measure. Build a simple coding system (10-15 codes), train operators to log every stop in real time, and Pareto the results weekly. Focus improvement on the top 2-3 causes. When those improve, re-Pareto and attack the new top causes. Combine with OEE tracking on your bottleneck to see how downtime reduction translates directly into throughput recovery.

Interactive Demo

Edit downtime events to see how categories, MTBF, MTTR, and availability metrics update. Identify the top loss drivers from your data.

⚡
Try It Yourself
Downtime Analyzer
▼
Enter downtime events and durations for each category. See which causes drive the most lost time and how they affect availability, MTBF, and MTTR.
720 hrs
200 hrs1200 hrs
CategoryEventsAvg MinTotal Lost Time
Mechanical Failure
360m
Changeover
300m
Material Shortage
300m
Operator Absence
270m
Quality Hold
120m
Minor Stops
100m
24.2 hrs
Total Downtime
96.6%
Availability
12.9 hrs
MTBF
27 min
MTTR
Ready for the full knowledge check? Test your understanding with guided scenarios and data export.
PROTake the Pro Knowledge Check →
Free forever · Every feature included

Stop reading, start modeling

Model your process flow, run simulations, optimize staffing with TOC math, and test your knowledge with 107 interactive checks — all in one platform.

Open Workbench → Pro Knowledge Check

Take this to a room

The running order

For a line team starting to measure. They should leave with a draft code list of 10 to 15 codes and agreement to log in real time.

7 beats · 10 min
  1. 1

    You cannot reduce what you did not code

    Every improvement conversation about downtime stalls in the same place: nobody knows what the minutes were.

    • We were down for a while is not data.
    • Recording at end of shift from memory is not data either.
    • The number you need is minutes by cause, logged as it happens.

    Ask the room How much downtime did our line have last week, by cause?

  2. 2

    Planned and unplanned are different problems

    Both are lost time, but you attack them with entirely different tools.

    • Planned: changeovers, PM, breaks, meetings. Minimise with SMED and structured handoffs.
    • Unplanned: breakdowns, quality holds, material shortages. Eliminate with TPM and error-proofing.
    • Calling a changeover downtime is fine. Treating it like a breakdown is not.
  3. 3

    Ten to fifteen codes, no more

    The number of codes is the design decision that makes or breaks the system.

    • Too few and you get no detail. Too many and use becomes inconsistent.
    • Start from the six big losses, then add two or three specific to this plant.
    • Fifty codes nobody remembers gives you an Other category at the top.
  4. 4

    Codes must be mutually exclusive

    If an operator cannot decide between two codes, the definitions are wrong - not the operator.

    • Every event fits exactly one code.
    • Laminate the list and post it at every station.
    • If they have to remember from memory they will guess or skip it.
  5. 5

    Weekly Pareto, top two

    The analysis is the same every week and takes ten minutes.

    • Sort minutes by code. Post the chart on the visual board.
    • The top two or three codes are the improvement targets.
    • Then deep dive: which machine, which component, which shift?
  6. 6

    Watch the trend, not the week

    A single week tells you nothing. Unplanned downtime as a percentage of available time, tracked weekly, tells you everything.

    • If the trend is flat, the countermeasures are not working.
    • Say that out loud rather than re-explaining the same Pareto.
    • Re-Pareto after each fix - the new top cause is rarely the old second.
  7. 7

    Convert to money

    Minutes do not move budgets. Lost production value does.

    • Turn downtime minutes into lost output at your own rate.
    • On the bottleneck, every recovered minute is a shipped unit.
    • Everywhere else it is not - which is itself worth saying.

    Ask the room Which machine do we code first, and who trains the operators on it?