Measuring a Schedule: OTIF, Tardiness, and Utilization

Ask two planners what makes a schedule “good” and you’ll often get two different answers, and both can be right for their own shop. One runs a contract machine shop where a single missed customer date is expensive and personal; the metric that matters most to them is whether every order ships on time. Another runs a make-to-stock line where the real cost is a churning, unstable plan that changes every time it’s regenerated; the metric that matters most to them is whether yesterday’s plan still looks like today’s. Neither is wrong. A schedule doesn’t have one score — it has several, and which ones a shop weighs most heavily is a business decision, not a universal ranking.

Here are the measurements that come up most, what each one actually captures, and where they pull against each other.

On-time delivery and OTIF

On-time delivery is the simplest of the group: of the orders due in a period, what percentage shipped on or before their due date. OTIF (On Time, In Full) tightens it further — an order only counts as a success if it shipped on time and at the full quantity ordered, closing the loophole where a partial, late-adjacent shipment gets counted as a win. OTIF is usually the metric a customer actually feels, which is why it tends to be the headline number a shop reports upward even when internal planning leans on more granular metrics day to day.

Lateness is the raw difference between when an order actually finished and when it was due — and it can be negative. An order that finishes three days early has a lateness of −3; an order that finishes two days late has a lateness of +2. Tardiness is lateness with the negative values clipped to zero: an order finishing early contributes 0 tardiness, never a negative “credit.” This distinction matters because a scheduler that’s allowed to net early finishes against late ones would happily let some orders run six days late as long as others finish six days early to “balance out” — which is nonsense from a customer’s point of view. Nobody’s late order gets forgiven because a different customer’s order arrived early.

Tardiness itself gets reported a few different ways depending on what a shop actually cares about:

  • Total tardiness — sum of days (or minutes) late across every order. Good for a plant-wide trend line, bad at spotting one catastrophically late order buried among many that are only slightly late.
  • Maximum tardiness — the single worst offender. Good for catching the one order that’s going to trigger an angry phone call; blind to whether everything else is a little late too.
  • Weighted tardiness — total tardiness with each order multiplied by a priority or value weight before summing, so a day late on a strategic account counts for more than a day late on a low-priority stock replenishment. This is usually the closest to what a business actually wants to minimize, and the hardest to compute by hand.

A shop that only tracks total tardiness can look fine on the dashboard while quietly burning its most important customer’s patience, because total tardiness doesn’t know which order was the important one.

Makespan

Makespan is the total span of time from when the first scheduled operation in a set starts to when the last one finishes. It’s a compression metric, not a due-date metric — a schedule can have a short makespan and still miss every due date if it’s compressed in the wrong place, or a long makespan and hit every date if there was slack to spare. Makespan matters most when the question is “how fast can this whole batch of work clear the floor,” which comes up in make-to-order shops racing to free up capacity for the next release, more than in shops managing a steady drip of individually-due orders.

Utilization — and why 100% is a warning sign, not a goal

Utilization measures how much of a resource’s available time was actually used. It’s intuitive to want it high — an idle machine looks like wasted money — and it’s one of the most commonly over-weighted metrics in scheduling, for a reason covered in more depth in Bottlenecks, Drum-Buffer-Rope, and Theory of Constraints: a resource running at 100% utilization has zero slack to absorb an ordinary bad day, and pushing every non-bottleneck resource toward 100% just builds work-in-process in front of whatever the real constraint is, without changing what actually ships. High utilization on the bottleneck is valuable. High utilization everywhere else is often a sign the schedule is optimizing something other than what the business actually needs.

Schedule adherence and stability

Adherence measures how closely a newly regenerated plan matches the plan that was already published and acted on — how many operations kept their previously-communicated start time versus how many moved. This is a different axis entirely from on-time performance: a plan can be perfectly on-time and still churn constantly, reshuffling tomorrow’s sequence every time it’s regenerated, which is its own real cost. A shop floor that gets handed a different dispatch order every morning stops trusting the schedule and starts working around it — see Plan Stability: Frozen Horizons, Nervousness, and When to Reschedule for the mechanism that manages this trade-off directly.

WIP and queue time

Work-in-process (WIP) — how many orders are in some partially-completed state at once — and the queue time each of them spends waiting for a resource to free up are less headline-grabbing than on-time delivery, but they’re often the earliest warning sign that something’s wrong. Rising WIP with flat throughput usually means work is being released faster than the true bottleneck can absorb it (the exact failure mode described in the bottleneck article’s cutting-welding-finishing example) — the schedule looks busy and is quietly getting less efficient at the same time.

A practical version of this shows up on almost any shop floor: a batch that finishes machining at 2pm and doesn’t start its next operation until 6pm wasn’t actually “in process” for those four hours in any meaningful sense — it was sitting in a queue behind other work. If queue time is quietly climbing week over week while run times stay flat, the honest read is rarely “the machines got slower.” It’s almost always “more work is being released into the plant than the tightest resource can clear,” and the fix lives upstream of the schedule, in how much work gets released, not in the sequence itself.

Changeover time as its own KPI

Total changeover time — minutes spent on setup driven purely by the transition between two jobs, as opposed to the fixed setup every job needs regardless of sequence — is worth tracking on its own line, not folded into generic “setup time.” See Sequence-Dependent Setup for the full mechanism; the KPI version of the story is simple: a shift’s worth of changeover time is capacity that never produced anything, and it’s one of the few costs on this list that a smarter sequence can reduce without buying anything or working anyone harder.

Every schedule is a weighted compromise

Here’s the part that surprises planners new to the craft: you cannot maximize all of the above at once, and treating any one of them as an unconditional goal breaks the others. Push hardest on zero tardiness and you’ll sometimes need to compress the plan, running some jobs earlier than optimal and increasing changeover or reducing stability to hit every date. Push hardest on plan stability and some genuinely better sequence will go unused because it would have moved too many already-committed start times. Push hardest on minimizing changeover and, as the sequence-dependent-setup article’s paint-line example shows, a due-today order can get buried behind a “cheaper” sequence that has no idea it was due today.

There is no metric-free way out of this. A real production schedule is the result of some set of priorities — implicit or explicit, in someone’s head or in a solver’s objective weights — deciding which of these trade-offs wins when they collide, because in a fully loaded plant, they collide constantly. A contract machine shop and a make-to-stock bakery can look at the identical set of KPIs and legitimately weigh them in opposite orders, and both can be running a well-managed operation. The craft isn’t finding the schedule that wins on every metric — that schedule usually doesn’t exist — it’s being honest and deliberate about which metric gives ground when they can’t all win together.

A trade-off, worked

Take one work center with five orders queued, three of them at risk of running late under the current sequence. A planner has two candidate resequences to choose from.

Sequence A clears all three at-risk orders first. Total tardiness drops to zero and maximum tardiness drops to zero — a clean sweep on the on-time metrics. But getting there means reordering all five operations relative to what was published yesterday, so adherence takes a real hit, and because the three at-risk orders don’t share a product family, the resequence also adds two extra changeovers that the previous plan didn’t have.

Sequence B only moves the single worst-late order to the front and leaves the other four in their previously published order. Total tardiness drops, but not to zero — the two lesser-late orders stay a little late. Adherence stays high (only one operation moved), and changeover time barely changes, because the reordering only touched one slot in an otherwise-unchanged sequence.

Neither sequence is objectively better. A shop where every late order draws a customer escalation will take Sequence A’s clean on-time sweep and accept the churn and the extra changeover as the price of it. A shop that got burned last quarter by a floor that stopped trusting a schedule that changed every morning will take Sequence B’s smaller, more stable move and accept that two orders stay marginally late. Both are legitimate answers to the same set of numbers — which is exactly the point: the KPIs don’t choose the sequence, they just make the trade-off visible enough that a planner (or a solver configured with the shop’s actual priorities) can choose it deliberately instead of by accident.

Where SmartFlow fits

SmartFlow APS surfaces these metrics on a KPI dashboard and role center, and a finite plan’s review panel shows the before/after picture on the KPIs that moved when a schedule change is proposed — so the trade-off described above is visible before a planner accepts it, not discovered afterward. See KPI Dashboard and Role Center for what’s on the dashboard today.

Ready to see finite scheduling on your own data?

Support: support@dynamicspro.ca · 416-843-6575

Get SmartFlow APS on AppSource