MalakarConsulting

HomeGuides › Supplier scorecard KPIs: OTIF, PPM and the metrics that actually change behavior

Supplier scorecard KPIs: OTIF, PPM and the metrics that actually change behavior

Updated September 21, 2026

A supplier scorecard is a small table that turns "I feel like this vendor is always late" into numbers both sides can look at. Done well, it takes an hour a month, gives you leverage in renegotiation, and shows which suppliers to grow and which to replace. Done badly, it is 30 metrics nobody reads.

This guide gives you six core metrics with exact formulas, an example of how to combine them into one score, suggested review rhythm, and the traps that make scorecards useless. The weights and thresholds are examples for you to adjust. There are no industry-standard targets that fit every business, so set yours from your own history and your customers' requirements.

Principles before metrics

  • Measure what you can verify from your own records (purchase orders, receipts, inspection results, invoices), not what the supplier reports about itself.
  • Use few metrics. Five to seven, each with a clear formula and a data source.
  • Agree on definitions with the supplier first. Is "on time" measured against the promised date on the PO, the date they confirmed, or the date you requested? Whichever you choose, write it down.
  • Measure by line or by PO consistently. Line-level metrics are more precise. Order-level metrics are easier to collect.

The six core metrics

MetricFormulaData source
On-time delivery (OTD)Lines received on or before the agreed date ÷ total lines due in the periodPO due dates, receiving log
In-full (fill rate)Lines received at or above the ordered quantity (within your tolerance) ÷ total lines duePO quantity, receiving log
OTIF (on time, in full)Lines that are both on time and in full ÷ total lines dueCombines the two above; a line counts only if it passes both tests
Quality: defect PPM(Defective units ÷ units received) × 1,000,000Incoming inspection, line rejects, returns
Lead-time reliabilityAverage promised lead time vs. actual, and the standard deviation of actualPO date, receipt date
Responsiveness and corrective actionDays to acknowledge a problem; days to close a corrective action requestEmail and CAR log

Two more that matter for many buyers: price variance (actual price paid against the agreed or quoted price, as a percentage) and cost to serve (expedites, rework and extra freight you paid because of this supplier). Add them if you can measure them without extra effort.

Why OTIF is stricter than OTD

A supplier can score 95 percent on time and 95 percent in full and still be well below 95 percent OTIF, because the two failures may land on different lines. Suppose out of 100 lines, 5 arrive late and a different 5 arrive short. On time is 95 percent, in full is 95 percent, and OTIF is 90 percent. If both misses fall on the same 5 lines, OTIF is 95 percent. Always compute OTIF line by line, not by multiplying the percentages.

Worked defect PPM example

You receive 48,000 units in a quarter and reject 96. PPM is 96 ÷ 48,000 × 1,000,000, which is 2,000 PPM, or 0.2 percent. PPM is useful for comparing suppliers with very different volumes. Count defects consistently: units rejected at receiving, units found defective on your line, and field returns each tell a different story, so track them in separate columns.

Combining metrics into one score

A single composite score makes ranking easy, but the weights are a business decision. Here is an example for a manufacturer that cares mostly about quality and delivery. Adjust to your priorities.

CategoryWeightScored 0-100 from
Quality35%PPM against your target; 100 at or below target, 0 at your reject limit
Delivery (OTIF)30%OTIF percentage, scaled between your floor and target
Cost20%Price variance and cost to serve
Responsiveness15%Corrective-action closure time and communication

Suppose a supplier scores 80 on quality, 90 on delivery, 70 on cost and 60 on responsiveness. The composite is 0.35×80 + 0.30×90 + 0.20×70 + 0.15×60 = 28 + 27 + 14 + 9 = 78. Set bands such as 85 and above preferred, 70 to 84 acceptable with monitoring, and below 70 on a formal improvement plan. Again, the bands are yours to define.

Scale each metric explicitly. For OTIF, for example, decide that 98 percent or better is 100 points and 80 percent or worse is 0 points, and interpolate between them. Without that, a scorecard drifts as people improvise.

Review rhythm

  • Monthly: update the numbers, flag anything that moved more than a few points.
  • Quarterly: share the scorecard with strategic suppliers and hold a short review: what went well, the worst three misses and their root cause, and agreed actions with dates.
  • Annually: revisit weights, targets and the supplier list. Decide who to grow, hold, develop or exit.

Use trends more than snapshots. A supplier at 85 and falling is a bigger concern than one at 75 and improving. If you keep a quality management system, scorecard reviews double as evidence for supplier evaluation (see the ISO 9001 documents guide).

Traps that make scorecards useless

  • Too many metrics. Nobody can act on twenty numbers.
  • No agreed definitions. When the supplier disputes the numbers, the review becomes an argument about data.
  • Ignoring volume. A 60 percent OTIF from a supplier with three lines a year is noise. Show the count next to the percentage and set a minimum sample size before you rank.
  • Charging your own mistakes to the supplier. If your purchase orders were late or your forecasts changed, the "late" delivery may be your fault. Measure against the date the supplier actually agreed to.
  • No consequences and no recognition. If the score never changes anything, it is a report, not a management tool. Link it to volume awards, payment terms or an improvement plan.
  • Scoring only performance. Add a risk column too: single-source parts, country concentration, and the supplier's financial health. The supplier verification guide covers the initial checks that feed it.

Setting one up in an afternoon

  1. Export the last 6 to 12 months of PO lines with promised date, received date, ordered and received quantity, and supplier.
  2. Pull rejects and returns by supplier from your quality log.
  3. Calculate OTD, in-full, OTIF and PPM per supplier per month in a spreadsheet.
  4. Choose weights and scoring scales, and calculate the composite.
  5. Rank suppliers, highlight the top and bottom, and pick two conversations to have this quarter.

Reviewed September 21, 2026. Formulas are standard supply chain practice; weights and targets are examples and should be set from your own data.

Frequently asked questions

What is OTIF and how do you calculate it?

OTIF stands for On Time, In Full. It is the number of order lines received on or before the agreed date and at or above the ordered quantity, divided by the total number of lines due. Calculate it line by line, since a line only counts if it passes both tests.

How do you calculate defect PPM?

Defect parts per million equals defective units divided by units received, multiplied by 1,000,000. For example, 96 rejects out of 48,000 units received is 2,000 PPM, or 0.2 percent.

How many KPIs should a supplier scorecard have?

Five to seven is enough for most small buyers: on-time delivery, in-full, OTIF, quality PPM, lead-time reliability, responsiveness, and optionally price variance. More metrics tend to reduce action, not improve it.

How often should I review supplier scorecards?

Update the numbers monthly, hold a short review with strategic suppliers quarterly, and revisit weights, targets and the supplier list annually.

Have purchase and quality records but no time to score them? Send your supplier list and records to the Supplier Scorecard & Vendor Performance Report and get a scored, ranked scorecard per supplier with red flags and talking points for your next review. — see what's included and order ($200) →