Field guide · Accuracy
WMS inventory accuracy:
measuring it so the number survives scrutiny
Accuracy is a measured quantity, and the measurement decides whether the number means anything. The same cycle count can honestly produce 99.97%, 98.18% or 96.4% — and only one of those predicts whether picking is about to short.
Short answer
WMS inventory accuracy measures how closely system records match physical stock. The figure only means something once you state its unit: accuracy by piece, by SKU, by location or by value produce different numbers from the same count. Location-level accuracy — does this location hold exactly this item in exactly this quantity — is the measure that predicts whether picking will short.
What inventory accuracy is actually measuring
Inventory accuracy is the proportion of inventory records that match physical reality. The definition is uncontroversial. The measurement is where it falls apart, because "proportion of records" requires you to say what a record is — and there are at least four defensible answers.
- Piece accuracy (net)
- Total counted eaches against total system eaches, with overs and unders allowed to cancel. Produces the highest number of the four, almost always. It answers a financial question: is the total valuation approximately right?
- Piece accuracy (absolute)
- The same calculation using the absolute value of each variance, so an over of 400 and an under of 400 sum to 800 rather than to zero. This is the honest version of piece accuracy and the one worth trending.
- Location accuracy
- Of the locations counted, how many held exactly the expected item in exactly the expected quantity and status. Binary per location: a location wrong by one each scores zero, the same as a location wrong by nine hundred. Harsh, and the closest proxy for whether a picker will find what the system promised.
- SKU or value accuracy
- Accuracy weighted by item or by extended cost. Useful for prioritising work and for finance, misleading as a general operational health measure, because a warehouse's service performance is not weighted by cost.
None of these is wrong. The failure mode is reporting one number without saying which it is — and it is almost always net piece accuracy, because that is what falls out of a standard variance report most easily, and it is the most flattering.
One count, three honest answers
Take a single cycle count pass. 500 locations counted, holding 120,000 eaches in total. 18 of those locations turn out to be wrong. Because some are over and some are under, the variances partly cancel: the net variance across all 18 is 40 eaches, while the gross variance — the sum of the absolute values — is 2,180 eaches.
Every number below comes from that one count. Nothing is estimated and nothing is benchmarked against anyone else.
The gap between the first and second line is the entire problem with net measurement. Netting collapsed 2,180 eaches of genuine error into 40, because an overstatement in one location was allowed to pay for an understatement in another. Those two errors are not related, they have different causes, and neither is fixed by the other's existence. Worse, the pair is a signature — equal and opposite variances usually mean a move that posted one leg, which is a specific, findable defect that netting deletes from the report.
The third line is the one a picker experiences. 96.4% location accuracy over 500 locations means roughly one location in twenty-eight will not hold what the system says. On a pick path of thirty lines, that is about one exception per path — which is exactly the kind of number that produces a floor convinced the system is unreliable while the weekly accuracy report reads 99.97%.
Both figures are arithmetically correct. Only one of them is answering the question anyone on the floor is asking.
Signs your accuracy number is lying to you
These are measurement defects, not inventory defects. Each one inflates the published figure without changing a single unit on the floor.
- Nobody can state the unit. If the person who publishes the figure cannot say whether it is net pieces, absolute pieces, locations or value, the number is not a measurement. This is the first thing to check and it resolves surprisingly often.
- Counts are not blind. A counter who can see the expected quantity is confirming, not counting. Anchoring is not a character flaw; it is what happens to everyone, which is why the blind-count flag is a configuration setting rather than a training topic.
- Variance is "resolved" by recounting until it matches. A recount is legitimate when the first count is in doubt. It is not legitimate as a default response to any non-zero variance, and a recount rate that rises with variance size is evidence of the second pattern rather than the first. See cycle count discrepancies for separating noise from repeatable failure.
- Tolerances absorb real error. A ±2 each tolerance on a pack of 12 does not smooth noise, it hides a third of a case. Tolerances set in pieces rather than as a percentage of the location's quantity systematically forgive small-quantity locations.
- The sampling frame excludes the hard locations. Reserve, bulk, floor-stacked, mixed-item, damaged and quarantine locations are the most likely to be wrong and the most often left out of the count programme because they are inconvenient.
- Only fast movers are counted. ABC-driven counting is sound practice, but if the accuracy figure is computed only from the A-class sample it describes a subset chosen for being well-managed.
- Accuracy recovers after every count and decays at a constant rate. The count is treating symptoms on a schedule. The decay slope is the actual measurement of interest, and almost nobody plots it.
- Reported accuracy is above 99% while short picks are routine. The two cannot both be true at location level. One of them is measuring something else.
What actually drives accuracy down
Assuming the measurement itself is sound, these are the mechanisms that put real error into the records. They are the same mechanisms that produce individual discrepancies, viewed as a population rather than as instances.
- Master data that no longer matches the physical world
- Case packs, UOM conversions, TI and HI, and location capacities that were correct at go-live and have not been maintained through vendor changes. Produces directional, proportional error that grows with receipt volume — see phantom inventory.
- Confirmation by keying rather than scanning
- Wherever a WMS allows a task to be confirmed by entering the expected location instead of scanning the actual one, the system records the plan rather than the event. Quantities stay right, placement drifts, and location accuracy falls while piece accuracy holds up — which is exactly the pattern that makes a net figure look healthy.
- Mixed-item and mixed-lot locations
- Any location holding more than one item or lot multiplies the ways a count can be wrong and makes blind counting substantially harder. Locations configured to permit mixing are worth reporting on separately; they usually carry error out of proportion to their share of the building.
- Partial and interrupted replenishment
- Replenishment tasks confirmed short, split across two pallets, or abandoned mid-task when a higher-priority interleave arrives. Each interruption is an opportunity for one leg to post without the other.
- Unrecorded physical work
- Stock moved to clear an aisle, consolidated at shift end, pulled forward for a large order, or set down "temporarily". The physical world changed and the system was never told. This is the category that responds to process discipline and to nothing else.
- Timing gaps between counting and posting
- A count taken at 06:00 and posted at 14:00 describes a location that has since been picked from twice. The resulting adjustment is not a correction, it is a new error — and it is indistinguishable from a real one in the variance report.
- Status and allocation drift
- Quantity correct, status wrong: damaged stock still flagged available, quarantined stock allocatable, allocations that never released after a cancelled order. Piece accuracy scores these as perfect. Picking does not.
How to audit your own accuracy measure
Before trying to improve accuracy, establish that the number you have describes something. This takes a day and frequently changes the priorities.
Write down the current formula
Exactly as computed, including the sampling frame, the tolerance, the treatment of recounts and whether variances are netted. If this cannot be reconstructed from documentation, reconstruct it from the report's own SQL or export logic. A formula nobody can state cannot be defended and should not be trended.
Recompute all three measures from the raw count file
Net pieces, absolute pieces, locations. Use the raw count detail, not the summary report — the summary has already applied the decisions you are trying to examine. The spread between the three is the size of the reporting problem.
Verify the count was blind
Check the configuration flag, and then check behaviour: the distribution of counted quantities under a blind count has a visible tail of wrong answers. A distribution where almost every count exactly equals system quantity, with a handful of large exceptions, is the fingerprint of confirmation rather than counting.
Examine the sampling frame
List every location type in the building and mark which are eligible for counting. Reserve, bulk, floor stack, mixed, damaged, quarantine, returns bench, staging. Anything ineligible is unmeasured, and unmeasured locations are where error accumulates precisely because nothing checks them.
Recompute with tolerances removed
Run the same file with zero tolerance. The difference between that and the published figure is the quantity of real variance the tolerance is currently forgiving. If it is large, the tolerance is a reporting decision rather than a measurement-noise allowance.
Join to the short-pick log
The strongest validity test available. Locations that shorted during the period should be over-represented in the variance set. If they are not, the count programme is not sampling where the errors are — and the accuracy figure, however carefully computed, is describing a different building than the one picking is failing in.
Segment the variance
By item class, vendor, zone, location type, UOM, receiver, shift and counter. A single building-wide accuracy figure is an average of several different problems. Segmentation is what turns it back into findable causes, and the segments almost always turn out to be very unequal.
Trend gross variance and the decay slope
Plot absolute variance over time, and plot how fast location accuracy falls between counts. A flat accuracy line with a steepening decay slope means the count programme is running harder to stand still — which is a process finding, not a counting finding.
The data and fields to inspect
Accuracy work needs the count detail, not the count summary. These are the fields that make recomputation possible.
- Cycle count header — count ID, count type (ABC, location, random, directed), scheduled date, count date, post date, counter ID, supervisor, sampling rule applied.
- Cycle count detail — location, item, LPN, system quantity at the moment of counting, counted quantity, variance, UOM, blind-count flag, recount sequence, tolerance applied, accepted or rejected.
- Tolerance configuration — by item class, value band, UOM or location type, and whether tolerance is expressed in pieces or as a percentage.
- Location master — location type, zone, pick or reserve, capacity, item-mixing permitted, LPN enforcement, count eligibility flag, count frequency class.
- Item master — base UOM, conversions, case pack, ABC class, standard cost, lot and serial control.
- Location inventory snapshot — taken at the same instant as the count extract, so the comparison is against the balance that actually existed at count time rather than at report time.
- Adjustment history — quantity, reason code, user, timestamp, and the link back to the count that generated it, so count-driven and non-count adjustments can be separated.
- Short-pick and exception log — location, item, expected, found, picker, timestamp, resolution. This is the ground truth against which any accuracy figure should be validated.
- Transaction history for counted locations — covering the window between count time and post time, so manufactured variances can be identified and excluded.
Record accuracy vs. financial accuracy
These are two different questions that share a word, and the example above is the cleanest demonstration of the gap between them.
Financial accuracy asks whether the total valuation on the balance sheet is right. It is a netted, aggregate, value-weighted question, and 99.97% is a perfectly good answer to it. Finance is not wrong to compute it that way — for their purpose, an over in aisle four genuinely does offset an under in aisle nine.
Record accuracy asks whether an individual record can be relied upon to direct physical work. It is binary, per location, and does not net, because a picker standing in aisle nine cannot be paid by a surplus in aisle four. 96.4% is the answer to that question, from the identical count.
Most disagreements about inventory accuracy between operations and finance are this, and neither side is being careless — and reconciling the two is a defined workflow rather than an argument. They are computing different measures, both correctly, and reporting them under one word. The fix is not to pick a winner — both numbers are needed — it is to publish both, labelled, so the 3.57-point gap becomes a visible and discussable quantity instead of a source of mutual suspicion.
One practical consequence: an accuracy improvement programme measured in net pieces can show progress while location accuracy is flat or falling, because netting rewards offsetting errors. If a programme's measure can improve without any location becoming more reliable, it is measuring the wrong thing. For the method that turns a segmented variance population into named causes, see inventory root cause analysis.
The operational takeaway
Publish the unit with the number, every time. "98.18% absolute piece accuracy, 96.40% location accuracy, 500 locations, blind, zero tolerance" is a measurement. "99.97%" on its own is a claim.
Trend gross variance rather than net, because netting deletes exactly the equal-and-opposite pairs that point at findable transaction defects. Validate against the short-pick log, because that is the only independent evidence you have about whether the records are usable. And treat the decay slope between counts as the real health metric — a building that returns to 96% one week after every count is not a building with a counting problem.
When the measure is sound and the number is still poor, the work moves to causes: tracing individual discrepancies, ruling out phantom stock, and auditing the WMS data as a whole.
Related WMSAudit guides
- Cycle count discrepancies Five tests that separate counting noise from a mechanism, and what each attribute pattern means.
- Inventory discrepancies The parent method: replay an item's transaction history until the balance stops reconciling.
- Inventory adjustment analysis Churn ratio, reason-code decomposition, and the five defect classes an adjustment log exposes.
- The WMS audit What a warehouse data audit examines, the exports it needs, and what it honestly cannot conclude.
- The WMSAudit engagement Scope, method, fixed fee and what the report contains.