Work measurement decomposes labor into timed elements and reassembles them into standard hours through performance ratings and allowances. This article dissects time study, work sampling, sample-size formulas, and activity diagrams, exposes the method’s s
Work measurement converts labor from an assumed quantity into a calibrated one. At its core, the work measurement method decomposes a job into repeated, observable elements, times those elements, adjusts the timings for the skill of the observed worker and the interruptions real work inevitably absorbs, and reassembles everything into a standard time — the defensible number behind staffing plans, cost standards, and pay structures. My concern in this article is less with the arithmetic, which is genuinely simple, than with the discipline surrounding it. Every credible capacity figure I have ever produced — whether for a labor pool or a land parcel — survives scrutiny only when the upstream parameter choices are explicit. Work measurement is no different. The method earns its authority through decomposition, transparent adjustment, and honest treatment of error; it forfeits that authority the moment performance ratings go unexamined or allowances get negotiated in silence. What follows dissects the method’s four instruments, traces the fault lines where it cracks, borrows a reading from capacity analysis in an adjacent field, and closes with a deployment sequence that keeps the numbers honest.
The work measurement method operates mainly inside the performance evaluation process, and its native habitat is processing and manufacturing work, where tasks repeat and observation is physically possible. Its function is to supply accurate, workable data and patterns for that evaluation — not impressions, not managerial intuition, but numbers a skeptic can trace back to a stopwatch and a sample.mbalib.com
The method’s first move is disassembly. A job, viewed whole, resists measurement; the same job broken into a work cycle — the sequence of elements a worker repeats — yields to it. Consider the example the source literature uses: a server at a cafeteria salad station. One complete work cycle includes fetching the plate, filling it, adding dressing, and delivering it to the diner. Four elements, each separately timed, each recorded on an observation sheet across multiple cycles.mbalib.com
The cycle logic has a statistical footing worth making explicit. Element times scatter from cycle to cycle — a plate slips, a dressing dispenser runs slow, a diner asks a question. Averaging across many cycles lets that scatter cancel out, and the amount of scatter itself tells the analyst how many cycles are needed: tight, consistent element times converge quickly, while volatile ones demand extended observation. The observation sheet also carries space for external elements — incidental occurrences that fall outside the cycle proper — and these records feed the later choice of allowance, since an allowance is supposed to reflect real interruption, not imagination.mbalib.com
The second move is adjustment, and this is where the method’s character shows. The average of observed element times describes one worker on one day — not a standard. Two parameters translate observation into standard. The first is the performance rating, a decimal expressing the observed worker’s skill relative to a “normal” worker. An analyst judging a worker to operate at 90 percent of normal pace multiplies observed time by 0.9; a worker judged at 110 percent gets a multiplier of 1.10. The product is the normal element time. The second parameter is the allowance, the share of working time conceded to rest, personal needs, and unavoidable interruption — a 10 percent allowance amounts to 6 minutes per hour. Dividing normal time by (1 − A), where A is the allowance expressed as a decimal, produces the standard time for each element; summing the element standard times yields the standard cycle time.mbalib.com
The formulas are objective in form and subjective in substance. The source itself flags the problem: performance ratings, being judgments, invite dispute.mbalib.com Two analysts watching the same worker can defensibly disagree, and the arithmetic downstream will faithfully carry either opinion to three decimal places. I have spent 16 years building capacity estimates, and the pattern is constant — precision at the end of a calculation cannot compensate for judgment at its beginning. A standard time is only as honest as the rating and allowance choices embedded upstream, which is why those choices deserve documentation as careful as the timing itself.
The toolkit answers four distinct questions: how long a task takes (time study), how time is distributed across activities (work sampling), how many observations a defensible estimate requires (sample size), and how workers and customers interact in time (activity diagrams). Each instrument has a proper use, and applying one where another belongs is among the most common failures in practice.
Time study exists to develop performance standards, and those standards propagate through the organization: they inform how many employees are needed, how tasks are assigned, what standard costs should be, how employee performance is judged, and how work-based pay plans are structured.mbalib.com The mechanics proceed in order. Identify the repetitive work cycle. Break it into discrete elements. Time each element with a stopwatch across enough cycles that variation stabilizes. Record everything on an observation sheet, including external and incidental occurrences — events outside the cycle proper that nonetheless inform the choice of allowances later. Sum the element times across cycles, divide by cycle count, and the average observed time per element emerges. From there, the rating and allowance adjustments described above convert observation into standard. The cafeteria example is worth dwelling on because it is unglamorous on purpose: plate, salad, dressing, delivery. If a method cannot handle a salad station cleanly, it has no business near an assembly line.mbalib.com
Time study answers how long; work sampling answers how much. The technique does not measure the duration of an activity but the proportion of time a worker spends across activities — and that difference defines its proper territory. Work sampling is most useful for designing and redesigning jobs where no prior procedure exists, particularly direct customer-contact activities that resist the pre-set structure time study requires.mbalib.com
Seven steps structure the method. Define activities in as few non-overlapping categories as possible. Design an observation form that is easy to apply and easy to analyze later. Determine the study’s length — long enough to constitute a random sample of behavior. Test the form to confirm the categories are well defined and the instrument workable. Determine the sample size and observation pattern using ordinary statistical methods. Conduct the study with trained observers. Analyze the data, which in its simplest form means counting observations per category and computing the percentage of time each represents.mbalib.com Proportion data suits design questions because redesign is fundamentally about reallocation: if 30 percent of a shift disappears into an activity nobody planned for, that is a design finding no stopwatch average would surface.
Two design decisions inside those steps deserve emphasis because they manage human behavior rather than statistics. The method recommends either telling observed workers they are not being evaluated — so they do not perform for the observer — or observing without their knowledge; and it recommends discarding the first day or two of data so workers can settle into routine.mbalib.com Both provisions acknowledge the same truth: measurement changes the measured.
How many observations suffice? The formula is N = Z²P(1 − P)/E², where N is the required sample size, Z is the standard normal deviate for the desired confidence level, P is the assumed proportion of time spent on the activity of interest expressed as a decimal, and E is the maximum allowable error, also as a decimal.mbalib.com
The parameters translate plainly. A manager wanting 95 percent confidence (Z ≈ 1.96) that the estimated proportion lies within ±5 percentage points of the truth (E = 0.05), assuming P = 0.5, needs roughly 385 observations. Larger samples buy accuracy and representativeness, and there is no shortcut around the arithmetic.mbalib.com
The awkward parameter is P, since it is exactly what the study is trying to discover. Three sources resolve the circularity: a small pilot study, a deliberately conservative assumption of 0.5 — which maximizes the required sample size and therefore guarantees adequacy — or a reasonable estimate drawn from experience.mbalib.com The conservative choice costs observation effort; the optimistic choice costs credibility. I know which trade I would take.
When the question involves how workers and customers occupy time together, charts replace stopwatches. Two graphical tools dominate. The worker–customer chart applies when the worker’s cycle time runs shorter than the customer’s service requirement, mapping the interaction on a shared time scale so that a single employee can serve more than one customer simultaneously.mbalib.com
The source’s bank example illustrates the logic with numbers worth keeping. A drive-through teller serves two lanes. A customer takes 15 seconds to enter the service area, 48 seconds to be served, and 9 seconds to leave — a customer cycle of 72 seconds. The teller’s own cycle, however, is only 48 seconds. The 24-second gap is structural slack: while the first customer occupies the exit phase, the second is already staged in the other lane. One teller, two lanes, no idle time — the chart makes the capacity visible.mbalib.com
When the situation involves multiple workers and multiple customers simultaneously, the simpler chart buckles, and the more elaborate activity chart takes over, mapping parallel interactions that a single time scale cannot hold.mbalib.com
A measurement method earns trust not by its successes but by the clarity with which it admits its failure modes. Work measurement has several, and most trace to upstream human choices rather than downstream arithmetic.
The first defect is rating subjectivity, and the source concedes it directly: performance ratings, being subjective, can cause disputes.mbalib.com The dispute is not incidental — a rating of 0.90 versus 1.10 swings normal time by more than 20 percent, and standard time inherits the swing. The remedies are procedural: multiple independent raters, calibration sessions where raters converge before the study begins, and documentation of the rating rationale alongside the timing sheets.
The second defect is the observer effect. People under observation behave differently, which contaminates the very quantity being measured. The method’s built-in responses — reassuring subjects they are not being evaluated, observing unobtrusively, discarding the first day or two of data — are honest admissions that the instrument disturbs the system.mbalib.com What they cannot fully remove is the deeper problem: a worker who knows a study is running, even without knowing its purpose, may pace differently for weeks.
The third defect is domain mismatch. Time study presupposes repetition. Knowledge work, exception handling, and variable-demand service resist cycle decomposition because they lack stable cycles to decompose. Work sampling partially covers this territory by measuring proportions rather than durations, but a proportion is not a standard — it tells you how time was spent, not how long a task should take. Applying time study to non-repetitive work produces a number with the form of a standard and the content of a guess.
The fourth defect is allowance politics. The allowance parameter A enters the formula ST = NT/(1 − A) with full arithmetic authority and no scientific protection. A 10 percent allowance lifts standard time by roughly 11 percent; a 15 percent allowance lifts it by roughly 18 percent. The difference between those two standard times is not measurement — it is negotiation, and it should be labeled as such. Allowances grounded in fatigue studies and ergonomics research deserve deference; allowances grounded in bargaining position deserve scrutiny.
The fifth defect is sampling bias. Observations taken on Tuesday mornings describe Tuesday mornings. Day-of-week, seasonal, and demand-cycle variation all distort proportions if the observation schedule ignores them. The sample size formula guarantees precision for the conditions sampled, not representativeness for the conditions that matter.
A governance note ties the defects together. Every parameter that invites dispute — rating, allowance, observation schedule — should carry a named owner and a documented basis before the study runs, not after the number lands. Retrofitting justification onto a disputed standard time is advocacy, not measurement, and everyone in the room recognizes the difference. Run root cause tracing down any of these defects and it terminates in the same place: parameter choice. Which worker, observed when, rated by whom, under what allowance, across which hours — the arithmetic amplifies these decisions rather than correcting them. A standard time is a conclusion; the parameters are the argument.
I work on urban systems, where the closest cousin of standard time is carrying capacity — the load a system sustains under stated conditions. The parallel is worth one section because it clarifies what a standard time is and is not.
A carrying capacity figure is an engineering ceiling, not an aspiration; nobody treats it as a challenge. Standard time deserves the same reading. It is the output a normal worker sustains at normal pace with normal interruptions — a rated condition, not a stretch target. Organizations that treat standard time as a floor to be beaten have misunderstood the instrument and usually end up gaming it.
The allowance parallels a reserve margin. No system, mechanical or human, runs at theoretical maximum indefinitely; the allowance is the deliberate inefficiency that keeps operation sustainable. Stripping allowances to make numbers look better works exactly as well in labor as it does in infrastructure — briefly.
The bank teller example carries a further lesson from capacity analysis. The 24-second gap between the customer cycle (72 seconds) and the teller cycle (48 seconds) is redundancy that absorbs variability. Parallel channels with slack are what keep systems stable under demand spikes; the same logic that lets one teller cover two lanes lets a road network survive an accident. Removing the slack raises average utilization and fragility together.
Work sampling, for its part, mirrors land-use surveying: nobody censuses every parcel, so surveyors sample and report proportions with stated confidence. The shared discipline is the error bar. A proportion without a confidence statement is a claim; with one, it is a measurement. The analogy ends there — labor is not land — but the epistemic standard transfers completely.
A disciplined deployment follows a sequence, and the sequence matters more than any single step.
First, define the decision the data must serve — staffing levels, standard costs, performance evaluation, or pay design — because each decision demands different precision and different instruments. Second, match the instrument to the question: time study for standards on repetitive work, work sampling for time allocation and job redesign, activity diagrams for interaction and capacity design. Third, run a pilot to estimate the proportion P before committing to a full sample; guessing wrong in either direction wastes either observations or credibility. Fourth, compute the sample size from the formula and schedule observations across the representative conditions — days, shifts, demand periods — rather than the convenient ones. Fifth, train observers and manage the observer effect deliberately, choosing between informed reassurance and unobtrusive observation and discarding warm-up data in either case. Sixth, apply the rating and allowance adjustments transparently, with the rating rationale and allowance basis documented beside the timings. Seventh, report the result with its conditions attached: which workers, which periods, which confidence level.
The boundary conditions deserve equal statement. The method assumes repetitive, observable, stable work. Where work is non-repetitive, where activity resists observation, or where the process itself is changing, the outputs degrade from measurement toward opinion — and an opinion formatted as a number is worse than a plain one, because it borrows credibility it did not earn.
The work measurement method is best understood as a contract between arithmetic and honesty. The arithmetic is genuinely simple: decompose the cycle, time the elements, adjust for pace and interruption, sample enough observations, chart the interactions. The honesty is harder — rating the worker openly, setting allowances on evidence, admitting the observer effect, reporting the error bar, and labeling negotiation as negotiation. Standards built this way support staffing plans, cost systems, evaluation, and pay with numbers that survive skeptical review. Standards built any other way produce the illusion of precision, which in 16 years of capacity work I have found to be the most expensive illusion an organization can buy. The stopwatch is cheap; the discipline is the price. Anyone who tells you otherwise is selling something, and the number they hand you will not survive the first serious audit.
Reference Block:
Source Reference Link: https://wiki.mbalib.com/wiki/工作测量法
Link Brief: An MBA Lib Encyclopedia entry explaining the work measurement method: time study for standard hours, work sampling for time allocation, the sample-size formula, and worker–customer and activity charts. This article adopts it as the primary methodological source for the four instruments, formulas, and worked examples.
Measurement is a craft best deepened one careful study at a time — keep pulling at the parameters, and the discipline will keep rewarding you.
Content Disclaimer: This article is for general reference only and does not constitute professional R&D guidance, production process advice or quality certification. All material performance data has specific test premises; readers should verify parameters against actual equipment and working conditions.

