Daloy Operations Score methodology

Most operating assessments are a scoring rubric someone made up over a weekend, wrapped in a nice chart. Here is ours, in full, including the parts we are not confident about yet.

What the number is

Your Operations Score runs from 20 to 100. It measures one thing: how close your business is to running on a system instead of on the people who remember how things work.

It is not a measure of how good your business is. A superb company with a heroic owner will score low, and that is exactly the point. The score tells you how much of your operation would still work on a Monday when the right person is not there.

The scale starts at 20 rather than 0 because you are running a real business that serves customers and makes payroll. Twenty means nothing systematic is in place. It does not mean nothing is there.

What it is, and what it is not

This is a structured diagnostic. It is not an industry benchmark, and it is not a statistically validated assessment. We will say that plainly for as long as it stays true, and we will say so here when it changes.

The five dimensions

Each one counts for 20 percent. We weight them equally because we have no evidence that would justify weighting them any other way, and we would rather tell you that than invent a hierarchy.

DimensionWhat we look at
How work flowsHow long a job takes end to end, how much is open at once, how often work gets redone, and what waits on one person
How standardized it isWhether the way work gets done is written down, current, actually followed, and owned by someone
How much leverage you get from systemsWhether there is one source of truth, how much gets re-typed by hand, how much completes untouched, and how your weekly numbers get made
How well change sticksWhether the last thing you changed is still happening, and whether there is a rhythm holding it in place
How you bring people inWhether hiring is structured, whether onboarding exists, how long people take to become productive, and who could cover whom

Four questions per dimension. Twenty in total.

The maturity levels

Every question is scored 0 to 4 against a written description of what each level looks like in a business. We do not ask you to rate yourself. We ask what you do, and we match your answer to a description.

LevelWhat it means
0Fully dependent on individual memory or intervention
1Informal and inconsistent
2Repeatable but not reliably followed
3Standardized, owned, and supported by systems
4Measured, resilient, and able to improve without founder intervention

That last detail matters more than it sounds. Anchoring answers to described behavior, rather than to how someone feels about their business, is the method the World Management Survey has used to score more than 13,000 companies across 35 countries. It is the reason their scores hold up.

We do not just take your word for it

Every question carries an evidence rating.

RatingWhat it meansEffect
ObservedWe saw it. The document, the screen, the exportNo cap. Can score the full 4
AttestedYou described it with a specific example, but we did not see itCapped at 3
AssertedYou told us it is true. No exampleCapped at 2

So you cannot score full marks on documentation by saying you have documentation. Someone has to open the folder.

This is deliberate, and it means your score is capped by what we could verify rather than by what is true. If you think a question is underscored, the fix is to show us. We will re-score it.

Every score ships with a confidence rating, which is simply the share of the twenty questions we were able to verify directly. A score without its confidence rating is marketing. We would rather give you the measurement.

Why we do not just average the five

We combine your five dimension scores using a geometric mean rather than a straight average. In plain terms: your weakest dimension pulls the number down harder than an average would let it.

A business scoring 80, 80, 80, 80 and 20 gets a 71 on a straight average. It gets a 65 from us.

That six point difference is the whole philosophy. Your business is a chain, not a portfolio. Beautiful documentation does not rescue you if every job waits three days on one person's approval. Output is governed by the constraint, so the score is too.

This is a documented and recommended approach for exactly this situation. The OECD and European Commission handbook on building composite indicators names the problem directly: a straight average lets a surplus in one area quietly offset a deficit in another, which is fine for a portfolio and wrong for a system.

The bands

ScoreBandWhat it means
20 to 39It runs on youThe business works because specific people remember how. Remove them and it stops
40 to 54It runs on habitsIt is consistent, but only because the same people do it the same way. Nothing is written
55 to 69It runs on a standardThe way work gets done is written, owned, and mostly followed
70 to 84It runs on numbersThe standard is measured, and the measurements change what you do
85 to 100It runs itself and gets betterImprovements land, stick, and compound without being pushed

A business can start in either of the first two bands when its operation still depends heavily on individual memory and intervention. The number is a starting position rather than a grade.

The bands are adapted from the CMMI maturity levels, the standard used for process maturity assessment for the last thirty years, translated out of engineering language.

What each dimension rests on

We are not going to claim we invented a science of operations. We picked five things other people have already shown matter, and built a way to measure them consistently.

How work flows rests on Little's Law, proven in 1961, which states that the amount of work you have open equals how fast you finish work multiplied by how long each job takes. It is a mathematical relationship, not a trend. If you have a lot open relative to what you finish, your jobs are slow, and no amount of effort changes that arithmetic. We often do this division live in the room.

How standardized it is carries the strongest evidence of the five, and it is usually the constraint. Researchers at Stanford and the LSE scored management practices at over 13,000 firms and found the scores predict productivity, profitability, growth and survival. Then they ran an actual randomized trial: 28 plants, half of them given five months of help fixing management practices. Productivity rose 11 percent, quality defects fell 32 percent, and profit rose by roughly 228,000 dollars per firm per year. This is not a consulting opinion. It is a controlled experiment.

How much leverage you get from systems uses published benchmarks rather than vendor claims. APQC's benchmarking data puts the median organization at 55 percent of invoices going out with no manual intervention, and top performers at 94 percent of orders processed without a human touching them at any point. Those are the numbers your touchless rate is scored against.

How well change sticks uses Prosci's benchmarking research, which found that 88 percent of projects with strong change management met their objectives against 13 percent of projects without it. We will note that this is Prosci's own practitioner research rather than independent peer review.

How you bring people in rests on a 2022 re-analysis of the personnel selection literature that corrected a statistical error running back to 1998. The corrected result: structured interviews, where every candidate gets the same questions scored against written criteria, are now the best validated predictor of job performance available, ahead of aptitude testing. Unstructured conversational interviews score less than half as well. Structure is not bureaucracy. It is the difference between predicting and guessing.

What we are not claiming

We would rather tell you this than have you find it.

We are not claiming your score predicts your revenue. It measures operating maturity. We have not run the study that would let us say more, and when we have, we will publish it.

We do not have industry benchmarks yet. When we have enough scored businesses to publish a distribution honestly, we will. Anyone showing you an industry benchmark today should be asked where it came from.

We are not using the statistic that 70 percent of change initiatives fail. It gets quoted constantly. It traces back to an author's self-described unscientific estimate from 1993, which he later disowned. There is no study underneath it.

We are not treating ISO certification as evidence of operational quality. The research on whether it improves performance is genuinely mixed. We score what we observe you doing.

Our scorers are human. Written descriptions reduce subjectivity, they do not remove it. Every third assessment is independently scored by a second person as a check.

What counts as improvement

Any score that gets re-run needs a rule for when a change is real and when it is measurement noise. Without one, the number just drifts upward and stops meaning anything.

Our current rule: a move of fewer than 11 points on the overall score, or fewer than 14 points on a single dimension, is reported as no measurable change. That threshold comes from a standard reliability calculation, and right now it uses an assumed reliability figure rather than a measured one. Once we have thirty assessments scored independently by two people, we will replace the assumption with the real number and update this page.

One thing worth knowing before we start. The overall score moves slowly, and it is supposed to. A single project that genuinely fixes one dimension will move the overall number by three to four points, which is inside our own noise floor. The dimension score will move by twenty, which is not.

So we will not tell you your Operations Score jumped after one project, because that would not be true. We will show you the dimension that moved, and we will re-score the whole thing once a year.

When we score it

WhenWhat you get
The Flow Check on our siteAn estimate. Five questions. Everything is capped, because we cannot verify anything through a form
The Daloy Operations ReviewYour real baseline. This is the one that counts
After a projectWe re-score the dimensions the work touched
Once a yearFull re-score, including the overall number

Questions about any of this are welcome, including the awkward ones. If you find something here we have got wrong, we would like to know, and we will correct it.