A Number With a Job It Was Given Rather Than Earned

The net promoter score was introduced with a specific empirical claim attached, which is that it predicted growth better than the satisfaction measures it was meant to replace (Reichheld, 2003). That claim is what moved it into thousands of companies, and it is the only thing that would justify the position it now holds on a quarterly dashboard. A score that does not forecast anything is a mood reading, useful for conversation and not for allocation, and every organization tracking one is implicitly asserting that it forecasts something.

Twenty years of testing have produced an unflattering answer, and the shape of the testing matters more than the volume of it. Almost all of the work asked whether the score predicts a stated intention, a self reported behavior, or a firm level financial figure. Very little of it asked whether the score predicts what a customer actually did, measured at the turnstile rather than in a follow up survey. Two studies have done that, in different sectors, and neither of them helped the metric. One of them also found something that has not appeared anywhere else and that lands directly on industrial distribution, which is that the score’s predictive power falls away as the customer’s experience with the supplier accumulates.

That finding is the subject of this article. Everything a manufacturer sells through distribution is sold to experienced repeat buyers, which means the population a brand cares most about is the population the score was measured to be worst at reading.

What Most of the Tests Actually Measured

The critical literature on the metric is large and it is worth being precise about what it examined, because a manufacturer defending its dashboard will reasonably ask whether the criticism has ever been tested rather than asserted.

Morgan and Rego (2006) linked American Customer Satisfaction Index survey data to Compustat and CRSP financial records across a 1994 to 2000 panel and tested 6 metrics against business performance: average satisfaction, top 2 box satisfaction, repurchase likelihood, the proportion of customers complaining, recommendation intentions of the kind the net promoter score is built from, and average recommendation behavior. Average satisfaction carried the greatest predictive value, repurchase likelihood carried some depending on the performance dimension being predicted, and the recommendation based metrics showed little or none. Their conclusion is stated without hedging, that recent prescriptions to focus customer feedback systems and metrics solely on customers’ recommendation intentions and behaviors are misguided.

Keiningham, Cooil, Andreassen, and Aksoy (2007a) ran the replication that should have settled the original claim, using Norwegian Customer Satisfaction Barometer data across 21 firms and more than 15,500 interviews, in the same industries the original work had cited as exemplars, and failed to reproduce any clear superiority of the recommendation measure over established loyalty measures. In the same year, Keiningham, Cooil, Aksoy, Andreassen, and Weiner (2007b) approached it from the customer level rather than the firm level, following more than 8,000 customers across retail banking, mass merchant retail, and internet service providers for 2 years, and found that no single metric performed best across the 3 outcomes they examined. Repurchase intention beat recommendation intention at predicting retention, recommendation intention was better at predicting recommendation, and repurchase intention and brand preference beat it again on share of wallet. The result from that paper worth carrying forward is not about which single item won. It is that multivariate models, using several metrics together, outperformed the best single metric models by 20 to 25 percent in adjusted R squared in banking and retail.

Van Doorn, Leeflang, and Tijs (2013) replicated the question again on 11,967 customers across 46 companies in banking, insurance, utilities, and telecommunications, comparing the metrics against current gross margin, current sales revenue growth, and future sales growth. They report that all the metrics performed equally well on current performance and equally poorly on future sales growth, with no significant differences among them, which leaves the recommendation measure neither superior nor inferior to the others and leaves all of them weak at the thing a forecast is for.

Read together, that body of work says the metric is not distinctively predictive. What it does not say, because none of these designs could, is how the metric performs against behavior that was recorded rather than reported. Every study above measured either an intention, a self report, or a company’s financial statements.

The general reason to expect trouble at exactly that step was established outside marketing altogether. Sheeran (2002), reviewing 10 meta analyses covering 422 studies, computed a sample weighted average correlation between stated intentions and later behavior of .53, which accounts for roughly 28 percent of the variance in what people subsequently do. Intentions are real and they are informative and they leave about three quarters of behavior unexplained. A metric built entirely from a stated intention inherits that ceiling before any of its own problems are considered.

Two Studies That Used Behavior

Tests against recorded behavior are rare rather than absent, and I have found 2. The scarcity is worth a moment on its own, because a literature that has spent 2 decades arguing about whether a metric predicts growth, while almost never checking it against a record of what customers did, has been conducting the argument one step removed from the question. Both of the exceptions reach the same conclusion from different data, and neither was designed to embarrass the metric.

Schmitt, Meyer, and Skiera (2012) worked with a financial services provider’s actual transaction records, computing customer lifetime value from the account data rather than from what customers said they would do, and tested intention to recommend against it. Their finding is narrow and awkward for the metric: intention to recommend was associated with higher contribution margin, and it did not enhance actual customer lifetime value or actual loyalty. The willingness to recommend and the economic value of the relationship came apart when the value was measured from the ledger.

Matzler, Strobl, Teichmann, and Aigner (2026) built the study this article is mostly about, and its design is the reason it is worth a manufacturer’s attention. They assembled 136 destination year observations from 38 European Alpine ski resorts, combining bi annual customer satisfaction survey data from 2011/12 through 2017/18, more than 120,000 responses in total, with archival access records from the resorts themselves. The outcome variable is not an intention to return and it is not a revenue line. It is how many times people actually came through the lift gates, recorded electronically, at the destination level, over 7 seasons. They analyzed it with panel regression, fixed effects, robust standard errors.

A ski resort is not an industrial distributor, and the transfer has to be argued rather than assumed. What makes the study relevant is not the sector but the design, since it is the closest thing available to a test of whether a recommendation score predicts repeat purchase in a setting where repeat purchase was counted rather than described.

Four Percent

The headline result is that the net promoter score does significantly predict skier visits, so the metric is not empty, and it is not superior to plain customer satisfaction. Comparing the 2 coefficients directly, the authors report an F test of F(1,37) = 4.55, significant at the 5 percent level, with satisfaction carrying a significantly stronger positive influence on visits than the recommendation score. On its own, by their adjusted R squared comparison, the net promoter score accounts for about 4 percent of the variance in how many visits a destination received.

Four percent is the number to carry out of this article, and it deserves to be read carefully rather than triumphantly. It does not mean the score is noise, since 4 percent of the variance in a real behavioral outcome across 38 resorts and 7 seasons is a genuine signal and most marketing metrics would be pleased to have it. What it means is that a quantity explaining 4 percent of an outcome cannot bear the weight of being the single number an organization manages by, and that the metric it was built to replace did better on the same data.

The Part That Matters in This Industry

The finding with the most direct application to industrial distribution is the one that has attracted the least attention, and it concerns how the score behaves as respondents accumulate experience with the thing they are rating.

Matzler and colleagues tested the interaction between the recommendation score and the destination experience of the sampled respondents, measured as prior visits, and report in their abstract that destination experience attenuates the predictive ability of the score. The consequence is stated plainly in the paper: if the sample exceeds destination experience levels of 2.3, the net promoter score does not significantly predict destination skier visits anymore. Below that level of accumulated experience the score retained its predictive power, and above it the relationship stopped being distinguishable from zero.

Two qualifications belong in the same paragraph as that result. The interaction is significant at the 10 percent level rather than the 5 percent level, which is a weaker threshold than the rest of the paper’s findings rest on and should be reported as such. And the result is close to unreplicated, since I have not located another study testing whether customer experience or tenure moderates the relationship between a recommendation score and subsequent recorded behavior. A single interaction at p below .10 in one sector is a lead worth following rather than a settled fact, and the reason to take it seriously is not its statistical strength but the mechanism it implies and the population it implicates.

Consider which customers an industrial manufacturer actually has. A plant that bought 60 drives from a brand across 12 years, a distributor that has carried a line since before the current channel manager arrived, a motor shop that has specified the same bearing for a decade, a maintenance group that has run the product long enough to know which revision was the bad one. There is almost no one in an industrial channel whose experience with a supplier consists of 2 transactions. The entire customer base sits above the threshold at which, in the one study that measured it against recorded behavior, the score stopped predicting anything.

A brand tracking its recommendation score across that population is not measuring a weak signal. It is measuring a signal that was observed to disappear in precisely the stratum it is being collected from, which is a different and worse problem than imprecision. An imprecise instrument reports the right direction with error around it. An instrument whose relationship to the outcome has gone to zero reports movement that is unrelated to the thing being managed, and it reports that movement with the same confident 2 decimal places it used when the relationship was intact.

Why Experience Would Do That

The finding is an interaction in a regression, and an interaction is not a mechanism, so the explanation has to be argued from something other than the coefficient. What follows is reasoning rather than measurement, offered because a result with no plausible mechanism behind it should not change anyone’s practice, and because the mechanism, if it is the right one, predicts where else the same failure would appear.

A customer with little experience of a supplier answers a recommendation question with a forecast assembled from very little. Reputation, the brand’s standing in the trade, the confidence of the last salesperson, and whatever the buyer’s peers have said are most of what is available, and the answer they give is largely a report on those inputs. That kind of answer will track future behavior reasonably well for exactly as long as it is the same set of inputs driving the behavior.

An experienced customer is in a different position. They have accumulated episodes: the quotation that took 6 days, the warranty argument in year 4, the applications engineer who answered on a Sunday, the configurator login that expired. Those episodes are what determine whether the product gets specified again, and a single question asking how likely they are to recommend the brand compresses all of it into one number, in which a strongly positive judgment about the product can fully absorb a deeply negative judgment about the experience of transacting. The compression is the problem, and it gets worse as the material being compressed accumulates.

There is a second mechanism, and it comes out of the part of a supplier relationship the customer cannot verify. Where a quality is a credence quality in the sense Darby and Karni (1973) established, meaning the buyer cannot confirm it even after years of use, the belief about it is free to move independently of the thing itself, and it moves on reputation. My own work in this channel found that operating at the level of the brand, where a manufacturer holds a position of meaning in the trade regardless of the accuracy or the age of that position, and where the proposition that nobody was ever fired for buying the dominant brand functions as a live constraint (Tolbert, 2022). A recommendation question asked of an experienced customer collects a mixture of accumulated episodes and inherited reputation, weighted in a way no one can recover from the number, and the reputation component is the part least connected to what the customer will do next.

The Lag Problem

There is a further result in the same paper, and it carries a practical implication sharp enough to be worth stating even where the paper is not perfectly consistent about the detail. The inconsistency is reported here rather than resolved quietly, since a reader who goes to the article will find it, and a piece that smoothed it over would deserve less trust than one that did not.

The authors examined lagged specifications, testing the score measured 1 and 2 years before the visits it was predicting. Their abstract states that the score with a 2 year lag demonstrates the strongest explanatory power. The body of the paper reports an F test comparing the 1 year and 2 year lagged coefficients, F(1,33) = 16.52, significant at the 1 percent level, in a direction that reads as favoring the 1 year lag, so the ranking between those 2 specifications is left open here rather than asserted.

Whichever lag wins, the implication is the same and it is uncomfortable for the way the metric is used. If the score’s relationship to behavior is strongest at a lag of a year or more, then the score read this quarter is not telling an organization about this quarter, and the quarter over quarter movement that dashboards are built to display is the part of the series carrying the least information. An instrument whose signal arrives a year late is a poor instrument for a monthly review and a reasonable one for an annual strategy cycle, and almost every implementation has it backwards.

The Strongest Case for Keeping It

An argument this one sided invites a reader to assume the other side has not been heard, so it is worth putting the defense of the metric as well as its defenders would put it. The defense is not weak, and a manufacturer who has run a recommendation program for a decade has usually arrived at some version of it independently.

The score has 3 properties that the instruments proposed to replace it routinely sacrifice. It is cheap, which means it can be administered often and to everyone rather than to a sample somebody had to design. It is comparable, across business units, across countries, and above all across time, and a company that has collected the same item the same way for 8 years owns a baseline that no newly designed instrument can produce retroactively. And it is legible to people who will never read a methods section, which is not a trivial property in an organization where the audience for the number includes a board.

The empirical defense is also stronger than the critical literature’s tone suggests. Matzler and colleagues (2026) did not find that the score fails to predict behavior. They found that it predicts behavior significantly and that satisfaction predicts it better, which is a finding about non superiority rather than about emptiness. Four percent of the variance in a recorded behavioral outcome, across 38 organizations and 7 seasons, is a real relationship, and a good many of the quantities on an executive dashboard have never been tested against a recorded outcome at all. A manufacturer defending its program can fairly point out that the metric has been examined more rigorously than most of what sits beside it, and has survived as something small and real rather than being refuted.

The concession that follows from taking that seriously is specific. If an organization wants 1 number for a board slide, the recommendation item is a defensible choice, on the condition that nobody allocates resources against its quarter over quarter movement. The failure in practice is almost never the collection of the item. It is the decision to treat its movement as information and to put budget behind that reading.

What the defense cannot rescue is the compression. Cheapness, comparability, and legibility are properties of how an instrument is administered and reported, and all 3 are available to a multidimensional instrument that reports its dimensions at their own levels rather than combining them into an index. The property that cannot be fixed by better administration is asking 1 question, because a single item has no way to report that a brand is admired and difficult to transact with at the same time. That is not a limitation of the recommendation question specifically. It is a limitation of any single item, which is why the answer is not a better one.

What Replaces It Is Not Exotic

A manufacturer persuaded by all of this reasonably wants to know what to use instead, and the literature’s answer is disappointing in the most useful way, since the replacement is not a cleverer single number.

On the same data, in the same models, plain customer satisfaction predicted actual visits significantly better than the recommendation score did (Matzler et al., 2026). Across a very different design, average satisfaction was the best performing metric of the 6 that Morgan and Rego (2006) tested. The measure the net promoter score was introduced to supersede outperformed it in both. That is not an argument for going back to a satisfaction score and tracking that instead, because the same critique of compression applies to any single item, and van Doorn and colleagues (2013) found every metric they tested equally poor at predicting future growth.

The argument is for more dimensions, and it has a number attached to it. Keiningham and colleagues (2007b) found that models using several metrics together beat the best single metric models by 20 to 25 percent in adjusted R squared in 2 of their 3 industries. A modest gain, and it is a gain in the right direction from the cheapest possible change, which is to stop collapsing a relationship into one question.

So what should an organization actually track, if the single score explains 4 percent and stops working on its longest standing customers? Several dimensions, reported at their own levels rather than combined, on a cadence matched to the lag at which they carry information, collected from more than one person inside each account. None of that is novel as measurement practice. It is only novel against the thing most companies are currently doing. The design argument behind it, and what an instrument has to do to survive a line review, is worked through in measuring channel loyalty, and what satisfaction scores miss.

If You Cannot Kill the Program

Most manufacturers reading this cannot retire a global recommendation program, because it is embedded in incentive plans, reported to a parent company, or attached to a contract with a measurement vendor. The useful question is what to do with a program that has to continue, and 5 moves are available without terminating anything.

Keep the item, and keep its wording frozen. The long time series is the most valuable thing the program has produced, and rewording the question to improve it destroys comparability with every prior administration. An instrument that retains the recommendation item unchanged stays legible to everyone already running the original, which is the design decision behind the index described below.

Add dimensions beside it rather than replacing it. The additional questions cost almost nothing to administer once a survey is already going out, and Keiningham and colleagues (2007b) put the gain from using several metrics together at 20 to 25 percent in adjusted R squared over the best single metric in 2 of their 3 industries.

Stop reporting a composite as the headline. Combining the dimensions into one index reintroduces the exact compression the additional questions were added to remove, and a brand admired for its products and avoided in daily transaction will produce a healthy composite while the relationship changes shape underneath it.

Split the respondents by role. The person who signs the agreement and the person who builds the quotation hold different parts of the answer, and a blended figure describes neither, which matters most in exactly the experienced accounts where the single score was observed to stop predicting.

Change the cadence to match the lag. If the relationship between the score and behavior is strongest a year or more out, then a monthly or quarterly review of the movement is reading the least informative part of the series, and an annual reading tied to a strategy cycle is reading the part that carries signal.

None of those 5 requires abandoning anything, and 4 of them cost only the willingness to report an existing result differently. The one that costs money is adding the dimensions, and it costs the least of anything in a research budget, since the expensive part of a survey program is reaching the respondents rather than asking them 4 more questions once reached.

The Same Compression, Measured in a Channel

The channel case is where the compression argument stops being theoretical, and I have measured it directly. In an exploratory study of one automation manufacturer’s distributors, drawn from its advisory council and key accounts, 159 respondents across 11 distributors rated the brand on 5 dimensions rather than 1. Willingness to recommend returned 50.31 and product quality 56.60, both strongly positive. Pricing returned -19.50 and the experience of doing business with the brand returned -34.59, with 54.7 percent of respondents falling in the detractor band on that last dimension. A single recommendation question run on that population would have reported a channel in good health.

Those were experienced partners, most of them long standing, which places them in the stratum Matzler and colleagues found the score least able to read. The divergence between a healthy recommendation figure and a deeply negative operational figure is what the single score conceals by construction, and the mechanism behind the concealment is the same compression that would explain why accumulated experience degrades the score’s predictive power. What that divergence does to a line over time, and why a brand does not see it coming, is worked through in why distributors stop selling your line.

In Finality

The net promoter score entered thousands of companies on a prediction claim, and the prediction claim has been tested more thoroughly than most metrics ever are. It survives as something real and small: significantly related to behavior, worse than the measure it replaced, accounting for a few percent of the outcome, carrying its information at a lag most implementations ignore, and, in the single study that has tested the question against recorded behavior, losing its predictive power entirely once customers have enough experience to have formed a judgment from episodes rather than from reputation.

An industrial manufacturer’s customers are all in that last group. The plant that has run the product for a decade, the distributor whose relationship with the brand predates the current channel manager, the shop that has specified the line since the last recession. They are the accounts that determine whether the business exists in 5 years, and they are the accounts about which one number is least likely to tell a brand anything it can act on.

The practical consequence is not that a brand should stop asking. It is that a brand asking its most experienced partners a single question, once a year, and watching the result move by 3 points, is watching the part of the signal that carries the least information, about the customers it can least afford to misread.


Talk to us

If you sell through distribution. The CHI reads your channel across all 5 dimensions at their own levels, split by the roles that decide what gets quoted, administered independently of your channel management, on a cadence built for the lag at which the reading means something. What it costs, and the terms.

If you carry the lines. The CHI Line Review runs the same instrument across your own top brands with your people as the respondents, which gives a defensible basis for the line card conversations most firms conduct on impression. For distributors, or for repair, service, and integration.


References

Darby, M. R., & Karni, E. (1973). Free competition and the optimal amount of fraud. The Journal of Law and Economics, 16(1), 67-88. https://doi.org/10.1086/466756

Keiningham, T. L., Cooil, B., Andreassen, T. W., & Aksoy, L. (2007a). A longitudinal examination of net promoter and firm revenue growth. Journal of Marketing, 71(3), 39-51. https://doi.org/10.1509/jmkg.71.3.039

Keiningham, T. L., Cooil, B., Aksoy, L., Andreassen, T. W., & Weiner, J. (2007b). The value of different customer satisfaction and loyalty metrics in predicting customer retention, recommendation, and share-of-wallet. Managing Service Quality, 17(4), 361-384. https://doi.org/10.1108/09604520710760526

Matzler, K., Strobl, A., Teichmann, K., & Aigner, G. (2026). Testing the predictive validity of the net promoter score in ski resorts: A longitudinal analysis. Review of Managerial Science, 20(9), 3265-3288. https://doi.org/10.1007/s11846-026-00986-2

Morgan, N. A., & Rego, L. L. (2006). The value of different customer satisfaction and loyalty metrics in predicting business performance. Marketing Science, 25(5), 426-439. https://doi.org/10.1287/mksc.1050.0180

Reichheld, F. F. (2003). The one number you need to grow. Harvard Business Review, 81(12), 46-54.

Schmitt, P., Meyer, S., & Skiera, B. (2012). An analysis of the link between customers’ intention to recommend a firm and the lifetime value of its customers. Recherche et Applications en Marketing (English Edition), 27(4), 121-142. https://doi.org/10.1177/205157071202700405

Sheeran, P. (2002). Intention-behavior relations: A conceptual and empirical review. European Review of Social Psychology, 12(1), 1-36. https://doi.org/10.1080/14792772143000003

Tolbert, C. L. (2022). A hermeneutic study of industrial distribution: The nuanced understanding of organizational fitness in the context of complex systems and memetic culture [Doctoral dissertation, Columbia International University].

van Doorn, J., Leeflang, P. S. H., & Tijs, M. (2013). Satisfaction as a predictor of future performance: A replication. International Journal of Research in Marketing, 30(3), 314-318. https://doi.org/10.1016/j.ijresmar.2013.04.002

The instrument this argument produced. Channel Health Index