The Survey That Went to the Person Who Answered the Phone

Most industrial voice of customer work is a satisfaction survey sent once a year to whichever contact is listed in the CRM, scored on a scale, reported as a number, and circulated with a slide that says customers rate us 8.2 out of 10 on technical support. The people who receive that slide are being asked to accept 3 separate claims at once, and the survey has tested none of them. The first claim is that the person who answered represents the customer, when in most industrial accounts no single person does. The second is that the question asked about something the respondent is able to judge. The third is that what a respondent says on a form corresponds to what that respondent’s organization will do at the next specification cycle.

None of those 3 claims is exotic and all of them have been examined in the research literature, mostly in settings other than this one. The purpose of this article is to bring the relevant parts of that literature together against the specific shape of an industrial manufacturer’s problem, which is that the products are technical, the buyers are organizations rather than people, the purchase is made by several functions with different stakes, and a good deal of what the manufacturer most wants to know is about performance the customer cannot independently confirm. The article covers how to ask, who to ask, and the part that rarely appears in a research proposal, which is the class of questions where more or better research will not produce an answer and something else has to.

Three Kinds of Quality, and Only One of Them a Customer Can Check

The distinction that organizes everything below came out of economics rather than marketing, and it is 55 years old. Nelson (1970) separated goods by how a buyer is able to learn about them, distinguishing those whose relevant attributes can be established by inspection before purchase from those that can only be judged through use afterward. Darby and Karni (1973), in a paper whose title is a fair summary of its concern, added the third category that makes the framework useful here. Credence qualities are those a buyer cannot verify even after purchase and consumption, and their argument connects that unverifiability directly to a seller’s incentive and opportunity to misrepresent, since a claim no one can check is a claim that costs nothing to make. Zeithaml (1981) carried the whole continuum into services and argued that services sit much further toward the experience and credence end than physical goods do, with consequences she traced for how buyers behave, including heavier reliance on word of mouth, greater use of price and visible evidence as quality cues, and higher perceived risk.

That literature has been applied extensively to consumer purchases and to professional services, and as far as I can find it has not been developed against industrial buying, which is where it does the most work. I say as far as I can find deliberately, because the searching behind that statement was thorough rather than exhaustive, and a manufacturer should treat it as an invitation to look rather than a settled claim about the state of a field.

The reason it does the most work here is that an industrial offering is a bundle in which the 3 kinds of quality are mixed and the mixture is not obvious to anyone. A drive’s frame size, its current rating, its communication options, and its price are search qualities, verifiable from a catalog page before anyone signs anything. Its ease of configuration, the responsiveness of the applications group, the clarity of the manual, and the behavior of the local rep are experience qualities, learned across the first year of using it and learned reliably. Its 10 year reliability in a dusty plant at 45 degrees ambient, the actual contribution of its output filtering to motor insulation life, the energy it saved against a baseline nobody recorded, and whether the failure that happened in year 4 was the drive, the installation, the power quality, or the load, are credence qualities, and a maintenance manager who has run the product for a decade is no better positioned to settle them than a manager who has never seen one.

What the Industrial Buyer Cannot Verify

The practical consequence is that a survey question can be well written, well administered, and returned by exactly the right person, and still measure something other than what the manufacturer believes it measures. Asked how satisfied they are with a drive’s reliability, a respondent reports a belief assembled from the failures they happen to remember, the stories told in their own building, the last conversation with a competitor’s salesperson, and the general reputation the brand carries in their industry. That belief is worth knowing, and it is not a measurement of reliability. Treating it as one puts a manufacturer in the position of improving a product attribute in response to a signal that was never reading the attribute.

The same structure explains a pattern most industrial manufacturers have seen and few have named, which is a brand whose field data is strong and whose reputation for reliability is weak, or the reverse. Where a quality is unverifiable, the belief about it is free to move independently of the thing itself, and it moves on whatever inputs are available, which are usually social. My own work in this channel found exactly that mechanism operating at the level of the brand, where a manufacturer can hold a position of meaning in the trade regardless of the accuracy or the age of the position, and where the old proposition that nobody was ever fired for buying the dominant brand functions as a live constraint on what a challenger can achieve with evidence alone (Tolbert, 2022). A challenger brand that wins every technical comparison and loses the specification is not usually losing to data. It is losing to a credence judgment that data does not reach.

The design implication is not that these questions should go unasked. It is that the answers have to be labeled correctly when they come back. A reading on a credence quality is a reading on reputation and belief, it belongs in the report as such, and the action it calls for is a communications and evidence problem rather than an engineering problem. A reading on an experience quality is a reading on the product and the organization around it, and it calls for the opposite response. A voice of customer program that does not sort its findings into those 2 piles will hand a manufacturer a list of improvements in which some items cannot be improved by the department they are handed to.

Asking for the Answer Instead of the Experience

The second common failure is in the form of the question rather than the subject of it. Ulwick (2002) made the argument in its most usable form for industrial work, observing that customers asked to propose solutions produce poor input, because proposing solutions is not their job and their proposals are bounded by the products they already know, while customers asked about the outcomes they are trying to achieve produce input a manufacturer can actually design against. Von Hippel (1986) had established the underlying limitation in a study whose language is worth quoting exactly, since it states the constraint better than a paraphrase: users steeped in the present, he wrote, “are thus unlikely to generate novel product concepts which conflict with the familiar.” His exception is the lead user, defined as a user whose present strong needs will become general in the marketplace months or years later, and the exception is narrow by construction. Most of a manufacturer’s customers are not lead users, which means most of a manufacturer’s customers cannot describe a product that does not exist.

What they can describe, in detail and with precision, is what happened. A maintenance supervisor cannot tell a manufacturer what the next generation of a product should be. The same supervisor can walk an interviewer through the last 3 times a unit came out of service, what was tried first, who was called, how long the part took, what the plant did in the meantime, and what argument was had afterward about whose fault it was. The difference between those 2 conversations is the difference between a question about preference, which invites the respondent to construct an answer, and a question about a specific episode, which asks them to report one.

Griffin and Hauser (1993) established what that kind of asking actually yields, in the study that remains the reference point for elicitation. Working within a segment, they estimated that roughly 20 to 30 interviews capture 90 to 95 percent of the needs the method is capable of voicing at all, and they compared group sessions against one on one interviews, finding that groups did not produce meaningfully more unique needs per person hour than individual interviews did. They also reported a result about satisfaction measurement that industrial manufacturers should sit with, which is that a monadic rating, where a respondent rates the brand they actually chose, carries a self selection bias, and that this monadic measure correlated poorly with market share at roughly .20 while a relative measure, where the respondent rates competing options against each other, correlated at roughly .83. A manufacturer asking its own customers how satisfied they are with its own products is running the weaker of the 2 designs.

There is one more form problem, and it appears whenever a study asks what a customer would pay. The most careful work on that question sits in environmental valuation rather than industrial marketing, so the transfer is an analogy and should be presented as one, but the direction is consistent and the magnitude is not small. List and Gallet (2001), across a set of studies that elicited both hypothetical and real payments for the same good, found hypothetical values averaging roughly 3 times actual values. Murphy, Allen, Stevens, and Weatherhead (2005), analyzing 28 studies and 83 observations, reported a median ratio of hypothetical to actual value of 1.35 with a severely skewed distribution, which is the same finding read through a statistic that outliers do not dominate. The 2 results are not in conflict, they are a mean and a median on skewed data, and together they say that stated willingness to pay runs high by an amount that varies enormously and cannot be predicted for any particular study. A quoted price in an interview is a data point about how a customer talks about price. It is not a forecast.

Who Is the Customer

An industrial purchase is made by an organization, and the research tradition that took that seriously begins with Webster and Wind (1972), whose buying center model treats the purchase as a group decision distributed across roles that want different things: users, influencers, buyers, deciders, and gatekeepers. Johnston and Bonoma (1981) examined the structure and interaction patterns of those groups empirically for capital equipment and industrial services. The concept is 50 years old, it is taught in every industrial marketing course, and it is routinely abandoned the moment a company runs a survey, because the survey goes to one contact.

The methodological cost of that abandonment was measured a long time ago in the same field. Phillips (1981), studying organizations in marketing channels with a design built to separate the trait being measured from artifacts of the informant and the method, found that informant reports converge weakly, and that a substantial share of the variance in what informants say reflects who is answering rather than the organizational reality being described. Kumar, Stern, and Anderson (1993) built on that body of evidence with procedures for selecting competent informants and for assessing agreement across several of them inside the same organization, and their paper opens by noting that despite long standing recommendations to use multiple informants, most published research was still relying on one. Their contribution is a method for doing it properly rather than a fresh demonstration of the error, and the error had already been demonstrated.

Translate that into an industrial account and the roles are concrete rather than abstract. The reliability engineer owns the failure history and the standards. The maintenance supervisor owns what actually happens at 2 in the morning and whether the part is on the shelf. Operations owns the production cost of the downtime and is usually the only party who can put a number on it. Purchasing owns the terms, the approved supplier list, and the pressure to consolidate. Plant engineering owns the specification for the next project, and corporate engineering may own a standard that overrides everything the plant thinks. Those 5 or 6 people describe different relationships with the same manufacturer, and any 2 of them will disagree about which problem is the important one. Averaging them produces a figure that describes nobody, and interviewing whichever one returns the call produces a study of that person.

The same split governs the distributor and the service firm, which in most industrial channels stand between the manufacturer and the plant and are not the same respondent as either. A brand studying only its distributors learns what the channel experiences and not what the plant experiences, and a brand studying only end users learns nothing about why its product is or is not being offered first when the plant calls. Both populations are part of the customer, and a study that collapses them has decided an empirical question by design.

Whether the Respondent Was a Person

Everything in the section above assumes the respondent is who they appear to be, and in panel sourced work that assumption has stopped holding. The Insights Association’s Global Data Quality Benchmarking Report for the first half of 2026, drawing on nearly 1.8 million survey records across 13 countries, found research agencies removing 28.5 percent of records through combined pre survey and in survey screening against 21.2 percent among sample suppliers. The finding that matters for anyone doing industrial work is the one stated most plainly in the report: B2B studies reported the highest combined pre survey and in survey respondent removal rates, along with exceptionally high post survey cleanout rates. A removal rate is a detection rate rather than a contamination rate, which means those figures describe what was caught and say nothing directly about what was not.

On the size of the underlying problem the best available non vendor source is a research brief from NORC at the University of Chicago, which reviewed the literature and reported estimates of fraud running 15 to 30 percent across the market research industry and reaching as high as 45 percent on some platforms. That band has to be quoted with the brief’s own qualification attached, since it pools deliberate fraud, bot generated responses, AI assisted responding, duplicated identities, and ordinary inattentive answering, and the brief states that these categories overlap and are difficult to separate. The honest reading is that a substantial and poorly bounded share of panel sourced survey data is not what it claims to be, with the boundaries of the problem contested and the direction not.

The newer finding concerns the part of a voice of customer report that industrial clients trust most, which is the verbatim. Wang, Mamaev, and Leckie (2026), in a preprint that has not yet been peer reviewed, tested machine generated text detectors against three modes of AI use on verified pre 2020 survey responses. Detectors performed acceptably against a response generated wholesale, reaching .79 to .93 AUROC. Against a human answer lightly revised by a model they fell to .55 to .74. Against persona grounded agentic completion, where a model answers in the sustained character of a specific respondent, they collapsed to .50 to .65, which is chance. Their own purpose built method reached .75 mean AUROC and identified 27.7 percent of AI assisted responses at a 5 percent false positive rate, which the authors themselves describe as short of usable screening. The prior work they cite found 34 percent of respondents on a major crowdsourcing platform self reporting that they used a language model on open ended questions (Zhang et al., 2025, as cited in Wang et al., 2026). A quoted verbatim in a panel sourced report now carries an unquantified probability of having been written by a model, and no reliable way to check.

Synthetic respondents are the same problem arriving by the front door rather than the back, and the strongest evidence on them is recent and peer reviewed. Peng and colleagues (2026), across 19 pre registered studies and 164 outcomes, built digital twins from more than 500 questions per individual and found them correlating with the humans they modeled at r = .20, showing lower within group variance than those humans in 154 of the 164 outcomes, and differing significantly from them on 105. They name 5 systematic distortions, of which the one that matters here is variance compression. A synthetic panel that reproduces a topline while under dispersing the spread has destroyed precisely the property an industrial study exists to read, which is the disagreement between the reliability engineer and purchasing about what the supplier’s real problem is. The aggregate can be right and the study can still be silent on the only question that was worth asking.

Three questions follow for a manufacturer commissioning this work, and they should be asked before the contract rather than after the report. Name the sampling frame, specifically, down to where the respondents came from and who holds the relationship with them. Ask how each respondent was established to be the person and the role they claimed. Ask, in writing, whether any portion of the delivered output was generated or augmented by a model. A supplier who cannot answer the first is selling a panel, and a supplier who will not answer the third has told you something.

The alternative is not a better detector, it is a sampling frame that never had the problem. An interview arranged through the manufacturer’s own account records or through the channel, with a named person in a named plant, on a scheduled call, in a conversation where the interviewer can hear a maintenance supervisor describe a failure they clearly lived through, has a provenance that a panel verbatim does not have and cannot acquire. That is a methodological property rather than a stylistic preference, and it is the reason this practice does not buy panel.

How Many, and the Number That Gets Sold

Sample size in qualitative work is where research proposals are least honest, including proposals I have written, and the literature has moved considerably in the last decade. The dishonesty is rarely a lie about the number. It is that the number arrives in the proposal before the study’s aim has been settled, carried in from a precedent that fit a different study, and once a client has seen a price attached to a count, the design gets fitted to the count rather than the count to the design.

The number most often quoted comes from Guest, Bunce, and Johnson (2006), who analyzed 60 interviews and found that no new themes appeared after the 12th, with the basic elements of their larger themes present as early as the 6th. That finding is real and it is also tightly scoped, and the scope conditions are the part that gets dropped when the number is quoted in a proposal. Their sample was relatively homogeneous, the research question was narrow and well defined, and the researchers already knew the context. Hagaman and Wutich (2017) reanalyzed data across more varied sites and found 20 to 40 interviews were needed before saturation. Hennink and Kaiser (2022), reviewing the studies that have tested the question empirically, report saturation falling within a range of about 9 to 17 interviews with a mean near 12 to 13, and 4 to 8 groups for focus group work, again in studies with fairly homogeneous populations and narrow objectives.

Set those conditions against an industrial voice of customer study and the mismatch is immediate. A study that speaks to 5 buying center roles across 3 plant types in 2 industries is not a homogeneous sample and its question is not narrow. Twelve interviews is a defensible number for a specific question inside a single role, and it is the wrong number for a study that intends to describe a market. Both statements are true at once, and a proposal that quotes 12 interviews for a market level question is selling the arithmetic of the first study and the promise of the second.

Two developments make the sample size conversation more useful than a count. Malterud, Siersma, and Guassora (2016) proposed information power in place of saturation, arguing that the required number falls as the sample carries more of what the study needs, and specifying 5 dimensions along which that varies: the breadth of the study aim, the specificity of the sample, whether established theory is being applied, the quality of the interview dialogue, and whether the analysis is case based or cross case. A narrow aim, a dense and well chosen sample, a theoretical frame, strong interviews, and a case oriented analysis together justify a small study, and a broad aim with a scattered sample justifies nothing of the kind. Braun and Clarke (2021b) go further and argue that saturation is conceptually incoherent for the kind of analysis described in the next section, since saturation belongs to a model in which themes are lying in the data waiting to be found, and that is not the model.

The practical rule that follows is one a manufacturer can hold a research supplier to. The number of interviews should be argued from the study’s aim, its sample, and its analysis rather than asserted from a precedent, and any supplier who quotes a number before the aim is settled has quoted a price rather than a design.

What Happens to the Transcripts

Analysis is where qualitative work either earns its cost or quietly becomes a quotation service. Braun and Clarke (2006) set out thematic analysis in the form most researchers learned it, and then spent the following 15 years distinguishing what they actually meant from what the field did with it. Braun and Clarke (2019) name reflexive thematic analysis as its own approach, built on the researcher’s subjectivity as an analytic resource rather than a contaminant, and on themes as actively generated through interpretation rather than discovered sitting in the transcripts. Braun and Clarke (2021a) draw the map, separating at least 3 variants that share a name, a coding reliability version, a codebook version, and the reflexive version, and argue that each has to be judged by criteria appropriate to its own assumptions, with much of the confusion in published work coming from studies that mix assumptions across variants.

That distinction matters commercially and not only academically. A coding reliability approach asks 2 analysts to agree, reports how often they did, and treats agreement as evidence of quality. It is well suited to counting how many respondents mentioned lead time. It is poorly suited to the question an industrial manufacturer is usually paying to answer, which is why the same complaint means different things in 2 plants, or what a group of people are doing when they describe a supplier as easy to work with. Reflexive thematic analysis is built for that second kind of question, it produces an interpretation rather than a tally, and the researcher’s judgment is doing visible work in the result, which is why the method requires the researcher to say where they stand rather than claim a neutrality nobody has.

What a manufacturer should expect from it is a set of themes that cut across the roles and the sites, each one stated as a claim rather than a label, each one traceable to the material it came from, and each one carrying the cases that do not fit. Negative cases are not an embarrassment in this method, they are the part that keeps a theme honest, and a report with no disconfirming material in it has usually been smoothed. The debrief exists for the same reason. Putting the themes in front of people who know the business, and letting them argue, is a test of whether the interpretation travels, and it regularly changes what the final report says.

When Research Will Not Settle It

Everything above concerns doing the work well, and doing it well is the smaller half of the value. The judgment a manufacturer is actually buying is the one that identifies which of its questions the work cannot answer at all, and identifies them before the money has been spent finding out. Five kinds come up repeatedly in industrial engagements, and each of them has somewhere else it needs to go.

The first kind is the credence question, which is the subject this article opened with. How reliable is our product, how long will the coating hold, did the filter actually extend the motor’s insulation life, is our mean time between failures better than theirs. A customer cannot settle any of those and neither can a study of customers. The answer lives in field failure data, in warranty records, in an instrumented trial, or in a third party test, and a manufacturer who commissions interviews on those questions has purchased a reputation study while believing it purchased an engineering one.

The second kind is the question about something the customer has not experienced. Von Hippel’s finding applies directly, and the practical form of it is that asking 30 plants whether they would adopt a technology none of them has run produces 30 constructions rather than 30 reports. The route around it is a lead user sample, a prototype in a real plant, or a paid trial, and each of those is a different budget line than interviews.

The third kind is the question whose answer is a number the customer does not possess. What does an hour of downtime on that line cost you is the standard example, and in most plants nobody has ever calculated it, so the number produced in an interview was produced during the interview. The same applies to what share of your spend goes to our category and how many of these do you replace a year. Those belong in the customer’s own systems, and getting them means asking for a record rather than an estimate, or accepting that the figure is an impression and reporting it as one.

The fourth kind is price. The stated preference evidence is clear enough about direction to make a simple rule defensible: a willingness to pay question tells you how a customer talks about your price, and only a real offer tells you what they will pay. A study can tell a manufacturer what a price signals, which of the competing brands are being used as the reference, and where the price is doing reputational damage out of proportion to its actual level. It cannot forecast take rate.

The fifth kind is the question that is not a research question at all. Two functions inside the manufacturer disagree about a direction, neither can win the argument internally, and a study is commissioned to settle it. What comes back will be read by both parties as support for their own position, because a qualitative report is rich enough to supply material to either, and the disagreement will survive the research intact. The tell is a scope conversation in which the sponsor keeps returning to a decision that has already been made or has already been blocked, and the honest response is to name it and stop, before anyone spends the money.

So what is a manufacturer buying when a voice of customer program is run properly, if a fifth of the questions it arrives with belong somewhere else? The sorting, in large part. A study that comes back having answered what interviews can answer, having labeled the credence findings as beliefs rather than measurements, and having routed the rest to the field data, the trial, or the internal decision that was masquerading as a research question, has done more for the manufacturer than a study that answered everything it was asked.

The Order the Work Runs In

The sequence matters because each stage sets the conditions under which the next one can be read. Run them out of order and the work still produces a report, which is what makes the error easy to miss, since nothing in a finished document shows whether the baseline existed at the time the interview guide was written.

The work starts with a baseline. What the stage produces is a written account of the business as it currently understands itself, its position, and the customers it believes it has. That account is not research and it is not treated as true. It is the thing the research gets compared against, and without it there is no way to tell whether a theme that came out of the interviews is news to the organization or something everyone already knew and nobody had written down. A finding that confirms what the company believed has a different value than a finding that contradicts it, and only a baseline lets anyone tell them apart.

Identification comes next, and it is a research decision rather than an administrative one. Which customers, which roles inside them, which segments, which of the channel partners between the manufacturer and the plant, and specifically which of them are not currently buying. A study composed entirely of current customers is a study of the people who already said yes, which carries the same limitation the customer advisory literature identified long ago, and it will not explain a loss.

Capture is semi-structured interviewing, conducted by someone with no stake in the answer, organized around episodes rather than opinions, with a guide firm enough to hold the study together and loose enough that a respondent can take the conversation somewhere the guide did not anticipate. The instrument in this practice is the Voice, and its scope is set by the aim and the sample under the information power criteria above rather than by a default count.

Analysis is reflexive thematic analysis, producing themes with their disconfirming cases attached. The report states what the themes are, what they rest on, where they are firm and where they are provisional, and which of the sponsor’s original questions the study could not answer and why. The debrief puts all of it in front of the people who have to act on it, and the argument that happens there is part of the method rather than a formality at the end of it.

In Finality

An industrial manufacturer asking its customers what they think is doing something sensible, and the reason it so often produces little is not that the customers were unhelpful or that the questions were badly worded. It is that a single instrument was pointed at 3 different kinds of question, one of which it answers well, one of which it answers only if the asking is built around episodes rather than preferences, and one of which it cannot answer at all because the customer has no way to know.

The program that is worth paying for is the one that separates them before the fieldwork rather than after, reaches the several people inside the account who hold different parts of the answer, can say where its respondents came from and how it knows they were people, argues its sample size from its aim instead of a precedent, interprets rather than tallies, and comes back willing to say which of the questions it was handed belong to the field data, the trial, or the conversation the company has been avoiding having with itself.

The last of those is the one that clients remember. A research supplier who takes a question off the table has given up revenue to do it, which is the clearest evidence available that the rest of the report can be trusted.


Talk to us

If you manufacture and sell through distribution. The Voice reads your customers and your channel through semi-structured interviews across the roles that actually hold the answer, analyzed by reflexive thematic analysis, reported with the credence findings labeled as what they are. It pairs with the CHI when the question is how widespread something is rather than why it is happening. What it costs, and the terms.

If you carry the lines or service the installed base. The same method reads your own customers and your own brands, and the debrief puts the disagreement inside your building on the table before leadership decides anything. For distributors, or for repair, service, and integration.


References

Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77-101. https://doi.org/10.1191/1478088706qp063oa

Braun, V., & Clarke, V. (2019). Reflecting on reflexive thematic analysis. Qualitative Research in Sport, Exercise and Health, 11(4), 589-597. https://doi.org/10.1080/2159676X.2019.1628806

Braun, V., & Clarke, V. (2021a). One size fits all? What counts as quality practice in (reflexive) thematic analysis? Qualitative Research in Psychology, 18(3), 328-352. https://doi.org/10.1080/14780887.2020.1769238

Braun, V., & Clarke, V. (2021b). To saturate or not to saturate? Questioning data saturation as a useful concept for thematic analysis and sample-size rationales. Qualitative Research in Sport, Exercise and Health, 13(2), 201-216. https://doi.org/10.1080/2159676X.2019.1704846

Darby, M. R., & Karni, E. (1973). Free competition and the optimal amount of fraud. The Journal of Law and Economics, 16(1), 67-88. https://doi.org/10.1086/466756

Griffin, A., & Hauser, J. R. (1993). The voice of the customer. Marketing Science, 12(1), 1-27. https://doi.org/10.1287/mksc.12.1.1

Guest, G., Bunce, A., & Johnson, L. (2006). How many interviews are enough? An experiment with data saturation and variability. Field Methods, 18(1), 59-82. https://doi.org/10.1177/1525822X05279903

Hagaman, A. K., & Wutich, A. (2017). How many interviews are enough to identify metathemes in multisited and cross-cultural research? Field Methods, 29(1), 23-41. https://doi.org/10.1177/1525822X16640447

Hennink, M., & Kaiser, B. N. (2022). Sample sizes for saturation in qualitative research: A systematic review of empirical tests. Social Science and Medicine, 292, 114523. https://doi.org/10.1016/j.socscimed.2021.114523

Insights Association. (2026). Global data quality benchmarking report, H1 2026. Insights Association.

Johnston, W. J., & Bonoma, T. V. (1981). The buying center: Structure and interaction patterns. Journal of Marketing, 45(3), 143-156. https://doi.org/10.1177/002224298104500312

Kumar, N., Stern, L. W., & Anderson, J. C. (1993). Conducting interorganizational research using key informants. Academy of Management Journal, 36(6), 1633-1651. https://doi.org/10.5465/256824

List, J. A., & Gallet, C. A. (2001). What experimental protocol influence disparities between actual and hypothetical stated values? Environmental and Resource Economics, 20(3), 241-254. https://doi.org/10.1023/A:1012791822804

Malterud, K., Siersma, V. D., & Guassora, A. D. (2016). Sample size in qualitative interview studies: Guided by information power. Qualitative Health Research, 26(13), 1753-1760. https://doi.org/10.1177/1049732315617444

Murphy, J. J., Allen, P. G., Stevens, T. H., & Weatherhead, D. (2005). A meta-analysis of hypothetical bias in stated preference valuation. Environmental and Resource Economics, 30(3), 313-325. https://doi.org/10.1007/s10640-004-3332-z

NORC at the University of Chicago, Center for Panel Survey Sciences. (2026). Fraudulent respondents and bots in nonprobability surveys: A literature review (CPSS research brief). NORC at the University of Chicago.

Nelson, P. (1970). Information and consumer behavior. Journal of Political Economy, 78(2), 311-329. https://doi.org/10.1086/259630

Peng, T., Gui, G., Brucks, M., Merlau, D. J., Fan, G. J., Ben Sliman, M., Johnson, E. J., Althenayyan, A., Bellezza, S., Donati, D., Fong, H., Friedman, E., Guevara, A., Hussein, M., Jerath, K., Kogut, B., Kumar, A., Lane, K., Li, H., … Toubia, O. (2026). Digital twins are funhouse mirrors: Five systematic distortions. Science Advances, 12(36), eaeh8260. https://doi.org/10.1126/sciadv.aeh8260

Phillips, L. W. (1981). Assessing measurement error in key informant reports: A methodological note on organizational analysis in marketing. Journal of Marketing Research, 18(4), 395-415. https://doi.org/10.1177/002224378101800401

Tolbert, C. L. (2022). A hermeneutic study of industrial distribution: The nuanced understanding of organizational fitness in the context of complex systems and memetic culture [Doctoral dissertation, Columbia International University].

Ulwick, A. W. (2002). Turn customer input into innovation. Harvard Business Review, 80(1), 91-97.

von Hippel, E. (1986). Lead users: A source of novel product concepts. Management Science, 32(7), 791-805. https://doi.org/10.1287/mnsc.32.7.791

Wang, Q., Mamaev, B., & Leckie, C. (2026). Towards detecting AI-assisted responses in online surveys (arXiv:2609.17317) [Preprint]. arXiv. https://arxiv.org/abs/2609.17317

Webster, F. E., Jr., & Wind, Y. (1972). A general model for understanding organizational buying behavior. Journal of Marketing, 36(2), 12-19. https://doi.org/10.1177/002224297203600204

Zeithaml, V. A. (1981). How consumer evaluation processes differ between goods and services. In J. H. Donnelly & W. R. George (Eds.), Marketing of services (pp. 186-190). American Marketing Association.

The instrument this argument produced. Voice of Customer