← All cases and insights Knowledge base →

ARTICLE · Research

“Let’s test the solution on customers” is four different procedures

Boris Kaptelov · 02.08.2026 · 9 min

The solution has already been invented. Then someone says, “Let’s test it on customers,” and everyone nods, because each person has understood the phrase in their own way.

One of them meant showing a description of the idea and watching the reaction. Another meant checking whether the promise is clearly worded. A third meant letting people click through a prototype. A fourth meant seeing whether they can complete the task without any prompting. These are four different procedures with four different limits on what can be concluded — and the money is usually paid for one of them while an answer to all four is expected.

Why the problem comes first and the solution second

The split between problem interviews and solution interviews comes from customer development practice. The logic is to separate two distinct sources of uncertainty: whether a meaningful problem exists, and whether the proposed concept fits it.

At the problem stage, the product is not mentioned at all. Not out of ritual purity, but because once you demonstrate it, the conversation irreversibly turns into an evaluation of your idea. You lose the one thing you came for: a description of the situation in the person’s own words and own logic.

But this is a rule of product practice, not a universal law. In research on an existing product, in studying how it is used, in comparing alternatives, contact with the solution is precisely the subject of study. The rule has a narrow meaning: do not mix reactions to what you have shown with data about first-hand experience. How the conversation about the customer’s job is structured before any demonstration is covered in our articles on the three schools of jobs-to-be-done and on questioning technique.

What exactly you are testing: four objects

It is useful to distinguish four objects.

The first is the product idea itself: what is being offered, to whom, in what situation, with what promised effect.

The second is the representation of value: how clearly, relevantly and believably the promise is worded.

The third is the form of interaction: whether a person can perform the required actions and understand what happens in response.

The fourth is how the solution works in a real context: compatibility with processes, the environment, other people, infrastructure and organizational constraints.

The problem is that a single stimulus touches several objects at once, and the sources of the reaction get mixed. An interactive prototype simultaneously conveys the idea, the visual style and the intended mechanics — and you do not know which of these the person actually liked. A physical sample adds sensory experience. A service pilot adds staff behavior, waiting, process organization and real-life exceptions.

Hence the rule: decide in advance which of the four objects you are testing, and choose a stimulus that touches that object and, as far as possible, none of the others.

A concept test tests the representation, not the product

A concept can be shown as text, a drawing, a storyboard, a set of attributes, a positioning statement, a mock-up or a sample. The method exists in both qualitative and quantitative form: a discussion reveals interpretations and the grounds for a reaction, while a standardized test measures predefined indicators on a larger sample.

The key point: the object of study is the reaction to the representation presented. Not to the future product. Change the wording, the level of detail, the picture, the brand, the price or the availability of a trial, and the answers change.

How much they change is an open question. A study by Dickinson and Wilby on a familiar consumer product found no substantial interaction between product trial and three positioning variants on the main measures. But this particular result cannot be transferred to other categories, especially sensory, complex or unfamiliar ones.

And a more general caveat: Peng and Finn’s review of concept testing practice found wide variation in study designs and limited public evidence of their reliability and validity. The name of the method in a vendor’s proposal guarantees nothing by itself. What matters is how the concept is shown, who the respondents are, the comparison context and the evaluation criteria.

A value proposition test tests understanding

A value proposition is how a company words the value it offers. Customer-perceived value is something else: it arises in use and in exchange.

An interview can establish whether the message is clear, relevant, believable, distinguishable from others’, and what trade-offs the person sees in it. It does not establish the value actually received, because that requires the experience of use.

There is a subtlety here that people rarely think about. The audience fills in what is implied: they understand terms in their own way and draw conclusions that the literal message does not contain. This has long been studied in research on advertising and labeling claims — Harris, and later Morris and colleagues, showed that listeners confidently infer from a message things it never said. So a respondent’s retelling of the message and their own inferred conclusions are two different data objects, and they must be recorded separately. The first speaks to the clarity of the wording; the second, to the promises you made without noticing.

A separate note on remarks like “that’s important,” “sounds interesting,” “I’d give it a try.” They describe a reaction within the research situation and are not equal to observed choice, to willingness to pay the price of switching, or to sustained use.

A prototype speaks only for what it contains

A prototype is a partial representation of a solution, built to explore specific questions before full implementation.

Its fidelity is not one-dimensional, and the split into “low fidelity” and “high fidelity” is useless without specifying which properties are actually reproduced. A prototype can be visually realistic and non-functional. Functional and incomplete. It can convey the form precisely without conveying the materials. It can reproduce the customer-facing part of a service without reproducing the internal operations.

The data extends to the characteristics the prototype actually represents, and to no others. A paper interface will show how people understand the structure and what sequence of actions they expect — and will not show how they react to system delays. A packaging mock-up will show shelf visibility and how the elements are interpreted — and will not show the experience of storage and use. A role-played service simulation will reveal the sequence of touchpoints — and will not reveal real workload or the variability of staff performance.

The practical takeaway: before the test begins, write a list of what this prototype does not test. It will turn out longer than expected and will spare you a wrong conclusion.

Usability testing is observation, not conversation

The difference is fundamental. Here the primary material is a person’s observed interaction with the product while pursuing set goals. Questions are needed to clarify expectations and interpretations, but they do not replace data on actions, errors, outcomes and time spent.

The international standard ISO 9241-11, on the ergonomics of human-system interaction, defines usability as the extent to which specified users achieve specified goals with effectiveness, efficiency and satisfaction in a specified context. It follows from the definition that the phrase “the product is easy to use” is incomplete without naming the users, the goals and the context. And the standard covers not only software but also physical products, environments and services.

And the key distinction. Usability is not the same as the appeal of the idea, market value or overall impression. A person can easily use something they do not need. A much-needed solution can have severe interaction problems. A good impression can coexist with execution errors. All of this can be measured in a single study, but it must be interpreted separately.

There is also the distinction between thinking aloud during the action and explaining afterwards. The first gives access to what the person holds in mind as the task unfolds. The second gives a coherent reconstruction, assembled with knowledge of the outcome. These are different data, and they must not be mixed in the report.

A pilot tests a configuration, not scalability

The fourth object on the list above is how the solution works in a real context. It has procedures of its own: product trial and pilot implementation.

This is where you get data no prototype can provide: compatibility with processes, staff behavior, real workload, real-life exceptions, data readiness on the client’s side. It is the most expensive and the most substantive test of all.

It has its own limit too, and that limit is regularly overestimated. A successful pilot is evidence that a specific configuration works under specific conditions: at this site, with this team, with this level of management attention. It does not prove scalability. Pilots often run with reinforced support and the most motivated participants — and those are precisely the two conditions that disappear first at rollout.

Formative and summative testing are different jobs

The last distinction — the simplest one, and the most frequently violated.

The distinction is not about stage but about function. Formative evaluation is run in order to change the design: what matters is finding problems, and small samples and rough stimuli will do. Summative evaluation characterizes the level achieved against predefined criteria: what matters is comparability. An informal test of an already finished product can be formative, and an early comparative measurement can serve a summative function. The stage determines nothing.

One study cannot do both well. When a test is commissioned to make a launch-or-not decision but conducted by formative logic, the result is a confident “yes” from a sample that was never assembled for that conclusion.

So here is the question worth asking before you start: are we looking for what to fix, or deciding whether it is ready? Everything else, including the price, depends on the answer.

In practice, we begin with a single conversation: what decision are you making, and what is at stake. If the stake is reversible and small, the honest answer is not to commission a test but to release and watch. If a production line, a packaging print run or a contract is on the line, we match the procedure to the object and write down in advance what it will not show. This takes anywhere from a few days to two months, depending on how many objects need testing.

Sources

  • Osterwalder, A. «Problem» vs Solution in Customer Interviews. Strategyzer, 17 января 2018.
  • International Organization for Standardization. ISO 9241-11:2018. Ergonomics of human-system interaction. Part 11: Usability: Definitions and concepts.
  • Ericsson, K. A., Simon, H. A. Protocol Analysis: Verbal Reports as Data. Revised ed. MIT Press, 1993 — различие проговаривания по ходу и объяснения после.
  • Peng, L., Finn, A. Concept Testing: The State of Contemporary Practice. Marketing Intelligence & Planning, 2008, 26(6), 649–674.
  • Dickinson, J. R., Wilby, C. P. Concept Testing With and Without Product Trial. Journal of Product Innovation Management, 1997, 14(2), 117–125.
  • Payne, A., Frow, P., Eggert, A. The Customer Value Proposition: Evolution, Development, and Application in Marketing. Journal of the Academy of Marketing Science, 2017, 45(4), 467–489.
  • Harris, R. J. Comprehension of Pragmatic Implications in Advertising. Journal of Applied Psychology, 1977, 62(5), 603–608.
  • Morris, L. A., Hastak, M., Mazis, M. B. Consumer Comprehension of Environmental Advertising and Labeling Claims. Journal of Consumer Affairs, 1995, 29(2), 328–350.
  • Theofanos, M. F., Quesenbery, W. Towards the Design of Effective Formative Test Reports. Journal of Usability Studies, 2005, 1(1), 27–45.

Shall we discuss your task?

Get in touch →