ARTICLE · Research
What research costs: six questions for your vendor
One proposal offers thirty in-depth interviews, four weeks, and a report with recommendations. Another offers fifteen interviews, six weeks, and a higher price. Choosing between them on these numbers is impossible, because the interview count describes neither the amount of work nor the reliability of the conclusions.
Below are six questions that separate a proposal with a method inside from a proposal with a price list.
How many interviews are needed
The short answer: there is no single number, and no standard exists. For a narrow question within one homogeneous group, nine to seventeen interviews are usually enough. When comparing several segments, interviews are counted for each group separately, so thirty interviews across five segments is five studies of six interviews each, not one large study.
A vendor's answer of "twelve, that's the standard" is wrong. Here is what the studies this figure was taken from actually showed.
| Study | What was measured | Result | Authors' caveat |
|---|---|---|---|
| Guest, Bunce, Johnson, 2006. Field Methods, vol. 18, no. 1, pp. 59–82 | 60 interviews, accumulation of codes and themes | Basic elements of themes after 6 interviews; claimed saturation at around 12 | The result applies to their corpus, their guide, and their coding approach |
| Hennink, Kaiser, Marconi, 2017. Qualitative Health Research, vol. 27, no. 4, pp. 591–608 | 25 interviews; code saturation and meaning saturation measured separately | Code saturation at around 9; meaning saturation at 16 to 24 | Hearing a theme and understanding it are different criteria |
| Hennink, Kaiser, 2022. Social Science & Medicine, vol. 292, art. 114523 | Systematic review of empirical tests | A range of 9–17 for homogeneous samples with a narrow objective | The authors explicitly limit generalization to other designs |
| Squire et al., 2024. Journal of Medical Internet Research, vol. 26, art. e52998 | Secondary analysis of five studies with 30–70 web interviews each | Near saturation after 15–23; no new codes at all after 30–67 | Structured and predominantly deductive designs reached the threshold earlier |
A substantive answer is built on information power, a concept proposed by Malterud and colleagues. The sufficiency of a corpus depends on five things: how narrow the question is, how specifically the cases are selected, whether the analysis draws on existing theory, how substantive the conversation is, and how the analysis is structured.
It is also worth knowing that justifying sample size is simply not customary in this field. Vasileiou and colleagues analyzed fifteen years of publications and found that sample size is almost always reported, while its sufficiency is justified far less often. So your question "why exactly this many" is not nitpicking — it is a rare practice.
What "interviews until saturation" means
The phrase sounds methodical and means nothing until it is stated what exactly is supposed to stop changing. Saturation is not one state but at least four different ones.
| Type of saturation | What stops changing | What it does not prove |
|---|---|---|
| Data saturation | New interviews bring no new information | That the analysis has reached depth |
| Code or thematic saturation | No new codes or categories emerge | That the meaning is understood and the connections explained |
| Meaning saturation | New cases no longer deepen the understanding of variations | That all contexts and rare cases have been exhausted |
| Theoretical saturation | The categories of the emerging theory require no further refinement | That statistical representativeness has been achieved |
The difference between code and meaning saturation is money. In the studies, the two are separated by seven to fifteen interviews.
So the question is simple: saturation of what, and how will you know it has been reached. The phrase "interviews were conducted until saturation" — without the type of saturation, the unit of counting, and the assessment procedure — explains nothing about data sufficiency.
Who is selected, and on what basis
This is where most of the quality hides, and this is precisely the section that is usually empty in proposals.
The sampling strategy determines which conclusions will be possible.
| Strategy | What it strengthens | Main limitation |
|---|---|---|
| Criterion sampling | The experience of people who have actually encountered the phenomenon | Everything depends on the precision of the criterion |
| Homogeneous sample | Depth within a single segment | Shows no variation beyond it |
| Maximum variation | The range of mechanisms and contrasts | Each subgroup ends up with too few cases |
| Typical cases | A description of the central practice | Typicality requires external grounding and does not equal the average |
| Extreme cases | Mechanisms visible at the margins | Does not describe ordinary frequency |
| Quota or cell-based sampling | Comparability of roles and segments | Small cells create the appearance of comparison without depth |
| Snowball | Access to closed and rare groups | Network ties amplify similarity and cut off the periphery |
| Convenience sample | Quick access to material | Availability becomes a hidden selection criterion |
The last two sneak into a project unnoticed, and both skew the result in the same direction: you systematically interview the people who are happy to talk to you.
And here is the arithmetic that explains the price better than any explanation. Crossing four buying-center roles, three stages of the customer journey, and two types of organizations yields twenty-four possible cells. Even a large overall corpus will contain one case or none in some of those cells. The budget grows not with the number of interviews but with the number of cells in which you need a conclusion.
A practical test: ask whom the vendor will not include and why. A coherent answer is what separates a researcher from a recruiter.
Are the people it didn't work for included
A corpus of current, satisfied customers explains why you suit the people you already suit.
Negative and deviant cases are needed not for completeness but to test the explanation: customers who left, those who declined at the selection stage, those who implemented with no effect, those who solved the problem without your product. One such case sharpens the mechanism more than a tenth confirmation of what is already understood.
That said, a negative case is informative only where the outcome was possible in principle. Including a company that could not have bought your product under any circumstances is not a contrast — it is a wasted interview.
In the budget this should be a separate line item, because recruiting lost and declined customers is more expensive and slower than recruiting current ones. If the line is not there, it is not in the work either.
What decision the research is for
The question it is wiser to start with, but which gets checked last, because everything else shows up in it.
If the vendor frames the goal as "understanding customers better," any result will fit it, which means none can be declared a failure. A workable framing sounds different: decide whether to launch the line in this segment; understand why customers churn in month nine; choose one of three product configurations.
Both the scope and the price depend on this. A narrow question can be answered with a smaller corpus. A broad one requires either more data or narrowing. A vendor who never asked about the decision will be selling volume.
What is said about the limits of the conclusions
The most revealing question.
If the promise is "a map of the problems the market cares about," or "an assessment of demand potential," or "a forecast of retention growth" — what is promised is something interviews do not deliver. Frequency within a target sample describes that sample. The phrase "six out of ten said" refers to ten people. Prevalence is measured by a survey with a justified sampling procedure; effect size, by an experiment or behavioral data.
A separate trap: a list of needs sorted by frequency of mention looks like a list of priorities and is not one. Why that is, and how to read such a list, we cover separately in the article on what gets called a need.
A good proposal itself contains a section on what the research will not show, and what the next step should be to get the rest.
What this means for the budget
Three things worth pinning down before you sign.
Count analytical groups, not interviews. It is their number, not the total number of meetings, that determines both the price and the kind of conclusion you will get.
Require separate line items for lost and declined customers and for cross-checking against your own data. Without the first, the corpus is biased; without the second, there is nothing to compare the narrative to.
Require a section on limits in the report itself, not in conversation. We include one in every project, and not out of caution: the client should understand what decision the result will let them make, before they pay.
If you have a proposal in hand and are not sure what is behind it, we review third-party research proposals as a separate short engagement — it takes a couple of days and usually pays for itself either through a lower budget or through walking away from a useless project.
Short answers
How many in-depth interviews are needed?
There is no single number. For a narrow question within a homogeneous group, nine to seventeen interviews are usually enough. When comparing segments, interviews are counted for each group separately.
Where did the figure of 12 interviews come from?
From the 2006 study by Guest, Bunce, and Johnson: in a corpus of sixty interviews, the basic elements of themes emerged after six, and claimed saturation was reached by around the twelfth. The authors specified that the result applied to their corpus, their guide, and their coding approach. The caveat got lost in the retellings.
What does "interviews until saturation" mean?
Nothing, until the proposal says saturation of what. Code saturation arrives at around the ninth interview; meaning saturation, after the sixteenth to twenty-fourth. The difference between them is the money in the budget.
What does the price of research depend on?
On the breadth of the question, the number of analytical groups in which a conclusion is needed, the difficulty of reaching respondents, and the depth of the analysis. Not on the number of interviews in the proposal.
Sources
- Guest, G., Bunce, A., Johnson, L. How Many Interviews Are Enough? An Experiment with Data Saturation and Variability. Field Methods, 2006, 18(1), 59–82.
- Hennink, M. M., Kaiser, B. N., Marconi, V. C. Code Saturation Versus Meaning Saturation: How Many Interviews Are Enough? Qualitative Health Research, 2017, 27(4), 591–608.
- Hennink, M., Kaiser, B. N. Sample Sizes for Saturation in Qualitative Research: A Systematic Review of Empirical Tests. Social Science & Medicine, 2022, 292, 114523.
- Squire, C. M., Giombi, K. C., Rupert, D. J., Amoozegar, J., Williams, P. Determining an Appropriate Sample Size for Qualitative Interviews to Achieve True and Near Code Saturation: Secondary Analysis of Data. Journal of Medical Internet Research, 2024, 26, e52998.
- Saunders, B., Sim, J., Kingstone, T., Baker, S., Waterfield, J., Bartlam, B., Burroughs, H., Jinks, C. Saturation in Qualitative Research: Exploring Its Conceptualization and Operationalization. Quality & Quantity, 2018, 52(4), 1893–1907.
- Malterud, K., Siersma, V. D., Guassora, A. D. Sample Size in Qualitative Interview Studies: Guided by Information Power. Qualitative Health Research, 2016, 26(13), 1753–1760.
- Vasileiou, K., Barnett, J., Thorpe, S., Young, T. Characterising and Justifying Sample Size Sufficiency in Interview-Based Studies. BMC Medical Research Methodology, 2018, 18, 148.
- Patton, M. Q. Qualitative Research & Evaluation Methods: Integrating Theory and Practice. 4th ed. Thousand Oaks, CA: SAGE, 2015.
- Mahoney, J., Goertz, G. The Possibility Principle: Choosing Negative Cases in Comparative Research. American Political Science Review, 2004, 98(4), 653–669.
Shall we discuss your task?
Get in touch →