ARTICLE · Research
“The analysis was done with AI” is not a description of a procedure
A vendor delivers a report on forty interviews. Done in a week instead of a month, at half the price, and it reads smoothly. Asked about the method, they answer that the transcription and analysis were done with artificial intelligence.
That answer says nothing about the quality of the work. Not because AI is bad, but because “analysis with AI” is not one operation — it is five different ones, with very different reliability. They line up in ascending order: the further from the original recording, the more interpretation has been added and the higher the cost of an early error.
The five operations: transcription, organizing the corpus, summarization, coding, and building themes and conclusions.
Let’s take them in order and see exactly what can be handed to the machine — and where the zone begins for which a human answers.
Transcription produces a draft, not a copy
Modern speech recognition systems work in many languages and remove the heaviest manual labor. The Whisper speech recognition model was trained on a large corpus of weakly labeled multilingual data and proved robust across recording conditions, but the original paper does not claim error-free performance for every language, accent, and subject domain.
Something else matters more: the errors are not distributed randomly. Their probability rises with poor audio, interruptions, noise, accents, fast speech, industry terminology, and proper names. Speaker separation is a task of its own, and a system can merge two people’s remarks into one or split a single remark in two.
Look at that list carefully. Terms, numbers, names, who said what, negation — these are exactly the elements on which conclusions are built. A lost “not” flips the meaning of a sentence to its opposite, and in a smooth transcript you cannot see it.
Hence the rule: any quote that goes into the report, and any number a calculation rests on, gets checked against the audio. Not the whole transcript — the places that carry weight.
Organizing the corpus is the only place where the gain is pure
The most underrated part of the work, and the safest to automate.
Search across the entire corpus, linking fragments, tagging by participant and stage, navigating between transcript and audio, version control — here the machine does not interpret; it keeps order. An error here does not distort a conclusion, it slows the work down, and this is the only one of the five operations where the savings are not paid for with risk.
The practical implication for choosing a vendor: if they run the project in a dedicated system rather than in chat threads and spreadsheets, you will later be able to verify any quote. If it lives in chat threads, there will be nothing to verify.
Summarization changes how the data is represented
An automatic summary decides what counts as important, merges similar wording, and smooths over whatever does not fit: contradictions, rare cases, the sequence of events over time.
For navigating the corpus, that is excellent. As a replacement for the primary text — no. Summaries across several interviews at once are especially sensitive: the result depends on which chunks went into processing and how the query was phrased.
There is one practical test. A summary that gives you no way to drill down to the source fragment cannot be verified. Some systems do provide such links: the qualitative analysis software MAXQDA states that its generated reports offer interactive jumps to the source fragments, and the research repository Dovetail offers answers with links to the underlying materials. That makes the work verifiable. It does not confirm that the interpretation is correct.
Coding against an existing codebook and coding from scratch are different tasks
“AI-assisted coding” hides at least four operations: applying a codebook defined in advance, finding fragments that fit a given code, inventing new codes from the corpus, and assembling codes into categories and themes.
The first task is close to machine labeling: there are clear definitions, there are examples, the output can be checked against a reference and agreement can be measured. Here models help and save noticeable time.
Beyond that, something different begins. Inductive coding requires deciding what counts as significant in the first place, how a fragment relates to the context of the whole conversation, and what central idea a theme expresses. Research from recent years shows that models reproduce some of the codes and themes of human analysis, but the result depends on the model, the phrasing of the task, fragment size, whether a codebook is provided, and what exactly the comparison is made against. Consistent weaknesses are also on record: shallower interpretation, loss of context, trouble with the links between categories, and divergence between repeated runs.
And the error profile is not the same as a human’s. In the study by Parfenova and colleagues, people handled complex statements better, while models did relatively better on simple fragments. In the comparison by Ayik and colleagues, four different tools partially matched human analysis and saved noticeable time, but none of them reproduced the depth of interpretation in full. There is no single “AI accuracy percentage” for qualitative analysis, and a vendor who quotes one is quoting a number pulled out of thin air.
Why agreement with a human proves nothing
The temptation is simple: give the model and an analyst the same corpus and see how closely they agree.
In a deductive design this works — there is a reference standard. In an inductive one it does not: several different coding systems can illuminate the same material equally plausibly. High agreement with one human coding proves neither completeness nor analytical value. Low agreement does not necessarily mean an error.
You have to evaluate on other grounds: is the result connected to the research question, is context preserved, does each theme hold together around one idea, are the themes distinguishable from one another, have the cases that do not fit the scheme been considered, and can you get from a conclusion down to specific fragments. What exactly to demand as that audit trail is something we cover in the article on the report.
Generating conclusions is furthest from the data
The chain looks like this: fragment — code — category — theme — explanation — conclusion for a decision. Interpretation is added at every step, and a selection error at the first step is amplified at the last.
A generative system will write a coherent, persuasive text even on an incomplete or contradictory basis. That is its strength as a writing tool and its main danger as an analysis tool. Smooth phrasing is not a sign of evidential weight — and smoothness is exactly what the reader of a report picks up.
It is telling how the developers themselves talk about their features. ATLAS.ti describes fully automatic inductive coding as a beta feature that requires subsequent review by the researcher. MAXQDA speaks of code suggestions and explanations for segments. The wording is chosen precisely: the system proposes analytical objects; it does not establish their truth.
Why “done with a language model” does not describe a procedure
The output of a language model depends on the provider, the model’s name and version, the system instructions, the prompt text, the examples supplied, temperature and other parameters, the amount of context passed in, the order of the fragments, and the date of the run. A model update changes the output without a single change to the prompt.
Compare that with how a conventional method is described: who selected the participants and by what criterion, what the interview guide was, who did the coding, against which codebook, how disagreements were resolved. All of this lets another person understand what was done and repeat it.
For the computational step to be described at the same level, you need to preserve the corpus or a frozen version of it, the text-splitting rules, the codebook, the full prompts with parameters, the model’s responses, the date, the tool version, and every manual decision made afterwards. A repeat run on part of the data will show stability, but it will not replace substantive verification.
That is the question worth putting to a vendor instead of asking whether they used AI. The other five questions to ask about a research proposal are in a separate article on the budget.
Confidentiality is defined by the contract, not by the model’s name
Sending interviews to an external service touches personal data, the client’s trade secrets, and obligations to research participants who were promised confidentiality.
What matters here is not which model was used, but where the data is physically processed, what the contract says about using it for training, how long it is stored, and what exactly was promised to the people who were interviewed. Consent to take part in a study is not consent to hand the recording to a third party.
What actually gets saved
Let’s sum up.
Transcription — saved almost entirely, with spot checks of the key passages. Organizing the material, search, corpus navigation — saved substantially. Coding against an existing codebook — saved noticeably, after a check on a sample. A first-pass summary for orientation — saved, as a navigation layer.
Building themes, explaining the mechanism, and drawing the conclusion for a decision — not saved. What gets cut here is not work but responsibility, and the result of that cut becomes visible only once the decision has already been made.
The acceleration in the early stages is real and worth having. It simply does not turn a one-week project into a report of the same quality as a one-month one — it frees the analyst’s time for the part that cannot be delegated. The difference between a good and a bad use of AI lies precisely in where the freed-up time went.
If you are holding a report that was done fast and cheap and you have doubts, we review such reports as a separate engagement: we check the key quotes against the recordings, test whether the conclusions trace down to the fragments, and tell you which decisions this report can support and which it cannot. It takes a few days and usually pays for itself with a single decision not taken.
Sources
- Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., Sutskever, I. Robust Speech Recognition via Large-Scale Weak Supervision. Proceedings of the 40th International Conference on Machine Learning, PMLR 202, 2023, 28492–28518.
- Than, N., Fan, L., Law, T., Nelson, L. K., McCall, L. Updating «The Future of Coding»: Qualitative Coding with Generative Large Language Models. Sociological Methods & Research, 2025, 54(3), 849–888.
- Parfenova, A., Marfurt, A., Pfeffer, J., Denzler, A. Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data Analysis. Findings of ACL: NAACL 2025, 6471–6484.
- De Paoli, S. Performing an Inductive Thematic Analysis of Semi-Structured Interviews With a Large Language Model. Social Science Computer Review, 2024, 42(4), 997–1019.
- Ayik, B., Gu, D., Zan, Y., Kim, S., Kim, W. L. Human vs. AI: Evaluating Thematic Analysis With ChatGPT, QInsights, ATLAS.ti AI, and MAXQDA AI Assist. Qualitative Inquiry, 2026.
- Schroeder, H., Aubin Le Quéré, M., Randazzo, C., Mimno, D., Schoenebeck, S. Large Language Models in Qualitative Research: Uses, Tensions, and Intentions. CHI 2025, Article 481.
- MAXQDA. AI Assist for Qualitative Data Analysis. Официальная документация, дата обращения 28.07.2026.
- ATLAS.ti. AI Coding. ATLAS.ti 26 User Manual. Официальная документация, дата обращения 28.07.2026.
- Dovetail. Research Repository. Официальная документация, дата обращения 28.07.2026.
Shall we discuss your task?
Get in touch →