top of page

Many Codes, Little Context: What Open-Source LLMs Actually Add to Qualitative Analysis

Healthcare interview analysis with open-source AI
Healthcare interview analysis with open-source AI

Qualitative analysis has never been short of labels

Large language models make coding look deceptively simple. Feed in a set of interview transcripts, and the model can quickly return phrases that sound analytically useful: patient education, physician support, communication barriers, emotional distress, treatment burden. The output is fluent, organized, and often plausible. Yet qualitative analysis has never been difficult because researchers cannot produce enough labels. The harder work begins after labeling: deciding whether a code preserves context, whether it repeats another code in different words, whether it can support theme development, and whether it actually helps explain participants’ experience.


Misra and colleagues examine that problem through a demanding empirical case: 34 semi-structured interviews with rural Appalachian patients living with Type 2 diabetes, focused on how they discuss their disease with healthcare providers. The interviews had already been coded by two trained qualitative researchers using a traditional inductive approach. The study then asked two open-source LLMs, Gemma2 and Llama3.1, to analyze the same material and compared their codes with the human-generated benchmark. The point was not whether the models could produce codes. They could. The more consequential question was how much of that output remained usable once context, duplication, and fit with researcher-developed themes were taken seriously.


From model performance to research workflow

The study sits at the intersection of thematic analysis and AI-assisted qualitative research. Thematic analysis is commonly used with interviews, focus groups, and semi-structured interviews to identify patterns, themes, and meanings in narrative material. Its strength lies in interpretation: researchers move between transcripts, codes, and emerging themes, while attending to both explicit statements and latent meanings. That process is slow and resource-intensive, and it can be shaped by researcher expertise, theoretical commitments, and bias. For this reason, qualitative researchers often rely on multiple coders and negotiated agreement to strengthen trustworthiness (Castleberry & Nolen, 2018; Terry et al., 2017; Braun & Clarke, 2019).

Open-source LLMs in qualitative analysis
Open-source LLMs in qualitative analysis

LLMs appear attractive in precisely this space. They can process large amounts of text quickly, generate candidate codes, and help researchers enter the material at an earlier stage. Healthcare interviews, however, introduce a constraint that generic discussions of AI often underplay: data privacy. The article notes that much existing work relies on tools such as ChatGPT, which require data to be processed online and therefore raise confidentiality concerns when sensitive healthcare material is involved. Misra and colleagues instead used Gemma2 and Llama3.1, both open-source models, so that patient interview data would not need to be exposed to internet-based systems.


That choice shifts the study away from a simple contest between models. It asks a methodological question: when researchers cannot sacrifice data ethics for convenience, can open-source LLMs still serve a useful role in qualitative analysis? The answer depends on where they are placed in the workflow.


Early coding is not the same as thematic analysis

The LLMs in this study were not treated as independent qualitative analysts. The authors describe their approach as informed by Braun and Clarke’s (2006) six-phase thematic analysis, but the models were used mainly at the front end of the process: summarizing text segments and proposing potential codes. Reviewing, refining, organizing, and interpreting themes remained the work of the research team. Model output therefore functioned as analytic material, not as a finished finding.


The study tested two prompting routes. In the deductive route, the models received the research question, an analytic direction, predefined transcripts, and examples of initial codes. This few-shot setup encouraged the models to approximate researcher coding. In the inductive route, the models received only the task, without examples or predefined categories, and generated codes through zero-shot prompting. The contrast matters because it separates two different forms of model assistance: coding with a supplied frame, and coding from the material alone.


The models produced a large amount of material. Llama3.1 generated 1,042 unique codes under the deductive framework and 829 under the inductive framework. Gemma2 generated 723 unique codes deductively and 715 inductively. Llama3.1 became noticeably more expansive once it was given the research question and examples; Gemma2 produced similar volumes across the two approaches. On the surface, these numbers suggest productivity. The later results show why productivity is the least interesting part of the story.

More codes narrowed quickly into fewer usable codes

Once the model outputs were evaluated by qualitative researchers, the apparent abundance of codes became less impressive. Deductive prompting performed better than inductive prompting, but not by enough to make the outputs analytically reliable on their own. Among Gemma2’s deductive codes, 45.6% had initial context, 29.5% were duplicative, and 27.1% fit researcher-identified themes. Among Llama3.1’s deductive codes, 43.7% had initial context, 36.2% were duplicative, and 26.8% fit. The inductive results were weaker: 19.2% of Gemma2’s inductive codes and 11.7% of Llama3.1’s inductive codes fit the human-developed themes.


These figures change how the model outputs should be read. A long list of codes does not mean a deep analysis has occurred. Much of the output remained at the level of candidate labeling, and only a smaller portion carried enough context or thematic fit to support interpretation.


The duplication problem is especially revealing. The study notes that researchers treated “Benefitted from Educational Program” and “Patient Education” as overlapping expressions under the broader context of “Group Support and Education.” “Positive Physician Relationship” closely resembled “Positive patient doctor relationship,” while “Open Communication with Healthcare Provider” paralleled “Open Communication with Doctor.” In cases like these, the model was not opening a new analytic direction; it was generating alternate phrasings for a similar idea.


For qualitative researchers, that matters. Manual coding is slow, but conceptual boundaries are adjusted during the act of coding. LLM output is fast, but it can leave researchers with a second layer of labor: cleaning near-synonyms, merging codes, deciding what counts as a new concept, and separating meaningful variation from verbal variation. Speed at the beginning can reappear as sorting work later.

LLM-generated codes and human interpretation
LLM-generated codes and human interpretation

Where the models lose the interview

The most important weakness in the study is not simply duplication. It is the thinness of context. Gemma2’s deductive codes reached only 27.1% good fit with researcher-developed themes, while Llama3.1 reached 26.8%. Under the inductive framework, the good-fit rates dropped to 19.2% for Gemma2 and 11.7% for Llama3.1. These numbers suggest that the models could detect many surface cues related to diabetes, providers, communication, support, and barriers, but had difficulty preserving the interpretive depth of the interviews.


In patient-provider communication, a phrase such as focusing only on “the numbers” is not merely a communication issue. It may register a patient’s sense of being reduced to clinical metrics. It may also point toward frustration in chronic disease management, limited appointment time, uneven continuity of care, or an asymmetrical clinical relationship. A model can detect the obvious terms. The analytic question is how those terms work inside a patient’s account of care.


That difference is where code generation separates from thematic interpretation. A theme is not a pile of related labels. It is an argument about how meanings are patterned across the material. Without stable contextual judgment, the model may remain close to the vocabulary of the transcript while missing the structure of the experience.


The useful role: a second reading, not a final reading

The human analysis identified two broad themes in patient-provider communication: barriers and facilitators. Barriers included disrespectful providers, providers talking down to patients, lack of continuity of care, attention to “the numbers” rather than the person, insufficient appointment time, and inaccessibility. Facilitators included positive and supportive providers, attentiveness, partnership in clinical care, providers with lived experience of diabetes, and participation in the DHSMP program, which improved diabetes education, health literacy, and empowerment.


The LLM outputs did not simply reproduce these themes, but they did bring some additional material into view. Researchers identified supplementary subthemes in the model-generated codes, including distrust of healthcare providers, financial struggles, perceptions that providers lacked diabetes knowledge or holistic understanding, lack of emotional support, and lack of open communication. The models also surfaced occasional expressions that appeared only once or twice and had not been developed into themes in the original analysis, such as patients feeling guilty about burdening providers, fear or hesitation about discussing diabetes, age-related discomfort in discussing diabetes, and feeling abandoned by providers.


This is where the models become methodologically useful. Their value is not that they interpret better than trained researchers. It is that they may disturb the completeness of the first reading. Qualitative research has long used multiple coders to strengthen rigor and trustworthiness, partly because different readers notice different things and can expose blind spots in a single interpretation (Lincoln & Guba, 1986). An LLM cannot replace that human process of negotiation, but it can add a supplementary pass through the data. That pass is noisy and uneven, yet it may return certain marginal experiences to the researcher’s attention.

AI coding for thematic analysis
AI coding for thematic analysis

What happens after the model finishes coding

After reading the study, the place of LLMs in qualitative analysis becomes fairly concrete. They are not systems for automatically producing themes, but they are not trivial text-processing tools either. Their strongest role appears earlier in the workflow: helping researchers open up the material, generate a pool of candidate codes, and bring forward expressions that might otherwise remain marginal. In sensitive healthcare research, open-source models such as Gemma2 and Llama3.1 also offer a practical advantage: researchers can experiment with LLM-assisted analysis without uploading patient interview data to internet-based tools.


The difficulty is that candidate codes are still far from themes. Misra et al.’s results repeatedly show that model outputs have to be examined again by human researchers: which codes actually carry context, which ones are merely repetitions of similar concepts, and which ones may relate to diabetes or self-management but still fail to answer the specific question of patient-provider communication. Gemma2 and Llama3.1 performed better under the deductive framework, but their good-fit rates were still only 27.1% and 26.8%; under the inductive framework, the rates were lower, at 19.2% and 11.7%. These figures are not just performance metrics. They point to a methodological risk: the readability of LLM-generated codes can easily be mistaken for interpretive adequacy.


In that sense, LLMs bring researchers back to the text rather than freeing them from it. They can expose peripheral experiences again, such as distrust of healthcare providers, financial strain, lack of emotional support, fear of discussing diabetes, or feeling abandoned by providers. Yet whether these expressions are strong enough to become subthemes, whether they remain isolated moments, and whether they should reshape the original thematic structure still depends on repeated checking against the interview transcripts. The central work of qualitative analysis still begins after code generation: returning to narrative, comparing context, judging meaning, and deciding which experiences deserve to be developed into themes.

 

 

 

References:

Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa

 

Braun, V., & Clarke, V. (2019). Reflecting on reflexive thematic analysis. Qualitative Research in Sport, Exercise and Health, 11(4), 589–597. https://doi.org/10.1080/2159676X.2019.1628806

 

Castleberry, A., & Nolen, A. (2018). Thematic analysis of qualitative research data: Is it as easy as it sounds? Currents in Pharmacy Teaching and Learning, 10(6), 807–815. https://doi.org/10.1016/j.cptl.2018.03.019

 

Lincoln, Y. S., & Guba, E. G. (1986). But is it rigorous? Trust worthiness and authenticity in naturalistic evaluation. New Directions for Program Evaluation, 1986(30), 73–84. https:// doi.org/10.1002/ev.1427

 

Misra, R., Dahal, R., Kirk, B., Khan, R., Dogan, G., Chataut, R., & Gyawali, P. (2026). Large Language Models in Qualitative Analysis: Comparing Traditional and Researcher-Interpreted Approaches. International Journal of Qualitative Methods, 25. https://doi.org/10.1177/16094069261426100

 

Terry, G., Hayfield, N., Clarke, V., & Braun, V., (2017). Thematic analysis. The SAGE Handbook of Qualitative Research in Psychology, 2(17-37), Article 25. https://doi.org/10.4135/ 9781526405555.n2.

 
 
 

Comments


This initiative is supported by the following organizations:

  • Twitter
  • LinkedIn
  • YouTube
logo_edited.png
bottom of page