What do I write when two studies on the same question reach opposite conclusions?

Treat the disagreement as your finding rather than an obstacle. State that the field is divided, then explain the division by comparing what the studies actually did: population, measure, time frame, analysis. Most contradictions in a literature dissolve into a boundary condition once you line the designs up.

Updated

The instinct is to hide the conflict, because a review that ends in "the evidence is mixed" feels like a failure to reach a conclusion. It is the opposite. A field where everything agrees needs no new study. Contradiction is where the research questions live, and an examiner reading your chapter is watching for whether you noticed.

Start by checking that the studies really do contradict each other. A surprising share of apparent conflicts are not conflicts at all. One study reports a statistically significant effect and another reports a non-significant one, which is not the same as one finding an effect and the other finding none. One measures attitudes and the other measures behavior. One follows participants for six months and the other for four years. Two results can both be correct and still point in opposite directions.

There is a published list of the differences worth checking. Carlson and colleagues name nation, design, intervention, subject population, study quality, enrollment size, and unusually favorable or influential outlier results as sources of variation between studies on one question (Carlson et al., 2023, p. 10). They also draw a line you should respect in your wording: saying the studies disagree describes the results, while saying the disagreement reflects differences in populations or methods is an interpretation you have to support from the designs themselves (Carlson et al., 2023, pp. 5-6).

When the conflict survives that check, put the designs next to each other and look for the difference that would explain it. Sample composition explains a great many. So does the outcome measure, the analysis strategy, the country, the decade, and whether the researchers had a stake in the result. Write the comparison out explicitly, because the reader cannot see the papers you are holding.

If a meta-analysis covers your question, read its heterogeneity reporting before you describe the conflict in your own words. Carlson and colleagues define statistical heterogeneity as variation among studies in the observed effects or in whether results reached significance, and they tell readers to inspect the Forest plot for whether the individual 95% confidence interval bars overlap and cross the line through the summary diamond (Carlson et al., 2023, p. 6). Where heterogeneity is substantial, a random effects model may fit better than a fixed effects one, and the authors may decide not to pool the studies at all (Carlson et al., 2023, p. 7). A refusal to pool is information about the field, and you can cite it as such.

The strongest paragraph you can write here does three things: it names the contradiction, it offers a candidate explanation grounded in the designs, and it says what evidence would decide the matter. That third move is often where your own study comes from.

Do conflicting studies belong in the literature review or the discussion?

Both, doing different jobs. In the review, the conflict is part of the state of the field and often the reason your study exists. In the discussion, you return to it and say what your own results imply about which account holds, or under what conditions each one does.

Keeping them separate matters because the two passages have different scope. The review states the disagreement as a fact about the literature and does not resolve it. The discussion is allowed to argue, because by then you have your own data on the table.

A common weakness is a review that resolves the conflict early, so the discussion has nothing left to do except repeat itself. Leave the question genuinely open in the chapter where it belongs open.

How do I explain contradictory results without declaring a winner?

Describe the conditions under which each result was obtained and let the reader see the difference. Write that an effect appears in one setting and not another, and name what separates them. You are allowed to say the evidence does not currently settle the question.

The phrasing that does this work is conditional rather than evaluative. "Reported in adolescent samples but not in adult ones" is a finding. "Alvarez is more convincing than Fenton" is an opinion you have not earned unless you can point to a specific methodological reason and state it.

If one study is genuinely stronger, say why in terms a reader can check: larger sample, preregistered analysis, longer follow-up, independent replication. That is a judgment supported by evidence rather than a preference.

Design strength is a matter of degree, and Ioannidis puts numbers on it. Under those modelling assumptions, an adequately powered randomized controlled trial testing an intervention with a 50 percent pre-study chance of working is true about 85 percent of the time, an underpowered early phase trial with bias can be true one time in four or less, a well powered exploratory epidemiological finding has about a one in five chance of being true when true relationships are outnumbered ten to one, and a meta-analysis that pools inconclusive studies to compensate for low power is probably false at a ratio of 1 to 3 (Ioannidis, 2005, pp. 699-700). Those figures are illustrative rather than universal, so use them to justify phrases like the evidence remains uncertain or the result requires replication, not to declare either study false. Ioannidis argues for the same reason that conclusions should reflect the totality of evidence rather than one team's significant result, and that even established findings can fail confirmation in large replication studies (Ioannidis, 2005, p. 701).

If your studies are qualitative, the comparison runs on context rather than effect sizes. Thomas and Harden keep each study's aims, methods, methodological quality, setting, and sample visible in the synthesis so a reader can judge whether a theme transfers (Thomas and Harden, 2008, pp. 7-8). They also recommend attending to agreement and disagreement between studies, deliberately looking for negative cases, and seeking maximum variability, while noting that the practical use of these principles is still uncertain (Thomas and Harden, 2008, p. 3). In their own review they compared boys and girls and found no theme belonging only to one group, but they report that thin contextual reporting in the source papers made that judgment hard (Thomas and Harden, 2008, pp. 7-8). If the papers you are comparing say little about their settings, say that too, because it limits what any contradiction between them can mean.