Cambridge, UK – September 12, 2026
A new analysis published in Nature Human Behaviour reveals that nearly half of social science research papers claim cause and effect relationships that their study designs cannot establish, with the practice of using causal language in correlational research tripling over the past two decades.
The study, conducted by researchers Christian Isch, Tobias Dörr, Neil Fasching, Grace Jennings, and Duncan J. Watts, analyzed 194,631 cross-sectional social science papers published between 1980 and 2024. The researchers found that 46.3% of these papers used causal language in their titles or abstracts—phrases like "leads to," "causes," or "affects"—despite using study designs that can only demonstrate associations, not causation.
Study Design and Scale
The research team faced a formidable task: manually reviewing nearly 200,000 articles would be impractical. Instead, they developed machine learning classifiers trained on expert human judgments to automatically detect causal language in article titles and abstracts. The classifiers were validated against human annotators to ensure accuracy.
To understand the real-world impact of such language, the researchers conducted a separate experiment with 1,105 college-educated adults. Participants read abstracts from cross-sectional studies, with some seeing the original abstract containing causal language, others seeing a rewritten version describing findings only as associations, and a third group receiving the original abstract plus a note explaining that cross-sectional designs cannot establish causality.
"Readers frequently indicated that abstracts with causal phrasing provided causal evidence, but methodological labels like 'cross-sectional' reduced this effect, though not enough to eliminate it entirely," the study authors write.
Key Findings
The analysis revealed several important patterns:
Prevalence Over Time: The annual rate of causal language has risen almost threefold since 2000, from 20% to 60% by 2024. This suggests the problem is not simply a legacy of older publishing norms but an ongoing trend in how social scientists communicate their findings.
Study Design Mismatch: Cross-sectional designs—common in social science research—reveal associations between variables at a single point in time but cannot establish that manipulating one variable will produce changes in another. For example, a study finding that people who exercise more report higher happiness cannot determine whether exercise causes happiness, whether happy people exercise more, or whether a third factor like good health influences both.
Reader Impact: The human-subjects experiment showed that causal language meaningfully affects how readers interpret evidence. Participants exposed to abstracts with causal phrasing were more likely to believe the research established cause and effect, even when the underlying design could not support such conclusions.
Why Causal Language Matters
The distinction between correlation and causation is not merely academic. Causal claims shape public policy, influence educational and clinical interventions, guide media coverage, and determine which social problems receive funding.
When a policy is justified by a causal claim that rests on correlational evidence, interventions may fail to produce the intended effects. Public trust in science can also erode when confident predictions do not materialize or when interventions based on scientific research do not deliver promised outcomes.
The study builds on earlier work concerned with scientific communication integrity. Tal Yarkoni's "generalizability crisis" argued that psychological research often makes claims disconnected from empirical evidence. Duncan Watts has similarly noted that sociological explanations frequently conflate what is intuitively understandable with what is causally established. A 2018 review by Isabelle Boutron and Philippe Ravaud documented systematic misrepresentation in biomedical research, showing that the problem is not unique to the social sciences.
Why Causal Overclaiming Persists
The researchers point to several factors that may drive the trend:
-
Publication and media incentives: Journals and media outlets tend to favor findings that sound decisive. Hedged language can seem weak or unremarkable, while causal claims attract more attention.
-
Pressure to demonstrate significance: Researchers face pressure to show the real-world significance of their work, which can lead to stronger claims about what the evidence means.
-
Complex research challenges: Social science research often deals with high-complexity problem spaces where causal identification is genuinely difficult. Researchers cannot randomly assign people to conditions like poverty, discrimination, or specific life experiences.
"The underlying measurement and design problems in social science are genuinely difficult," the authors note. "In that environment, a modest correlational result can be nudged, sentence by sentence, into a causal story."
The Role of AI in Scientific Communication
The study also examined how large language models—increasingly used to summarize and disseminate research—respond to causal language. The researchers found that AI systems trained or prompted with overreaching causal statements inherited and reproduced the exaggerated interpretations.
Given that language models increasingly mediate public access to scientific knowledge by summarizing papers, answering questions, and generating secondary content, this mechanism could amplify causal overclaiming well beyond the original readership of any single article.
Potential Solutions
The researchers suggest several approaches to address the issue:
-
Pre-publication detection: Automated systems could screen manuscripts for causal language that exceeds what the study design justifies, flagging papers for editorial review.
-
Author guidance: Journals could provide clearer guidance on language appropriate to different study designs, helping authors describe their findings accurately.
-
AI guardrails: Systems summarizing research could be programmed to preserve the evidential caution of the original work, preventing the amplification of overstated claims.
-
Reader education: Media and science communicators could help audiences understand the limitations of different research designs and what claims each can legitimately support.
Implications for Research and Policy
The study's message is ultimately corrective rather than condemnatory. The social sciences can produce valuable knowledge even from imperfect designs, provided the language used to describe that knowledge stays within the bounds of what the evidence supports.
The researchers note that faithful science communication depends on the claim strength of a summary matching the strength of the underlying evidence. This standard, codified in norms for the science of science communication, which the new study suggests is routinely violated.
As scientific publishing continues to accelerate and AI-mediated reading becomes more common, the gap between what research actually demonstrates and what it claims to demonstrate may widen further. The new analysis provides both a warning about the current state of research communication and a roadmap for addressing it using the same computational tools that helped uncover the problem.
Sources: Isch, C., Dörr, T., Fasching, N., Jennings, G., & Watts, D. J. (2026). Quantifying the prevalence and impact of overreaching causal claims in social science. Nature Human Behaviour. . Scienmag, September 12, 2026. Nature Human Behaviour, September 1, 2026.
FIRAT Editorial Board
Institutional Research Desk · Foresight Institute of Research and Translation
The collective editorial and research translation board of FIRAT, synthesising peer-reviewed evidence, policy briefs, and division milestones across our seven foundational research pillars.

