
Overview: The ability to communicate statistical results to domain experts and stakeholders is an important goal of the undergraduate statistics and data science curriculum. However, as large language models (LLMs) have become more accessible, a major concern is that students are offloading important cognitive tasks to generative AI. Using a corpus of over 1,600 undergraduate students’ data analysis reports from 2021 to 2025, we show how students’ writing style and verb usage have become more similar to that of LLMs. This shift is most pronounced in the first and fifth quintiles of students’ reports, which roughly map onto the introduction and conclusion sections, respectively. At the same time, we demonstrate that students’ writing style has become more similar to that of statistics experts with the addition of LLMs. We end by discussing the implications of our findings for statistics and data science educators. In particular, we propose alternative modes of assessment that still emphasize statistical thinking, such as targeted writing assignments for structuring a report introduction.
Colando, S.*, Franke, E.*, Weinberg, G., and Reinhart, A. Analyzing students’ statistics writing before and after the emergence of large language models, submitted, 2026.
Versions of this work have also been presented at various statistics education conferences:
USCOTS 2025 Research Satellite as a poster
ECOTS 2026 as a breakout session
ICOTS 12 as both a main topic session talk and a contributed session talk
* denotes equal contribution.