A Comparative Analysis of the Lexical Diversity of the Introduction Section in AI-Generated vs Human-Authored Linguistic Articles
Abstract
Lexical diversity is a core indicator of academic writing quality. Despite the rapid growth of AI-assisted academic writing, limited work has compared lexical diversity in advanced linguistic articles by examining AI-generated and human-written introductions through matched quantitative indices. This gap is important because lexical diversity can help educators, researchers, and journal stakeholders understand how AI-generated academic prose differs from human-written prose in vocabulary use and information density. Accordingly, this study aimed to compare the lexical diversity and lexical density of five paired introduction sections written by ChatGPT-4 and human authors, to determine whether GPT-4 matched or exceeded human writing in lexical behaviour, to examine whether RTTR and the Maas index remained consistent across three preprocessing conditions, and to evaluate whether lexical density functioned as a clear discriminator between the two text types. For the five paired introductions used in this study, ChatGPT-generated texts tended to present greater lexical density and overall greater lexical diversity than human-written texts. However, lexical diversity results varied across RTTR and Maas index values, suggesting that authorship differences should be interpreted through multiple measures rather than a single index.
Keywords: academic discourse; comparative linguistics; large language models; lexical richness; machine-generated text
Full Text:
PDFReferences
Aldosari, L. A., & Altuwairesh, N. (2026). Assessing Legal Translations Generated by GPT-4 Turbo Using MQM: A Comparative Study. 3L: Language, Linguistics, Literature® The Southeast Asian Journal of English Language Studies, 32(1), 172-186.
http://doi.org/10.17576/3L-2026-3201-11
Alheety, A. A., khalaf, M. K., & Mohammed, H. J. (2026). Syntactic Complexity in AI-Generated vs. Human-Authored Linguistic and Literary Texts. Arab World English Journal (AWEJ) Special Issue on CALL 17. (12) 109-125. DOI: https://dx.doi.org/10.24093/awej/call12.7
Biber, D., Conrad, S., & Reppen, R. (1998). Corpus linguistics: Investigating language structure and use. Cambridge University Press.
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S. M., Nori, H., Palangi, H., Ribeiro, M., & Zhang, Y. (2023). Sparks of Artificial General Intelligence: Early experiments with GPT-4. ArXiv, abs/2303.12712.
Dwivedi, Y. K., Kshetri, N., Hughes, L., Slade, E. L., Jeyaraj, A., Kar, A. K., Baabdullah, A. M., Koohang, A., Raghavan, V., Ahuja, M., Albanna, H., Albashrawi, M. A., Al-Busaidi, A. S., Balakrishnan, J., Barlette, Y., Basu, S., Bose, I., Brooks, L., Buhalis, D., … Wright, R. (2023).
Opinion Paper: "So what if ChatGPT wrote it?" Multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy. Int. J. Inf. Manag., 71, 102642.
Guo, B., Zhang, X., Wang, Z., Jiang, M., Nie, J., Ding, Y., Yue, J., & Wu, Y. (2023). How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection. ArXiv, abs/2301.07597.
Hamat, A. (2024). The language of AI and human poetry: A comparative lexicometric study. 3L: Language, Linguistics, Literature® The Southeast Asian Journal of English Language Studies, 30(2), 1-20.
Herbold, S., Hautli-Janisz, A., Heuer, U., Kikteva, Z., & Trautsch, A. (2023). AI, write an essay for me: A large-scale comparison of human-written versus ChatGPT-generated essays. ArXiv, abs/2304.14276.
Hout, R. V., & Vermeer, A. (2007). Lexical richness and the language proficiency of native and non-native speakers. John Benjamins Publishing Company.
Hyland, K. (2004). Disciplinary discourses: Social interactions in academic writing. University of Michigan Press.
Iskender, A. (2023). Holy or unholy? ChatGPT and the future of universities. World Journal of English Language, 13(4), 1-13.
Jarvis, S., & Daller, M. (2013). Vocabulary knowledge: Human ratings and automated measures. John Benjamins Publishing Company.
Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274.
Khalil, M., & Er, E. (2023). Will ChatGPT get you a PhD? On the creative abilities of large language models. arXiv preprint arXiv:2302.05287.
Laufer, B., & Nation, P. (1995). Vocabulary size and use: Lexical richness in L2 written production. Applied Linguistics, 16(3), 307-322.
McEnery, T., & Hardie, A. (2012). Corpus linguistics: Method, theory and practice. Cambridge University Press.
McNamara, D. S., Graesser, A. C., Kurby, C. A., & Louwerse, M. M. (2014). The Coh-Metrix Common Core Text Ease and Readability Assessor. Cambridge University Press.
Nation, I. S. (2013). Learning vocabulary in another language. Cambridge University Press.
Nor, N. F. M., & Aziz, J. (2026). ChatGPT as a Tool in Developing Research and Language Skills in a Research Methodology Course. 3L: Language, Linguistics, Literature® The Southeast Asian Journal of English Language Studies, 32(1), 187-200.
OpenAI. (2023). GPT-4 technical report. arXiv.
Ray, P. P. (2023). ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet of Things and Cyber-Physical Systems, 3, 121-154.
Schmitt, N. (2010). Researching vocabulary: A vocabulary research manual. Palgrave Macmillan.
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., & Anderson, R. (2023). The Curse of Recursion: Training on Generated Data Makes Models Forget. ArXiv, abs/2305.17493.
Tweedie, F. J., & Baayen, R. H. (1998). How variable may a constant be? Measures of lexical richness in perspective. Computers and the Humanities, 32(5), 323-352.
Wu, T., He, S., Liu, J., Sun, S., Liu, K., Han, Q.-L., & Tang, Y. (2023). A brief overview of ChatGPT: The history, status quo and future. IEEE/CAA Journal of Automatica Sinica, 10(5), 1122-1136.
Zhai, X. (2022). ChatGPT user experience: Implications for education. SSRN Electronic Journal, 1-18. https://doi.org/10.2139/ssrn.4312418
DOI: http://dx.doi.org/10.17576/3L-2026-3203-14
Refbacks
- There are currently no refbacks.
eISSN : 2550-2247
ISSN : 0128-5157