This is an outdated version published on 2026-07-31. Read the most recent version.

Implementation of mDeBERTa v3 for Zero-Shot Classification of AI-Generated Text Scientific Abstracts

Authors

  • Aldo Universitas Jenderal Achmad Yani Yogyakarta
  • Aris Wahyu Murdiyanto Universitas Jenderal Achmad Yani Yogyakarta https://orcid.org/0000-0002-9829-753X
  • Ulfi Saidata Aesyi Universitas Jenderal Achmad Yani Yogyakarta

DOI:

https://doi.org/10.56705/ijodas.v7i2.457

Keywords:

AI Text Detection;, IndoT5, mDeBERTa v3, Scientific Abstract, Zero-Shot Classification

Abstract

Human evaluators face difficulties distinguishing original scientific abstracts from AI generated text, as the formal language produced by AI appears highly neat and convincing. Previous studies have proven that human reviewers correctly identified only 68% of AI-generated abstracts, while 14% of original human texts were mistakenly classified as machine-generated. This study evaluates the performance of the mDeBERTa v3 model using a zero-shot classification approach based on Natural Language Inference (NLI) to classify human-written and AI-synthesized Indonesian scientific abstracts without model retraining. A balanced dataset of 2,274 abstracts was constructed from SINTA 3 journal scraping published before 2018 as the human class, and IndoT5-base-paraphrase outputs as the AI class. Mann-Whitney U statistical analysis on seven linguistic features revealed significant differences (p = 0.00) across all features. A critical anomaly was discovered, AI texts from IndoT5-base-paraphrase exhibit higher sentence length variation (mean SD = 15.44) than human texts (mean SD = 7.99), contradicting the theoretical assumption that AI text is more uniform. This phenomenon results from context window exhaustion in IndoT5 when processing information-dense scientific abstracts, producing semantic hallucinations such as substituting the research object "chili pepper" with "pickles" and distorting location names. Testing three NLI instruction scenarios revealed that a single instruction targeting sentence fluctuation anomalies (Scenario 1) achieved the highest Recall of 76.52%, but triggered 790 False Positives, limiting overall accuracy to 53.52%. Adding instructional complexity degraded performance due to AI formal vocabulary bias. These results indicate that zero-shot NLI effectively tracks AI mechanical anomalies with a single targeted instruction.

Downloads

Download data is not yet available.

References

References:[1] J. G. Meyer et al., “ChatGPT and large language models in academia: opportunities and challenges,” BioData Min., vol. 16, no. 1, pp. 1–11, 2023, doi: 10.1186/s13040-023-00339-9.

[2] D. R. E. Cotton, P. A. Cotton, and J. R. Shipway, “Chatting and cheating: Ensuring academic integrity in the era of ChatGPT,” Innov. Educ. Teach. Int., vol. 61, no. 2, pp. 228–239, 2024, doi: 10.1080/14703297.2023.2190148.

[3] F. M. Howard, A. Li, M. F. Riffon, and E. Garrett-mayer, “Characterizing the Increase in Artificial Intelligence Content Detection in Oncology Scienti fi c Abstracts From 2021 to 2023,” 2023, doi: 10.1200/CCI.24.00077.

[4] C. A. Gao et al., “Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers,” npj Digit. Med., vol. 6, p. 75, 2023, doi: 10.1038/s41746-023-00819-6.

[5] O. Marsh, M. Yates, A. Rowe, and M. Song, “Adversarial Attacks and Robustness in LLM- Generated Text Detection Systems,” 2025.

[6] A. Alikhanov et al., “AI Generated Text Detection,” 2026, [Online]. Available: http://arxiv.org/abs/2601.03812

[7] P. He, J. Gao, and W. Chen, “Debertav3: Improving Deberta Using Electra-Style Pre-Training With Gradient-Disentangled Embedding Sharing,” 11th Int. Conf. Learn. Represent. ICLR 2023, no. Mlm, pp. 1–16, 2023.

[8] W. Yin, J. Hay, and D. Roth, “Benchmarking Zero-shot Text Classification : Datasets , Evaluation and Entailment Approach,” pp. 3914–3923, 2019.

[9] E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn, “DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature,” Proc. Mach. Learn. Res., vol. 202, pp. 24950–24962, 2023.

[10] W. X. Zhao et al., “A Survey of Large Language Models,” Mar. 2025, [Online]. Available: http://arxiv.org/abs/2303.18223

[11] V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, and S. Feizi, “Can AI-Generated Text be Reliably Detected?,” pp. 1–37, 2025, [Online]. Available: http://arxiv.org/abs/2303.11156

[12] M. Hahn, “Theoretical Limitations of Self-Attention in Neural Sequence Models,” vol. 8, pp. 156–171, 2020.

[13] J. Han, M. Kamber, and J. Pei, Data Mining: Concepts and Techniques, 3rd ed. Morgan Kaufmann, 2011.

[14] L. Lukman et al., “Proposal of the S-score for measuring the performance of researchers, institutions, and journals in Indonesia,” Sci. Ed., vol. 5, no. 2, pp. 135–141, 2018, doi: 10.6087/KCSE.138.

[15] H. de Sompel and M. L. Nelson, “Reminiscing about 15 years of interoperability efforts,” D-Lib Mag., vol. 21, no. 11/12, 2015, doi: 10.1045/november2015-vandesompel.

[16] A. Radford et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019.

[17] A. Field, Discovering statistics using IBM SPSS statistics. Sage publications limited, 2024.

[18] M. Laurer, W. Van Atteveldt, A. Casas, and K. Welbers, “Less Annotating , More Classifying : Addressing the Data Scarcity Issue of Supervised Machine Learning with Deep Transfer Learning and,” 2023, doi: 10.1017/pan.2023.20.

[19] N. F. Liu, K. Lin, J. Hewitt, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the Middle : How Language Models Use Long Contexts,” 2023.

[20] A. Conneau et al., “Unsupervised cross-lingual representation learning at scale,” in Proceedings of the 58th annual meeting of the association for computational linguistics, 2020, pp. 8440–8451.

[21] A. Zhang, Z. C. Lipton, M. Li, and A. J. Smola, Dive into Deep Learning. Cambridge University Press, 2023.

[22] K. M. Sujon, R. Hassan, K. Choi, and M. A. Samad, “Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models,” J. Big Data, vol. 12, no. 1, 2025, doi: 10.1186/s40537-025-01313-4.

[23] S. Sathyanarayanan, “Confusion Matrix-Based Performance Evaluation Metrics,” African J. Biomed. Res., vol. 27, no. 4, pp. 4023–4031, 2024, doi: 10.53555/ajbr.v27i4s.4345.

Published

2026-07-31

Versions

How to Cite

Implementation of mDeBERTa v3 for Zero-Shot Classification of AI-Generated Text Scientific Abstracts. (2026). Indonesian Journal of Data and Science, 7(2). https://doi.org/10.56705/ijodas.v7i2.457