Zero-Shot Detection of IndoT5-Synthesized Indonesian Scientific Abstracts Using mDeBERTa v3

Authors

  • Aldo Universitas Jenderal Achmad Yani Yogyakarta
  • Aris Wahyu Murdiyanto Universitas Jenderal Achmad Yani Yogyakarta https://orcid.org/0000-0002-9829-753X
  • Ulfi Saidata Aesyi Universitas Jenderal Achmad Yani Yogyakarta

DOI:

https://doi.org/10.56705/ijodas.v7i2.457

Keywords:

AI Text Detection;, IndoT5, mDeBERTa v3, Scientific Abstract, Zero-Shot Classification

Abstract

 Introduction: Distinguishing human-written scientific abstracts from AI-synthesized text remains challenging, particularly when machine-generated language appears fluent and formally structured. This study evaluates mDeBERTa v3 in a zero-shot Natural Language Inference (NLI) setting for detecting Indonesian scientific abstracts specifically synthesized using IndoT5-base-paraphrase. Method: A balanced dataset of 2,274 abstracts comprising 1,137 human-written abstracts from SINTA 3 journals and 1,137 IndoT5-synthesized counterparts was analyzed. Seven linguistic features were examined using the Mann–Whitney U test, followed by zero-shot mDeBERTa v3 classification using one-, three-, and five-aspect NLI instruction scenarios. A Random Forest classifier using the same linguistic features was included as a supervised baseline. Results and Discussion: All seven linguistic features differed significantly between classes (p < 0.001), with AI texts showing substantially higher sentence-length variation than human texts. The targeted one-aspect NLI scenario achieved the highest recall of 76.52% but only 53.52% accuracy because 790 human abstracts were misclassified as AI. Increasing instruction complexity further reduced recall. In contrast, Random Forest achieved 91.21% accuracy and an F1-score of 0.9130, confirming that the identified linguistic anomalies are strong learnable signals. Conclusion: Zero-shot mDeBERTa v3 can detect generator-specific structural artifacts but remains insufficiently precise for standalone academic-integrity screening and should be supplemented by supervised methods and human review.

Downloads

Download data is not yet available.

References

[1] J. G. Meyer et al., “ChatGPT and large language models in academia: opportunities and challenges,” BioData Min., vol. 16, no. 1, pp. 1–11, 2023, doi: https://doi.org/10.1186/s13040-023-00339-9.

[2] D. R. E. Cotton, P. A. Cotton, and J. R. Shipway, “Chatting and cheating: Ensuring academic integrity in the era of ChatGPT,” Innov. Educ. Teach. Int., vol. 61, no. 2, pp. 228–239, 2024, doi: https://doi.org/10.1080/14703297.2023.2190148.

[3] F. M. Howard, A. Li, M. F. Riffon, and E. Garrett-mayer, “Characterizing the Increase in Artificial Intelligence Content Detection in Oncology Scienti fi c Abstracts From 2021 to 2023,” 2023, doi: https://doi.org/10.1200/CCI.24.00077.

[4] P. C. Theocharopoulos, P. Anagnostou, A. Tsoukala, S. V. Georgakopoulos, S. K. Tasoulis, and V. P. Plagianakos, “Detection of Fake Generated Scientific Abstracts,” Apr. 2023, doi: https://doi.org/10.1109/BigDataService58306.2023.00011.

[5] C. A. Gao et al., “Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers,” npj Digit. Med., vol. 6, p. 75, 2023, doi: https://doi.org/10.1038/s41746-023-00819-6.

[6] O. Marsh, M. Yates, A. Rowe, and M. Song, “Adversarial Attacks and Robustness in LLM- Generated Text Detection Systems,” 2025.

[7] H. Desaire, A. E. Chua, M. Isom, R. Jarosova, and D. Hua, “Distinguishing academic science writing from humans or ChatGPT with over 99% accuracy using off-the-shelf machine learning tools,” Cell Reports Phys. Sci., vol. 4, no. 6, 2023.

[8] A. Alikhanov et al., “AI Generated Text Detection,” Jan. 2026

[9] P. He, J. Gao, and W. Chen, “Debertav3: Improving Deberta Using Electra-Style Pre-Training With Gradient-Disentangled Embedding Sharing,” 11th Int. Conf. Learn. Represent. ICLR 2023, no. Mlm, pp. 1–16, 2023.

[10] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio, Eds., Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4171–4186. doi: https://doi.org/10.18653/v1/N19-1423.

[11] W. Yin, J. Hay, and D. Roth, “Benchmarking Zero-shot Text Classification : Datasets , Evaluation and Entailment Approach,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China, 2019, pp. 3914–3923.

[12] E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn, “DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature,” Proc. Mach. Learn. Res., vol. 202, pp. 24950–24962, 2023.

[13] W. X. Zhao et al., “A Survey of Large Language Models,” Mar. 2026.

[14] C. Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” J. Mach. Learn. Res., vol. 21, pp. 1–67, Sep. 2023.

[15] V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, and S. Feizi, “Can AI-Generated Text be Reliably Detected?,” pp. 1–37, Jan. 2025.

[16] M. Hahn, “Theoretical Limitations of Self-Attention in Neural Sequence Models,” Trans. Assoc. Comput. Linguist., vol. 8, pp. 156–171, Dec. 2020, doi: https://doi.org/10.1162/tacl_a_00306.

[17] J. Han, M. Kamber, and J. Pei, Data Mining: Concepts and Techniques, 3rd ed. Morgan Kaufmann, 2011.

[18] L. Lukman et al., “Proposal of the S-score for measuring the performance of researchers, institutions, and journals in Indonesia,” Sci. Ed., vol. 5, no. 2, pp. 135–141, 2018, doi: https://doi.org/10.6087/KCSE.138.

[19] H. de Sompel and M. L. Nelson, “Reminiscing about 15 years of interoperability efforts,” D-Lib Mag., vol. 21, no. 11/12, 2015, doi: https://doi.org/10.1045/november2015-vandesompel.

[20] A. Radford et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019.

[21] M. Freitag and Y. Al-Onaizan, “Beam Search Strategies for Neural Machine Translation,” in Proceedings of the First Workshop on Neural Machine Translation, T. Luong, A. Birch, G. Neubig, and A. Finch, Eds., Stroudsburg, PA, USA: Association for Computational Linguistics, Aug. 2017, pp. 56–60. doi: https://doi.org/10.18653/v1/W17-3207.

[22] N. S. Keskar, B. McCann, L. R. Varshney, C. Xiong, and R. Socher, “CTRL: A Conditional Transformer Language Model for Controllable Generation,” pp. 1–18, Sep. 2019.

[23] R. Paulus, C. Xiong, and R. Socher, “A Deep Reinforced Model for Abstractive Summarization,” CoRR, vol. abs/1705.0, Nov. 2017.

[24] Y. Wu et al., “Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation,” CoRR, vol. abs/1609.0, Oct. 2016.

[25] A. K. Uysal and S. Gunal, “The impact of preprocessing on text classification,” Inf. Process. Manag., vol. 50, no. 1, pp. 104–112, 2014, doi: https://doi.org/10.1016/j.ipm.2013.08.006.

[26] B. Krawczyk, “Learning from imbalanced data : open challenges and future directions,” Prog. Artif. Intell., vol. 5, no. 4, pp. 221–232, 2016, doi: https://doi.org/10.1007/s13748-016-0094-0.

[27] A. Field, Discovering statistics using IBM SPSS statistics. Sage publications limited, 2024.

[28] M. Mars, “From Word Embeddings to Pre-Trained Language Models: A State-of-the-Art Walkthrough,” Appl. Sci., vol. 12, no. 17, 2022, doi: https://doi.org/10.3390/app12178805.

[29] M. Laurer, W. van Atteveldt, A. Casas, and K. Welbers, “Less Annotating, More Classifying: Addressing the Data Scarcity Issue of Supervised Machine Learning with Deep Transfer Learning and BERT-NLI,” Polit. Anal., vol. 32, no. 1, pp. 84–100, Jan. 2024, doi: https://doi.org/10.1017/pan.2023.20.

[30] N. F. Liu, K. Lin, J. Hewitt, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the Middle : How Language Models Use Long Contexts,” 2023.

[31] I. V Pantic and S. Mugosa, “Artificial intelligence strategies based on random forests for detection of AI-generated content in public health,” Public Health, vol. 242, pp. 382–387, May 2025, doi: https://doi.org/10.1016/j.puhe.2025.03.029.

[32] S. K. Aityan, W. Claster, K. S. Emani, S. Rais, and T. Tran, “A Lightweight Approach to Detection of AI-Generated Texts Using Stylometric Features,” Jan. 2026.

[33] P. He, X. Liu, J. Gao, and W. Chen, “Deberta: Decoding-Enhanced Bert With Disentangled Attention,” ICLR 2021 - 9th Int. Conf. Learn. Represent., 2021.

[34] A. Conneau et al., “Unsupervised cross-lingual representation learning at scale,” in Proceedings of the 58th annual meeting of the association for computational linguistics, 2020, pp. 8440–8451.

[35] A. Zhang, Z. C. Lipton, M. Li, and A. J. Smola, Dive into Deep Learning. Cambridge: Cambridge University Press, 2023.

[36] K. M. Sujon, R. Hassan, K. Choi, and M. A. Samad, “Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models,” J. Big Data, vol. 12, no. 1, 2025, doi: https://doi.org/10.1186/s40537-025-01313-4.

[37] S. Sathyanarayanan, “Confusion Matrix-Based Performance Evaluation Metrics,” African J. Biomed. Res., vol. 27, no. 4, pp. 4023–4031, 2024, doi: https://doi.org/10.53555/ajbr.v27i4s.4345

Downloads

Published

2026-07-31 — Updated on 2026-07-31

Versions

How to Cite

Zero-Shot Detection of IndoT5-Synthesized Indonesian Scientific Abstracts Using mDeBERTa v3. (2026). Indonesian Journal of Data and Science, 7(2), 333-349. https://doi.org/10.56705/ijodas.v7i2.457