Zero-Shot Detection of IndoT5-Synthesized Indonesian Scientific Abstracts Using mDeBERTa v3

Authors

  • Aldo Universitas Jenderal Achmad Yani Yogyakarta
  • Aris Wahyu Murdiyanto Universitas Jenderal Achmad Yani Yogyakarta https://orcid.org/0000-0002-9829-753X
  • Ulfi Saidata Aesyi Universitas Jenderal Achmad Yani Yogyakarta

DOI:

https://doi.org/10.56705/ijodas.v7i2.457

Keywords:

AI Text Detection;, IndoT5, mDeBERTa v3, Scientific Abstract, Zero-Shot Classification

Abstract

Human evaluators struggle to distinguish original scientific abstracts from AI-generated text, as AI-produced formal language appears neat and convincing; prior studies report reviewers correctly identify only 68% of AI-generated abstracts while misclassifying 14% of human texts. This study presents an exploratory, generator-specific evaluation of mDeBERTa v3 using zero-shot Natural Language Inference (NLI) classification, applied to Indonesian scientific abstracts synthesized via IndoT5-base-paraphrase rather than AI-generated text in general. A balanced 2,274-abstract dataset paired human abstracts (SINTA 3 journals) with IndoT5-base-paraphrase outputs as the AI class. Mann-Whitney U analysis on seven linguistic features revealed significant differences (p < 0.001) across all. A critical anomaly emerged: AI texts showed higher sentence-length variation (SD = 15.44) than human texts (SD = 7.99), contradicting the assumption that AI text is more uniform, attributable to context-window exhaustion in IndoT5 producing semantic hallucinations when synthesizing dense abstracts. Testing three NLI scenarios showed a single instruction targeting this fluctuation achieved the highest Recall (76.52%) but with 790 false positives among 1,137 human abstracts, limiting accuracy to 53.52%; added complexity further degraded AI-class recall due to vocabulary overlap. A Random Forest classifier trained on the same features achieved 91.21% accuracy (F1 = 0.9130), substantially outperforming the zero-shot approach and confirming the anomaly as a strong, learnable signal. These results indicate zero-shot NLI can partially track a generator's mechanical artifacts through a single targeted instruction, but remains insufficiently precise to separate machine-error fluctuation from natural human variation, and is not recommended for standalone academic-integrity screening without further refinement

Downloads

Download data is not yet available.

References

[1] J. G. Meyer et al., “ChatGPT and large language models in academia: opportunities and challenges,” BioData Min., vol. 16, no. 1, pp. 1–11, 2023, doi: https://doi.org/10.1186/s13040-023-00339-9.

[2] D. R. E. Cotton, P. A. Cotton, and J. R. Shipway, “Chatting and cheating: Ensuring academic integrity in the era of ChatGPT,” Innov. Educ. Teach. Int., vol. 61, no. 2, pp. 228–239, 2024, doi: https://doi.org/10.1080/14703297.2023.2190148.

[3] F. M. Howard, A. Li, M. F. Riffon, and E. Garrett-mayer, “Characterizing the Increase in Artificial Intelligence Content Detection in Oncology Scienti fi c Abstracts From 2021 to 2023,” 2023, doi: https://doi.org/10.1200/CCI.24.00077.

[4] P. C. Theocharopoulos, P. Anagnostou, A. Tsoukala, S. V. Georgakopoulos, S. K. Tasoulis, and V. P. Plagianakos, “Detection of Fake Generated Scientific Abstracts,” Apr. 2023, doi: https://doi.org/10.1109/BigDataService58306.2023.00011.

[5] C. A. Gao et al., “Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers,” npj Digit. Med., vol. 6, p. 75, 2023, doi: https://doi.org/10.1038/s41746-023-00819-6.

[6] O. Marsh, M. Yates, A. Rowe, and M. Song, “Adversarial Attacks and Robustness in LLM- Generated Text Detection Systems,” 2025.

[7] H. Desaire, A. E. Chua, M. Isom, R. Jarosova, and D. Hua, “Distinguishing academic science writing from humans or ChatGPT with over 99% accuracy using off-the-shelf machine learning tools,” Cell Reports Phys. Sci., vol. 4, no. 6, 2023.

[8] A. Alikhanov et al., “AI Generated Text Detection,” Jan. 2026

[9] P. He, J. Gao, and W. Chen, “Debertav3: Improving Deberta Using Electra-Style Pre-Training With Gradient-Disentangled Embedding Sharing,” 11th Int. Conf. Learn. Represent. ICLR 2023, no. Mlm, pp. 1–16, 2023.

[10] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio, Eds., Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4171–4186. doi: https://doi.org/10.18653/v1/N19-1423.

[11] W. Yin, J. Hay, and D. Roth, “Benchmarking Zero-shot Text Classification : Datasets , Evaluation and Entailment Approach,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China, 2019, pp. 3914–3923.

[12] E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn, “DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature,” Proc. Mach. Learn. Res., vol. 202, pp. 24950–24962, 2023.

[13] W. X. Zhao et al., “A Survey of Large Language Models,” Mar. 2026.

[14] C. Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” J. Mach. Learn. Res., vol. 21, pp. 1–67, Sep. 2023.

[15] V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, and S. Feizi, “Can AI-Generated Text be Reliably Detected?,” pp. 1–37, Jan. 2025.

[16] M. Hahn, “Theoretical Limitations of Self-Attention in Neural Sequence Models,” Trans. Assoc. Comput. Linguist., vol. 8, pp. 156–171, Dec. 2020, doi: https://doi.org/10.1162/tacl_a_00306.

[17] J. Han, M. Kamber, and J. Pei, Data Mining: Concepts and Techniques, 3rd ed. Morgan Kaufmann, 2011.

[18] L. Lukman et al., “Proposal of the S-score for measuring the performance of researchers, institutions, and journals in Indonesia,” Sci. Ed., vol. 5, no. 2, pp. 135–141, 2018, doi: https://doi.org/10.6087/KCSE.138.

[19] H. de Sompel and M. L. Nelson, “Reminiscing about 15 years of interoperability efforts,” D-Lib Mag., vol. 21, no. 11/12, 2015, doi: https://doi.org/10.1045/november2015-vandesompel.

[20] A. Radford et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019.

[21] M. Freitag and Y. Al-Onaizan, “Beam Search Strategies for Neural Machine Translation,” in Proceedings of the First Workshop on Neural Machine Translation, T. Luong, A. Birch, G. Neubig, and A. Finch, Eds., Stroudsburg, PA, USA: Association for Computational Linguistics, Aug. 2017, pp. 56–60. doi: https://doi.org/10.18653/v1/W17-3207.

[22] N. S. Keskar, B. McCann, L. R. Varshney, C. Xiong, and R. Socher, “CTRL: A Conditional Transformer Language Model for Controllable Generation,” pp. 1–18, Sep. 2019.

[23] R. Paulus, C. Xiong, and R. Socher, “A Deep Reinforced Model for Abstractive Summarization,” CoRR, vol. abs/1705.0, Nov. 2017.

[24] Y. Wu et al., “Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation,” CoRR, vol. abs/1609.0, Oct. 2016.

[25] A. K. Uysal and S. Gunal, “The impact of preprocessing on text classification,” Inf. Process. Manag., vol. 50, no. 1, pp. 104–112, 2014, doi: https://doi.org/10.1016/j.ipm.2013.08.006.

[26] B. Krawczyk, “Learning from imbalanced data : open challenges and future directions,” Prog. Artif. Intell., vol. 5, no. 4, pp. 221–232, 2016, doi: https://doi.org/10.1007/s13748-016-0094-0.

[27] A. Field, Discovering statistics using IBM SPSS statistics. Sage publications limited, 2024.

[28] M. Mars, “From Word Embeddings to Pre-Trained Language Models: A State-of-the-Art Walkthrough,” Appl. Sci., vol. 12, no. 17, 2022, doi: https://doi.org/10.3390/app12178805.

[29] M. Laurer, W. van Atteveldt, A. Casas, and K. Welbers, “Less Annotating, More Classifying: Addressing the Data Scarcity Issue of Supervised Machine Learning with Deep Transfer Learning and BERT-NLI,” Polit. Anal., vol. 32, no. 1, pp. 84–100, Jan. 2024, doi: https://doi.org/10.1017/pan.2023.20.

[30] N. F. Liu, K. Lin, J. Hewitt, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the Middle : How Language Models Use Long Contexts,” 2023.

[31] I. V Pantic and S. Mugosa, “Artificial intelligence strategies based on random forests for detection of AI-generated content in public health,” Public Health, vol. 242, pp. 382–387, May 2025, doi: https://doi.org/10.1016/j.puhe.2025.03.029.

[32] S. K. Aityan, W. Claster, K. S. Emani, S. Rais, and T. Tran, “A Lightweight Approach to Detection of AI-Generated Texts Using Stylometric Features,” Jan. 2026.

[33] P. He, X. Liu, J. Gao, and W. Chen, “Deberta: Decoding-Enhanced Bert With Disentangled Attention,” ICLR 2021 - 9th Int. Conf. Learn. Represent., 2021.

[34] A. Conneau et al., “Unsupervised cross-lingual representation learning at scale,” in Proceedings of the 58th annual meeting of the association for computational linguistics, 2020, pp. 8440–8451.

[35] A. Zhang, Z. C. Lipton, M. Li, and A. J. Smola, Dive into Deep Learning. Cambridge: Cambridge University Press, 2023.

[36] K. M. Sujon, R. Hassan, K. Choi, and M. A. Samad, “Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models,” J. Big Data, vol. 12, no. 1, 2025, doi: https://doi.org/10.1186/s40537-025-01313-4.

[37] S. Sathyanarayanan, “Confusion Matrix-Based Performance Evaluation Metrics,” African J. Biomed. Res., vol. 27, no. 4, pp. 4023–4031, 2024, doi: https://doi.org/10.53555/ajbr.v27i4s.4345

Published

2026-07-31 — Updated on 2026-07-31

Versions

How to Cite

Zero-Shot Detection of IndoT5-Synthesized Indonesian Scientific Abstracts Using mDeBERTa v3. (2026). Indonesian Journal of Data and Science, 7(2). https://doi.org/10.56705/ijodas.v7i2.457