Handling Imbalanced Data in Concept Map Proposition Quality Classification: A Comparative Study of Resampling and Class-Weighted SVM
DOI:
https://doi.org/10.56705/ijodas.v7i1.482Keywords:
Concept Map, Proposition Quality Classification, Imbalanced Data, Support Vector Machine, TF-IDF, Class-Weighted SVMAbstract
Introduction: Concept maps are widely used in education to represent students’ knowledge structures, while automatic proposition-level quality classification can support faster and more consistent assessment. However, imbalanced class distributions may cause classification models to favor majority classes and reduce performance on minority categories. Method: This study compared several imbalance-handling techniques for concept map proposition quality classification using Support Vector Machine (SVM) with TF-IDF features. The dataset comprised 691 propositions grouped into four quality classes, with class 3 dominating the distribution. The evaluated approaches included baseline training, Random OverSampler, SMOTE, ADASYN, SMOTE-Tomek, SMOTE-ENN, and cost-sensitive SVM with class_weight='balanced'. Data were divided using an 80:20 stratified split, with balancing applied only to the training set. Results and Discussion: The cost-sensitive SVM achieved the best overall performance with 0.8417 accuracy, 0.7098 macro precision, 0.7320 macro recall, 0.7149 macro F1-score, 0.8455 weighted F1-score, and 0.2158 MAE. Among resampling approaches, ADASYN and Random OverSampler were the most competitive, whereas SMOTE-based variants produced lower macro-level performance. Conclusion: Class-weighted SVM is more effective than the evaluated resampling techniques for sparse, imbalanced concept map proposition data, highlighting the importance of macro-level and ordinal-aware metrics in educational text classification.
Downloads
References
[1] D. D. Prasetya, A. Pinandito, Y. Hayashi, and T. Hirashima, “Analysis of quality of knowledge structure and students’ perceptions in extension concept mapping,” Res. Pract. Technol. Enhanc. Learn., vol. 17, no. 1, p. 14, Dec. 2022, doi: https://doi.org/10.1186/s41039-022-00189-9.
[2] D. D. Prasetya, T. Widiyaningtyas, and T. Hirashima, “Interrelatedness patterns of knowledge representation in extension concept mapping,” Res. Pract. Technol. Enhanc. Learn., vol. 20, p. 009, May 2024, doi: https://doi.org/10.58459/rptel.2025.20009.
[3] T. Evans and I. Jeong, “Concept maps as assessment for learning in university mathematics,” Educ. Stud. Math., vol. 113, no. 3, pp. 475–498, Jul. 2023, doi: https://doi.org/10.1007/s10649-023-10209-0.
[4] C. Cischke and S. T. Mueller, “Concept Mapping Assessments as a Tool for Judgment of Learning.” Jul. 29, 2022, doi: https://doi.org/10.31234/osf.io/69bjx.
[5] N. Rismayanti, D. D. Prasetya, T. Widiyaningtyas, and T. Hirashima, “Comparative Study of Random Forest and Ordinal Regression in Concept Map Quality Assessment: The Role of TF-IDF, BERT, and SMOTE-based Balancing,” Ilk. J. Ilm., vol. 17, no. 3, pp. 336–345, 2025, doi: http://dx.doi.org/10.33096/ilkom.v17i3.2906.336-345.
[6] M. S. Mohosheu, M. A. al Noman, A. Newaz, Al-Amin, and T. Jabid, “A Comprehensive Evaluation of Sampling Techniques in Addressing Class Imbalance Across Diverse Datasets,” 2024 6th Int. Conf. Electr. Eng. Inf. Commun. Technol., 2024, doi: https://doi.org/10.1109/iceeict62016.2024.10534464.
[7] N. Habbat, “Sentiment analysis of imbalanced datasets using BERT and ensemble stacking for deep learning,” Eng. Appl. Artif. Intell., vol. 126, 2023, doi: https://doi.org/10.1016/j.engappai.2023.106999.
[8] F. De Engenharia, “Measuring the performance of ordinal classification,” Int. J. Pattern Recognit. Arti¯cial Intell., vol. 25, no. 8, pp. 1173–1195, 2011, doi: https://doi.org/10.1142/S0218001411009093.
[9] D. Yilmaz Eroglu and M. S. Pir, “Hybrid Oversampling and Undersampling Method (HOUM) via Safe-Level SMOTE and Support Vector Machine,” Appl. Sci., vol. 14, no. 22, p. 10438, Nov. 2024, doi: https://doi.org/10.3390/app142210438.
[10] A. Tuppad and S. D. Patil, “Data Pre-processing Issues in Medical Data Classification,” 2023 Int. Conf. …, 2023, [Online]. Available: https://ieeexplore.ieee.org/abstract/document/10275855/.
[11] J. Zhao, K. S. Chong, W. Shu, and ..., “A Data Pre-Processing Module for Improved-Accuracy Machine-Learning-based Micro-Single-Event-Latchup Detection,” 2023 IEEE 9th Int. …, 2023, [Online]. Available: https://ieeexplore.ieee.org/abstract/document/10207447/.
[12] G. Ketepalli and P. Bulla, “Data Preparation and Pre-processing of Intrusion Detection Datasets using Machine Learning,” 2023 Int. Conf. …, 2023, [Online]. Available: https://ieeexplore.ieee.org/abstract/document/10134025/.
[13] K. K. Agustiningsih, E. Utami, and M. A. Alsyaibani, “Sentiment Analysis of COVID-19 Vaccines in Indonesia on Twitter Using Pre-Trained and Self-Training Word Embeddings,” J. Ilmu Komput. dan Inf., vol. 15, no. 1, pp. 39–46, 2022, doi: https://doi.org/10.21609/jiki.v15i1.1044.
[14] A. Al Tawil, “Comparative Analysis of Machine Learning Algorithms for Email Phishing Detection Using TF-IDF, Word2Vec, and BERT,” Comput. Mater. Contin., vol. 81, no. 2, pp. 3395–3412, 2024, doi: https://doi.org/10.32604/cmc.2024.057279.
[15] S. M. M. Hossain, “TF-IDF feature-based spam filtering of mobile SMS using a machine learning approach,” Applied Intelligence for Industry 4.0. pp. 162–175, 2023, [Online]. Available: https://api.elsevier.com/content/abstract/scopus_id/85161154224.
[16] C. A. N. Agustina, “The Implementation of TF-IDF and Word2Vec on Booster Vaccine Sentiment Analysis Using Support Vector Machine Algorithm,” Procedia Computer Science, vol. 234. pp. 156–163, 2024, doi: https://doi.org/10.1016/j.procs.2024.02.162.
[17] S. M. M. Hossain, K. M. A. Kamal, A. Sen, and I. H. Sarker, TF-IDF Feature-Based Spam Filtering of Mobile SMS Using a Machine Learning Approach. 2023.
[18] G. Popoola, “Sentiment Analysis of Financial News Data using TF-IDF and Machine Learning Algorithms,” 2024 IEEE 3rd International Conference on AI in Cybersecurity, ICAIC 2024. 2024, doi: https://doi.org/10.1109/ICAIC60265.2024.10433843.
[19] E. A. Aldhahri, A. A. Almazroi, M. H. Alkinani, N. Ayub, E. A. Alghamdi, and N. F. Janbi, “Smart Farming: Enhancing Urban Agriculture through Predictive Analytics and Resource Optimization,” IEEE Access, vol. 13, no. October 2024, pp. 72375–72388, 2025, doi: https://doi.org/10.1109/ACCESS.2025.3530006.
[20] N. T. Singh, “An Innovative URL-Based System Approach with ML Based Prevention,” Proceedings - International Conference on Computing, Power, and Communication Technologies, IC2PCT 2024. pp. 1498–1503, 2024, doi: https://doi.org/10.1109/IC2PCT60090.2024.10486712.
[21] G. Husain et al., “SMOTE vs. SMOTEENN: A Study on the Performance of Resampling Algorithms for Addressing Class Imbalance in Regression Models,” Algorithms, vol. 18, no. 1, p. 37, Jan. 2025, doi: https://doi.org/10.3390/a18010037.
[22] X. Ning, “Classification of Sulfadimidine and Sulfapyridine in Duck Meat by Surface Enhanced Raman Spectroscopy Combined with Principal Component Analysis and Support Vector Machine,” Anal. Lett., vol. 53, no. 10, pp. 1513–1524, 2020, doi: https://doi.org/10.1080/00032719.2019.1710524.
[23] O. Hamidi, “Analysis of the Response of Urban Water Consumption to Climatic Variables: Case Study of Khorramabad City in Iran,” Adv. Meteorol., vol. 2021, 2021, doi: https://doi.org/10.1155/2021/6615152.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Rahmadani Rahmadani, Nurul Rismayanti, Lukman Syafie, A Sinra, Roesman Ridwan Raja

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Authors retain copyright and full publishing rights to their articles. Upon acceptance, authors grant Indonesian Journal of Data and Science a non-exclusive license to publish the work and to identify itself as the original publisher.
Self-archiving. Authors may deposit the submitted version, accepted manuscript, and version of record in institutional or subject repositories, with citation to the published article and a link to the version of record on the journal website.
Commercial permissions. Uses intended for commercial advantage or monetary compensation are not permitted under CC BY-NC 4.0. For permissions, contact the editorial office at ijodas.journal@gmail.com.
Legacy notice. Some earlier PDFs may display “Copyright © [Journal Name]” or only a CC BY-NC logo without the full license text. To ensure clarity, the authors maintain copyright, and all articles are distributed under CC BY-NC 4.0. Where any discrepancy exists, this policy and the article landing-page license statement prevail.










