Handling Imbalanced Data in Concept Map Proposition Quality Classification: A Comparative Study of Resampling and Class-Weighted SVM

Authors

  • Nurul Rismayanti Universitas Muslim Indonesia
  • Lukman Syafie Universiti Kuala Lumpur
  • A Sinra UPTP. Wil. Makassar 1 Badan Pendapatan Daerah Prov. Sulsel
  • Rahmadani Rahmadani Universitas Muslim Indonesia
  • Roesman Ridwan Raja Kyushu Institute of Technology

DOI:

https://doi.org/10.56705/ijodas.v7i1.482

Keywords:

Concept Map, Proposition Quality Classification, Imbalanced Data, Support Vector Machine, TF-IDF, Class-Weighted SVM

Abstract

Introduction: Concept maps are widely used in education to represent students’ knowledge structures, while automatic proposition-level quality classification can support faster and more consistent assessment. However, imbalanced class distributions may cause classification models to favor majority classes and reduce performance on minority categories. Method: This study compared several imbalance-handling techniques for concept map proposition quality classification using Support Vector Machine (SVM) with TF-IDF features. The dataset comprised 691 propositions grouped into four quality classes, with class 3 dominating the distribution. The evaluated approaches included baseline training, Random OverSampler, SMOTE, ADASYN, SMOTE-Tomek, SMOTE-ENN, and cost-sensitive SVM with class_weight='balanced'. Data were divided using an 80:20 stratified split, with balancing applied only to the training set. Results and Discussion: The cost-sensitive SVM achieved the best overall performance with 0.8417 accuracy, 0.7098 macro precision, 0.7320 macro recall, 0.7149 macro F1-score, 0.8455 weighted F1-score, and 0.2158 MAE. Among resampling approaches, ADASYN and Random OverSampler were the most competitive, whereas SMOTE-based variants produced lower macro-level performance. Conclusion: Class-weighted SVM is more effective than the evaluated resampling techniques for sparse, imbalanced concept map proposition data, highlighting the importance of macro-level and ordinal-aware metrics in educational text classification.

Downloads

Download data is not yet available.

References

[1] D. D. Prasetya, A. Pinandito, Y. Hayashi, and T. Hirashima, “Analysis of quality of knowledge structure and students’ perceptions in extension concept mapping,” Res. Pract. Technol. Enhanc. Learn., vol. 17, no. 1, p. 14, Dec. 2022, doi: https://doi.org/10.1186/s41039-022-00189-9.

[2] D. D. Prasetya, T. Widiyaningtyas, and T. Hirashima, “Interrelatedness patterns of knowledge representation in extension concept mapping,” Res. Pract. Technol. Enhanc. Learn., vol. 20, p. 009, May 2024, doi: https://doi.org/10.58459/rptel.2025.20009.

[3] T. Evans and I. Jeong, “Concept maps as assessment for learning in university mathematics,” Educ. Stud. Math., vol. 113, no. 3, pp. 475–498, Jul. 2023, doi: https://doi.org/10.1007/s10649-023-10209-0.

[4] C. Cischke and S. T. Mueller, “Concept Mapping Assessments as a Tool for Judgment of Learning.” Jul. 29, 2022, doi: https://doi.org/10.31234/osf.io/69bjx.

[5] N. Rismayanti, D. D. Prasetya, T. Widiyaningtyas, and T. Hirashima, “Comparative Study of Random Forest and Ordinal Regression in Concept Map Quality Assessment: The Role of TF-IDF, BERT, and SMOTE-based Balancing,” Ilk. J. Ilm., vol. 17, no. 3, pp. 336–345, 2025, doi: http://dx.doi.org/10.33096/ilkom.v17i3.2906.336-345.

[6] M. S. Mohosheu, M. A. al Noman, A. Newaz, Al-Amin, and T. Jabid, “A Comprehensive Evaluation of Sampling Techniques in Addressing Class Imbalance Across Diverse Datasets,” 2024 6th Int. Conf. Electr. Eng. Inf. Commun. Technol., 2024, doi: https://doi.org/10.1109/iceeict62016.2024.10534464.

[7] N. Habbat, “Sentiment analysis of imbalanced datasets using BERT and ensemble stacking for deep learning,” Eng. Appl. Artif. Intell., vol. 126, 2023, doi: https://doi.org/10.1016/j.engappai.2023.106999.

[8] F. De Engenharia, “Measuring the performance of ordinal classification,” Int. J. Pattern Recognit. Arti¯cial Intell., vol. 25, no. 8, pp. 1173–1195, 2011, doi: https://doi.org/10.1142/S0218001411009093.

[9] D. Yilmaz Eroglu and M. S. Pir, “Hybrid Oversampling and Undersampling Method (HOUM) via Safe-Level SMOTE and Support Vector Machine,” Appl. Sci., vol. 14, no. 22, p. 10438, Nov. 2024, doi: https://doi.org/10.3390/app142210438.

[10] A. Tuppad and S. D. Patil, “Data Pre-processing Issues in Medical Data Classification,” 2023 Int. Conf. …, 2023, [Online]. Available: https://ieeexplore.ieee.org/abstract/document/10275855/.

[11] J. Zhao, K. S. Chong, W. Shu, and ..., “A Data Pre-Processing Module for Improved-Accuracy Machine-Learning-based Micro-Single-Event-Latchup Detection,” 2023 IEEE 9th Int. …, 2023, [Online]. Available: https://ieeexplore.ieee.org/abstract/document/10207447/.

[12] G. Ketepalli and P. Bulla, “Data Preparation and Pre-processing of Intrusion Detection Datasets using Machine Learning,” 2023 Int. Conf. …, 2023, [Online]. Available: https://ieeexplore.ieee.org/abstract/document/10134025/.

[13] K. K. Agustiningsih, E. Utami, and M. A. Alsyaibani, “Sentiment Analysis of COVID-19 Vaccines in Indonesia on Twitter Using Pre-Trained and Self-Training Word Embeddings,” J. Ilmu Komput. dan Inf., vol. 15, no. 1, pp. 39–46, 2022, doi: https://doi.org/10.21609/jiki.v15i1.1044.

[14] A. Al Tawil, “Comparative Analysis of Machine Learning Algorithms for Email Phishing Detection Using TF-IDF, Word2Vec, and BERT,” Comput. Mater. Contin., vol. 81, no. 2, pp. 3395–3412, 2024, doi: https://doi.org/10.32604/cmc.2024.057279.

[15] S. M. M. Hossain, “TF-IDF feature-based spam filtering of mobile SMS using a machine learning approach,” Applied Intelligence for Industry 4.0. pp. 162–175, 2023, [Online]. Available: https://api.elsevier.com/content/abstract/scopus_id/85161154224.

[16] C. A. N. Agustina, “The Implementation of TF-IDF and Word2Vec on Booster Vaccine Sentiment Analysis Using Support Vector Machine Algorithm,” Procedia Computer Science, vol. 234. pp. 156–163, 2024, doi: https://doi.org/10.1016/j.procs.2024.02.162.

[17] S. M. M. Hossain, K. M. A. Kamal, A. Sen, and I. H. Sarker, TF-IDF Feature-Based Spam Filtering of Mobile SMS Using a Machine Learning Approach. 2023.

[18] G. Popoola, “Sentiment Analysis of Financial News Data using TF-IDF and Machine Learning Algorithms,” 2024 IEEE 3rd International Conference on AI in Cybersecurity, ICAIC 2024. 2024, doi: https://doi.org/10.1109/ICAIC60265.2024.10433843.

[19] E. A. Aldhahri, A. A. Almazroi, M. H. Alkinani, N. Ayub, E. A. Alghamdi, and N. F. Janbi, “Smart Farming: Enhancing Urban Agriculture through Predictive Analytics and Resource Optimization,” IEEE Access, vol. 13, no. October 2024, pp. 72375–72388, 2025, doi: https://doi.org/10.1109/ACCESS.2025.3530006.

[20] N. T. Singh, “An Innovative URL-Based System Approach with ML Based Prevention,” Proceedings - International Conference on Computing, Power, and Communication Technologies, IC2PCT 2024. pp. 1498–1503, 2024, doi: https://doi.org/10.1109/IC2PCT60090.2024.10486712.

[21] G. Husain et al., “SMOTE vs. SMOTEENN: A Study on the Performance of Resampling Algorithms for Addressing Class Imbalance in Regression Models,” Algorithms, vol. 18, no. 1, p. 37, Jan. 2025, doi: https://doi.org/10.3390/a18010037.

[22] X. Ning, “Classification of Sulfadimidine and Sulfapyridine in Duck Meat by Surface Enhanced Raman Spectroscopy Combined with Principal Component Analysis and Support Vector Machine,” Anal. Lett., vol. 53, no. 10, pp. 1513–1524, 2020, doi: https://doi.org/10.1080/00032719.2019.1710524.

[23] O. Hamidi, “Analysis of the Response of Urban Water Consumption to Climatic Variables: Case Study of Khorramabad City in Iran,” Adv. Meteorol., vol. 2021, 2021, doi: https://doi.org/10.1155/2021/6615152.

Downloads

Published

2026-03-17

How to Cite

Handling Imbalanced Data in Concept Map Proposition Quality Classification: A Comparative Study of Resampling and Class-Weighted SVM. (2026). Indonesian Journal of Data and Science, 7(1), 1-11. https://doi.org/10.56705/ijodas.v7i1.482