Weakly Supervised Sentiment Analysis of Gold Price Discussions Using Conventional Machine Learning and IndoBERT

Authors

  • M Rizki Hardika Universitas Internasional Semen Indonesia
  • Brina Miftahurrohmah Universitas Internasional Semen Indonesia

DOI:

https://doi.org/10.56705/ijodas.v7i2.445

Keywords:

Sentiment Analysis, Gold Price, Weak Supervision, IndoBERT, Platform X

Abstract

Introduction: Gold price movements attract substantial public and investor attention because gold serves as both a safe-haven asset and a hedging instrument. This study investigates Indonesian public sentiment toward gold price discussions on Platform X using a weakly supervised sentiment-analysis framework. Method: A total of 7,283 Indonesian-language tweets containing the keyword “harga emas” were collected during 2023–2025, with 4,429 tweets retained after preprocessing. Sentiment labels were generated using a domain-specific lexicon and validated through manual annotation. Naïve Bayes, K-Nearest Neighbor, Support Vector Machine, and IndoBERT were evaluated using the same train–test partition. Results and Discussion: Manual validation achieved a Cohen’s Kappa of 0.8718, indicating almost perfect inter-annotator agreement, while the lexicon-based labels achieved 70.62% accuracy against the manually annotated reference. IndoBERT achieved the highest performance on weakly supervised labels with 98.31% accuracy and a 98.16% macro F1-score, outperforming SVM, Naïve Bayes, and KNN. However, its accuracy decreased to 69.49% when evaluated against manually annotated data, demonstrating that downstream performance remains strongly influenced by weak-label quality. Conclusion: Weak supervision provides an efficient and scalable approach for large-scale Indonesian financial sentiment annotation, while contextual models such as IndoBERT offer superior classification performance; however, reliable manual validation remains essential to mitigate label noise and improve generalizability.

Downloads

Download data is not yet available.

References

[1] T. I. Tanin, A. Sarker, R. Brooks, and H. X. Do, “Does oil impact gold during COVID-19 and three other recent crises?,” Energy Economics, vol. 108, p. 105938, Apr. 2022, https://doi.org/10.1016/j.eneco.2022.105938.

[2] A. A. Sikiru and A. A. Salisu, “Assessing the hedging potential of gold and other precious metals against uncertainty due to epidemics and pandemics,” Quality & Quantity, vol. 56, no. 4, pp. 2199–2214, Aug. 2022, https://doi.org/10.1007/s11135-021-01214-7.

[3] I. N. Agustin, “Can Social Responsible Investment and Gold be a Good Diversifier for Indonesia Sharia Investors?,” Jurnal Keuangan dan Perbankan, vol. 26, no. 1, pp. 146–160, Feb. 2022, https://doi.org/10.26905/jkdp.v26i1.6929.

[4] D. G. Baur and T. K. McDermott, “Is gold a safe haven? International evidence,” Journal of Banking & Finance, vol. 34, no. 8, pp. 1886–1898, Aug. 2010, https://doi.org/10.1016/j.jbankfin.2009.12.008.

[5] A. Algaba, D. Ardia, K. Bluteau, S. Borms, and K. Boudt, “Econometrics Meets Sentiment: An Overview Of Methodology And Applications,” Journal of Economic Surveys, vol. 34, no. 3, pp. 512–547, Jul. 2020, https://doi.org/10.1111/joes.12370.

[6] L. Barbaglia, S. Consoli, S. Manzan, L. Tiozzo Pezzoli, and E. Tosetti, “Sentiment analysis of economic text: A lexicon‐based approach,” Economic Inquiry, vol. 63, no. 1, pp. 125–143, Jan. 2025, https://doi.org/10.1111/ecin.13264.

[7] M. Baker and J. Wurgler, “Investor Sentiment in the Stock Market,” Journal of Economic Perspectives, vol. 21, no. 2, pp. 129–151, Apr. 2007, https://doi.org/10.1257/jep.21.2.129.

[8] X. Huang and H. Song, “Investor Sentiment Combined with Multisource Information to Predict Stock Prices: An Analysis of China’s A-Share Market,” Scientific Programming, vol. 2021, pp. 1–9, Dec. 2021, https://doi.org/10.1155/2021/9094032.

[9] E. Blankespoor, “Firm communication and investor response: A framework and discussion integrating social media,” Accounting, Organizations and Society, vol. 68–69, pp. 80–87, Jul. 2018, https://doi.org/10.1016/j.aos.2018.03.009.

[10] A. Giachanou and F. Crestani, “Like It or Not,” ACM Computing Surveys, vol. 49, no. 2, pp. 1–41, Jun. 2017, https://doi.org/10.1145/2938640.

[11] U. Naseem, I. Razzak, M. Khushi, P. W. Eklund, and J. Kim, “COVIDSenti: A Large-Scale Benchmark Twitter Data Set for COVID-19 Sentiment Analysis,” IEEE Transactions on Computational Social Systems, vol. 8, no. 4, pp. 1003–1015, Aug. 2021, https://doi.org/10.1109/TCSS.2021.3051189.

[12] Simon Kemp, “Digital 2025: Indonesia,” DataReportal – Global Digital Insights, Feb. 25, 2025.

[13] K. Kowsari, K. Jafari Meimandi, M. Heidarysafa, S. Mendu, L. Barnes, and D. Brown, “Text Classification Algorithms: A Survey,” Information, vol. 10, no. 4, p. 150, Apr. 2019, https://doi.org/10.3390/info10040150.

[14] Daniel Jurafsky and James H. Martin, Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, with Language Models, 3rd ed. 2026.

[15] B. Liu, “Sentiment Analysis and Opinion Mining,” Synthesis Lectures on Human Language Technologies, vol. 5, no. 1, pp. 1–167, May 2012, https://doi.org/10.2200/S00416ED1V01Y201204HLT016.

[16] A. F. A. H. Alnuaimi and T. H. K. Albaldawi, “An overview of machine learning classification techniques,” BIO Web of Conferences, vol. 97, p. 00133, Apr. 2024, https://doi.org/10.1051/bioconf/20249700133.

[17] Y. Zhang, Q. Li, and Y. Xin, “Research on eight machine learning algorithms applicability on different characteristics data sets in medical classification tasks,” Frontiers in Computational Neuroscience, vol. 18, Jan. 2024, https://doi.org/10.3389/fncom.2024.1345575.

[18] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North, 2019, pp. 4171–4186, https://doi.org/10.18653/v1/N19-1423.

[19] B. Wilie et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” in Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, 2020, pp. 843–857, https://doi.org/10.18653/v1/2020.aacl-main.85.

[20] S. H. Ramadhani and M. I. Wahyudin, “Analisis Sentimen Terhadap Vaksinasi Astra Zeneca pada Twitter Menggunakan Metode Naïve Bayes dan K-NN,” Jurnal JTIK (Jurnal Teknologi Informasi dan Komunikasi), vol. 6, no. 4, pp. 526–534, Feb. 2022, https://doi.org/10.35870/jtik.v6i4.530.

[21] N. Wulandari, Y. Cahyana, R. Rahmat, and H. Hikmayanti, “Sentiment Analysis on the Relocation of the National Capital (IKN) on Social Media X Using Naive Bayes and K-Nearest Neighbor (KNN) Methods,” Journal of Applied Informatics and Computing, vol. 9, no. 3, pp. 724–731, Jun. 2025, https://doi.org/10.30871/jaic.v9i3.9552.

[22] Syahril Dwi Prasetyo, Shofa Shofiah Hilabi, and Fitri Nurapriani, “Analisis Sentimen Relokasi Ibukota Nusantara Menggunakan Algoritma Naïve Bayes dan KNN,” Jurnal KomtekInfo, pp. 1–7, Jan. 2023, https://doi.org/10.35134/komtekinfo.v10i1.330.

[23] E. Hokijuliandy, H. Napitupulu, and Firdaniza, “Application of SVM and Chi-Square Feature Selection for Sentiment Analysis of Indonesia’s National Health Insurance Mobile Application,” Mathematics, vol. 11, no. 17, p. 3765, Sep. 2023, https://doi.org/10.3390/math11173765.

[24] N. F. Rozy, N. Amini, N. Hakiem, and S. H. Afrizal, “Analysis of Multi-Class Sentiment on Indonesian Twitter Using Support Vector Machine Classification Algorithm with Particle Swarm Optimization,” in Proceedings of the 2023 7th International Conference on Advances in Artificial Intelligence, Oct. 2023, pp. 62–67, https://doi.org/10.1145/3633598.3633609.

[25] A. M. van der Veen and E. Bleich, “The advantages of lexicon-based sentiment analysis in an age of machine learning,” PLOS ONE, vol. 20, no. 1, p. e0313092, Jan. 2025, https://doi.org/10.1371/journal.pone.0313092.

[26] A. Ratner, S. H. Bach, H. Ehrenberg, J. Fries, S. Wu, and C. Ré, “Snorkel,” Proceedings of the VLDB Endowment, vol. 11, no. 3, pp. 269–282, Nov. 2017, https://doi.org/10.14778/3157794.3157797.

[27] B. Frenay and M. Verleysen, “Classification in the Presence of Label Noise: A Survey,” IEEE Transactions on Neural Networks and Learning Systems, vol. 25, no. 5, pp. 845–869, May 2014, https://doi.org/10.1109/TNNLS.2013.2292894.

[28] C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge University Press, 2008.

[29] R. Feldman and J. Sanger, The Text Mining Handbook. Cambridge University Press, 2006.

[30] Bo Han and Timothy Baldwin, Lexical Normalisation of Short Text Messages: Makn Sens a #twitter. Portland, Oregon, USA: Association for Computational Linguistics, 2011.

[31] M. F. Porter, “An algorithm for suffix stripping,” Program, vol. 14, no. 3, pp. 130–137, Mar. 1980, https://doi.org/10.1108/eb046814.

[32] J. Cohen, “A Coefficient of Agreement for Nominal Scales,” Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, Apr. 1960, https://doi.org/10.1177/001316446002000104.

[33] J. R. Landis and G. G. Koch, “The Measurement of Observer Agreement for Categorical Data,” Biometrics, vol. 33, no. 1, p. 159, Mar. 1977, https://doi.org/10.2307/2529310.

[34] D. Berrar, “Performance Measures for Binary Classification,” in Encyclopedia of Bioinformatics and Computational Biology, Elsevier, 2019, pp. 546–560.

[35] D. Chicco and G. Jurman, “The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation,” BMC Genomics, vol. 21, no. 1, p. 6, Dec. 2020, https://doi.org/10.1186/s12864-019-6413-7.

[36] M. A. Pienaar and K. Naidoo, “Classification and predictive models using supervised machine learning: A conceptual review,” Southern African Journal of Critical Care, p. e2937, May 2025, https://doi.org/10.7196/SAJCC.2025.v411.2937.

[37] C. C. Aggarwal, Machine Learning for Text. Cham: Springer International Publishing, 2018.

[38] C. Cortes and V. Vapnik, “Support-vector networks,” Machine Learning, vol. 20, no. 3, pp. 273–297, Sep. 1995, https://doi.org/10.1007/BF00994018.

Downloads

Published

2026-07-31

How to Cite

Weakly Supervised Sentiment Analysis of Gold Price Discussions Using Conventional Machine Learning and IndoBERT. (2026). Indonesian Journal of Data and Science, 7(2), 274-290. https://doi.org/10.56705/ijodas.v7i2.445