Comparative Evaluation of Machine Learning Models for Heavy Crude Oil Viscosity Prediction Using Repeated Nested Cross-Validation and Independent Holdout Testing
DOI:
https://doi.org/10.56705/ijodas.v7i2.455Keywords:
Heavy Crude Oil Viscosity, Machine Learning, Gradient Boosting, Repeated Nested Cross-Validation, Independent Holdout Testing, Comparative EvaluationAbstract
Introduction: Accurate prediction of heavy crude oil viscosity is important for reservoir engineering, production planning, and flow assurance because viscosity strongly affects fluid mobility and transport behavior. This study comparatively evaluates established machine learning models under a rigorous validation protocol rather than proposing a new predictive framework. Method: A published Middle Eastern heavy crude-oil dataset containing 196 development measurements and 47 independent holdout measurements was used. Linear Regression, Support Vector Regression, Random Forest, Gradient Boosting, and the Beggs–Robinson correlation were evaluated using repeated nested cross-validation with five outer folds repeated twice and five inner folds. Preprocessing and hyperparameter selection were embedded within the validation pipeline, while the untouched holdout set was used only for final evaluation. Results and Discussion: Gradient Boosting achieved the best internal performance with R² = 0.99313 and RMSE = 11.41 cP. On the independent holdout set, it achieved R² = 0.99308, RMSE = 8.43 cP, MAE = 6.64 cP, and MAPE = 0.78%, outperforming Random Forest and Support Vector Regression. Residual diagnostics showed no detectable heteroscedasticity, while permutation importance identified temperature and C7+ as the dominant predictors. Conclusion: Gradient Boosting provides highly accurate viscosity predictions within the sampled domain; however, the absence of row-level oil identifiers and external reservoir data limits conclusions regarding oil-disjoint and field-level generalization.
Downloads
References
[1] M. A. Oloso, M. G. Hassan, M. B. Bader-El-Den, and J. M. Buick, “Ensemble SVM for characterisation of crude oil viscosity,” J. Pet. Explor. Prod. Technol., vol. 8, no. 2, pp. 531–546, Jun. 2018, doi: https://doi.org/10.1007/s13202-017-0355-x.
[2] F. Souas, A. Safri, and A. Benmounah, “A review on the rheology of heavy crude oil for pipeline transportation,” Pet. Res., vol. 6, no. 2, pp. 116–136, Jun. 2021, doi: https://doi.org/10.1016/j.ptlrs.2020.11.001.
[3] E. Bahonar, M. Chahardowli, Y. Ghalenoei, and M. Simjoo, “New correlations to predict oil viscosity using data mining techniques,” J. Pet. Sci. Eng., vol. 208, p. 109736, Jan. 2022, doi: https://doi.org/10.1016/j.petrol.2021.109736.
[4] U. Sinha, B. Dindoruk, and M. Soliman, “Machine learning augmented dead oil viscosity model for all oil types,” J. Pet. Sci. Eng., vol. 195, p. 107603, Dec. 2020, doi: https://doi.org/10.1016/j.petrol.2020.107603.
[5] M. Ghanavati, M.-J. Shojaei, and A. R. S. A., “Effects of Asphaltene Content and Temperature on Viscosity of Iranian Heavy Crude Oil: Experimental and Modeling Study,” Energy & Fuels, vol. 27, no. 12, pp. 7217–7232, Dec. 2013, doi: https://doi.org/10.1021/ef400776h.
[6] H. S. Naji, “Comparative Study of the C7+ Characterization Methods: An Object-Oriented Approach,” Arab. J. Sci. Eng., vol. 36, no. 7, pp. 1423–1446, Nov. 2011, doi: https://doi.org/10.1007/s13369-011-0118-9.
[7] C. H. Whitson, “Effect of C7+ Properties on Equation-of-State Predictions,” Soc. Pet. Eng. J., vol. 24, no. 06, pp. 685–696, Dec. 1984, doi: https://doi.org/10.2118/11200-PA.
[8] M. Baghani, “Energy optimization in electrochemical oxidation for wastewater treatment using interpretable machine learning,” Chem. Eng. J. Adv., vol. 25, 2026, doi: https://doi.org/10.1016/j.ceja.2026.101044.
[9] Y. Bouslihim, “Predicting soil organic matter from color indices: economic and technical feasibility in semi-arid agricultural soils,” Carbon Res., vol. 5, no. 1, 2026, doi: https://doi.org/10.1007/s44246-025-00240-6.
[10] O. Alomair, E. Folad, and A. Elsharkawy, “Effects of asphaltene and resin contents on crude oil viscosity: Experimental and modeling studies,” Fuel, vol. 405, p. 136575, 2026, doi: https://doi.org/10.1016/j.fuel.2025.136575.
[11] Y. He et al., “rediction of High-Pressure Physical Properties of crude oil Using Explainable Machine Learning Models,” IEEE Access, vol. 13, pp. 4912–4931, 2025, doi: https://doi.org/10.1109/ACCESS.2024.3522595.
[12] K. Peiro Ahmady Langeroudy, P. Kharazi Esfahani, and M. R. Khorsand Movaghar, “Enhanced intelligent approach for determination of crude oil viscosity at reservoir conditions,” Sci. Rep., vol. 13, no. 1, p. 1666, Jan. 2023, doi: https://doi.org/10.1038/s41598-023-28770-2.
[13] J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Ann. Stat., vol. 29, no. 5, Oct. 2001, doi: https://doi.org/10.1214/aos/1013203451.
[14] C. Nadeau and Y. Bengio, “Inference for the {Generalization} {Error},” Mach. Learn., vol. 52, no. 3, pp. 239–281, Sep. 2003, doi: https://doi.org/10.1023/A:1024068626366.
[15] T. G. Dietterich, “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms,” Neural Comput., vol. 10, no. 7, pp. 1895–1923, Oct. 1998, doi: https://doi.org/10.1162/089976698300017197.
[16] C. Strobl, A.-L. Boulesteix, T. Kneib, T. Augustin, and A. Zeileis, “Conditional variable importance for random forests,” BMC Bioinformatics, vol. 9, no. 1, p. 307, Dec. 2008, doi: https://doi.org/10.1186/1471-2105-9-307.
[17] B. Dou, “Machine Learning Methods for Small Data Challenges in Molecular Science,” Chemical Reviews, vol. 123, no. 13. pp. 8736–8780, 2023, doi: https://doi.org/10.1021/acs.chemrev.3c00189.
[18] P. T. Zakowicz, “Machine learning helps predict early onset psychosis with serum protein biomarkers, neuropsychometry, and clinicodemographic data,” Sci. Rep., vol. 16, no. 1, 2026, doi: https://doi.org/10.1038/s41598-025-33765-2.
[19] O. Karal, “Performance comparison of different kernel functions in SVM for different k value in k-fold cross-validation,” Proc. - 2020 Innov. Intell. Syst. Appl. Conf. ASYU 2020, 2020, doi: https://doi.org/10.1109/ASYU50717.2020.9259880.
[20] Y. Nie, “Deep Melanoma classification with K-Fold Cross-Validation for Process optimization,” IEEE Med. Meas. Appl. MeMeA 2020 - Conf. Proc., 2020, doi: https://doi.org/10.1109/MeMeA49120.2020.9137222.
[21] J. M. Dahr, “A Hard Voting Ensemble Model of the Logistic Regression, Support Vector Machine and Random Forest for Network Intrusion Detection,” Inform. Slov., vol. 49, no. 27, pp. 43–56, 2025, doi: https://doi.org/10.31449/inf.v49i27.8573.
[22] Z. Liu, “Prediction of landfill gases concentration based on Grey Wolf Optimization – Support Vector Regression during landfill excavation process,” Waste Manag., vol. 198, pp. 128–136, 2025, doi: https://doi.org/10.1016/j.wasman.2025.02.040.
[23] J. C. M. Sánchez, “Improving wheat yield prediction through variable selection using Support Vector Regression, Random Forest, and Extreme Gradient Boosting,” Smart Agricultural Technology, vol. 10.. 2025, doi: https://doi.org/10.1016/j.atech.2025.100791.
[24] C. C. Onyekwena, “Support vector machine regression to predict gas diffusion coefficient of biochar-amended soil,” Appl. Soft Comput., vol. 127, 2022, doi: https://doi.org/10.1016/j.asoc.2022.109345.
[25] H. Su, “Support Vector Regression-Based Reduced- Reference Perceptual Quality Model for Compressed Point Clouds,” IEEE Trans. Multimed., vol. 26, pp. 6238–6249, 2024, doi: https://doi.org/10.1109/TMM.2023.3347638.
[26] V. L. Yerrabolu, “Performance Comparison of Random Forest Regressor and Support Vector Regression for Solar Energy Prediction,” Iop Conference Series Earth and Environmental Science, vol. 1375, no. 1. 2024, doi: https://doi.org/10.1088/1755-1315/1375/1/012013.
[27] A. Primantara, “Bagging System Performance Analysis Using Artificial Neural Network, Random Forest Regression, Linear Regression, and Support Vector Regression,” Digest of Technical Papers IEEE International Conference on Consumer Electronics. pp. 618–622, 2024, doi: https://doi.org/10.1109/ISCT62336.2024.10791247.
[28] M. Y. Shams, “Water quality prediction using machine learning models based on grid search method,” Multimed. Tools Appl., vol. 83, no. 12, pp. 35307–35334, 2024, doi: https://doi.org/10.1007/s11042-023-16737-4.
[29] Y. Boutahri, “Machine learning-based predictive model for thermal comfort and energy optimization in smart buildings,” Results Eng., vol. 22, 2024, doi: https://doi.org/10.1016/j.rineng.2024.102148.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Enggie Hendrawan Saputra, Ilham Ari Elbaith Zaeni, Didik Dwi Prasetya, Azlan Mohd Zain, Welly Antonius, I Made Wirawan

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Authors retain copyright and full publishing rights to their articles. Upon acceptance, authors grant Indonesian Journal of Data and Science a non-exclusive license to publish the work and to identify itself as the original publisher.
Self-archiving. Authors may deposit the submitted version, accepted manuscript, and version of record in institutional or subject repositories, with citation to the published article and a link to the version of record on the journal website.
Commercial permissions. Uses intended for commercial advantage or monetary compensation are not permitted under CC BY-NC 4.0. For permissions, contact the editorial office at ijodas.journal@gmail.com.
Legacy notice. Some earlier PDFs may display “Copyright © [Journal Name]” or only a CC BY-NC logo without the full license text. To ensure clarity, the authors maintain copyright, and all articles are distributed under CC BY-NC 4.0. Where any discrepancy exists, this policy and the article landing-page license statement prevail.










