Comparative Evaluation of Machine Learning Models for Heavy Crude Oil Viscosity Prediction Using Repeated Nested Cross-Validation and Independent Holdout Testing
DOI:
https://doi.org/10.56705/ijodas.v7i2.455Keywords:
Heavy Crude Oil Viscosity, Machine Learning, Gradient Boosting, Repeated Nested Cross-Validation, Independent Holdout Testing, Comparative EvaluationAbstract
Accurate prediction of heavy crude oil viscosity supports reservoir engineering, production planning, and flow assurance. This study presents a rigorous comparative evaluation rather than a new machine-learning framework. The published dataset, derived from Kamel et al. and reproduced by Li et al., represents 28 Middle Eastern heavy crude-oil samples through 196 development measurements and 47 independent holdout measurements at 20–80 °C. The supplied workbook contains API gravity, temperature, C1, C2, C3, C4–C6, C7+, and viscosity, with a development viscosity range of 632.88–1267.65 cP; however, it does not include row-level oil or reservoir identifiers. Linear Regression, the Beggs–Robinson correlation, SVR, Random Forest, and Gradient Boosting were evaluated. Imputation, scaling, and hyperparameter selection were embedded in repeated nested cross-validation with five outer folds repeated twice and five inner folds. Final models were evaluated once on the untouched 47-record holdout, with bootstrap confidence intervals, corrected pairwise tests, residual diagnostics, and held-out permutation importance. Gradient Boosting achieved an internal R² of 0.99313 (95% CI: 0.99077–0.99549) and RMSE of 11.41 cP (10.03–12.80). On the independent holdout, it achieved R² = 0.99308 (bootstrap 95% CI: 0.98772–0.99637), RMSE = 8.43 cP, MAE = 6.64 cP, and MAPE = 0.78%. Corrected comparisons showed lower RMSE than Random Forest and SVR. Holdout diagnostics detected no statistically significant heteroscedasticity for Gradient Boosting, and extreme-value sensitivity produced similar performance. Temperature and C7+ were the dominant held-out predictors. These results are encouraging within the sampled domain, but the absent oil identifiers prevent group-disjoint validation, and no external reservoir dataset supports field-level generalization.
Downloads
References
[1] M. A. Oloso, M. G. Hassan, M. B. Bader-El-Den, and J. M. Buick, “Ensemble SVM for characterisation of crude oil viscosity,” J. Pet. Explor. Prod. Technol., vol. 8, no. 2, pp. 531–546, Jun. 2018, doi: https://doi.org/10.1007/s13202-017-0355-x.
[2] F. Souas, A. Safri, and A. Benmounah, “A review on the rheology of heavy crude oil for pipeline transportation,” Pet. Res., vol. 6, no. 2, pp. 116–136, Jun. 2021, doi: https://doi.org/10.1016/j.ptlrs.2020.11.001.
[3] E. Bahonar, M. Chahardowli, Y. Ghalenoei, and M. Simjoo, “New correlations to predict oil viscosity using data mining techniques,” J. Pet. Sci. Eng., vol. 208, p. 109736, Jan. 2022, doi: https://doi.org/10.1016/j.petrol.2021.109736.
[4] U. Sinha, B. Dindoruk, and M. Soliman, “Machine learning augmented dead oil viscosity model for all oil types,” J. Pet. Sci. Eng., vol. 195, p. 107603, Dec. 2020, doi: https://doi.org/10.1016/j.petrol.2020.107603.
[5] M. Ghanavati, M.-J. Shojaei, and A. R. S. A., “Effects of Asphaltene Content and Temperature on Viscosity of Iranian Heavy Crude Oil: Experimental and Modeling Study,” Energy & Fuels, vol. 27, no. 12, pp. 7217–7232, Dec. 2013, doi: https://doi.org/10.1021/ef400776h.
[6] H. S. Naji, “Comparative Study of the C7+ Characterization Methods: An Object-Oriented Approach,” Arab. J. Sci. Eng., vol. 36, no. 7, pp. 1423–1446, Nov. 2011, doi: https://doi.org/10.1007/s13369-011-0118-9.
[7] C. H. Whitson, “Effect of C7+ Properties on Equation-of-State Predictions,” Soc. Pet. Eng. J., vol. 24, no. 06, pp. 685–696, Dec. 1984, doi: https://doi.org/10.2118/11200-PA.
[8] M. Baghani, “Energy optimization in electrochemical oxidation for wastewater treatment using interpretable machine learning,” Chem. Eng. J. Adv., vol. 25, 2026, doi: https://doi.org/10.1016/j.ceja.2026.101044.
[9] Y. Bouslihim, “Predicting soil organic matter from color indices: economic and technical feasibility in semi-arid agricultural soils,” Carbon Res., vol. 5, no. 1, 2026, doi: https://doi.org/10.1007/s44246-025-00240-6.
[10] O. Alomair, E. Folad, and A. Elsharkawy, “Effects of asphaltene and resin contents on crude oil viscosity: Experimental and modeling studies,” Fuel, vol. 405, p. 136575, 2026, doi: https://doi.org/10.1016/j.fuel.2025.136575.
[11] Y. He et al., “rediction of High-Pressure Physical Properties of crude oil Using Explainable Machine Learning Models,” IEEE Access, vol. 13, pp. 4912–4931, 2025, doi: https://doi.org/10.1109/ACCESS.2024.3522595.
[12] K. Peiro Ahmady Langeroudy, P. Kharazi Esfahani, and M. R. Khorsand Movaghar, “Enhanced intelligent approach for determination of crude oil viscosity at reservoir conditions,” Sci. Rep., vol. 13, no. 1, p. 1666, Jan. 2023, doi: https://doi.org/10.1038/s41598-023-28770-2.
[13] J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Ann. Stat., vol. 29, no. 5, Oct. 2001, doi: https://doi.org/10.1214/aos/1013203451.
[14] C. Nadeau and Y. Bengio, “Inference for the {Generalization} {Error},” Mach. Learn., vol. 52, no. 3, pp. 239–281, Sep. 2003, doi: https://doi.org/10.1023/A:1024068626366.
[15] T. G. Dietterich, “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms,” Neural Comput., vol. 10, no. 7, pp. 1895–1923, Oct. 1998, doi: https://doi.org/10.1162/089976698300017197.
[16] C. Strobl, A.-L. Boulesteix, T. Kneib, T. Augustin, and A. Zeileis, “Conditional variable importance for random forests,” BMC Bioinformatics, vol. 9, no. 1, p. 307, Dec. 2008, doi: https://doi.org/10.1186/1471-2105-9-307.
[17] B. Dou, “Machine Learning Methods for Small Data Challenges in Molecular Science,” Chemical Reviews, vol. 123, no. 13. pp. 8736–8780, 2023, doi: https://doi.org/10.1021/acs.chemrev.3c00189.
[18] P. T. Zakowicz, “Machine learning helps predict early onset psychosis with serum protein biomarkers, neuropsychometry, and clinicodemographic data,” Sci. Rep., vol. 16, no. 1, 2026, doi: https://doi.org/10.1038/s41598-025-33765-2.
[19] O. Karal, “Performance comparison of different kernel functions in SVM for different k value in k-fold cross-validation,” Proc. - 2020 Innov. Intell. Syst. Appl. Conf. ASYU 2020, 2020, doi: https://doi.org/10.1109/ASYU50717.2020.9259880.
[20] Y. Nie, “Deep Melanoma classification with K-Fold Cross-Validation for Process optimization,” IEEE Med. Meas. Appl. MeMeA 2020 - Conf. Proc., 2020, doi: https://doi.org/10.1109/MeMeA49120.2020.9137222.
[21] J. M. Dahr, “A Hard Voting Ensemble Model of the Logistic Regression, Support Vector Machine and Random Forest for Network Intrusion Detection,” Inform. Slov., vol. 49, no. 27, pp. 43–56, 2025, doi: https://doi.org/10.31449/inf.v49i27.8573.
[22] Z. Liu, “Prediction of landfill gases concentration based on Grey Wolf Optimization – Support Vector Regression during landfill excavation process,” Waste Manag., vol. 198, pp. 128–136, 2025, doi: https://doi.org/10.1016/j.wasman.2025.02.040.
[23] J. C. M. Sánchez, “Improving wheat yield prediction through variable selection using Support Vector Regression, Random Forest, and Extreme Gradient Boosting,” Smart Agricultural Technology, vol. 10.. 2025, doi: https://doi.org/10.1016/j.atech.2025.100791.
[24] C. C. Onyekwena, “Support vector machine regression to predict gas diffusion coefficient of biochar-amended soil,” Appl. Soft Comput., vol. 127, 2022, doi: https://doi.org/10.1016/j.asoc.2022.109345.
[25] H. Su, “Support Vector Regression-Based Reduced- Reference Perceptual Quality Model for Compressed Point Clouds,” IEEE Trans. Multimed., vol. 26, pp. 6238–6249, 2024, doi: https://doi.org/10.1109/TMM.2023.3347638.
[26] V. L. Yerrabolu, “Performance Comparison of Random Forest Regressor and Support Vector Regression for Solar Energy Prediction,” Iop Conference Series Earth and Environmental Science, vol. 1375, no. 1. 2024, doi: https://doi.org/10.1088/1755-1315/1375/1/012013.
[27] A. Primantara, “Bagging System Performance Analysis Using Artificial Neural Network, Random Forest Regression, Linear Regression, and Support Vector Regression,” Digest of Technical Papers IEEE International Conference on Consumer Electronics. pp. 618–622, 2024, doi: https://doi.org/10.1109/ISCT62336.2024.10791247.
[28] M. Y. Shams, “Water quality prediction using machine learning models based on grid search method,” Multimed. Tools Appl., vol. 83, no. 12, pp. 35307–35334, 2024, doi: https://doi.org/10.1007/s11042-023-16737-4.
[29] Y. Boutahri, “Machine learning-based predictive model for thermal comfort and energy optimization in smart buildings,” Results Eng., vol. 22, 2024, doi: https://doi.org/10.1016/j.rineng.2024.102148.
Published
Issue
Section
License
Copyright (c) 2026 Enggie Hendrawan Saputra, Ilham Ari Elbaith Zaeni, Didik Dwi Prasetya, Azlan Mohd Zain, Welly Antonius, I Made Wirawan

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Authors retain copyright and full publishing rights to their articles. Upon acceptance, authors grant Indonesian Journal of Data and Science a non-exclusive license to publish the work and to identify itself as the original publisher.
Self-archiving. Authors may deposit the submitted version, accepted manuscript, and version of record in institutional or subject repositories, with citation to the published article and a link to the version of record on the journal website.
Commercial permissions. Uses intended for commercial advantage or monetary compensation are not permitted under CC BY-NC 4.0. For permissions, contact the editorial office at ijodas.journal@gmail.com.
Legacy notice. Some earlier PDFs may display “Copyright © [Journal Name]” or only a CC BY-NC logo without the full license text. To ensure clarity, the authors maintain copyright, and all articles are distributed under CC BY-NC 4.0. Where any discrepancy exists, this policy and the article landing-page license statement prevail.










