Comparative Evaluation of Machine Learning Models for Heavy Crude Oil Viscosity Prediction Using Repeated Nested Cross-Validation and Independent Holdout Testing

Authors

  • Enggie Hendrawan Saputra Universitas Negeri Malang
  • Ilham Ari Elbaith Zaeni Universitas Negeri Malang
  • Didik Dwi Prasetya Universitas Negeri Malang
  • Azlan Mohd Zain Universiti Teknologi Malaysia
  • Welly Antonius PT Kilang Pertamina Internasional
  • I Made Wirawan Universitas Negeri Malang

DOI:

https://doi.org/10.56705/ijodas.v7i2.455

Keywords:

Heavy Crude Oil Viscosity, Machine Learning, Gradient Boosting, Repeated Nested Cross-Validation, Independent Holdout Testing, Comparative Evaluation

Abstract

Accurate prediction of heavy crude oil viscosity supports reservoir engineering, production planning, and flow assurance. This study presents a rigorous comparative evaluation rather than a new machine-learning framework. The published dataset, derived from Kamel et al. and reproduced by Li et al., represents 28 Middle Eastern heavy crude-oil samples through 196 development measurements and 47 independent holdout measurements at 20–80 °C. The supplied workbook contains API gravity, temperature, C1, C2, C3, C4–C6, C7+, and viscosity, with a development viscosity range of 632.88–1267.65 cP; however, it does not include row-level oil or reservoir identifiers. Linear Regression, the Beggs–Robinson correlation, SVR, Random Forest, and Gradient Boosting were evaluated. Imputation, scaling, and hyperparameter selection were embedded in repeated nested cross-validation with five outer folds repeated twice and five inner folds. Final models were evaluated once on the untouched 47-record holdout, with bootstrap confidence intervals, corrected pairwise tests, residual diagnostics, and held-out permutation importance. Gradient Boosting achieved an internal R² of 0.99313 (95% CI: 0.99077–0.99549) and RMSE of 11.41 cP (10.03–12.80). On the independent holdout, it achieved R² = 0.99308 (bootstrap 95% CI: 0.98772–0.99637), RMSE = 8.43 cP, MAE = 6.64 cP, and MAPE = 0.78%. Corrected comparisons showed lower RMSE than Random Forest and SVR. Holdout diagnostics detected no statistically significant heteroscedasticity for Gradient Boosting, and extreme-value sensitivity produced similar performance. Temperature and C7+ were the dominant held-out predictors. These results are encouraging within the sampled domain, but the absent oil identifiers prevent group-disjoint validation, and no external reservoir dataset supports field-level generalization.

Downloads

Download data is not yet available.

References

[1] M. A. Oloso, M. G. Hassan, M. B. Bader-El-Den, and J. M. Buick, “Ensemble SVM for characterisation of crude oil viscosity,” J. Pet. Explor. Prod. Technol., vol. 8, no. 2, pp. 531–546, Jun. 2018, doi: https://doi.org/10.1007/s13202-017-0355-x.

[2] F. Souas, A. Safri, and A. Benmounah, “A review on the rheology of heavy crude oil for pipeline transportation,” Pet. Res., vol. 6, no. 2, pp. 116–136, Jun. 2021, doi: https://doi.org/10.1016/j.ptlrs.2020.11.001.

[3] E. Bahonar, M. Chahardowli, Y. Ghalenoei, and M. Simjoo, “New correlations to predict oil viscosity using data mining techniques,” J. Pet. Sci. Eng., vol. 208, p. 109736, Jan. 2022, doi: https://doi.org/10.1016/j.petrol.2021.109736.

[4] U. Sinha, B. Dindoruk, and M. Soliman, “Machine learning augmented dead oil viscosity model for all oil types,” J. Pet. Sci. Eng., vol. 195, p. 107603, Dec. 2020, doi: https://doi.org/10.1016/j.petrol.2020.107603.

[5] M. Ghanavati, M.-J. Shojaei, and A. R. S. A., “Effects of Asphaltene Content and Temperature on Viscosity of Iranian Heavy Crude Oil: Experimental and Modeling Study,” Energy & Fuels, vol. 27, no. 12, pp. 7217–7232, Dec. 2013, doi: https://doi.org/10.1021/ef400776h.

[6] H. S. Naji, “Comparative Study of the C7+ Characterization Methods: An Object-Oriented Approach,” Arab. J. Sci. Eng., vol. 36, no. 7, pp. 1423–1446, Nov. 2011, doi: https://doi.org/10.1007/s13369-011-0118-9.

[7] C. H. Whitson, “Effect of C7+ Properties on Equation-of-State Predictions,” Soc. Pet. Eng. J., vol. 24, no. 06, pp. 685–696, Dec. 1984, doi: https://doi.org/10.2118/11200-PA.

[8] M. Baghani, “Energy optimization in electrochemical oxidation for wastewater treatment using interpretable machine learning,” Chem. Eng. J. Adv., vol. 25, 2026, doi: https://doi.org/10.1016/j.ceja.2026.101044.

[9] Y. Bouslihim, “Predicting soil organic matter from color indices: economic and technical feasibility in semi-arid agricultural soils,” Carbon Res., vol. 5, no. 1, 2026, doi: https://doi.org/10.1007/s44246-025-00240-6.

[10] O. Alomair, E. Folad, and A. Elsharkawy, “Effects of asphaltene and resin contents on crude oil viscosity: Experimental and modeling studies,” Fuel, vol. 405, p. 136575, 2026, doi: https://doi.org/10.1016/j.fuel.2025.136575.

[11] Y. He et al., “rediction of High-Pressure Physical Properties of crude oil Using Explainable Machine Learning Models,” IEEE Access, vol. 13, pp. 4912–4931, 2025, doi: https://doi.org/10.1109/ACCESS.2024.3522595.

[12] K. Peiro Ahmady Langeroudy, P. Kharazi Esfahani, and M. R. Khorsand Movaghar, “Enhanced intelligent approach for determination of crude oil viscosity at reservoir conditions,” Sci. Rep., vol. 13, no. 1, p. 1666, Jan. 2023, doi: https://doi.org/10.1038/s41598-023-28770-2.

[13] J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Ann. Stat., vol. 29, no. 5, Oct. 2001, doi: https://doi.org/10.1214/aos/1013203451.

[14] C. Nadeau and Y. Bengio, “Inference for the {Generalization} {Error},” Mach. Learn., vol. 52, no. 3, pp. 239–281, Sep. 2003, doi: https://doi.org/10.1023/A:1024068626366.

[15] T. G. Dietterich, “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms,” Neural Comput., vol. 10, no. 7, pp. 1895–1923, Oct. 1998, doi: https://doi.org/10.1162/089976698300017197.

[16] C. Strobl, A.-L. Boulesteix, T. Kneib, T. Augustin, and A. Zeileis, “Conditional variable importance for random forests,” BMC Bioinformatics, vol. 9, no. 1, p. 307, Dec. 2008, doi: https://doi.org/10.1186/1471-2105-9-307.

[17] B. Dou, “Machine Learning Methods for Small Data Challenges in Molecular Science,” Chemical Reviews, vol. 123, no. 13. pp. 8736–8780, 2023, doi: https://doi.org/10.1021/acs.chemrev.3c00189.

[18] P. T. Zakowicz, “Machine learning helps predict early onset psychosis with serum protein biomarkers, neuropsychometry, and clinicodemographic data,” Sci. Rep., vol. 16, no. 1, 2026, doi: https://doi.org/10.1038/s41598-025-33765-2.

[19] O. Karal, “Performance comparison of different kernel functions in SVM for different k value in k-fold cross-validation,” Proc. - 2020 Innov. Intell. Syst. Appl. Conf. ASYU 2020, 2020, doi: https://doi.org/10.1109/ASYU50717.2020.9259880.

[20] Y. Nie, “Deep Melanoma classification with K-Fold Cross-Validation for Process optimization,” IEEE Med. Meas. Appl. MeMeA 2020 - Conf. Proc., 2020, doi: https://doi.org/10.1109/MeMeA49120.2020.9137222.

[21] J. M. Dahr, “A Hard Voting Ensemble Model of the Logistic Regression, Support Vector Machine and Random Forest for Network Intrusion Detection,” Inform. Slov., vol. 49, no. 27, pp. 43–56, 2025, doi: https://doi.org/10.31449/inf.v49i27.8573.

[22] Z. Liu, “Prediction of landfill gases concentration based on Grey Wolf Optimization – Support Vector Regression during landfill excavation process,” Waste Manag., vol. 198, pp. 128–136, 2025, doi: https://doi.org/10.1016/j.wasman.2025.02.040.

[23] J. C. M. Sánchez, “Improving wheat yield prediction through variable selection using Support Vector Regression, Random Forest, and Extreme Gradient Boosting,” Smart Agricultural Technology, vol. 10.. 2025, doi: https://doi.org/10.1016/j.atech.2025.100791.

[24] C. C. Onyekwena, “Support vector machine regression to predict gas diffusion coefficient of biochar-amended soil,” Appl. Soft Comput., vol. 127, 2022, doi: https://doi.org/10.1016/j.asoc.2022.109345.

[25] H. Su, “Support Vector Regression-Based Reduced- Reference Perceptual Quality Model for Compressed Point Clouds,” IEEE Trans. Multimed., vol. 26, pp. 6238–6249, 2024, doi: https://doi.org/10.1109/TMM.2023.3347638.

[26] V. L. Yerrabolu, “Performance Comparison of Random Forest Regressor and Support Vector Regression for Solar Energy Prediction,” Iop Conference Series Earth and Environmental Science, vol. 1375, no. 1. 2024, doi: https://doi.org/10.1088/1755-1315/1375/1/012013.

[27] A. Primantara, “Bagging System Performance Analysis Using Artificial Neural Network, Random Forest Regression, Linear Regression, and Support Vector Regression,” Digest of Technical Papers IEEE International Conference on Consumer Electronics. pp. 618–622, 2024, doi: https://doi.org/10.1109/ISCT62336.2024.10791247.

[28] M. Y. Shams, “Water quality prediction using machine learning models based on grid search method,” Multimed. Tools Appl., vol. 83, no. 12, pp. 35307–35334, 2024, doi: https://doi.org/10.1007/s11042-023-16737-4.

[29] Y. Boutahri, “Machine learning-based predictive model for thermal comfort and energy optimization in smart buildings,” Results Eng., vol. 22, 2024, doi: https://doi.org/10.1016/j.rineng.2024.102148.

Published

2026-07-31

How to Cite

Comparative Evaluation of Machine Learning Models for Heavy Crude Oil Viscosity Prediction Using Repeated Nested Cross-Validation and Independent Holdout Testing. (2026). Indonesian Journal of Data and Science, 7(2). https://doi.org/10.56705/ijodas.v7i2.455