Comparative Evaluation of Machine Learning Models for Heavy Crude Oil Viscosity Prediction Using Repeated Nested Cross-Validation and Independent Holdout Testing

Authors

  • Enggie Hendrawan Saputra Universitas Negeri Malang
  • Ilham Ari Elbaith Zaeni Universitas Negeri Malang
  • Didik Dwi Prasetya Universitas Negeri Malang
  • Azlan Mohd Zain Universiti Teknologi Malaysia
  • Welly Antonius PT Kilang Pertamina Internasional
  • I Made Wirawan Universitas Negeri Malang

DOI:

https://doi.org/10.56705/ijodas.v7i2.455

Keywords:

Heavy Crude Oil Viscosity, Machine Learning, Gradient Boosting, Repeated Nested Cross-Validation, Independent Holdout Testing, Comparative Evaluation

Abstract

Introduction: Accurate prediction of heavy crude oil viscosity is important for reservoir engineering, production planning, and flow assurance because viscosity strongly affects fluid mobility and transport behavior. This study comparatively evaluates established machine learning models under a rigorous validation protocol rather than proposing a new predictive framework. Method: A published Middle Eastern heavy crude-oil dataset containing 196 development measurements and 47 independent holdout measurements was used. Linear Regression, Support Vector Regression, Random Forest, Gradient Boosting, and the Beggs–Robinson correlation were evaluated using repeated nested cross-validation with five outer folds repeated twice and five inner folds. Preprocessing and hyperparameter selection were embedded within the validation pipeline, while the untouched holdout set was used only for final evaluation. Results and Discussion: Gradient Boosting achieved the best internal performance with R² = 0.99313 and RMSE = 11.41 cP. On the independent holdout set, it achieved R² = 0.99308, RMSE = 8.43 cP, MAE = 6.64 cP, and MAPE = 0.78%, outperforming Random Forest and Support Vector Regression. Residual diagnostics showed no detectable heteroscedasticity, while permutation importance identified temperature and C7+ as the dominant predictors. Conclusion: Gradient Boosting provides highly accurate viscosity predictions within the sampled domain; however, the absence of row-level oil identifiers and external reservoir data limits conclusions regarding oil-disjoint and field-level generalization.

Downloads

Download data is not yet available.

References

[1] M. A. Oloso, M. G. Hassan, M. B. Bader-El-Den, and J. M. Buick, “Ensemble SVM for characterisation of crude oil viscosity,” J. Pet. Explor. Prod. Technol., vol. 8, no. 2, pp. 531–546, Jun. 2018, doi: https://doi.org/10.1007/s13202-017-0355-x.

[2] F. Souas, A. Safri, and A. Benmounah, “A review on the rheology of heavy crude oil for pipeline transportation,” Pet. Res., vol. 6, no. 2, pp. 116–136, Jun. 2021, doi: https://doi.org/10.1016/j.ptlrs.2020.11.001.

[3] E. Bahonar, M. Chahardowli, Y. Ghalenoei, and M. Simjoo, “New correlations to predict oil viscosity using data mining techniques,” J. Pet. Sci. Eng., vol. 208, p. 109736, Jan. 2022, doi: https://doi.org/10.1016/j.petrol.2021.109736.

[4] U. Sinha, B. Dindoruk, and M. Soliman, “Machine learning augmented dead oil viscosity model for all oil types,” J. Pet. Sci. Eng., vol. 195, p. 107603, Dec. 2020, doi: https://doi.org/10.1016/j.petrol.2020.107603.

[5] M. Ghanavati, M.-J. Shojaei, and A. R. S. A., “Effects of Asphaltene Content and Temperature on Viscosity of Iranian Heavy Crude Oil: Experimental and Modeling Study,” Energy & Fuels, vol. 27, no. 12, pp. 7217–7232, Dec. 2013, doi: https://doi.org/10.1021/ef400776h.

[6] H. S. Naji, “Comparative Study of the C7+ Characterization Methods: An Object-Oriented Approach,” Arab. J. Sci. Eng., vol. 36, no. 7, pp. 1423–1446, Nov. 2011, doi: https://doi.org/10.1007/s13369-011-0118-9.

[7] C. H. Whitson, “Effect of C7+ Properties on Equation-of-State Predictions,” Soc. Pet. Eng. J., vol. 24, no. 06, pp. 685–696, Dec. 1984, doi: https://doi.org/10.2118/11200-PA.

[8] M. Baghani, “Energy optimization in electrochemical oxidation for wastewater treatment using interpretable machine learning,” Chem. Eng. J. Adv., vol. 25, 2026, doi: https://doi.org/10.1016/j.ceja.2026.101044.

[9] Y. Bouslihim, “Predicting soil organic matter from color indices: economic and technical feasibility in semi-arid agricultural soils,” Carbon Res., vol. 5, no. 1, 2026, doi: https://doi.org/10.1007/s44246-025-00240-6.

[10] O. Alomair, E. Folad, and A. Elsharkawy, “Effects of asphaltene and resin contents on crude oil viscosity: Experimental and modeling studies,” Fuel, vol. 405, p. 136575, 2026, doi: https://doi.org/10.1016/j.fuel.2025.136575.

[11] Y. He et al., “rediction of High-Pressure Physical Properties of crude oil Using Explainable Machine Learning Models,” IEEE Access, vol. 13, pp. 4912–4931, 2025, doi: https://doi.org/10.1109/ACCESS.2024.3522595.

[12] K. Peiro Ahmady Langeroudy, P. Kharazi Esfahani, and M. R. Khorsand Movaghar, “Enhanced intelligent approach for determination of crude oil viscosity at reservoir conditions,” Sci. Rep., vol. 13, no. 1, p. 1666, Jan. 2023, doi: https://doi.org/10.1038/s41598-023-28770-2.

[13] J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Ann. Stat., vol. 29, no. 5, Oct. 2001, doi: https://doi.org/10.1214/aos/1013203451.

[14] C. Nadeau and Y. Bengio, “Inference for the {Generalization} {Error},” Mach. Learn., vol. 52, no. 3, pp. 239–281, Sep. 2003, doi: https://doi.org/10.1023/A:1024068626366.

[15] T. G. Dietterich, “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms,” Neural Comput., vol. 10, no. 7, pp. 1895–1923, Oct. 1998, doi: https://doi.org/10.1162/089976698300017197.

[16] C. Strobl, A.-L. Boulesteix, T. Kneib, T. Augustin, and A. Zeileis, “Conditional variable importance for random forests,” BMC Bioinformatics, vol. 9, no. 1, p. 307, Dec. 2008, doi: https://doi.org/10.1186/1471-2105-9-307.

[17] B. Dou, “Machine Learning Methods for Small Data Challenges in Molecular Science,” Chemical Reviews, vol. 123, no. 13. pp. 8736–8780, 2023, doi: https://doi.org/10.1021/acs.chemrev.3c00189.

[18] P. T. Zakowicz, “Machine learning helps predict early onset psychosis with serum protein biomarkers, neuropsychometry, and clinicodemographic data,” Sci. Rep., vol. 16, no. 1, 2026, doi: https://doi.org/10.1038/s41598-025-33765-2.

[19] O. Karal, “Performance comparison of different kernel functions in SVM for different k value in k-fold cross-validation,” Proc. - 2020 Innov. Intell. Syst. Appl. Conf. ASYU 2020, 2020, doi: https://doi.org/10.1109/ASYU50717.2020.9259880.

[20] Y. Nie, “Deep Melanoma classification with K-Fold Cross-Validation for Process optimization,” IEEE Med. Meas. Appl. MeMeA 2020 - Conf. Proc., 2020, doi: https://doi.org/10.1109/MeMeA49120.2020.9137222.

[21] J. M. Dahr, “A Hard Voting Ensemble Model of the Logistic Regression, Support Vector Machine and Random Forest for Network Intrusion Detection,” Inform. Slov., vol. 49, no. 27, pp. 43–56, 2025, doi: https://doi.org/10.31449/inf.v49i27.8573.

[22] Z. Liu, “Prediction of landfill gases concentration based on Grey Wolf Optimization – Support Vector Regression during landfill excavation process,” Waste Manag., vol. 198, pp. 128–136, 2025, doi: https://doi.org/10.1016/j.wasman.2025.02.040.

[23] J. C. M. Sánchez, “Improving wheat yield prediction through variable selection using Support Vector Regression, Random Forest, and Extreme Gradient Boosting,” Smart Agricultural Technology, vol. 10.. 2025, doi: https://doi.org/10.1016/j.atech.2025.100791.

[24] C. C. Onyekwena, “Support vector machine regression to predict gas diffusion coefficient of biochar-amended soil,” Appl. Soft Comput., vol. 127, 2022, doi: https://doi.org/10.1016/j.asoc.2022.109345.

[25] H. Su, “Support Vector Regression-Based Reduced- Reference Perceptual Quality Model for Compressed Point Clouds,” IEEE Trans. Multimed., vol. 26, pp. 6238–6249, 2024, doi: https://doi.org/10.1109/TMM.2023.3347638.

[26] V. L. Yerrabolu, “Performance Comparison of Random Forest Regressor and Support Vector Regression for Solar Energy Prediction,” Iop Conference Series Earth and Environmental Science, vol. 1375, no. 1. 2024, doi: https://doi.org/10.1088/1755-1315/1375/1/012013.

[27] A. Primantara, “Bagging System Performance Analysis Using Artificial Neural Network, Random Forest Regression, Linear Regression, and Support Vector Regression,” Digest of Technical Papers IEEE International Conference on Consumer Electronics. pp. 618–622, 2024, doi: https://doi.org/10.1109/ISCT62336.2024.10791247.

[28] M. Y. Shams, “Water quality prediction using machine learning models based on grid search method,” Multimed. Tools Appl., vol. 83, no. 12, pp. 35307–35334, 2024, doi: https://doi.org/10.1007/s11042-023-16737-4.

[29] Y. Boutahri, “Machine learning-based predictive model for thermal comfort and energy optimization in smart buildings,” Results Eng., vol. 22, 2024, doi: https://doi.org/10.1016/j.rineng.2024.102148.

Downloads

Published

2026-07-31

How to Cite

Comparative Evaluation of Machine Learning Models for Heavy Crude Oil Viscosity Prediction Using Repeated Nested Cross-Validation and Independent Holdout Testing. (2026). Indonesian Journal of Data and Science, 7(2), 247-160. https://doi.org/10.56705/ijodas.v7i2.455