Skip to main content

Refine your search

Aurinko.

Doctoral defence of Saifullah Khan, MSc, 9.10.2026: Computational intelligence for improving performance of machine learning models via statistics – Correlations, imputation, outlier detection, and augmentation

The doctoral dissertation in the field of Computer Science will be examined at the Faculty of Science, Forestry and Technology, Kuopio campus.

What is the topic of your doctoral research? Why is it important to study the topic?

This doctoral research examines how computationally inexpensive and interpretable statistical methods can improve machine learning by enhancing data before and during model training. The work focuses on correlation analysis, data augmentation, missing-data imputation, feature reduction, outlier and noise control, and bounded activation functions. This is important because even advanced machine-learning models can perform poorly when their input data are incomplete, noisy, weakly relevant or highly variable. Improving data quality can increase prediction accuracy and generalisation without always requiring larger or more complex models, while also reducing computation and data-collection costs. This directly reflects the thesis objective of using statistical methods to reduce noise, enhance significant features and improve prediction performance.

What are the key findings or observations of your doctoral research?

The dissertation shows that relatively simple statistical techniques can substantially improve machine-learning performance. Polynomial Abet (PA) data augmentation achieved an average RRMSE of 0.464% versus 1.088% for polynomial feature transformation; a 57.40% improvement while doing so with 14.02% less computation time. Seasonal Mean Imputation (SMI) improved wind-power forecasting by using seasonal information when replacing missing values, improving SMAPE by 102.74% compared to mean imputation in the reported LSTM experiment for the evaluation metric of “ratio of Error Percentage Improvement in prediction to missing data percentage.” Skewness Driven Distribution Imputation (SDDI) offered a computationally simple, distribution-aware alternative for missing-data handling, while correlation-based feature selection showed that much smaller datasets can retain nearly the same predictive accuracy e.g. reduction of data by 60.5% while accuracy changed by only 3%. 

The thesis also introduces Tanned-ReLU, a bounded and smooth activation function that improved prediction performance across diverse multivariate datasets. The main novelty is the systematic use of statistics not merely for describing data, but for actively shaping what machine-learning models learn i.e. externally through preprocessing and internally through activation control e.g. Tanned-ReLU reduced RRMSE by up to 82.44% versus P-ReLU and also improved on L-ReLU and ReLU in the tested datasets.

How can the results of your doctoral research be utilised in practice?

The results can be used wherever machine-learning systems must work with incomplete, noisy, or highly variable data. The proposed methods can support more reliable forecasting in areas such as energy production, wireless networks, environmental monitoring, and other sensor-based applications. Correlation-based feature reduction can help organisations collect and process only the most useful variables, potentially reducing sensor, storage, integration, and computation costs. The imputation methods of SMI and SDDI can restore missing measurements, while PA can improve learning from limited datasets. Tanned-ReLU can be used as an alternative activation function in neural networks processing volatile multivariate data. The methods provide practical ways to improve ML performance without automatically increasing model complexity.

What are the key research methods and materials used in your doctoral research?

The dissertation is based on four peer-reviewed studies using empirical machine-learning experiments on real-world datasets. We developed and evaluated statistical methods for data augmentation, missing-data imputation, and activation function of neural networks. The materials included nuclear-power time series from 31 US states, operational wind-power data from wind power plant in Pakistan, 5G Vehicle-to-Everything quality-of-service data from Germany, and additional multivariate datasets for wind power (Finnish grid), and other small open source datasets, including IoT based water quality monitoring. 

The proposed methods were compared with established alternatives using models such as LSTM, GRU and XGBoost at the time. Performance was assessed with measures including RMSE, RRMSE, SMAPE, prediction accuracy, correlation, and computation time. Repeated testing and cross-validation were used where appropriate.

Is there something else about your doctoral dissertation you would like to share in the press release?

A central message of the dissertation is that more attention should be given to the pre-processing stage of machine learning. While research into developing increasingly complex models is important, the quality and structure of the data provided to those models strongly influence their performance, yet data pre-processing often receives comparatively less attention. 

This dissertation shows that carefully examining and improving data through statistical methods such as correlation analysis, imputation, augmentation, and noise reduction can lead to meaningful improvements in prediction, efficiency, and interpretability. The work therefore emphasizes that understanding how data are prepared before learning begins is just as important as improving the machine-learning model itself.

The doctoral dissertation of Saifullah Khan, MSc, entitled Computational intelligence for improving performance of machine learning models via statistics - Correlations, imputation, outlier detection, and augmentation will be examined at the Faculty of Science, Forestry and Technology, Kuopio campus. The opponent will be Associate Professor Juho Kannala, Aalto University, and the custos will be Professor Pekka Toivanen, University of Eastern Finland. Language of the public defence is English. 

For further information, please contact: 

Saifullah Khan, [email protected], tel. +358 50 327 0452