TY - GEN
T1 - Comparative Analysis of Model Selection Criteria for Symbolic Regression Using Genetic Programming
AU - Ramlan, Fitria Wulandari
AU - Kronberger, Gabriel
AU - O’Riordan, Colm
AU - McDermott, James
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
PY - 2026
Y1 - 2026
N2 - Symbolic regression (SR) using genetic programming (GP) can generate a diverse set of candidate models that balance accuracy and complexity, particularly when configured with multi-objective optimisation, which produces a Pareto front of non-dominated solutions. However, selecting a single model from this population remains challenging, especially when relying only on training data. This study evaluates the effectiveness of model selection criteria in SR, which include Mean Squared Error (MSE), Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), Description Length (DL), and PySR Score Metric (PSM). These criteria are evaluated using their training scores on 20 real-world regression datasets from the PMLB collection using PySR. We calculate the Spearman rank correlation coefficient (ρ) between each metric and test MSE to evaluate how well the metric ranks generalisable models. The results show that no single metric performs reliably across all datasets. Metrics that focus mainly on accuracy often lead to overfitting, while simplicity-based metrics can underfit. PSM aims to balance accuracy and complexity, but its performance is inconsistent, sometimes helpful, but often unstable across datasets. This study provides practical insights into the behaviour of model selection metrics in SR and offers guidance for selecting models that generalise well without overfitting.
AB - Symbolic regression (SR) using genetic programming (GP) can generate a diverse set of candidate models that balance accuracy and complexity, particularly when configured with multi-objective optimisation, which produces a Pareto front of non-dominated solutions. However, selecting a single model from this population remains challenging, especially when relying only on training data. This study evaluates the effectiveness of model selection criteria in SR, which include Mean Squared Error (MSE), Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), Description Length (DL), and PySR Score Metric (PSM). These criteria are evaluated using their training scores on 20 real-world regression datasets from the PMLB collection using PySR. We calculate the Spearman rank correlation coefficient (ρ) between each metric and test MSE to evaluate how well the metric ranks generalisable models. The results show that no single metric performs reliably across all datasets. Metrics that focus mainly on accuracy often lead to overfitting, while simplicity-based metrics can underfit. PSM aims to balance accuracy and complexity, but its performance is inconsistent, sometimes helpful, but often unstable across datasets. This study provides practical insights into the behaviour of model selection metrics in SR and offers guidance for selecting models that generalise well without overfitting.
KW - Genetic Programming
KW - Information Criteria
KW - Model Selection
KW - Multi-Objective Optimisation
KW - Symbolic Regression
UR - https://www.scopus.com/pages/publications/105029536387
U2 - 10.1007/978-3-032-15635-8_6
DO - 10.1007/978-3-032-15635-8_6
M3 - Conference contribution
AN - SCOPUS:105029536387
SN - 9783032156341
T3 - Communications in Computer and Information Science
SP - 91
EP - 108
BT - Computational Intelligence - 17th International Joint Conference, IJCCI 2025, Proceedings
A2 - Marcelloni, Francesco
A2 - Madani, Kurosh
A2 - van Stein, Niki
A2 - Filipe, Joaquim
PB - Springer
T2 - 17th International Joint Conference on Computational Intelligence, IJCCI 2025
Y2 - 22 October 2025 through 24 October 2025
ER -