A Multi-Algorithm Approach using Spatial Land use Data to Predict Phthalate Ester Concentrations in the Surface Water of U-Tapao Canal, Southern Thailand

Howard IC, Okpara KA and Howard CC

Published on: 2026-06-06

Abstract

Phthalate esters (PAEs) are widespread endocrine-disrupting chemicals present in the surface freshwaters of the world, yet no established modelling approach exists for the spatial prediction of PAE concentrations across the study area. Hence this, this work introduces a novel machine learning (ML)-based framework for spatial PAE prediction - dibutyl phthalate (DBP), di(2-ethylhexyl) phthalate (DEHP), diisononyl phthalate (DiNP), along with their sum (PAEs)) in the U-Tapao Canal, southern Thailand based on empirical surface water quality data measured by [1]. Empirically-derived PAE concentrations at 17 georeferenced sample locations (PAEs = 1.44-12.08 g/L) are used in combination with spatial predictors derived from a GIS land use map of the catchment. Four ML algorithms (Random Forest (RF), eXtreme Gradient Boosting (XGBoost), Long Short-Term Memory (LSTM) neural network, and Support Vector Regression (SVR)) are applied to the data-set and assessed in comparison to a multiple linear regression (MLR) baseline through a purpose-built leave-one-out cross validation (LOOCV) framework optimally adapted for small sample sizes. The XGBoost algorithm recorded the closest fit to the empirical data-set (R= 0.96; RMSE= 0.35 g/L), with RF yielding a strong predictive skill (R= 0.94). Subsequent analysis of the AI algorithms' feature importances revealed that (in order of importance): 1. Distance to industrial areas, 2. Land use class, and 3. DEHP concentration was most explanatory for PAEs prediction. A combined two-stage architecture-using a binary classifier to predict the presence or absence of PAEs in a water sample, fed into a subsequent linear regressor for PAE concentration-proved most effective for handling ND data-sets without bias introduced through data imputation. Despite limited empirical surface water measurements, these results indicate that robust ML models can be constructed to predict PAE hotspots and relative contaminant levels, and lays a supportive road-mapping framework for other emerging contaminant classes in tropical watersheds.

Keywords

Hazard index; Concentration addition; Endocrine disruptors; Phthalate esters; Mixture toxicity; Response addition; Toxic equivalency

Introduction

Phthalate esters (PAEs) are among the most extensively produced industrial plasticisers, with global annual production exceeding six million metric tonnes [2]. As PAEs are widely incorporated in a vast array of consumer and industrial products ranging from plasticizers in PVC and cosmetics to food packaging and medical devices, they have become ubiquitous contaminants of water resources, originating from wastewater discharges, agricultural runoff, and industrial processes [3]. Among other properties, PAEs are recognised endocrine disruptors and the long-chain derivatives of phthalates such as DEHP, DBP and DiNP have documented adverse effects including impacts on fertility, foetal development and the risk of neoplasia in aquatic biota [4].

Rapid industrialisation and agricultural intensification in Southeast Asia, coupled with the development of a lagging municipal wastewater treatment capacity, have led to the accumulation of significant quantities of PAEs loading into surface waters. In the study by [1], PAE contamination was observed in the eutrophic U-Tapao Canal of Songkhla Province, South Thailand, a canal serving as a link for land uses such as agriculture, industry, urbanization, and waste management prior to being discharged into the Songkhla Lake Basin. Concentration of observed PAEs was varied between 1.44 g/L and 12.08 g/L at 17 georeferenced sampling sites and DEHP was found at 100% detection frequency indicating anthropogenic persistence at the sampling site.  DBP and DiNP showed uneven spatial detection (47–65%), suggesting discrete land-use-driven point source inputs.

Despite mounting evidence of PAE pollution in tropical water bodies, predicting the spatial concentration distribution of these contaminants from readily available geospatial data remains an unmet scientific need. Traditional water quality assessments are expensive, spatially coarse, and retrospective. ML provides a potential alternative to capture the non-linear relationship between predictor and response variables and enable interpolation in the unsampled areas to produce a spatial continuum prediction model based on data [5,6]. The widely used ML algorithms for prediction of water quality parameters were an ensemble of models like Random Forest (RF), XGBoost; and deep learning models such as LSTM [7,8]. Nevertheless, to the best of the authors' knowledge, these approaches have not previously been applied to PAEs in a small-sample tropical canal system.

Key methodological challenges addressed by this work include: (i) handling non-detect (ND) values in the response variable; (ii) engineering spatially meaningful features from GIS land use data; (iii) reliable model selection under small sample sizes (n = 17); and (iv) providing interpretable post-hoc explanations of model behaviour. This paper addresses these challenges by developing an ML-based PAE prediction framework for the U-Tapao Canal through five objectives: (1) constructing spatially relevant predictor variables from the [1] dataset and a GIS land use map; (2) training and comparing RF, XGBoost, LSTM, and SVR against an MLR baseline; (3) implementing a dual-stage detection–concentration architecture to eliminate the need for ND imputation; (4) quantifying predictor importance via SHapley Additive exPlanations (SHAP); and (5) producing spatial hotspot predictions across the canal catchment.

For Full-length paper, please go through this pdf file: