DOI: 10.1002/cpe.70960 ISSN: 1532-0626

SurroTune: A Surrogate‐Assisted Auto‐Tuning Framework for Spark Lambda Architectures With SHAP Validation

Kahina Kessi, Ibtisam Ferrahi, Hamza Nemouchi

ABSTRACT

Configuring Apache Spark for distributed data processing remains a critical challenge in Big Data engineering, particularly in resource‐constrained containerized deployments where elastic scaling tools are unavailable. This paper presents SurroTune, a surrogate‐assisted auto‐tuning framework for Lambda architectures implemented on Spark, comprising batch and streaming processing layers. The framework operates through five sequential stages: benchmark collection, surrogate model training, Bayesian Optimization, multi‐objective Pareto optimization, and SHAP‐based explainability. Evaluated on 1620 real execution traces collected on a reproducible containerized cluster (Hadoop 3.2.1, Spark 3.5.0), SurroTune achieves for batch execution time prediction (CV , stable across five random seeds: ) and for streaming P99 latency prediction (CV ). Exhaustive empirical evaluation identifies configurations that reduce batch execution time by up to 47.1% compared to a conservative baseline (95% bootstrap CI: [25.7%, 66.1%], ). Surrogate‐guided Bayesian Optimization reaches within 5% of the predicted optimum after only 1 surrogate query, whereas Random Search requires approximately 10 real executions to reach equivalent proximity. For the Speed Layer, all nine tested scenarios satisfy a 5,000 ms SLA. We derive an interpretable analytical model for the Speed Layer: (CV ). Its contribution ranking aligns closely with SHAP attributions derived independently from the nonlinear surrogate, providing dual cross‐validation of the learned performance structure. All experimental artifacts are released for reproducibility. Our results are specific to the tested environment; the methodology is designed for extension to larger clusters.