DOI: 10.3390/computers15080505 ISSN: 2073-431X

A Trust-Aware Extension to a Reinforcement Learning Hyper-Heuristic Framework for Multi-Objective Scientific Workflow Scheduling

Hadeel Amjed Saeed, Sufyan T. Faraj Al-Janabi, Esam Taha Yassen, Omar A. Aldhaibani

A reinforcement learning hyper-heuristic framework for multi-objective scientific workflow scheduling selects among five meta-heuristic optimisers and tunes their control parameters under a Nash Social Welfare reward over makespan, cost, security, and resource utilisation. In its base form it treats security as a static virtual-machine attribute and admits all candidates unconditionally. This paper contributes the mechanism design required to integrate two security-realism layers into the scheduling loop without redesigning the reward: a five-stage zero-trust admission pipeline, a bounded non-stationary per-machine dynamic trust signal, a state-vector augmentation that exposes trust to the agent, and a coupling that attenuates the effective security level seen by the security utility. The two layers act at distinct timescales: admission is a provisioning-time gate on a machine’s structural compliance, whereas the trust signal evolves per decision epoch for the machines already admitted, so static admission and dynamic trust coexist by construction. We evaluate three hyper-heuristic agents on 20 Pegasus workflow instances under both a trust suite and a trust-free baseline. The base framework establishes a sharp separation between the hyper-heuristic and direct task-to-machine RL families; we treat this as an inherited property and ask a different question: can the two security-realism layers be integrated without disturbing it? Across 20 Pegasus instances and three HH-RL agents, the family-level separation is preserved. Under a reproducible evaluation protocol—five independently seeded repeats of the full paired comparison, 100 greedy inference episodes per (agent, workflow, suite) cell, with per-workflow deltas averaged across repeats before testing—the trust extension imposes a small, heterogeneous absorption cost: the median per-workflow shift in Nash reward is −0.24, −0.24, and −0.05 for PDQN, DQNHH, and QLHH respectively, an order of magnitude below the absolute reward levels. The shift is statistically significant for DQNHH (two-sided Wilcoxon p = 0.0014, rank-biserial r = −0.77), marginal for PDQN (p = 0.058), and absent for QLHH (p = 0.18). The security utility stays above 0.91 on every instance, and the family-level scaling robustness is preserved intact. The contribution is therefore a drop-in mechanism whose cost is bounded and small relative to the between-family separation—with a robust workflow-level heterogeneity: the parameterised agent converts the trust signal into consistent gains on the largest DAGs (mean +1.28 on Sipht_1000 and Inspiral_1000 across the five repeats) while paying a small cost on typical instances.

More from our Archive