DOI: 10.1063/5.0341722 ISSN: 1070-6631

Reinforcement learning of fixed-budget moving-rod stirring for passive scalar mixing in Navier–Stokes–Brinkman flows

Hosung Lee, Dasan Kim, Byeongoh Hwang, Myungjoo Kang

Stirring-driven mixing is a fundamental mechanism for homogenizing passive scalars in fluids, but designing effective stirring trajectories remains challenging when the flow is generated by a constrained physical actuator rather than prescribed directly. We present a reinforcement-learning framework for passive scalar mixing in a two-dimensional circular container stirred by a moving rod. The fluid is modeled by the incompressible Navier–Stokes equations with Brinkman penalization for the moving rod and stationary wall, and the scalar evolves according to an advection–diffusion equation. The reinforcement-learning policy controls only the direction of rod motion, while the rod speed and total path length are fixed. This fixed-budget formulation enables a controlled comparison between learned and hand-designed stirring protocols by ensuring that all methods use the same actuation distance and speed. The learned policy is trained using proximal policy optimization and evaluated against hand-designed deterministic trajectories, including centered circular, off-center circular, W-shaped, and figure-eight paths, as well as smooth random trajectories. Mixing performance is quantified using normalized scalar variance, an H−1-type mix-norm proxy, early-stage scalar concentration snapshots, vorticity and strain-rate fields, mechanism-oriented diagnostics, stochastic rollout statistics, and sensitivity tests. In deterministic evaluation, the learned trajectory achieves a final normalized scalar variance of J/J0=0.090, compared with J/J0=0.391 for the strongest hand-designed deterministic baseline, W-path. The learned trajectory also obtains the lowest final normalized H−1 proxy, H−1/H0−1=0.030. In stochastic evaluation over 20 rollouts, the learned stochastic policy gives mean and median final variances of 0.395 and 0.446, comparable to 0.394 and 0.438 for smooth random trajectories. Thus, the stochastic comparison is interpreted as a robustness check rather than as evidence of stochastic superiority. The results show that reinforcement learning can discover effective moving-rod trajectories for scalar homogenization under fixed-speed, fixed-path length, actuator-constrained stirring.

More from our Archive