Skip to main navigation Skip to search Skip to main content

Deep reinforcement learning-based two-stage coevolutionary framework for multimodal multi-objective optimization

  • Qianlong DANG
  • , Xiaochuan GAO
  • , Shuwei HOU
  • , Zhengxin HUANG
  • , Guanghui ZHANG*
  • , Junhu RUAN
  • *Corresponding author for this work

Research output: Journal PublicationsJournal Article (refereed)peer-review

Abstract

In multimodal multi-objective optimization, different evolutionary operators with different characteristics are essential for reproducing offspring and searching for Pareto optimal solution sets. However, many existing multimodal multi-objective evolutionary algorithms (MMOEAs) are designed with a single evolutionary operator for reproduction, leading to their insufficient adaptability in solving different types of multimodal multi-objective optimization problems. To address the above problems, this paper proposes a deep reinforcement learning (DRL) based two-stage coevolutionary framework for MMOEA (MMOEA-DRL). Specifically, a DRL-based adaptive operator selection strategy is proposed, in which genetic algorithms and differential evolutionary operators are used as candidate evolutionary operators. Moreover, an offline strategy is employed to train a deep Q network (DQN) model, which determines the potential future payoffs of different evolutionary operators based on the current state of the population and the environment, and selects the more promising evolutionary operators. In addition, the proposed adaptive operator selection strategy is integrated into a two-stage coevolutionary framework. The DQN provides adaptive evolutionary operators to guide population evolution, and the large amount of diversified data generated by the population during the search process feeds back into the training of the DQN, which in turn enhances its strategy learning capability and decision accuracy. Experimental comparisons with MO_Ring_PSO_SCD, MMODE_CSCD, MMODE_ICD, MOEA/D-SS, and CMMO on CEC2019 and IDMP show that MMOEA-DRL obtains an optimal value of 64.71% and outperforms the nearest-performing algorithm by 55.74% for the performance indicator in the decision space. Furthermore, MMOEA-DRL still has the best distribution results on a real-world map-based distance minimization problem.

Original languageEnglish
Article number114836
JournalApplied Soft Computing
Volume193
Early online date18 Feb 2026
DOIs
Publication statusPublished - May 2026

Bibliographical note

Publisher Copyright:
© 2026 Elsevier B.V.

Funding

This work was supported in part by the National Nature Science Foundation of China under Grant 12501711, in part by Key Research and Development Program of Shaanxi under Grant 2024NC-XCZX-08, in part by the Postdoctoral Research Project of Shaanxi Province Grant 2025BSHEDZZ012, in part by the Natural Science Foundation of Shaanxi under Grant 2024JC-YBQN-0687 and 2024JC-YBMS-0591, in part by the China Postdoctoral Science Foundation under Grant 2024M762632, and in part by the Social Science Foundation of Shaanxi under Grant 2024R027.

Keywords

  • Adaptive operator selection
  • Deep Q network
  • Deep reinforcement learning
  • Evolutionary operator
  • Multimodal multi-objective optimization

Fingerprint

Dive into the research topics of 'Deep reinforcement learning-based two-stage coevolutionary framework for multimodal multi-objective optimization'. Together they form a unique fingerprint.

Cite this