EvolveRL is a research project on self-improving language-model agents. Its framework, AERL (Adversarial Evolutionary Reinforcement Learning), replaces manual prompt engineering with an evolutionary loop: candidate prompts are generated, attacked by adversarial models, scored by an automated judge, and mutated into the next generation.
Approach
The framework has four components. An Evolutionary Prompt Writer/Improver generates and mutates candidates, borrowing crossover and mutation operators from evolutionary algorithms. Evolutionary Models are parallel LLM instances that differ in prompt and configuration, in the manner of population-based training. Adversarial Models craft scenarios that probe for weaknesses: ambiguous code specifications, contradictory function signatures, adverse market conditions. A Judge scores every output on domain-specific metrics and acts as the fitness function: top performers survive, the rest are discarded.
Configurations therefore emerge from measured performance rather than human judgment. Adversaries that repeatedly break candidates can be preserved, raising difficulty over generations.