A toy greedy policy, not the trained Qwen2.5-3B agent. It reproduces the failure SurviveCity documented: reward can climb while agents starve.
01 OpenEnv Multi-Agent RL
SurviveCity
Reward design is part of the system—not a score added afterward.
- Figured out
- Agents optimized the rubric in ways that failed the actual survival objective.
- Broke
- The apparent policy-collapse signal was actually starvation.
- Built better
- Debug the environment and reward signals before blaming the policy.
OutcomeAgent survival rate went from 15% to 60%, four times higher; 3 reward-hacking exploits were closed.
