1UC Berkeley 2Johns Hopkins University
IEEE Conference on Decision and Control (CDC), 2026

Multi-agent reinforcement learning (MARL) is increasingly used to design learning-enabled agents that interact in shared environments. However, training MARL algorithms in general-sum games remains challenging: learning dynamics can become unstable, and convergence guarantees typically hold only in restricted settings such as two-player zero-sum or fully cooperative games. Moreover, when agents have heterogeneous and potentially conflicting preferences, it is unclear what system-level objective should guide learning. In this paper, we propose a new MARL pipeline called Near-Potential Policy Optimization (NePPO) for computing approximate Nash equilibria in mixed cooperative–competitive environments. The core idea is to learn a player-independent potential function such that the Nash equilibrium of a cooperative game with this potential as the common utility approximates a Nash equilibrium of the original game. To this end, we introduce a novel MARL objective such that minimizing this objective yields the best possible potential function candidate and consequently an approximate Nash equilibrium of the original game. We develop an algorithmic pipeline that minimizes this objective using zeroth-order gradient descent and returns an approximate Nash equilibrium policy. We empirically show the superior performance of this approach compared to popular baselines such as IPPO and MAPPO.
NePPO computes approximate Nash equilibria of a general-sum Markov game by learning a single, player-independent potential function and solving the cooperative game in which every player maximizes . If reproduces the change in each player's own utility under a unilateral deviation from that cooperative equilibrium, then the cooperative solution is an approximate equilibrium of the original game.
Let be a Nash equilibrium of the cooperative game with common utility , and let be player 's best response to under its own utility . For each player, measures the mismatch between the change in potential and the change in the player's own value along its best-response deviation. Each is non-negative, and if then is an -approximate Nash equilibrium of the original game (Theorem 3.1). NePPO therefore minimizes over . Unlike the defining condition of a Markov near-potential function, this only asks to be accurate around rather than uniformly over policy space — which is why NePPO can recover equilibria even in zero-sum games, where no global potential function exists.
is restricted to a parameterized family and the max over players is smoothed with a log-sum-exp, . Because depends on only through the solutions of two nested RL problems, its gradient is estimated with a two-point zeroth-order estimator. Each iteration:
The algorithm returns . Both solvers are modular: any cooperative MARL method can serve as the CoopGameSolver and any single-agent RL method as the RLSolver. The parameterization of is a design handle — restricting it to a structured class biases which equilibrium is selected without modifying any player's reward.
On the general-sum game of eq. (8), the candidate potential is the convex combination . The objective can be evaluated in closed form and is minimized (at zero) for every , even though the game is not a potential game.




Eq. (7) defines a family of two-player, two-action games indexed by : is the general-sum game above, is zero-sum, and for the unique equilibrium is fully mixed. Here the potential is the quadratic over the players' mixed strategies, with . NePPO attains the lowest max regret at every and recovers exact equilibria at , including the zero-sum case. IPPO cycles at the mixed equilibria (top figure), and MAPPO converges to non-equilibrium profiles because it maximizes the summed reward.
| NePPO | IPPO | MAPPO | |
|---|---|---|---|
| 0.0 | 0.000 | 0.000 | 0.500 |
| 0.2 | 0.000 | 0.000 | 0.800 |
| 0.4 | 0.036 | DNC | 1.100 |
| 0.6 | 0.052 | DNC | 1.400 |
| 0.8 | 0.034 | DNC | 1.700 |
| 1.0 | 0.000 | 0.011 | N/A |
A partially observable, general-sum Markov game from the Multi-Particle Environment suite with six agents: two heroes that collect food while avoiding capture, and four adversaries that pursue them, one of which sees the heroes and can broadcast to the rest. The potential is the discounted sum of a state-dependent softmax mixture of the agents' rewards, , so the learned potential can weight players differently in different states. Regret is measured per agent by training a PPO best response against the others' frozen policies.
| NePPO | IPPO | MAPPO | MADDPG | |
|---|---|---|---|---|
| Max regret | 11.14 | 23.90 | 51.78 | DNC |

@inproceedings{kalanther2026neppo,
author = {Kalanther, Addison and Bharvirkar, Sanika and Sastry, Shankar and Maheshwari, Chinmay},
title = {NePPO: Near-Potential Policy Optimization for General-Sum Multi-Agent Reinforcement Learning},
booktitle = {Proceedings of the 2026 65th IEEE Conference on Decision and Control (CDC)},
year = {2026},
month = {Dec},
organization = {IEEE},
note = {To appear}
}