- Laboratoire d’informatique Sorbonne Université - CNRS UMR 7606

Le LIP6 soutient la campagne Octobre Rose de prévention contre le cancer du sein

CHOUAKI Fares

Fares ChouakiDoctorant à Sorbonne Université (ATER, Sorbonne Université)
Équipe : SMA
Date d’arrivée : 11/10/2023
    Sorbonne Université - LIP6
    Boîte courrier 169
    Couloir 25-26, Étage 4, Bureau 420
    4 place Jussieu
    75252 PARIS CEDEX 05

01 44 27 88 40
Fares.Chouaki (at) nulllip6.fr
https://perso.lip6.fr/Fares.Chouaki/

Direction de recherche : Aurélie BEYNIER

Multiagent ethical behaviors in multi-objective reinforcement learning

Most decisions taken within a society affect several groups at once, and each of them judges the outcome against its own set of criteria. Reducing these criteria to a single quantity does not resolve the tension between them, it merely hides it. This thesis studies sequential decision problems in which the criteria are instead treated as distinct objectives. Solving these problems requires decisions to be taken in sequence, under dynamics that are not known in advance, and in pursuit of several conflicting objectives. We tackle these problems using multi-objective reinforcement learning, with the aim of learning good compromise solutions for complex trade-offs between the objectives.

We first consider the case where the trade-off is known in advance but expressed as a non-linear aggregation of the objectives. We look for a policy that realises this trade-off on every single run, rather than one that only realises it when its returns are averaged over many runs. We introduce MO-CDQN, a value-based distributional algorithm for this setting. It learns the full multivariate return distribution of every state-action pair, and conditions action selection on the return accumulated so far. We establish the soundness of the algorithm theoretically, and show empirically that it outperforms existing methods. An extension of the algorithm optimises a whole portfolio of aggregations from a single shared stream of experience.

We then remove the assumption that the trade-off is known, and ask how to recover a set of policies covering the achievable compromises so that a decision maker can choose among them afterwards. We introduce MO-CEM-RL, the first method to combine an evolution strategy with multi-objective reinforcement learning. A population of preference-conditioned policies is evolved by the Cross-Entropy Method while part of it is refined by gradient updates. The resulting solutions dominate those found by the leading population-based alternative on continuous-control benchmarks.

Next, we ask whether a set of policies obtained cheaply, for simple linear aggregations, can be reused at execution time to construct an efficient policy for a complex non-linear aggregation function revealed only afterwards. We formalise the notion of reuse as switching between members of this pre-trained set, either between executions or within one, and establish what each granularity can guarantee. We also give switching strategies for both cases. Each of them is proved to achieve a return value at least as high as the one obtained by following any of the previously pre-trained policies.

Finally, we turn to problems where the objectives are the concern of a team of cooperating agents. We argue that fairness between objectives must be assessed at each execution rather than on average, since a policy that appears balanced over many runs can be severely imbalanced on any given one. We then leverage multi-agent communication to propose a fully decentralised algorithm that optimise fairness under this evaluation criterion.

Taken together, these contributions provide a first set of algorithms for learning policies that optimise complex, non-linear aggregations of objectives. These policies are reliable, in the sense that the compromise they strike between the objectives holds at every execution rather than only on average. They also extend to problems involving several agents that must cooperate to satisfy the objectives.


Publications 2025-2026