- Computer Science Laboratory Sorbonne Université - CNRS UMR 7606

LIP6 supports the Pink October campaign for breast cancer awareness.

BHAN Milan

Postdoc at Sorbonne University
Team : LFI

Supervision : Marie-Jeanne LESOT
Co-supervision : VITTAUT Jean-Noël

Explainable Natural Language Processing : Methods, Evaluation, and Practical Applications

This thesis focuses on Explainable Artificial Intelligence (XAI) for Natural Language Processing (NLP), with the objective of making the predictions of deep neural networks applied to text more intelligible. Modern NLP models achieve remarkable performance but remain largely opaque due to their billions of parameters, which earns them the designation of "black boxes". My work explores several explanation paradigms organized along three axes: generation, evaluation, and practical applications. The generation axis examines how to produce high-quality explanations. The TIGTEC method generates counterfactual explanations satisfying several desirable properties, such as sparsity, plausibility, and diversity. The notion of Counterfactual Feature Importance assigns an importance score to each token modification to improve their intelligibility. CT-CBM transforms a specialized NLP classifier into an explainable-by-design model, whose concept bottleneck layer is generated and targeted automatically. The evaluation axis measures explanation quality according to two protocols. CLS-A is an attribution method for attention-based classifiers, whose human evaluation shows that it improves both the accuracy and speed of users. NeuroFaith automatically measures the faithfulness of self-generated natural language explanations produced by large language models by comparing them to their internal activity, and exploits the linear structure of this faithfulness to improve it through intervention on latent activations. The application axis finally shows how XAI methods can serve practical objectives beyond explainability itself. Attribution methods and counterfactual generation are combined to detect and neutralize toxic content in text. Self-AMPLIFY automatically generates post-hoc explanatory justifications to enrich the in-context learning prompts of small language models, improving their performance on complex NLP tasks.


Phd defence : 06/11/2026

Jury members :

Vincent GUIGUE, Professeur AgroParisTech, MMIP, Palaiseau [Rapporteur]
Céline HUDELOT, Professeure CentraleSupélec, Gif-sur-Yvette [Rapporteur]
Alexandre ALLAUZEN, Professeur Université Paris-Dauphine, LAMSADE, Paris
Thomas GUYET, Chargé de recherche INRIA, AIstroSight, Lyon
Benjamin PIWOWARSKI, Chargé de recherche CNRS, MLIA, Paris
Benoît SAGOT, Directeur de recherche Inria, ALMAnaCH Paris

Departure date : 06/30/2026

2023-2026 Publications