The Laboratory is committed to responsible and trustworthy artificial intelligence, and the team is dedicated to developing advanced machine unlearning methodologies, inspired by the latest advances in model interpretability and deep learning research. We aim to apply these cutting-edge technologies to address complex challenges in various application fields, with a particular focus on privacy compliance and knowledge control in large language models.
In our laboratory, we combine theory and practice to design solutions that can radically transform the way trained models manage knowledge. We use innovative localization and routing techniques to identify where a concept is expressed in a model and selectively suppress it, thereby removing specific information while preserving the model’s overall utility across diverse architectures and tasks.
We are committed to ongoing research and development of methodologies that address current privacy needs and anticipate future directions in AI governance and regulation. Explore our site to learn more about our work, our publications, and how our research is helping to build more trustworthy and compliant intelligent systems.
Machine Unlearning involves the exploration and innovation of methods for selectively removing knowledge from trained models, such as Gradient Ascent, Preference Optimization, and neuron-level interventions. This pivotal area of research is directed towards making a model behave as though it had never seen certain data during training, without the prohibitive cost of retraining from scratch. Such advancements are crucial to enabling large language models to comply with privacy regulations, such as the GDPR’s right to be forgotten, while remaining efficient enough to run on modest hardware.
Potential topics:
I. Concept Localization & Interpretability (Focus: Finding Where Knowledge Lives)
Variance-Based Neuron Selection: Identifying concept-carrying neurons with a single untrained forward pass. Concept Expression Analysis: Locating where knowledge is expressed in a model, rather than merely where it is retrieved. Layer Asymmetry Exploitation: Leveraging the concentration of semantic content in later transformer blocks.
II. Efficient Intervention & Routing (Focus: Making Forgetting Practical)
Causal Routing: Attaching lightweight, input-conditional gating modules that suppress concept neurons at query time. Frozen-Base Unlearning: Changing model behavior without altering a single base model weight, ensuring full attribution of any change. Parameter-Efficient Suppression: Removing concepts using a tiny fraction of the model’s parameters, drastically cutting memory and compute costs.
III. Verification, Privacy & Trust (Focus: Proving Knowledge Is Gone)
Right to Be Forgotten: Aligning trained models with GDPR and CCPA data erasure obligations. Adversarial Robustness: Stress-testing unlearning against extraction, probing, and membership inference attacks. Utility–Forgetting Trade-offs: Ensuring that removing knowledge does not collapse the model’s general capabilities.
Involved Researchers:
- Federico Fontana
- Bardh Prenkaj

