Francesco
Gentile

TL;DR PhD student at the University of Trento decoding the internal logic of vision models through mechanistic interpretability.

Index.01
Personal Profile

Profile.

Academic background, research motivations, and the journey toward model transparency.

Currently, I am a PhD student at the University of Trento working within the MHUG Lab under the supervision of Prof. Elisa Ricci. My work is based at the intersection of deep learning and model transparency, focusing on how we can better understand the decisions made by complex neural architectures.

My current research focuses on weight-based mechanistic interpretability within vision models. I am specifically interested in decomposing the weights of pre-trained models into simpler, more manageable components. Then, by analyzing how these subcomponents compose and interact, I aim to provide a more granular view of the computations performed by vision architectures.

Prior to my PhD, I completed both my Bachelor's in Computer Science and my Master's in Artificial Intelligence at the University of Trento. During these years, my research interests were quite broad, covering areas such as Temporal Graph Learning, Topological Deep Learning, and Scene Understanding. These projects provided a foundation for my current focus on the internal structure and compositionality of deep learning models.

Research Focus

01 Mechanistic Interpretability
Focus
02 Causality
Focus
03 Vision Models
Focus
04 Temporal Graph Learning
Foundational
05 Topological DL
Foundational
06 Scene Understanding
Foundational
Spotlight.02

Research Spotlight.

Active Focus

Investigating how deep neural networks represent complex visual concepts. My current research focuses on reverse-engineering vision transformers (ViTs) to understand which features they learn, how they process visual information, and how to make such understandings actionable.

Computer Vision and Pattern Recognition (CVPR) · 2026

From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition

"As vision-language models are deployed at scale, understanding their internal mechanisms becomes increasingly critical. Existing interpretability methods predominantly rely on activations, making them dataset-dependent, vulnerable to data bias, and often restricted to coarse head-level explanations. We introduce SITH (Semantic Inspection of Transformer Heads), a fully data-free, training-free framework that directly analyzes CLIP's vision transformer in weight space. For each attention head, we decompose its value-output matrix into singular vectors and interpret each one via COMP (Coherent Orthogonal Matching Pursuit), a new algorithm that explains them as sparse, semantically coherent combinations of human-interpretable concepts. We show that SITH yields coherent, faithful intra-head explanations, validated through reconstruction fidelity and interpretability experiments. This allows us to use SITH for precise, interpretable weight-space model edits that amplify or suppress specific concepts, improving downstream performance without retraining. Furthermore, we use SITH to study model adaptation, showing how fine-tuning primarily reweights a stable semantic basis rather than learning entirely new features."

Folio.03

Selected Projects.

Course Work NO.01

Continuous-Time Dynamic Graphs for Information Cascade Prediction

In this work, we propose a novel approach to information cascade prediction that leverages continuous-time dynamic graphs to capture the complex interactions and dynamics among cascades on a global scale; second, we introduce a mechanism for selecting the most informative neighbors when making predictions in memory-based models for continuous-time dynamic graphs.

#Topological DL#Temporal Graphs#Information Cascade Prediction
Coming Soon
Thesis NO.02

Reviving Graph Neural Networks for Human-Object Interaction Detection

This thesis addresses the challenges of Human-Object Interaction (HOI) detection by introducing three novel methodologies for complex scene understanding. It explores a graph-based approach to structure entity relationships, a hypergraph model to capture higher-order interactions beyond simple pairs, and a vision-language framework that leverages large language models to guide visual attention.

#HOI Detection#Topological DL#Scene Understanding
Coming Soon
Course Work NO.03

Estimation of Distribution using ENergy-based models

This work introduces Estimation of Distribution using ENergy-based models (EDEN), a novel Estimation of Distribution Algorithm (EDA) for black-box optimization. EDEN leverages a neural network equipped with hypergraph convolutions to approximate a population's fitness landscape as an energy-based probability model. Candidate solutions are subsequently generated by sampling this model using modified Langevin dynamics with adaptive noise.

#Optimization#Energy-based Models#Topological DL
Coming Soon
Index.04
Bulletins & Field Notes

Latest News.

NO.01
Visit us at CVPR Poster Session 1 - Poster 267 to discuss about SITH!
Bulletin
NO.02
Starting my position as Junior Researcher at FBK!
Bulletin