| ||||
| ||||
![]() Title:Explaining Visual-Language Foundation Models for Histopathology: a Patch-Level Approach Conference:IEEE CBMS 2026 Tags:Explainable AI, Foundation models and Histopathology Abstract: Visual–language foundation models have recently become the state of the art in computational histopathology, enabling zero-shot classification and region-level interpretation via text–image similarity. However, it remains unclear whether these models rely on features that are semantically meaningful to human experts at the tile/patch level. In this work, we assess the alignment between model-derived saliency maps and specialist annotations for three visual-language foundation models: CONCH, PathGen, and MUSK. Using the model-agnostic P-IBISA method, we generate attribution maps for histopathology patches from the WSSS4LUAD and BCSS datasets and compare them to ground-truth semantic segmentation masks. Faithfulness is measured using the Confidence Increase metric, while spatial correspondence is evaluated via the DICE score. Results show that P-IBISA saliencies consistently achieve higher faithfulness than ground-truth annotations, indicating that the highlighted regions are more influential to the models’ predictions than human-labeled regions. Additionally, localization analysis reveals low overlap between saliency maps and expert annotations, suggesting that the models rely on features that do not fully correspond to human-interpretable tissue regions. These findings highlight a gap between model reasoning and human understanding, motivating future work toward integrating segmentation-aware regularization into multimodal foundation models for histopathology. Explaining Visual-Language Foundation Models for Histopathology: a Patch-Level Approach ![]() Explaining Visual-Language Foundation Models for Histopathology: a Patch-Level Approach | ||||
| Copyright © 2002 – 2026 EasyChair |
