BaTCAVe: Trustworthy Explanations for Robot Behaviors

Summary
We introduce BaTCAVe, a Bayesian extension of concept activation vectors that attaches calibrated uncertainty estimates to concept-based explanations of robot decisions, making post-hoc interpretation of neural network policies trustworthy across both simulation platforms and real-world robotic systems.
BibTeX
@inproceedings{sagar2025batcave,
title={BaTCAVe: Trustworthy Explanations for Robot Behaviors},
author={Sagar, Som and Taparia, Aditya and Mankodiya, Harsh and Bidare, Pranav and Zhou, Yifan and Senanayake, Ransalu},
booktitle={2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
pages={13867--13874},
year={2025}
}