Failures Are Fated, But Can Be Faded: Characterizing and Mitigating Unwanted Behaviors in Large-Scale Vision and Language Models

Summary
We introduce a deep reinforcement learning framework that maps the failure landscape of large vision and language models, then restructures that landscape with limited human feedback to mitigate accuracy failures, social biases, and misalignment.
BibTeX
@inproceedings{sagar2024failures,
title={Failures are fated, but can be faded: characterizing and mitigating unwanted behaviors in large-scale vision and language models},
author={Sagar, Som and Taparia, Aditya and Senanayake, Ransalu},
booktitle={Proceedings of the 41st International Conference on Machine Learning},
pages={42999--43023},
year={2024}
}