Best! This research proposes selecting attributes for vision-language models directly from the target images themselves, rather than relying on large language models (LLMs) to generate descriptors based only on class names. Previous methods generated attributes conditioned solely on the label, which often led to misrepresentations, especially under distribution shifts.
This paper shows that descriptors generated without consulting images carry little visual evidence, and selecting attributes conditioned on images significantly improves accuracy and interpretability.