GIPCOL: Graph-Injected Soft Prompting for Compositional Zero-Shot Learning

Xu, Guangyue; Chai, Joyce; Kordjamshidi, Parisa

doi:10.1109/WACV57701.2024.00567

Citation Details

GIPCOL: Graph-Injected Soft Prompting for Compositional Zero-Shot Learning

Pre-trained vision-language models (VLMs) have achieved promising success in many fields, especially with prompt learning paradigm. In this work, we propose GIPCOL (Graph-Injected Soft Prompting for Compositional Learning) to better explore the compositional zero-shot learning (CZSL) ability of VLMs within the prompt-based learning framework. The soft prompt in GIPCOL is structured and consists of the prefix learnable vectors, attribute label and object label. In addition, the attribute and object labels in the soft prompt are designated as nodes in a compositional graph. The compositional graph is constructed based on the compositional structure of the objects and attributes extracted from the training data and consequently feeds the updated concept representation into the soft prompt to capture this compositional structure for a better prompting for CZSL. With the new prompting strategy, GIPCOL achieves state-of-the-art AUC results on all three CZSL benchmarks, including MIT-States, UT-Zappos, and C-GQA datasets in both closed and open settings compared to previous non-CLIP as well as CLIP-based methods. We analyze when and why GIPCOL operates well given the CLIP backbone and its training data limitations, and our findings shed light on designing more effective prompts for CZSL. more »

Award ID(s):: 2028626

PAR ID:: 10547204

Author(s) / Creator(s):: Xu, Guangyue; Chai, Joyce; Kordjamshidi, Parisa

Publisher / Repository:: IEEE

Date Published:: 2024-01-03

ISBN:: 979-8-3503-1892-0

Page Range / eLocation ID:: 5762 to 5771

Format(s):: Medium: X

Location:: Waikoloa, HI, USA

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.1109/WACV57701.2024.00567

More Like this