NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

GIPCOL: Graph-Injected Soft Prompting for Compositional Zero-Shot Learning

https://doi.org/10.1109/WACV57701.2024.00567

Xu, Guangyue; Chai, Joyce; Kordjamshidi, Parisa (January 2024, IEEE)

Pre-trained vision-language models (VLMs) have achieved promising success in many fields, especially with prompt learning paradigm. In this work, we propose GIPCOL (Graph-Injected Soft Prompting for Compositional Learning) to better explore the compositional zero-shot learning (CZSL) ability of VLMs within the prompt-based learning framework. The soft prompt in GIPCOL is structured and consists of the prefix learnable vectors, attribute label and object label. In addition, the attribute and object labels in the soft prompt are designated as nodes in a compositional graph. The compositional graph is constructed based on the compositional structure of the objects and attributes extracted from the training data and consequently feeds the updated concept representation into the soft prompt to capture this compositional structure for a better prompting for CZSL. With the new prompting strategy, GIPCOL achieves state-of-the-art AUC results on all three CZSL benchmarks, including MIT-States, UT-Zappos, and C-GQA datasets in both closed and open settings compared to previous non-CLIP as well as CLIP-based methods. We analyze when and why GIPCOL operates well given the CLIP backbone and its training data limitations, and our findings shed light on designing more effective prompts for CZSL.
more » « less
Full Text Available
MetaReVision: Meta-Learning with Retrieval for Visually Grounded Compositional Concept Acquisition

https://doi.org/10.18653/v1/2023.findings-emnlp.818

Xu, Guangyue; Kordjamshidi, Parisa; Chai, Joyce (January 2023, Association for Computational Linguistics)

Humans have the ability to learn novel compositional concepts by recalling and generalizing primitive concepts acquired from past experiences. Inspired by this observation, in this paper, we propose MetaReVision, a retrievalenhanced meta-learning model to address the visually grounded compositional concept learning problem. The proposed MetaReVision consists of a retrieval module and a metalearning module which are designed to incorporate retrieved primitive concepts as a supporting set to meta-train vision-language models for grounded compositional concept recognition. Through meta-learning from episodes constructed by the retriever, MetaReVision learns a generic compositional representation that can be fast updated to recognize novel compositional concepts. We create CompCOCO and CompFlickr to benchmark the grounded compositional concept learning. Our experimental results show that MetaReVision outperforms other competitive baselines and the retrieval module plays an important role in this compositional learning process.
more » « less
Full Text Available
Zero-Shot Compositional Concept Learning

https://doi.org/10.18653/v1/2021.metanlp-1.3

Xu, Guangyue; Kordjamshidi, Parisa; Chai, Joyce (July 2021, Proceedings of the 1st Workshop on Meta Learning and Its Applications to Natural Language Processing)

Full Text Available

Search for: All records