X -Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs

Sobal, Vlad; Ibrahim, Mark; Balestriero, Randall; Cabannes, Vivien; Bouchacourt, Diane; Astolfi, Pietro; Cho, Kyunghyun; LeCun, Yann

Citation Details

This content will become publicly available on April 24, 2026

X -Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs

Learning good representations involves capturing the diverse ways in which data samples relate. Contrastive loss - an objective matching related samples - underlies methods from self-supervised to multimodal learning. Contrastive losses, however, can be viewed more broadly as modifying a similarity graph to indicate how samples should relate in the embedding space. This view reveals a shortcoming in contrastive learning: the similarity graph is binary, as only one sample is the related positive sample. Crucially, similarities \textit{across} samples are ignored. Based on this observation, we revise the standard contrastive loss to explicitly encode how a sample relates to others. We experiment with this new objective, called X -Sample Contrastive, to train vision models based on similarities in class or text caption descriptions. Our study spans three scales: ImageNet-1k with 1 million, CC3M with 3 million, and CC12M with 12 million samples. The representations learned via our objective outperform both contrastive self-supervised and vision-language models trained on the same data across a range of tasks. When training on CC12M, we outperform CLIP by on both ImageNet and ImageNet Real. Our objective appears to work particularly well in lower-data regimes, with gains over CLIP of on ImageNet and on ImageNet Real when training with CC3M. Finally, our objective seems to encourage the model to learn representations that separate objects from their attributes and backgrounds, with gains of - \% over CLIP on ImageNet9. We hope the proposed solution takes a small step towards developing richer learning objectives for understanding sample relations in foundation models. more »

Award ID(s):: 1922658

PAR ID:: 10649825

Author(s) / Creator(s):: Sobal, Vlad; Ibrahim, Mark; Balestriero, Randall; Cabannes, Vivien; Bouchacourt, Diane; Astolfi, Pietro; Cho, Kyunghyun; LeCun, Yann

Publisher / Repository:: The International Conference on Learning Representations (ICLR 2025)

Date Published:: 2025-04-24

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
This content will become publicly available on April 24, 2026
Conference Paper:
The DOI is not currently available.

More Like this