Diffusion-based Text-to-Image (T2I) models have achieved impressive success in generating high-quality images from textual prompts. While large language models (LLMs) effectively leverage Direct Preference Optimization (DPO) for fine-tuning on human preference data without the need for reward models, diffusion models have not been extensively explored in this area. Current preference learning methods applied to T2I diffusion models immediately adapt existing techniques from LLMs. However, this direct adaptation introduces an estimated loss specific to T2I diffusion models. This estimation can potentially lead to suboptimal performance through our empirical results. In this work, we propose Direct Score Preference Optimization (DSPO), a novel algorithm that aligns the pretraining and fine-tuning objectives of diffusion models by leveraging score matching, the same objective used during pretraining. It introduces a new perspective on preference learning for diffusion models. Specifically, DSPO distills the score function of human-preferred image distributions into pretrained diffusion models, fine-tuning the model to generate outputs that align with human preferences. We theoretically show that DSPO shares the same optimization direction as reinforcement learning algorithms in diffusion models under certain conditions. Our experimental results demonstrate that DSPO outperforms preference learning baselines for T2I diffusion models in human preference evaluation tasks and enhances both visual appeal and prompt alignment of generated images.
more »
« less
This content will become publicly available on July 6, 2027
Iterative Dual-Model Alignment for Story Evaluation
Large language models (LLMs) can both evaluate and explain text quality; however, most existing evaluators operate as static classifiers and lack the ability to refine their reasoning through interaction. We propose an \textbf{Iterative Alpha--Beta Learning} framework that jointly trains two complementary 8B models: an Alpha () classifier that assesses pairwise story engagement, and a Beta () generator that produces structured, rubric-guided comparative explanations. The two models co-evolve within a closed feedback loop: provides probabilistic preference signals to guide ’s Direct Preference Optimization (DPO), while ’s improved explanations are reintegrated to retrain via a KL-based contrastive objective. This dual optimization enables mutual learning: gains interpretability and robustness from ’s textual rationales, while acquires stronger alignment and discriminative precision from ’s confidence deltas. Experiments on human-annotated story-pair datasets HANNA show that the proposed system consistently outperforms strong single-model baselines in both accuracy and explanation quality across multiple iterative rounds.
more »
« less
- Award ID(s):
- 2048001
- PAR ID:
- 10688182
- Publisher / Repository:
- Association for Computational Linguistics (ACL 2026)
- Date Published:
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
As large language models increasingly drive real-world applications, aligning them with human values becomes paramount. Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique, translating preference data into reward models when oracle human values remain inaccessible. In practice, RLHF mostly relies on approximate reward models, which may not consistently guide the policy toward maximizing the underlying human values. We propose Policy-Interpolated Learning for Aligned Feedback (PILAF), a novel response sampling strategy for preference labeling that explicitly aligns preference learning with maximizing the underlying oracle reward. PILAF is theoretically grounded, demonstrating optimality from both an optimization and a statistical perspective. The method is straightforward to implement and demonstrates strong performance in iterative and online RLHF settings where feedback curation is critical.more » « less
-
Large language models (LLMs) have achieved impressive performance but face high computational costs and latency, limiting their deployment in resource-constrained settings. In contrast, small-scale LLMs (SLMs) are more efficient yet struggle to capture evolving real-world knowledge. Retrieval-augmented generation (RAG) helps by integrating external knowledge, but imperfect retrieval can introduce distracting noise that misleads SLMs. We propose {\name}, a robust RAG framework for SLMs via Margin-aware Preference Optimization. {\name} employs multi-turn prompting for detailed reasoning, rejection sampling for high-quality explanations, and contrastive preference selection to refine responses by maximizing the likelihood gap between preferred and non-preferred outputs.more » « less
-
Preference learning, or the task of aligning generative models to preference comparison data, has yet to reach the conceptual maturity of classification, density estimation, etc. To close this gap, this work presents a framework to understand preference learning starting from the sampling distribution of pairwise preference data. First, we prove that the only evaluation of a generative model that respects both preferences and prevalences in the data distribution is a form of win rate, justifying win rate as the focal point to understand preference learning. We then analyze preference learning methods as win rate optimization (WRO) or non-WRO. We present novel instances of WRO beyond existing examples (RLHF, NLHF) and identify two key theoretical benefits of all such methods. We prove that common non-WRO methods like DPO and SFT on preferred samples lack these properties and suggest ways to mitigate such theoretical limitations. We also show that WRO underperforms in practice due optimization difficulties and that optimization success predicts performance better than choices which affect the objective's solution. Our analysis highlights best practices for existing methods and provides recommendations for future research, guided by the principle that one should either align non-WRO methods more closely with WRO or improve the optimization of WRO objectives.more » « less
-
Corneous proteins are an important component of the tetrapod integument. Duplication and diversification of keratins and associated proteins are linked with the origin of most novel integumentary structures like mammalian hair, avian feathers, and scutes covering turtle shells. Accordingly, the loss of integumentary structures often coincides with the loss of genes encoding keratin and associated proteins. For example, many hair keratins in dolphins and whales have become pseudogenes. The adhesive setae of geckos and anoles are composed of both intermediate filament keratins (IF-keratins, formerly known as alpha-keratins) and corneous beta-proteins (CBPs, formerly known as beta-keratins) and recent whole genome assemblies of two gecko species and an anole uncovered duplications in seta-specific CBPs in each of these lineages. While anoles evolved adhesive toepads just once, there are two competing hypotheses about the origin(s) of digital adhesion in geckos involving either a single origin or multiple origins. Using data from three published gecko genomes, I examine CBP gene evolution in geckos and find support for a hypothesis where CBP gene duplications are associated with the repeated evolution of digital adhesion. Although these results are preliminary, I discuss how additional gecko genome assemblies, combined with phylogenies of keratin and associated protein genes and gene duplication models, can provide rigorous tests of several hypotheses related to gecko CBP evolution. This includes a taxon sampling strategy for sequencing and assembly of gecko genomes that could help resolve competing hypotheses surrounding the origin(s) of digital adhesion.more » « less
An official website of the United States government
