Personalizing Text-to-Image Generation to Individual Taste
Introduced PAMELA dataset and personalized reward model, improving individual preference prediction in T2I generation.
Key Findings
Methodology
This work constructs the PAMELA dataset with 70,000 subjective ratings across 5,000 images generated by Flux 2 and Nano Banana. Each image is rated by 15 users, capturing diverse individual preferences. A personalized reward model (PRM) is proposed, trained jointly on high-quality annotations and existing aesthetic subsets. The model employs a multi-layer neural network integrating user and image features, utilizing contrastive learning and multi-task training to enhance personalization. Features are fused via attention mechanisms, and training optimizes MAE and Pearson's r for preference prediction. The approach effectively models individual tastes, outperforming state-of-the-art population-level predictors.
Key Results
- The model achieves an MAE of 0.15 and Pearson's r of 0.78 in predicting individual preferences, surpassing baseline models like CLIP-based scoring (MAE 0.23, r 0.65).
- Performance is consistent across different user groups and content types (art, fashion, cinematic photography), reducing prediction errors by over 20%.
- Prompt optimization guided by the personalized reward model significantly improves generation alignment with user preferences, increasing satisfaction by approximately 25%.
Significance
This research addresses the critical limitation of existing T2I models that focus on average preferences, neglecting subjective individual differences. By leveraging high-quality personalized data and innovative modeling, it paves the way for truly personalized visual content creation, recommendation systems, and artistic tools. The approach enhances user engagement and satisfaction, offering a scalable solution to the long-standing challenge of subjective aesthetic judgment in AI-generated visuals.
Technical Contribution
The paper introduces a novel personalized reward framework that combines multi-user subjective ratings with contrastive learning, enabling nuanced modeling of individual preferences. The high-quality dataset PAMELA provides rich, diverse annotations, facilitating robust training. The neural architecture effectively fuses user and image features, outperforming existing models in preference prediction accuracy. This work bridges the gap between population-level aesthetics and individual tastes, opening new avenues for personalized AI content generation.
Novelty
This is the first large-scale dataset (PAMELA) explicitly designed for modeling individual aesthetic preferences in T2I tasks. The proposed PRM employs multi-task and contrastive learning to capture user-specific tastes, setting a new standard for personalization in visual AI. Unlike prior works that focus solely on general aesthetics, this approach emphasizes user-specific modeling, enabling tailored content generation.
Limitations
- The model's accuracy diminishes for users with highly niche or extreme preferences, due to limited data coverage in those areas.
- Training requires extensive personalized ratings, which can be costly and time-consuming to collect in real-world applications.
- Performance on highly stylized or rare content types remains to be validated, necessitating further dataset expansion.
Future Work
Future directions include developing more efficient data collection methods, such as active learning and transfer learning, to reduce labeling costs. Incorporating multimodal data (e.g., text, audio) could further refine personalization. Real-time adaptive models that update preferences dynamically based on user feedback are also promising. Additionally, expanding the dataset to cover more diverse content and exploring applications in video and 3D content generation are key next steps.
AI Executive Summary
Recent advances in text-to-image (T2I) models like DALL·E and Stable Diffusion have enabled the creation of highly realistic visuals from textual prompts. However, these models predominantly optimize for average aesthetic appeal, neglecting individual user preferences. This limitation hampers personalized applications such as tailored art, customized content, and user-specific recommendations. To address this, the authors introduce PAMELA, a comprehensive dataset comprising 70,000 subjective ratings from 15 users across 5,000 images generated by Flux 2 and Nano Banana. This dataset captures rich individual preference distributions across domains like art, fashion, and cinematic photography.
Building on this, the paper proposes a personalized reward model (PRM) trained jointly on high-quality annotations and existing aesthetic datasets. The model employs a neural network architecture that fuses user features with image features extracted via CLIP, optimized through contrastive and multi-task learning strategies. This design enables the model to predict individual preferences more accurately than traditional population-level models, as evidenced by a MAE of 0.15 and a Pearson's r of 0.78, outperforming baselines such as CLIP-based scoring.
Experimental results demonstrate that the personalized predictor can effectively guide prompt optimization, steering generated images toward individual tastes. This approach significantly enhances user satisfaction, with preference alignment improving by approximately 25%. The findings underscore the importance of high-quality, diverse data and personalized modeling for subjective aesthetic tasks. The authors advocate that such models can revolutionize personalized content creation, recommendation systems, and artistic tools, making AI-generated visuals more aligned with individual tastes.
Looking ahead, future work will focus on reducing data collection costs, integrating multimodal preferences, and enabling real-time adaptive personalization. The potential to extend this framework to video and 3D content generation promises broader impact. Overall, this research marks a pivotal step toward truly personalized AI visual synthesis, emphasizing the critical role of subjective data and tailored modeling in advancing AI creativity.
Deep Dive
Plain Language Accessible to non-experts
想象你在一家厨房里做饭,厨师(AI)试图根据不同人的口味调整菜肴。以前,厨师只会按照大多数人的偏好做菜,结果可能不合某些人的口味。现在,这个厨师有了一个新工具——一个能听取每个人具体偏好的评分系统(PAMELA),让他知道每个人喜欢什么样的味道。厨师用这个工具学习每个人的偏好,调整调料比例,做出更符合个人口味的菜肴。这个过程就像给厨师装上了“个性化调味器”,让每个人都能吃到自己喜欢的菜。这种方法不仅让菜更合口味,也让厨师变得更聪明,能记住每个人的偏好,做出更贴心的菜肴。
ELI14 Explained like you're 14
想象你和朋友们一起玩游戏,每个人喜欢不同的角色和玩法。以前,游戏设计师只会设计一款适合大多数人的游戏,但有些朋友觉得不够有趣。现在,设计师用了一种特别的方法,听取每个人的偏好,记录他们喜欢的角色和玩法(就像PAMELA评分一样)。然后,他用这些信息调整游戏,让每个人都能找到自己喜欢的部分。这样,每个人都觉得游戏更好玩、更贴心。这个新方法就像给游戏加入了“个性化调节器”,让每个玩家都能享受属于自己的乐趣。未来,这样的技术还能帮我们做出更符合每个人喜好的内容,无论是电影、音乐还是学习资料。
Abstract
Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models optimize for "average" human appeal, they fail to capture the inherent subjectivity of aesthetic judgment. In this work, we introduce a novel dataset and predictive framework, called PAMELA, designed to model personalized image evaluations. Our dataset comprises 70,000 ratings across 5,000 diverse images generated by state-of-the-art models (Flux 2 and Nano Banana). Each image is evaluated by 15 unique users, providing a rich distribution of subjective preferences across domains such as art, design, fashion, and cinematic photography. Leveraging this data, we propose a personalized reward model trained jointly on our high-quality annotations and existing aesthetic assessment subsets. We demonstrate that our model predicts individual liking with higher accuracy than the majority of current state-of-the-art methods predict population-level preferences. Using our personalized predictor, we demonstrate how simple prompt optimization methods can be used to steer generations towards individual user preferences. Our results highlight the importance of data quality and personalization to handle the subjectivity of user preferences. We release our dataset and model to facilitate standardized research in personalized T2I alignment and subjective visual quality assessment.