ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
ELLA integrates LLM with diffusion models using TSC for enhanced semantic alignment.
Xiwei Hu, Rui Wang, Yixiao Fang et al.
ELLA integrates LLM with diffusion models using TSC for enhanced semantic alignment.
Xiwei Hu, Rui Wang, Yixiao Fang et al.
Enhanced robotic manipulation data collection via compositional generalization, achieving a 77.5% success rate.
Jensen Gao, Annie Xie, Ted Xiao et al.
mmPlace transforms intermediate frequency signals into range-azimuth heatmaps, employs rotation-based concatenation, achieving 87.37% recall@1 in challenging scenarios.
Chengzhen Meng, Yifan Duan, Chenming He et al.
Proposes Claim-Conditioned Probability (CCP) for token-level uncertainty quantification, improving factual accuracy detection across 7 models and 4 languages.
Ekaterina Fadeeva, Aleksandr Rubashevskii, Artem Shelmanov et al.
Proposed CAT model enhances multimodal reasoning with clue aggregation and GPT-based preference optimization, boosting AVQA performance.
Qilang Ye, Zitong Yu, Rui Shao et al.
TextMonkey: OCR-free large model, improves document understanding, scores 561 on OCRBench.
Yuliang Liu, Biao Yang, Qiang Liu et al.
RL-based H2O enables real-time humanoid teleoperation via RGB camera, achieving dynamic motion imitation with success rate over 72.5%.
Tairan He, Zhengyi Luo, Wenli Xiao et al.
Using GLM regression, report length ≥947 words significantly increases citation counts, confirming the impact of review depth on publication influence.
Abdelghani Maddi, Luis Miotti
This survey reviews controllable generation in text-to-image diffusion models, analyzing DDPMs' control mechanisms.
Pu Cao, Feng Zhou, Qing Song et al.
TTPXHunter fine-tunes SecureBERT with data augmentation to extract 193 TTPs, achieving 97.09% F1 on real reports.
Nanda Rani, Bikash Saha, Vikas Maurya et al.
Design2Code benchmarks multimodal models on real webpages, revealing gaps in visual element recall and layout accuracy, with GPT-4o leading but still limited.
Chenglei Si, Yanzhe Zhang, Ryan Li et al.
MathScale uses topic extraction and concept graphs to generate 2 million math QA pairs, boosting LLMs' reasoning by 43%.
Zhengyang Tang, Xingxing Zhang, Benyou Wang et al.
This survey reviews over 90 papers on cultural representation in LLMs, proposing proxies like demographic and semantic features, analyzing probing methods, and highlighting research gaps.
Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania et al.
CoAT method enhances zero-shot GUI agent prediction, supported by AITZ dataset.
Jiwen Zhang, Jihao Wu, Yihua Teng et al.
Proposes Uplift-guided Budget Allocation (UBA) for target user attacks, estimating heterogeneous treatment effects via path number proxies to optimize fake user budgets.
Wenjie Wang, Changsheng Wang, Fuli Feng et al.
Caduceus model enhances DNA sequence modeling with bi-directionality and RC equivariance.
Yair Schiff, Chia-Hsiang Kao, Aaron Gokaslan et al.
Proposes Wukong, a stacking FM-based architecture, establishing a recommendation scaling law with performance beyond 100 GFLOP/sample.
Buyun Zhang, Liang Luo, Yuxin Chen et al.
Proposes exploration-based trajectory optimization (ETO) for LLM agents, leveraging failure trajectories with contrastive learning, outperforming baselines by significant margins.
Yifan Song, Da Yin, Xiang Yue et al.
Introduces measure pre-conditioning via γ-convergence to enhance stability and convergence in parametric ML models and transfer learning.
Joaquín Sánchez García
Proposed Key-Point-Driven Data Synthesis (KPDDS), creating 800K+ math reasoning QA pairs, significantly boosting model performance.
Yiming Huang, Xiao Liu, Yeyun Gong et al.