WithAnyone: Towards Controllable and ID Consistent Image Generation
WithAnyone model achieves controllable and ID-consistent image generation using contrastive loss and the MultiID-2M dataset.
Hengyuan Xu, Wei Cheng, Peng Xing et al.
WithAnyone model achieves controllable and ID-consistent image generation using contrastive loss and the MultiID-2M dataset.
Hengyuan Xu, Wei Cheng, Peng Xing et al.
ToolPRM uses fine-grained step scoring to improve structured function calling inference, outperforming coarse reward models.
Jianghao Lin, Yuanyuan Shi, Xin Peng et al.
xLLM employs decoupled architecture with adaptive scheduling and multi-layer pipeline optimization, achieving 1.7× throughput over MindIE and 2.2× over vLLM-Ascend on Qwen models.
Tongxuan Liu, Tao Peng, Peijun Yang et al.
AEPO introduces dynamic entropy balancing and gradient regulation, enhancing stability and exploration in multi-turn web agent RL with only 1K samples.
Guanting Dong, Licheng Bao, Zhongyuan Wang et al.
SHaRe-SSM excels in ultra-long sequences with 52.1x energy efficiency improvement.
Kartikay Agrawal, Abhijeet Vikram, Vedant Sharma et al.
Proposes Generative Universal Verifier, trained on ViVerBench, improving visual verification by 8.3 points with OmniVerifier-7B and TTS strategies.
Xinchen Zhang, Xiaoying Zhang, Youbin Wu et al.
Confidence estimation based on activation signals improves LLM trustworthiness, achieving 95% accuracy with reduced latency.
Zhiqi Huang, Vivek Datla, Chenyang Zhu et al.
FusionNet and Transformer-based models achieved top PSNR of 26.35 in NTIRE 2025 low-light enhancement, demonstrating multi-model fusion effectiveness.
Xiaoning Liu, Zongwei Wu, Florin-Alexandru Vasluianu et al.
LIBERO-Plus systematically analyzes VLA model robustness under seven perturbations, revealing performance drops from 95% to below 30%.
Senyu Fei, Siyin Wang, Junhao Shi et al.
Mask-GRPO introduces reinforcement learning into masked generative models, significantly improving text-to-image generation with a 0.73 score on GenEval and FID of 8.32, surpassing SOTA.
Yifu Luo, Xinhao Hu, Keyu Fan et al.
DepthVLA integrates a pretrained depth module into a mixture-of-transformers framework, significantly improving spatial reasoning and manipulation success rates (e.g., 78.5% in real-world tasks).
Tianyuan Yuan, Yicheng Liu, Chenhao Lu et al.
Utilizing hypernetwork and adapters architecture, this study enhances perspective adaptation in hate speech detection with fewer parameters.
Daniil Ignatev, Denis Paperno, Massimo Poesio
This study shows that AI models are more persuasive when defending positions aligned with their prior beliefs, using sequential and simultaneous debate protocols, with models favoring sycophantic strategies in conflict scenarios.
María Victoria Carro, Denise Alejandra Mester, Facundo Nieto et al.
EgoSocial leverages multimodal cues to improve proactive intervention detection in OLMMs, boosting timing accuracy by 45.6% on Phi-4.
Xijun Wang, Tanay Sharma, Achin Kulshrestha et al.
ExtremBench benchmark reveals LLM discrepancies in solving extremal problems.
Binxin Gao, Jingjun Han
DriveVLA-W0 employs world modeling to predict future images, providing dense self-supervision that significantly enhances autonomous driving models' scalability, outperforming baselines on large datasets.
Yingyan Li, Shuyao Shang, Weisong Liu et al.
AnyUp is a universal feature upsampling method that generalizes to any feature type at inference, outperforming state-of-the-art.
Thomas Wimmer, Prune Truong, Marie-Julie Rakotosaona et al.
Omni-Detective pipeline and Omni-Captioner model enhance multimodal fine-grained perception, outperforming SOTA on key benchmarks.
Ziyang Ma, Ruiyang Xu, Zhenghao Xing et al.
Introduced Bellman-Wasserstein Distance (BWD) to assess offline RL dataset quality, significantly improving prediction accuracy.
Arip Asadulaev, Fakhri Karray, Martin Takac
ReRe employs RLVR with constrained beam search and combined rewards to significantly boost recommendation ranking performance.
Junfei Tan, Yuxin Chen, An Zhang et al.