Context-Aware Autoregressive Models for Multi-Conditional Image Generation
ContextAR model achieves efficient multi-conditional image generation, surpassing diffusion models.
Yixiao Chen, Zhiyuan Ma, Guoli Jia et al.
ContextAR model achieves efficient multi-conditional image generation, surpassing diffusion models.
Yixiao Chen, Zhiyuan Ma, Guoli Jia et al.
Agent4SR uses LLMs for low-knowledge, high-impact attacks, enhancing manipulation of recommender systems.
Shengkang Gu, Jiahao Liu, Dongsheng Li et al.
Proposed LSTAN-GERPE leverages multi-layer spatio-temporal attention, RoPE, and graph eigenvector embedding to significantly improve traffic forecasting accuracy.
Xiao Wang, Shun-Ren Yang
Proposes dLLM-Cache, a training-free adaptive caching method combining long-interval prompt caching and partial response updates, reducing FLOPs by up to 9.1× and inference latency.
Zhiyuan Liu, Yicun Yang, Yaojie Zhang et al.
SpatialCrafter uses video diffusion models to reconstruct 3D scenes from sparse views, enhancing reconstruction accuracy.
Songchun Zhang, Huiyao Xu, Sitong Guo et al.
Introduces CrafText benchmark with 3924 instructions, evaluating instruction following in dynamic, multimodal environments with advanced RL algorithms.
Zoya Volovikova, Gregory Gorbov, Petr Kuderov et al.
LifelongAgentBench systematically evaluates LLM agents' lifelong learning; introducing experience replay and self-consistency improves success rates from 19% to 78%.
Junhao Zheng, Xidi Cai, Qiuke Li et al.
CoLM enables progressive scaling and elastic inference; CoLM-Air reaches about 3× faster 1M-token prefilling.
Kaitao Song, Xiaohua Wang, Xu Tan et al.
211 student-made applied-math problems benchmark LLMs on asymptotics and approximation.
James V. Roggeveen, Erik Y. Wang, Will Flintoft et al.
Introduced MedCaseReasoning dataset with 14,489 clinical cases; fine-tuning improves diagnostic accuracy by 29% and reasoning recall by 41%.
Kevin Wu, Eric Wu, Rahul Thapa et al.
EgoDex leverages Apple Vision Pro to collect 829 hours of egocentric manipulation videos, enabling advanced imitation learning for dexterous robots.
Ryan Hoque, Peide Huang, David J. Yoon et al.
DPSeg integrates dual prompts and cost volume learning, significantly improving open-vocabulary semantic segmentation accuracy.
Ziyu Zhao, Xiaoguang Li, Linjia Shi et al.
Proposes Critique-Guided Distillation (CGD), leveraging teacher critiques to improve reasoning, achieving +7% average gains on benchmarks.
Berkcan Kapusuzoglu, Supriyo Chakraborty, Zain Sarwar et al.
Proposes ASRC-SNN with adaptive skip connections, achieving state-of-the-art long-term temporal modeling, outperforming traditional RSNN by 2-3% accuracy.
Shang Xu, Jiayu Zhang, Ziming Wang et al.
Time-R1 employs a three-stage RL fine-tuning framework to endow a 3B-parameter LLM with comprehensive temporal reasoning, outperforming models over 200 times larger in future event prediction and creative scenario generation.
Zijia Liu, Peixuan Han, Haofei Yu et al.
HAPO employs history-aware policy optimization to reduce response length by 33-59% with only 2-5% accuracy loss in LLMs.
Chengyu Huang, Zhengxin Zhang, Claire Cardie
SpecBranch combines hybrid speculative drafting with rollback-aware branch parallelism, achieving 1.8-4.5× speedup and halving rollback tokens in large language model inference.
Yuhao Shen, Junyi Shen, Quan Kong et al.
Ready2Unlearn optimizes models during training for future unlearning readiness.
Hanyu Duan, Yi Yang, Ahmed Abbasi et al.
TartanGround is a large-scale multi-modal simulation dataset with 1.44 million samples, supporting perception and navigation tasks for ground robots in diverse environments.
Manthan Patel, Fan Yang, Yuheng Qiu et al.
Proposes Prior Depth Anything, integrating sparse measurements and depth prediction to produce dense, metric depth maps with strong zero-shot generalization.
Zehan Wang, Siyu Chen, Lihe Yang et al.