ResearchCube: Multi-Dimensional Trade-off Exploration for Research Ideation
ResearchCube explores multi-dimensional trade-offs in research ideation using a 3D space.
Zijian Ding, Fenghai Li, Ziyi Wang et al.
ResearchCube explores multi-dimensional trade-offs in research ideation using a 3D space.
Zijian Ding, Fenghai Li, Ziyi Wang et al.
Developed a scalable problem reduction library with 100+ problem types and 200+ rules, integrated via AI agents for automated contribution and verification.
Xi-Wei Pan, Shi-Wen An, Jin-Guo Liu
Proposes TCER (Triviality Corrected Endogenous Reward), leveraging relative information gain to mitigate bias in open-ended text generation, improving diversity and quality.
Xinda Wang, Zhengxu Hou, Yangshijie Zhang et al.
METRO uses large language models to autonomously extract strategies from expert dialogues, improving performance by 9%-10%.
Haofu Yang, Jiaji Liu, Chen Huang et al.
WM-DAgger synthesizes OOD recovery data using World Models, achieving a 93.3% success rate in soft bag pushing with five demonstrations.
Anlan Yu, Zaishu Chen, Peili Song et al.
Proposes Polyglot Score to evaluate 10 multilingual teacher models; finds data quality outweighs model size in effectiveness.
Lester James V. Miranda, Ivan Vulić, Anna Korhonen
Introduces Bottleneck Tokens (BToks) to address structural issues in multimodal retrieval, achieving a score of 59.0 on MMEB-V2.
Siyu Sun, Jing Ren, Zhaohe Liao et al.
LaDA-Band uses Discrete Masked Diffusion for vocal-to-accompaniment generation, enhancing acoustic authenticity and global coherence.
Qi Wang, Zhexu Shen, Meng Chen et al.
PLOVIS leverages pre-trained open-vocabulary image segmentation (DeCLIP) to generate pseudo labels directly from 3D point clouds, enabling effective training under scarce data conditions.
Takahiko Furuya
WebForge automates the end-to-end creation of realistic, reproducible, and scalable web environments using a four-stage pipeline, enabling multi-dimensional capability profiling.
Peng Yuan, Yuyang Yin, Yuxuan Cai et al.
Introduces Energy-oriented Diffusion Bridge (E-Bridge), reducing sampling steps and improving image restoration quality via low-energy geodesic trajectories.
Jinhui Hou, Zhiyu Zhu, Junhui Hou
CFMS integrates multimodal perception with hierarchical symbolic reasoning, boosting complex table understanding by 15% accuracy on benchmarks.
Qixian Huang, Hongqiang Lin, Tong Fu et al.
WARPED generates wrist-view data from monocular RGB videos, reducing data collection time by 5-8x for robot learning.
Harry Freeman, Chung Hee Kim, George Kantor
OmniUMI enables physically grounded robot learning via human-aligned multimodal interaction, enhancing contact-rich manipulation performance.
Shaqi Luo, Yuanyuan Li, Youhao Hu et al.
AffordGen generates diverse demonstrations using 3D generative models to enhance robot manipulation generalization.
Jiawei Zhang, Kaizhe Hu, Yingqian Huang et al.
CARE framework evaluates LLM alignment with community linguistic behaviors through reaction tone analysis.
Nuan Wen, Xuezhe Ma
BLUEmed combines multi-agent debate with hybrid retrieval-augmented generation, achieving 69.13% accuracy in clinical terminology error detection.
Saukun Thika You, Nguyen Anh Khoa Tran, Wesley K. Marizane et al.
DeepShapeMatchingKit accelerates functional map solving by 33x using vectorized reformulation, addressing 3D shape matching bottlenecks.
Yizheng Xie, Lennart Bastian, Congyue Deng et al.
SMFormer integrates foundation models and data augmentation to enhance self-supervised stereo matching.
Yun Wang, Zhengjie Yang, Jiahao Zheng et al.
FatigueFusion generates fatigue-driven motion via latent space fusion, applicable to various synthesis tasks.
Iliana Loi, Konstantinos Moustakas