Attention in Large Language Models Yields Efficient Zero-Shot Re-Rankers
ICR leverages attention pattern changes in LLMs for efficient zero-shot re-ranking, reducing latency by over 60%.
Shijie Chen, Bernal Jiménez Gutiérrez, Yu Su
ICR leverages attention pattern changes in LLMs for efficient zero-shot re-ranking, reducing latency by over 60%.
Shijie Chen, Bernal Jiménez Gutiérrez, Yu Su
ExACT combines R-MCTS and Exploratory Learning, achieving 6%-30% improvement on VisualWebArena.
Xiao Yu, Baolin Peng, Vineeth Vajipey et al.
Proposed VideoCLIP-XL with TPCM, achieving significant improvements in long description video understanding, trained on over 2 million video-long description pairs.
Jiapeng Wang, Chengyu Wang, Kunzhe Huang et al.
UniSumEval evaluates 9 LLMs on fine-grained, multi-dimensional summarization across 9 domains, addressing gaps in existing benchmarks.
Yuho Lee, Taewon Yun, Jason Cai et al.
Divergence-based calibration (DC-PDD) outperforms existing methods in detecting training data, improving AUC by 8.6%.
Weichao Zhang, Ruqing Zhang, Jiafeng Guo et al.
Proposes LSCS architecture combining model parameters, explicit memory, knowledge graphs, and text storage for lifelong experience absorption and accurate recall.
Yu Wang, Chi Han, Tongtong Wu et al.
FRAMES dataset evaluates LLMs' factuality, retrieval, and reasoning via multi-hop questions, showing a 50%+ performance boost with multi-step retrieval.
Satyapriya Krishna, Kalpesh Krishna, Anhad Mohananey et al.
LogicPro synthesizes complex logical reasoning data using program-guided learning, enhancing model performance.
Jin Jiang, Yuchen Yan, Yang Liu et al.
Proposes SciLead, an LLM-based system for automated scientific leaderboard construction, achieving 55% F1 in full scenario, addressing manual errors.
Furkan Şahinuç, Thy Thy Tran, Yulia Grishina et al.
Leveraging large code models for conversational no-code programming of cobots; evaluated on industrial assembly tasks with 85% accuracy in instruction sequences.
Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen
Proposes Agent Workflow Memory (AWM), which induces and stores reusable workflows, boosting web navigation success rates by 24.6% and 51.1% on Mind2Web and WebArena respectively.
Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried et al.
PingPong benchmark evaluates role-playing language models using user emulation and multi-model evaluation, validating over 40 models.
Ilya Gusev
E2LLM combines chunk-based soft prompts with pre-trained encoders and decoders, overcoming the 'impossible triangle' of performance, efficiency, and compatibility for long-context tasks.
Zihan Liao, Jun Wang, Hang Yu et al.
DOLCE framework parameterizes problems with λ and k to identify retrieval and holistic understanding tasks in long contexts.
Zi Yang
This study evaluates LLMs' ability to generate novel research ideas, finding they outperform experts in novelty (p<0.05) but are slightly weaker in feasibility.
Chenglei Si, Diyi Yang, Tatsunori Hashimoto
xLAM models outperform GPT-4 on the Berkeley Function-Calling Leaderboard, enhancing AI agent task performance.
Jianguo Zhang, Tian Lan, Ming Zhu et al.
Introduces MMMU-Pro, a more rigorous benchmark with filtering, option augmentation, and vision-only input, causing performance drops of 16.8%-26.9%, enhancing evaluation robustness.
Xiang Yue, Tianyu Zheng, Yuansheng Ni et al.
OLMoE uses sparse Mixture-of-Experts, activating only 1B of 7B parameters, outperforming larger models.
Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld et al.
TinyAgent fine-tunes small open-source LLMs with high-quality datasets, surpassing GPT-4-Turbo in function calling, deploying efficiently on edge devices with tool retrieval and quantization.
Lutfi Eren Erdogan, Nicholas Lee, Siddharth Jha et al.
Symbolic working memory enhances language models for complex rule application, significantly improving multi-step reasoning accuracy.
Siyuan Wang, Zhongyu Wei, Yejin Choi et al.