The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
Sparse attention enhances long-sequence processing in Transformer LLMs using Quest and Vertical-Slash methods.
Piotr Nawrot, Robert Li, Renjie Huang et al.
Sparse attention enhances long-sequence processing in Transformer LLMs using Quest and Vertical-Slash methods.
Piotr Nawrot, Robert Li, Renjie Huang et al.
Proposes HOTET, combining Transformer embeddings and hypernetworks for efficient multi-distribution OT mapping.
Mingchen Jiang, Peng Xu, Xichen Ye et al.
Introduced Bias-Eli-W PnP estimator, significantly reducing localization errors on KITTI dataset.
Guangyang Zeng, Yuan Shen, Ziyang Hong et al.
By building a large math problem dataset, integrating code execution, and developing candidate solution selection, the model achieves state-of-the-art performance.
Ivan Moshkov, Darragh Hanley, Ivan Sorokin et al.
LiLaVe method extracts correctness signals from LLM hidden states, enhancing task efficiency and accuracy.
Bartosz Piotrowski, Witold Drzewakowski, Konrad Staniszewski et al.
ManipDreamer enhances robotic manipulation models with action trees and visual guidance, improving PSNR to 21.05.
Ying Li, Xiaobao Wei, Xiaowei Chi et al.
Proposes an efficient zero-shot subject-driven video generation framework using only 1% compute resources.
Daneul Kim, Jingxu Zhang, Wonjoon Jin et al.
Proposes MR. Video based on MapReduce, achieving 10%+ accuracy boost on LVBench for long video understanding.
Ziqi Pang, Yu-Xiong Wang
π0.5 leverages multi-source heterogeneous data to enable robots to perform long-horizon, multi-step tasks in unseen environments with high success rates.
Physical Intelligence, Kevin Black, Noah Brown et al.
Introduces the e-Partitioning principle as a necessary and sufficient condition for FDR control, improving existing methods like eBH, BY with flexible, simultaneous control.
Jelle Goeman, Rianne de Heide, Aldo Solari
SPECI introduces a hierarchical skill prompt framework with dynamic skill codebook and mode approximation, achieving state-of-the-art continual robot manipulation performance.
Jingkai Xu, Xiangli Nie
Eagle 2.5 employs Automatic Degrade Sampling and Image Area Preservation to enhance long-video understanding, achieving 72.4% on Video-MME with 8B parameters, comparable to GPT-4o.
Guo Chen, Zhiqi Li, Shihao Wang et al.
Proposes HAF-VT, a multimodal hierarchical attention model, boosting cross-domain sequential recommendation by 3-5% on key metrics.
Wangyu Wu, Zhenhong Chen, Siqi Song et al.
A large language model-based automatic nugget extraction and assignment framework for RAG evaluation, validated against human annotations with high correlation.
Ronak Pradeep, Nandan Thakur, Shivani Upadhyay et al.
Using prompts, ChatGPT detects and categorizes LSP translation errors with 64.1% accuracy, demonstrating strong potential for automated evaluation.
Joachim Minder, Guillaume Wisniewski, Natalie Kübler
Proposes new methods to enhance spatial reasoning in MLLMs, significantly improving performance on spatial tasks.
Huanyu Zhang, Chengzu Li, Wenshan Wu et al.
WindVE leverages CPU-NPU collaboration with linear regression-based queue management to boost concurrent vector embedding by 22.3%, reducing costs and improving throughput.
Jinqi Huang, Xuebing Yu, Yi Xiong et al.
Reformulates Expected Free Energy planning as variational inference, integrating goal and information gain, enabling scalable resource-aware decision-making.
Bert de Vries, Wouter Nuijten, Thijs van de Laar et al.
MARFT introduces a modular framework with Flex-MG and PPO-based algorithms, enhancing LaMAS fine-tuning with 15% performance gains on DeepScaler and DeepCoder.
Junwei Liao, Muning Wen, Jun Wang et al.
Twin-Co framework combines multi-turn dialogue and internal optimization, boosting image-text alignment with CLIP score reaching 0.338.
Jianhui Wang, Yangfan He, Yan Zhong et al.