Set the Clock: Temporal Alignment of Pretrained Language Models
Proposes temporal alignment of LLaMa2 via finetuning and prompting, achieving up to 62% performance boost on 2022 data.
Bowen Zhao, Zander Brumbaugh, Yizhong Wang et al.
Proposes temporal alignment of LLaMa2 via finetuning and prompting, achieving up to 62% performance boost on 2022 data.
Bowen Zhao, Zander Brumbaugh, Yizhong Wang et al.
SelectIT leverages LLM's intrinsic uncertainty through multi-granularity self-reflection to select high-quality instruction tuning data, boosting model performance.
Liangxin Liu, Xuebo Liu, Derek F. Wong et al.
Star-Searcher employs hierarchical path planning and viewpoint clustering, reducing path length by 15%, search time by 20%, achieving 100% target completeness.
Yiming Luo, Zixuan Zhuang, Neng Pan et al.
This study systematically reviews LLMs like GPT and BERT in health, urban planning, climate, and disaster management, highlighting their transformative potential.
Pravneet Kaur, Gautam Siddharth Kashyap, Ankit Kumar et al.
Proposed CDD detects data contamination via output distribution peakedness, improving detection accuracy by 21.8%-30.2%.
Yihong Dong, Xue Jiang, Huanyu Liu et al.
NaVid is a video-based vision-language navigation model that uses monocular RGB input to achieve state-of-the-art performance without maps or depth data, demonstrating strong Sim2Real transfer.
Jiazhao Zhang, Kunyu Wang, Rongtao Xu et al.
Noiseless Kernel Ridge Regression (KRR) achieves minimax optimal rates, revealing super-smoothness effects and saturation phenomena based on eigenvalue decay and target smoothness.
Jihao Long, Xiaojun Peng, Lei Wu
RoboEXP employs an action-conditioned scene graph (ACSG) via interactive exploration, integrating multimodal data for complex scene understanding.
Hanxiao Jiang, Binghao Huang, Ruihai Wu et al.
GROS is a general robust estimator aggregation method in metric spaces, ensuring sub-Gaussian tail behavior and high breakdown point.
Alejandro Cholaquidis, Emilien Joly, Leonardo Moreno
MACRec uses multi-agent collaboration to reframe recommendation for 4 tasks.
Zhefan Wang, Yuanqing Yu, Wendi Zheng et al.
Introduces EES method to enhance robot task success by estimating, extrapolating, and situating skill competence.
Nishanth Kumar, Tom Silver, Willie McClinton et al.
Proposes tinyBenchmarks using 100 examples with IRT and clustering to estimate LLM performance on benchmarks within 2% error.
Felipe Maia Polo, Lucas Weber, Leshem Choshen et al.
Proposes QontOT, a quantum-based framework for contextual optimal transport, outperforming classical neural OT in predicting distribution variations.
Nicola Mariella, Albert Akhriev, Francesco Tacchino et al.
CriticBench benchmarks 17 LLMs' critique and correction abilities across five reasoning domains, revealing linear relationships and size-dependent knowledge consistency.
Zicheng Lin, Zhibin Gou, Tian Liang et al.
MobileLLM employs deep-thin transformer architecture with parameter sharing, achieving state-of-the-art accuracy for sub-billion models on mobile devices.
Zechun Liu, Changsheng Zhao, Forrest Iandola et al.
LLMBind framework integrates multimodal tasks with dual-pathway mechanism, achieving superior performance and expandability.
Bin Zhu, Munan Ning, Peng Jin et al.
Proposes VPDD, a discrete diffusion-based framework leveraging large-scale actionless human videos for multi-task robot policy transfer, outperforming state-of-the-art methods.
Haoran He, Chenjia Bai, Ling Pan et al.
Proposes Sliver sliding window data stream paradigm, boosting live recommendation timeliness and accuracy, with a 6.76% CTR increase demonstrated.
Fengqi Liang, Baigong Zheng, Liqin Zhao et al.
OlympiadBench, with 8,476 high-level bilingual scientific problems, evaluates GPT-4V's reasoning, scoring 17.97%, highlighting current AI limitations in complex science tasks.
Chaoqun He, Renjie Luo, Yuzhuo Bai et al.
Introduces PGI and GELAN to improve gradient reliability and parameter efficiency, boosting YOLOv9 detection performance.
Chien-Yao Wang, I-Hau Yeh, Hong-Yuan Mark Liao