NExT-GPT: Any-to-Any Multimodal LLM
NExT-GPT achieves any-to-any modality input-output with multimodal adapters and diffusion decoders, tuning only 1% of parameters.
Shengqiong Wu, Hao Fei, Leigang Qu et al.
NExT-GPT achieves any-to-any modality input-output with multimodal adapters and diffusion decoders, tuning only 1% of parameters.
Shengqiong Wu, Hao Fei, Leigang Qu et al.
SignRound optimizes LLM quantization using SignSGD, achieving 6.91%-33.22% accuracy improvement at 2 bits.
Wenhua Cheng, Weiwei Zhang, Haihao Shen et al.
phi-1.5 model achieves comparable common sense reasoning with 1.3B parameters as models 5x larger.
Yuanzhi Li, Sébastien Bubeck, Ronen Eldan et al.
Controlled experiment shows InstructGPT significantly reduces content diversity, increasing similarity by ~8% and decreasing lexical diversity by ~15%.
Vishakh Padmakumar, He He
Introduces a space-efficient streaming SDP solver using spectral sketching, achieving ˜O(m² + n²) space with O(√n log(1/ε)) passes.
Zhao Song, Mingquan Ye, Lichen Zhang
Proposes 3D Implicit Transporter for temporally consistent keypoint detection, improving non-rigid object understanding.
Chengliang Zhong, Yuhang Zheng, Yupeng Zheng et al.
Proposes a subspace-constrained randomized Kaczmarz (SCRK) method accelerating convergence for low-rank systems, leveraging external knowledge for robustness.
Jackie Lok, Elizaveta Rebrova
MADLAD-400 is a manually audited multilingual monolingual dataset covering 419 languages, enabling high-quality pretraining for multilingual models.
Sneha Kudugunta, Isaac Caswell, Biao Zhang et al.
SayNav uses large language models for dynamic navigation in new environments, improving success rate by 8%.
Abhinav Rajvanshi, Karan Sikka, Xiao Lin et al.
ImageBind-LLM achieves multi-modality instruction tuning via image-text alignment, supporting audio, 3D point clouds, and video.
Jiaming Han, Renrui Zhang, Wenqi Shao et al.
SyncDreamer employs a multiview-synchronized diffusion framework to generate consistent multi-view images from a single view, achieving superior 3D reconstruction.
Yuan Liu, Cheng Lin, Zijiao Zeng et al.
This study benchmarks large-scale SNN inference on SATA and SpikeSim, revealing actual energy efficiency is far below estimates due to hardware bottlenecks.
Abhiroop Bhattacharjee, Ruokai Yin, Abhishek Moitra et al.
EGIC combines OASIS-C and ORP for low-bit-rate image compression, outperforming SOTA methods.
Nikolai Körber, Eduard Kromer, Andreas Siebert et al.
Proposed CoALA framework to organize language agents, enhancing reasoning and decision-making.
Theodore R. Sumers, Shunyu Yao, Karthik Narasimhan et al.
Lightweight EOAT combined with CMA-ES optimization achieves 100% success in robotic Lego assembly/disassembly.
Ruixuan Liu, Yifan Sun, Changliu Liu
CIEM method evaluates VLM hallucination by generating contrastive Q&A pairs, significantly improving model accuracy.
Hongyu Hu, Jiyuan Zhang, Minyi Zhao et al.
Kernel-based analysis of Ermakov-Zolotukhin quadrature derives error formulas, improving theoretical guarantees for DPP-based quadratures.
Ayoub Belhadji
Proposes Sequential Dexterity, using bi-directional RL to chain dexterous policies with transition feasibility, achieving 20% success boost in long-horizon tasks.
Yuanpei Chen, Chen Wang, Li Fei-Fei et al.
This paper analyzes how detoxification methods (fine-tuning and RLHF) influence language models' prompt dependence using attribution entropy metrics.
Daniel Scalena, Gabriele Sarti, Malvina Nissim et al.
The study explores adversarial attacks on aligned language models using methods like perplexity detection.
Neel Jain, Avi Schwarzschild, Yuxin Wen et al.