cs.AI 2309.05519

NExT-GPT: Any-to-Any Multimodal LLM

NExT-GPT achieves any-to-any modality input-output with multimodal adapters and diffusion decoders, tuning only 1% of parameters.

Shengqiong Wu, Hao Fei, Leigang Qu et al.

2023-09-11 46
cs.MM 2309.03905

ImageBind-LLM: Multi-modality Instruction Tuning

ImageBind-LLM achieves multi-modality instruction tuning via image-text alignment, supporting audio, 3D point clouds, and video.

Jiaming Han, Renrui Zhang, Wenqi Shao et al.

2023-09-08 20
cs.NE 2309.03388

Are SNNs Truly Energy-efficient? $-$ A Hardware Perspective

This study benchmarks large-scale SNN inference on SATA and SpikeSim, revealing actual energy efficiency is far below estimates due to hardware bottlenecks.

Abhiroop Bhattacharjee, Ruokai Yin, Abhishek Moitra et al.

2023-09-07 45
cs.AI 2309.02427

Cognitive Architectures for Language Agents

Proposed CoALA framework to organize language agents, enhancing reasoning and decision-making.

Theodore R. Sumers, Shunyu Yao, Karthik Narasimhan et al.

2023-09-06 5