cs.CV 2308.07498

DREAMWALKER: Mental Planning for Continuous Vision-Language Navigation

DREAMWALKER employs explicit world models with environment graphs and scene synthesizers, combined with Monte Carlo Tree Search, to enable strategic mental planning in continuous vision-language navigation, achieving over 78% success rate.

Hanqing Wang, Wei Liang, Luc Van Gool et al.

2023-08-15 121 citations 34
cs.CL 2308.07134

Language is All a Graph Needs

InstructGLM uses natural language prompts for graph structure description, surpassing GNNs in node classification with instruction fine-tuning.

Ruosong Ye, Caiqi Zhang, Runhui Wang et al.

2023-08-14 44
cs.CV 2308.06571

ModelScope Text-to-Video Technical Report

ModelScopeT2V uses diffusion with spatio-temporal blocks, 1.7B parameters, achieving superior text-to-video synthesis with high temporal coherence.

Jiuniu Wang, Hangjie Yuan, Dayou Chen et al.

2023-08-12 35
cs.AI 2308.03688

AgentBench: Evaluating LLMs as Agents

AgentBench benchmarks 29 LLMs across 8 environments, revealing performance gaps and guiding future improvements.

Xiao Liu, Hao Yu, Hanchen Zhang et al.

2023-08-08 44
quant-ph 2308.01582

Quantum speedups for stochastic optimization

Quantum variance reduction algorithms outperform classical methods in low-dimensional stochastic optimization, achieving quadratic speedups.

Aaron Sidford, Chenyi Zhang

2023-08-03 40