cs.CL 2410.17196

VoiceBench: Benchmarking LLM-Based Voice Assistants

VoiceBench benchmark evaluates LLM-based voice assistants across diverse real-world scenarios, revealing performance gaps in robustness and safety, with detailed multi-metric analysis.

Yiming Chen, Xianghu Yue, Chen Zhang et al.

2024-10-23 45
cs.CV 2410.16512

TIPS: Text-Image Pretraining with Spatial awareness

TIPS combines spatial awareness via synthetic captions and self-supervised masked modeling, boosting dense image understanding.

Kevis-Kokitsi Maninis, Kaifeng Chen, Soham Ghosh et al.

2024-10-22 27
cs.CL 2410.16464

Beyond Browsing: API-Based Web Agents

Proposed API-calling and Hybrid Agents, API Agent improved by 24.0% on WebArena, achieving 38.9% success rate.

Yueqi Song, Frank Xu, Shuyan Zhou et al.

2024-10-22 47