cs.RO 2305.11307

Semantic Anomaly Detection with Large Language Models

Leveraging large language models (LLMs) for semantic anomaly detection in vision-based policies, achieving over 90% accuracy in simulated autonomous driving scenarios.

Amine Elhafsi, Rohan Sinha, Christopher Agia et al.

2023-05-19 56
cs.LG 2305.10924

Structural Pruning for Diffusion Models

Diff-Pruning uses Taylor expansion to cut 50% FLOPs in diffusion models, with only 10-20% training cost, maintaining quality.

Gongfan Fang, Xinyin Ma, Xinchao Wang

2023-05-18 33
cs.CV 2305.10855

TextDiffuser: Diffusion Models as Text Painters

TextDiffuser combines Transformer layout prediction with latent diffusion models, enabling high-quality, controllable text image synthesis.

Jingye Chen, Yupan Huang, Tengchao Lv et al.

2023-05-18 42
eess.AS 2305.10790

Listen, Think, and Understand

LTU model integrates audio perception and reasoning using the OpenAQA-5M dataset for audio understanding.

Yuan Gong, Hongyin Luo, Alexander H. Liu et al.

2023-05-18 40
cs.CL 2305.10403

PaLM 2 Technical Report

PaLM 2 combines UL2-style objectives, compute-optimal scaling, and multilingual data to improve reasoning and translation with smaller, faster models.

Rohan Anil, Andrew M. Dai, Orhan Firat et al.

2023-05-18 26
cs.RO 2305.09765

OpenVR: Teleoperation for Manipulation

OpenVR links Quest 2 + Unity to Panda teleoperation, reaching 71.3 Hz simulation and 36.2 Hz hardware.

Abraham George, Alison Bartsch, Amir Barati Farimani

2023-05-17 48