cs.CL 2211.05826

The CRINGE Loss: Learning what language not to model

CRINGE loss leverages contrastive negative generation with iterative self-labeling, significantly improving safety and coherence in language models, outperforming baselines.

Leonard Adolphs, Tianyu Gao, Jing Xu et al.

2022-11-11 43
cs.CV 2211.05783

Unifying Flow, Stereo and Depth Estimation

Unified dense correspondence model using Transformer surpasses SOTA in optical flow, stereo, and depth tasks with shared parameters and no cost volume.

Haofei Xu, Jing Zhang, Jianfei Cai et al.

2022-11-11 59
cs.LG 2211.05244

Deep Learning for Time Series Anomaly Detection: A Survey

Proposes a comprehensive taxonomy of deep learning models for time series anomaly detection, covering forecasting, reconstruction, representation, and hybrid methods.

Zahra Zamanzadeh Darban, Geoffrey I. Webb, Shirui Pan et al.

2022-11-10 32
cs.LG 2211.05102

Efficiently Scaling Transformer Inference

Proposes multi-dimensional partitioning and communication optimization for TPU v4, achieving 29ms/token latency and 76% MFU on 540B models.

Reiner Pope, Sholto Douglas, Aakanksha Chowdhery et al.

2022-11-10 51
cs.SD 2211.02250

Real-Time Target Sound Extraction

Waveformer combines dilated causal convolution and Transformer for real-time target sound extraction, improving SI-SNRi by up to 3.3dB.

Bandhav Veluri, Justin Chan, Malek Itani et al.

2022-11-04 54