cs.CV 2303.11797

CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Proposes CAT-Seg, a cost aggregation-based open-vocabulary semantic segmentation framework, fine-tuning CLIP encoders to improve pixel-level recognition of seen and unseen classes, outperforming state-of-the-art.

Seokju Cho, Heeseong Shin, Sunghwan Hong et al.

2023-03-21 297 citations 49
cs.SD 2303.11510

ICASSP 2023 Deep Noise Suppression Challenge

E3Net-based deep noise suppression with speaker embedding improves signal quality by 0.145 score points and WAcc to 85.3% in ICASSP 2023 challenge.

Harishchandra Dubey, Ashkan Aazami, Vishak Gopal et al.

2023-03-21 49
cs.CV 2303.11328

Zero-1-to-3: Zero-shot One Image to 3D Object

Zero-1-to-3 leverages pre-trained diffusion models for zero-shot single-image 3D view synthesis and reconstruction, with strong generalization.

Ruoshi Liu, Rundi Wu, Basile Van Hoorick et al.

2023-03-21 47
cs.CL 2303.11315

Context-faithful Prompting for Large Language Models

Enhance LLMs' contextual faithfulness using opinion-based prompts and counterfactual demonstrations, significantly reducing memorization ratio.

Wenxuan Zhou, Sheng Zhang, Hoifung Poon et al.

2023-03-21 10
cs.CL 2303.13375

Capabilities of GPT-4 on Medical Challenge Problems

GPT-4 surpasses USMLE passing scores by over 20 points without domain-specific tuning, demonstrating strong reasoning and calibration.

Harsha Nori, Nicholas King, Scott Mayer McKinney et al.

2023-03-21 48
cs.LG 2303.09331

Model Based Explanations of Concept Drift

Proposes a model-based explanation framework for concept drift, leveraging spatial feature changes and explanation tools like LIME and SHAP.

Fabian Hinder, Valerie Vaquet, Johannes Brinkrolf et al.

2023-03-16 59 citations 55
cs.CV 2303.09295

DIRE for Diffusion-Generated Image Detection

Proposes DIRE, leveraging diffusion model reconstruction error for high-accuracy detection of diffusion-generated images, outperforming existing methods.

Zhendong Wang, Jianmin Bao, Wengang Zhou et al.

2023-03-16 44
cs.CL 2303.08774

GPT-4 Technical Report

GPT-4 is a multimodal Transformer trained with predictive and RLHF methods, achieving human-level performance and scalable predictability.

OpenAI, Josh Achiam, Steven Adler et al.

2023-03-16 36