DIRE for Diffusion-Generated Image Detection
Proposes DIRE, leveraging diffusion model reconstruction error for high-accuracy detection of diffusion-generated images, outperforming existing methods.
Zhendong Wang, Jianmin Bao, Wengang Zhou et al.
Proposes DIRE, leveraging diffusion model reconstruction error for high-accuracy detection of diffusion-generated images, outperforming existing methods.
Zhendong Wang, Jianmin Bao, Wengang Zhou et al.
Introduces MATH 401 dataset, evaluates GPT-4, ChatGPT on arithmetic tasks, showing GPT-4's superior performance with 83.54% accuracy.
Zheng Yuan, Hongyi Yuan, Chuanqi Tan et al.
ART enables automatic multi-step reasoning and tool use in LLMs, improving unseen task performance by over 22%.
Bhargavi Paranjape, Scott Lundberg, Sameer Singh et al.
SelfCheckGPT detects factual hallucinations via multi-sample response consistency without external data.
Potsawee Manakul, Adian Liusie, Mark J. F. Gales
GPT-4 is a multimodal Transformer trained with predictive and RLHF methods, achieving human-level performance and scalable predictability.
OpenAI, Josh Achiam, Steven Adler et al.
ViperGPT uses code generation with GPT-3 Codex to perform visual reasoning without training, achieving state-of-the-art zero-shot results.
Dídac Surís, Sachit Menon, Carl Vondrick
Introduces 'Tuned Lens'—layer-wise affine transformations—improving intermediate decoding accuracy and interpretability in transformers.
Nora Belrose, Igor Ostrovsky, Lev McKinney et al.
Proposes TCSF, a compressed-domain TSG framework using I-frames, motion vectors, residuals, achieving superior performance with lower complexity.
Xiang Fang, Daizong Liu, Pan Zhou et al.
MobileVOS combines contrastive learning and knowledge distillation to enable real-time video object segmentation on mobile devices, with 32x fewer parameters and 5x faster inference.
Roy Miles, Mehmet Kerim Yucel, Bruno Manganelli et al.
SR-init assesses layer redundancy by measuring accuracy drop after stochastic re-initialization, enabling interpretable pruning with significant parameter reduction.
Hui Tang, Yao Lu, Qi Xuan
AVLMaps fuses audio, visual, and language features into a 3D map, boosting robot multimodal goal navigation by 50%.
Chenguang Huang, Oier Mees, Andy Zeng et al.
Using Flamingo as a success detector framed as VQA, achieving cross-domain generalization with minimal human annotations.
Yuqing Du, Ksenia Konyushkova, Misha Denil et al.
Negative-entropy FTRL achieves three-world adaptation: adversarial O(√dT logT log|D|T) and corrupted-stochastic O(d logT/Δmin plus corruption terms).
Fang Kong, Canzhe Zhao, Shuai Li
Retinexformer integrates one-stage Retinex framework with Transformer, achieving significant PSNR gains over 4dB on 13 datasets.
Yuanhao Cai, Hao Bian, Jing Lin et al.
VLOSS enhances open-world segmentation using omni-supervised data, surpassing MaskCLIP by 2% on LVIS v1.
Bowen Dong, Jiaxi Gu, Jianhua Han et al.
This paper introduces LLM-GROP, integrating large language models with task and motion planning to improve multi-object rearrangement success rate by over 20%.
Yan Ding, Xiaohan Zhang, Chris Paxton et al.
GigaGAN introduces a new architecture for text-to-image synthesis, generating 512px images in just 0.13 seconds.
Minguk Kang, Jun-Yan Zhu, Richard Zhang et al.
Proposed a Text-Visual Prompting (TVP) framework, achieving a 9.79% improvement on Charades-STA dataset.
Yimeng Zhang, Xin Chen, Jinghan Jia et al.
Visual ChatGPT integrates visual foundation models for image and language interaction.
Chenfei Wu, Shengming Yin, Weizhen Qi et al.
Magnushammer employs contrastive learning with Transformer to improve premise retrieval, achieving 59.5% success on PISA with 4x fewer parameters.
Maciej Mikuła, Szymon Tworkowski, Szymon Antoniak et al.