Towards Generic Anomaly Detection and Understanding: Large-scale Visual-linguistic Model (GPT-4V) Takes the Lead
GPT-4V excels in multi-modal anomaly detection, particularly in zero/one-shot scenarios.
Yunkang Cao, Xiaohao Xu, Chen Sun et al.
GPT-4V excels in multi-modal anomaly detection, particularly in zero/one-shot scenarios.
Yunkang Cao, Xiaohao Xu, Chen Sun et al.
FETV benchmark introduces multi-dimensional, temporal-aware fine-grained evaluation for open-source T2V models, revealing poor correlation of existing metrics with human judgment.
Yuanxin Liu, Lei Li, Shuhuai Ren et al.
Introduced a variable selection method using Maximum Mean Discrepancy for interpretable distribution comparison.
Kensuke Mitsuzawa, Motonobu Kanagawa, Stefano Bortoli et al.
RoboGen uses generative models to automatically create diverse tasks and environments, enhancing robotic skill learning efficiency.
Yufei Wang, Zhou Xian, Feng Chen et al.
RoboFlamingo leverages open-source VLMs for robot control, achieving 96.4% success rate, outperforming SOTA by a large margin.
Xinghang Li, Minghuan Liu, Hanbo Zhang et al.
Copilot4D combines VQVAE and discrete diffusion for point-cloud forecasting, cutting Chamfer distance by over 65% at 1 s and 50% at 3 s.
Lunjun Zhang, Yuwen Xiong, Ze Yang et al.
nvblox leverages GPU acceleration for incremental Signed Distance Field (SDF) mapping, achieving 177× faster surface reconstruction and 31× faster distance field computation, significantly enhancing robotic path planning.
Alexander Millane, Helen Oleynikova, Emilie Wirbel et al.
This study reveals vulnerabilities of BERTScore, BLEURT, and COMET under adversarial attacks, proposing methods to enhance their robustness.
Yichen Huang, Timothy Baldwin
Leveraging large language models (LLMs) for graph augmentation significantly improves recommendation accuracy under data sparsity.
Wei Wei, Xubin Ren, Jiabin Tang et al.
ChipNeMo combines DAPT, instruction fine-tuning, and RAG to enhance chip design tasks, outperforming GPT-4 in key benchmarks.
Mingjie Liu, Teodor-Dumitru Ene, Robert Kirby et al.
Using LoRA, successfully reversed Llama 2-Chat 70B's safety training, reducing refusal rate to 1%.
Simon Lermen, Charlie Rogers-Smith, Jeffrey Ladish
RecInterpreter aligns sequential recommender embeddings to LLaMA with a linear adapter and residual prompts on MovieLens100K and Steam.
Zhengyi Yang, Jiancan Wu, Yanchen Luo et al.
FollowBench introduces a multi-level, fine-grained benchmark for LLM instruction following, covering five constraint types, revealing models' limitations at higher difficulty levels with detailed metrics.
Yuxin Jiang, Yufei Wang, Xingshan Zeng et al.
SimMMDG splits features into shared and specific parts, using contrastive learning and translation for robust multi-modal domain generalization, outperforming baselines on EPIC-Kitchens and HAC.
Hao Dong, Ismail Nejjar, Han Sun et al.
Study reveals vision-language models struggle with spatial reasoning, achieving only 56% accuracy on What'sUp benchmark.
Amita Kamath, Jack Hessel, Kai-Wei Chang
Proposes high-quality open-source diffusion models for T2V (1024×576) and I2V with content preservation, advancing video synthesis.
Haoxin Chen, Menghan Xia, Yingqing He et al.
Proposes a restoration gap metric and guidance strategy to improve diffusion-based behavior synthesis, achieving 8-10% performance gains on offline benchmarks.
Kyowoon Lee, Seongun Kim, Jaesik Choi
Survey on embedding techniques in recommender systems, covering matrix, sequential, and graph structures.
Maolin Wang, Xinjian Zhao, Wanyu Wang et al.
ProtoConcepts extends prototype networks with multiple visualizations, using geometric prototype balls, achieving comparable accuracy and significantly improved interpretability.
Chiyu Ma, Brandon Zhao, Chaofan Chen et al.
Proposed Spatial-Temporal Diversification Network (STDN) enhances video domain generalization via space-time content diversity, using spatial grouping and multi-scale relation modeling.
Kun-Yu Lin, Jia-Run Du, Yipeng Gao et al.