Exploring CLIP for Assessing the Look and Feel of Images
Utilizing CLIP for image perception assessment, achieving high correlation.
Jianyi Wang, Kelvin C. K. Chan, Chen Change Loy
Utilizing CLIP for image perception assessment, achieving high correlation.
Jianyi Wang, Kelvin C. K. Chan, Chen Change Loy
Proposes MHST, a multi-modal transformer model integrating text, layout, and visual features, achieving significant improvements on the TAT-DQA dataset for complex document VQA.
Fengbin Zhu, Wenqiang Lei, Fuli Feng et al.
PanGu-Coder uses two-stage function-level training, reaching 17.07% HumanEval pass@1 with 317M parameters.
Fenia Christopoulou, Gerasimos Lampouras, Milan Gritta et al.
Proposes VAT with 4D Convolutional Swin Transformer for few-shot segmentation, achieving state-of-the-art results.
Sunghwan Hong, Seokju Cho, Jisu Nam et al.
CodeT leverages pre-trained models to generate test cases, using dual execution agreement to significantly improve code solution selection accuracy.
Bei Chen, Fengji Zhang, Anh Nguyen et al.
RAVOS combines deep motion modeling with ROI prediction, achieving 86.1 J&F at 42 FPS for video object segmentation.
Bo Miao, Mohammed Bennamoun, Yongsheng Gao et al.
Diffsound uses a discrete diffusion model for non-autoregressive text-to-sound generation, achieving 5x faster inference and higher quality (MOS 3.56 vs 2.786).
Dongchao Yang, Jianwei Yu, Helin Wang et al.
Proposes Actor-Critic-based improper RL algorithms for combining multiple controllers to optimize unknown MDPs, with proven convergence rates.
Mohammadi Zaki, Avinash Mohan, Aditya Gopalan et al.
This work reveals SGD leverages Fourier spectral gaps to gradually amplify sparse features, approaching the computational limit in learning k-sparse parity problems.
Boaz Barak, Benjamin L. Edelman, Surbhi Goel et al.
Construct highly symmetric line systems using twisted spherical functions of finite groups, improving kissing number lower bounds for d=10,11,14.
Mikhail Ganzhinov
Proposes ST-P3, an end-to-end vision-based autonomous driving model using spatial-temporal features, achieving 85.2% mAP on nuScenes.
Shengchao Hu, Li Chen, Penghao Wu et al.
XMem employs a multi-store memory framework inspired by human cognition, achieving state-of-the-art long-video object segmentation with low GPU usage.
Ho Kei Cheng, Alexander G. Schwing
Introduced Confident Adaptive Language Modeling (CALM) to reduce computation while maintaining performance, achieving up to 3x speedup.
Tal Schuster, Adam Fisch, Jai Gupta et al.
RMT enhances long-sequence processing with memory, surpassing Transformer-XL.
Aydar Bulatov, Yuri Kuratov, Mikhail S. Burtsev
Proposes a unified 2D-3D molecular pretraining model, boosting property prediction accuracy by 8.3%.
Jinhua Zhu, Yingce Xia, Lijun Wu et al.
BM3 introduces a self-supervised multi-modal recommendation framework using dropout-based contrastive views, achieving 2-9x faster training and outperforming state-of-the-art on large datasets.
Xin Zhou, Hongyu Zhou, Yong Liu et al.
Language models can self-evaluate answer accuracy using P(True) and P(IK) methods effectively.
Saurav Kadavath, Tom Conerly, Amanda Askell et al.
Using CBF-based decentralized controllers, the study analyzes stability's role in balancing safety and liveness, demonstrating PCCA's superior performance with 15% faster arrival and no deadlocks.
Mrdjan Jankovic, Mario Santillo, Yan Wang
Combining pretrained large models' in-context learning with scratchpad prompting significantly enhances length generalization.
Cem Anil, Yuhuai Wu, Anders Andreassen et al.
Introduces PPD to train CNFs on manifolds, avoiding ODE solving, enabling high-dimensional generation.
Heli Ben-Hamu, Samuel Cohen, Joey Bose et al.