RedCaps: web-curated image-text data created by the people, for the people
RedCaps leverages Reddit data, creating 12M image-text pairs, improving vision-language models' performance.
Karan Desai, Gaurav Kaul, Zubin Aysola et al.
RedCaps leverages Reddit data, creating 12M image-text pairs, improving vision-language models' performance.
Karan Desai, Gaurav Kaul, Zubin Aysola et al.
DyTox achieves continual learning with dynamic token expansion, excelling on ImageNet1000.
Arthur Douillard, Alexandre Ramé, Guillaume Couairon et al.
Neural Schrödinger-Föllmer flows enable finite-time, low-variance Bayesian inference in high dimensions via stochastic control and neural network parametrization.
Francisco Vargas, Andrius Ovsianas, David Fernandes et al.
Swin Transformer V2 combines post-norm, cosine attention, and Log-CPB to scale vision models to 3B parameters and 84.0% ImageNet-V2 accuracy.
Ze Liu, Han Hu, Yutong Lin et al.
XLS-R model leverages wav2vec 2.0 for cross-lingual speech representation learning using 128 languages.
Arun Babu, Changhan Wang, Andros Tjandra et al.
LayoutLM model using Transformer architecture enhances Document AI task accuracy.
Lei Cui, Yiheng Xu, Tengchao Lv et al.
Proposed GRI combines offline expert demonstrations with online exploration, boosting urban autonomous driving scores by 17%.
Raphael Chekroun, Marin Toromanoff, Sascha Hornauer et al.
Proposed Mask-guided Spectral-wise Transformer (MST) achieves superior hyperspectral image reconstruction, outperforming SOTA with 6dB PSNR gain and reduced parameters by 54%.
Yuanhao Cai, Jing Lin, Xiaowan Hu et al.
iBOT employs online tokenizer-based masked image modeling, achieving 82.3% linear probing accuracy on ImageNet-1K.
Jinghao Zhou, Chen Wei, Huiyu Wang et al.
Proposed CLUE uses contrastive learning to scale user representations, outperforming task-specific models with significant transferability.
Kyuyong Shin, Hanock Kwak, Su Young Kim et al.
Proposes RSE-RL, combining VAE and SAC to recursively enhance images, outperforming traditional filters with 3.5dB PSNR gain on CelebA.
Chandrajit Bajaj, Yi Wang, Yunhao Yang
Proposes EN and PALM algorithms for on-the-fly co-occurrence rectification, enabling scalable large-vocabulary topic inference with high robustness.
Moontae Lee, Sungjun Cho, Kun Dong et al.
Comparative analysis of DeepONet and FNO, with practical extensions, highlighting robustness in complex geometries and noisy data.
Lu Lu, Xuhui Meng, Shengze Cai et al.
GRCN employs adaptive graph refinement with prototype networks to improve multimedia recommendation with implicit feedback, achieving over 9% improvement in Recall@10.
Wei Yinwei, Wang Xiang, Nie Liqiang et al.
Proposes MISLID, an adaptive fixed-confidence Top-m algorithm for misspecified linear bandits, matching the theoretical lower bound.
Clémence Réda, Andrea Tirinzoni, Rémy Degenne
Leveraging large pre-trained language models like BERT to enhance NLP tasks through fine-tuning, prompting, and text generation.
Bonan Min, Hayley Ross, Elior Sulem et al.
Proposes tighter online confidence intervals for RKHS elements, reducing width growth from O(√γn) to near logarithmic, enhancing regret bounds.
Sattar Vakili, Jonathan Scarlett, Tara Javidi
Scatterbrain unifies sparse and low-rank attention approximation using LSH and kernel features, reducing error by 2.1× over baselines.
Beidi Chen, Tri Dao, Eric Winsor et al.
MEST framework employs Elastic Mutation and Soft Memory Bound to enable accurate, fast sparse training on edge devices, boosting Top-1 accuracy on ImageNet by over 2% at 90% sparsity.
Geng Yuan, Xiaolong Ma, Wei Niu et al.
Introduces Linear State-Space Layers (LSSL), combining RNN, convolution, and continuous-time models, excelling at long dependencies.
Albert Gu, Isys Johnson, Karan Goel et al.