PP-OCR: A Practical Ultra Lightweight OCR System
PP-OCR is an ultra-lightweight OCR system with a 3.5M model for 6622 Chinese characters.
Yuning Du, Chenxia Li, Ruoyu Guo et al.
PP-OCR is an ultra-lightweight OCR system with a 3.5M model for 6622 Chinese characters.
Yuning Du, Chenxia Li, Ruoyu Guo et al.
Introduces 3D-FUTURE, a large-scale dataset with 20,240 synthetic indoor images and 9,992 detailed textured furniture models for multi-task 3D scene understanding.
Huan Fu, Rongfei Jia, Lin Gao et al.
This study applies classical machine learning and Seq2Seq models (LSTM, Transformer) to Minangkabau sentiment analysis and machine translation, constructing the first relevant corpora.
Fajri Koto, Ikhwan Koto
GraphCodeBERT leverages data flow structures with a graph-guided attention mechanism, achieving state-of-the-art results in code understanding tasks.
Daya Guo, Shuo Ren, Shuai Lu et al.
Survey of transfer learning methods in deep reinforcement learning, analyzing algorithms, applications, and challenges.
Zhuangdi Zhu, Kaixiang Lin, Anil K. Jain et al.
TadGAN uses GANs for time series anomaly detection, achieving the highest average F1 score.
Alexander Geiger, Dongyu Liu, Sarah Alnegheimish et al.
Using Hopf oscillator with delay feedback, optimal performance occurs at delay ≈ 1.6×input period, maximizing memory capacity.
Felix Köster, Dominik Ehlert, Kathy Lüdge
GeDi uses smaller LMs as generative discriminators to guide large LMs, enhancing safety and control.
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann et al.
SpanBasedSP predicts span trees to improve compositional generalization, boosting accuracy from 61.0 to 88.9 on key datasets.
Jonathan Herzig, Jonathan Berant
Graph Neural Network-based object importance prediction accelerates large-scale planning, outperforming baseline partial grounding strategies.
Tom Silver, Rohan Chitnis, Aidan Curtis et al.
Introduced IndoNLU benchmark with 12 tasks, trained IndoBERT models, outperforming multilingual models by significant margins.
Bryan Wilie, Karissa Vincentio, Genta Indra Winata et al.
Network dissection systematically identifies semantic roles of individual units in CNNs and GANs, revealing object detectors and controllable scene features.
David Bau, Jun-Yan Zhu, Hendrik Strobelt et al.
PPG introduces phased training with separate policy and value updates, boosting sample efficiency by ~30% on Procgen benchmarks.
Karl Cobbe, Jacob Hilton, Oleg Klimov et al.
A comprehensive 57-task benchmark shows GPT-3 (175B) improves 20% over random, revealing significant knowledge gaps.
Dan Hendrycks, Collin Burns, Steven Basart et al.
KILT benchmark unifies Wikipedia-based knowledge tasks with dense retrieval and Seq2Seq models, outperforming task-specific approaches.
Fabio Petroni, Aleksandra Piktus, Angela Fan et al.
WaveGrad uses diffusion-based gradient estimation, generating high-fidelity speech with only six iterations.
Nanxin Chen, Yu Zhang, Heiga Zen et al.
This paper introduces a deep CNN denoiser integrated into a half-quadratic splitting framework, significantly improving image deblurring, super-resolution, and demosaicing performance.
Kai Zhang, Yawei Li, Wangmeng Zuo et al.
Proposes DeepAL framework combining Bayesian sampling and batch strategies, reducing labeling costs by 30% while maintaining high accuracy.
Pengzhen Ren, Yun Xiao, Xiaojun Chang et al.
Validation of contact point angular momentum control in LIP-based model improves Cassie Blue's bipedal locomotion.
Yukai Gong, Jessy Grizzle
MTOP introduces a multilingual semantic parsing dataset with 100k utterances across 6 languages and 11 domains, achieving +6.3 Slot F1 with XLM-R models.
Haoran Li, Abhinav Arora, Shuohui Chen et al.