Behavior Sequence Transformer for E-commerce Recommendation in Alibaba
This paper introduces the Behavior Sequence Transformer (BST) for e-commerce CTR prediction, achieving significant online CTR gains.
Qiwei Chen, Huan Zhao, Wei Li et al.
This paper introduces the Behavior Sequence Transformer (BST) for e-commerce CTR prediction, achieving significant online CTR gains.
Qiwei Chen, Huan Zhao, Wei Li et al.
MP-DQN fixes P-DQN and wins on all three tasks: 0.987, 0.789, 0.913.
Craig J. Bester, Steven D. James, George D. Konidaris
Proposes a randomized block proximal distributed algorithm for high-dimensional nonsmooth convex optimization, ensuring convergence with explicit rates.
Francesco Farina, Giuseppe Notarstefano
DeepGLO combines global matrix factorization with local TCNs, addressing scale variance and global dependencies in high-dimensional time series forecasting.
Rajat Sen, Hsiang-Fu Yu, Inderjit Dhillon
Proposed a 'loss prediction module' for active learning, achieving a 0.42% accuracy increase on CIFAR-10.
Donggeun Yoo, In So Kweon
S4L combines self-supervised and semi-supervised learning, achieving new SOTA on ILSVRC-2012 with only 10% labels.
Xiaohua Zhai, Avital Oliver, Alexander Kolesnikov et al.
Proposes conformalized quantile regression (CQR), ensuring finite-sample, distribution-free coverage with shorter prediction intervals.
Yaniv Romano, Evan Patterson, Emmanuel J. Candès
Proposes jackknife+ for distribution-free predictive intervals with coverage ≥1−2α, robust to model instability.
Rina Foygel Barber, Emmanuel J. Candes, Aaditya Ramdas et al.
Proposes weakly-supervised attention for GNNs, achieving over 60% performance improvement on complex graph classification tasks.
Boris Knyazev, Graham W. Taylor, Mohamed R. Amer
FANTrack employs deep CNNs for 3D multi-object data association, achieving 77.72% MOTA on KITTI.
Erkan Baser, Venkateshwaran Balasubramanian, Prarthana Bhattacharyya et al.
Scaling Jigsaw and Colorization to 100M images reveals self-supervised learning's potential to rival supervised methods in specific tasks.
Priya Goyal, Dhruv Mahajan, Abhinav Gupta et al.
DRIT++ achieves diverse image-to-image translation via disentangled representations, evaluated using Fréchet Inception Distance and perceptual distance.
Hsin-Ying Lee, Hung-Yu Tseng, Qi Mao et al.
Proposed a billion-scale semi-supervised learning method using teacher/student paradigm; ResNet-50 achieved 81.2% accuracy on ImageNet.
I. Zeki Yalniz, Hervé Jégou, Kan Chen et al.
SuperGLUE introduces more challenging tasks and diverse formats, pushing NLP benchmarks toward deeper understanding.
Alex Wang, Yada Pruksachatkun, Nikita Nangia et al.
Introduces Centered Kernel Alignment (CKA) for measuring neural network representation similarity, overcoming CCA's high-dimensional limitations.
Simon Kornblith, Mohammad Norouzi, Honglak Lee et al.
The paper explores information-theoretic assumptions in batch reinforcement learning, providing theoretical results and sample complexity lower bounds.
Jinglin Chen, Nan Jiang
Proposes LU and LDL^H based algorithms for high-performance sampling of general DPPs, reducing complexity from O(n^3) to near O(n^2).
Jack Poulson
Introduced EigenPooling method to enhance graph classification performance, validated on 6 benchmark datasets.
Yao Ma, Suhang Wang, Charu C. Aggarwal et al.
HAWQ leverages Hessian spectrum for automatic mixed-precision quantization, achieving 8× activation compression on ResNet20 with only 0.15% accuracy loss.
Zhen Dong, Zhewei Yao, Amir Gholami et al.
Introduces an exact GPU-accelerated algorithm for Convolutional NTK, achieving 77% accuracy on CIFAR-10, outperforming previous kernel methods by 10%.
Sanjeev Arora, Simon S. Du, Wei Hu et al.