Learning Neural Templates for Recommender Dialogue System
Proposes NTRD framework using neural templates and slot filling for recommender dialogue, outperforming SOTA with 1.806% ReR@1 and zero-shot capabilities.
Zujie Liang, Huang Hu, Can Xu et al.
Proposes NTRD framework using neural templates and slot filling for recommender dialogue, outperforming SOTA with 1.806% ReR@1 and zero-shot capabilities.
Zujie Liang, Huang Hu, Can Xu et al.
CLIPort combines CLIP's semantic understanding with Transporter's spatial precision for multi-task robotic manipulation, achieving over 90% success rates.
Mohit Shridhar, Lucas Manuelli, Dieter Fox
Introduces sequential forecast calibration tests based on e-values, enabling real-time monitoring with statistical guarantees.
Sebastian Arnold, Alexander Henzi, Johanna F. Ziegel
Constructed Paint4Poem dataset for classical Chinese poem visualization; evaluated AttnGAN and MirrorGAN, showing style mimicry but limited semantic reflection.
Dan Li, Shuai Wang, Jie Zou et al.
Pix2Seq transforms object detection into sequence generation using a Transformer, achieving competitive COCO results with 43.0 AP from minimal task-specific assumptions.
Ting Chen, Saurabh Saxena, Lala Li et al.
SPLADE v2 enhances sparse lexical representations with max pooling and distillation, achieving over 9% NDCG@10 improvement on TREC DL 2019 for efficient first-stage retrieval.
Thibault Formal, Carlos Lassance, Benjamin Piwowarski et al.
ShapeMap 3-D employs GelSight tactile sensors and depth cameras, using Gaussian process spatial graphs for efficient dense shape mapping with high accuracy.
Sudharshan Suresh, Zilin Si, Joshua G. Mangelson et al.
RnG-KBQA combines contrastive ranking and generation, greatly improving zero-shot and generalization in KBQA.
Xi Ye, Semih Yavuz, Kazuma Hashimoto et al.
ActionCLIP transforms video action recognition into video-text matching, achieving 83.8% top-1 accuracy on Kinetics-400 with zero-shot transfer.
Mengmeng Wang, Jiazheng Xing, Yong Liu
ThriftyDAgger improves imitation learning efficiency with budget-aware novelty and risk gating, boosting performance by 58%.
Ryan Hoque, Ashwin Balakrishna, Ellen Novoseller et al.
HM3D is a large-scale, high-fidelity 3D indoor dataset with 1000 real building reconstructions, outperforming prior datasets and boosting embodied AI navigation.
Santhosh K. Ramakrishnan, Aaron Gokaslan, Erik Wijmans et al.
DAQA combines symbolic and deep learning methods for interpretability and generalizability.
Kwabena Nuamah
ObjectFolder integrates visual, auditory, and tactile implicit neural representations for multisensory object modeling, enabling recognition, retrieval, 3D reconstruction, and robotic grasping.
Ruohan Gao, Yen-Yu Chang, Shivani Mall et al.
RAFT-inspired multi-scale ConvGRU architecture achieves high-precision, real-time stereo matching with 29% error reduction on Middlebury.
Lahav Lipson, Zachary Teed, Jia Deng
This study introduces structured 'subject-style' prompts, analyzing 5493 images to optimize text prompts for better image quality in VQGAN+CLIP.
Vivian Liu, Lydia B. Chilton
Study shows language models can encode perceptual color structure without grounding, validated using CIELAB color space.
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich et al.
This study tracks neural language models' (NLMs) learning trajectories, revealing highly consistent stages across architectures and data, driven by an underlying inductive bias.
Leshem Choshen, Guy Hacohen, Daphna Weinshall et al.
This study analyzes Transformer attention weights in NMT, revealing biases towards uninformative tokens but demonstrating interpretability and improvement methods.
Javier Ferrando, Marta R. Costa-jussà
Proposed multi-topic, knowledge-enhanced art description framework using ResNet, BERT, Wikipedia, achieving significant improvements in diversity and accuracy.
Zechen Bai, Yuta Nakashima, Noa Garcia
Introduces vSLIP model with hierarchical planning enabling bipedal robots to navigate height-constrained environments autonomously.
Zhongyu Li, Jun Zeng, Shuxiao Chen et al.