GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.LG 2410.02628

Inverse Entropic Optimal Transport Solves Semi-supervised Learning via Data Likelihood Maximization

Proposes EBiEOT framework combining semi-supervised data via inverse entropic OT for conditional distribution maximization.

Mikhail Persiianov, Arip Asadulaev, Nikita Andreev et al.

2024-10-04 49
cs.CV 2410.02331

Self-eXplainable AI for Medical Image Analysis: A Survey and New Outlooks

Self-eXplainable AI with attention mechanisms enhances transparency in medical image analysis, covering 200 papers.

Junlin Hou, Sicen Liu, Yequan Bie et al.

2024-10-03 44
cs.LG 2410.02321

Convergence of Score-Based Discrete Diffusion Models: A Discrete-Time Analysis

Introduces Girsanov-based convergence bounds for discrete diffusion models in CTMC framework; KL divergence nearly linear in dimension d.

Zikun Zhang, Zixiang Chen, Quanquan Gu

2024-10-03 36
cs.LG 2410.02145

Active Learning of Deep Neural Networks via Gradient-Free Cutting Planes

Gradient-free cutting plane method for deep ReLU networks, providing convergence guarantees and improving sample efficiency in active learning.

Erica Zhang, Fangzhao Zhang, Mert Pilanci

2024-10-03 51
cs.SD 2410.02084

Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset

Generates symbolic music using LLM-enhanced MetaScore dataset, providing text and tag controls.

Weihan Xu, Julian McAuley, Taylor Berg-Kirkpatrick et al.

2024-10-03 25
cs.CL 2410.02052

ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning

ExACT combines R-MCTS and Exploratory Learning, achieving 6%-30% improvement on VisualWebArena.

Xiao Yu, Baolin Peng, Vineeth Vajipey et al.

2024-10-03 28
cs.RO 2410.02048

FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors

FeelAnyForce uses a multi-head Transformer with RGB and depth images to estimate 3D contact forces, achieving 4% MAE on unseen objects.

Amir-Hossein Shahidzadeh, Gabriele Caddeo, Koushik Alapati et al.

2024-10-03 50
cs.CV 2410.01962

LS-HAR: Language Supervised Human Action Recognition with Salient Fusion, Construction Sites as a Use-Case

LS-HAR leverages language-guided skeleton and visual features with Salient Fusion, achieving 89.5% accuracy on NTU-RGB+D XSub dataset.

Mohammad Mahdavian, Mohammad Loni, Ted Samuelsson et al.

2024-10-03 32
cond-mat.stat-mech 2410.01764

Integrable Matrix Probabilistic Diffusions and the Matrix Stochastic Heat Equation

Introduces matrix stochastic heat equation (MSHE), derives explicit invariant measure in 1D, and demonstrates classical integrability in weak-noise regime via inverse scattering.

Alexandre Krajenbrink, Pierre Le Doussal

2024-10-03 45
cs.CV 2410.01744

Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks

Leopard model excels in multi-image text tasks, surpassing Llama-3.2 using 1.2M open-source data.

Mengzhao Jia, Wenhao Yu, Kaixin Ma et al.

2024-10-03 18
cs.CV 2410.01699

Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding

Proposed training-free Probabilistic Speculative Jacobi Decoding (SJD) accelerates autoregressive text-to-image generation by ~2×, maintaining quality and diversity.

Yao Teng, Han Shi, Xian Liu et al.

2024-10-03 53
cs.CV 2410.01341

Cognition Transferring and Decoupling for Text-supervised Egocentric Semantic Segmentation

CTDN method for text-supervised egocentric semantic segmentation, significantly improving accuracy.

Zhaofeng Shi, Heqian Qiu, Lanxiao Wang et al.

2024-10-02 22
cs.CL 2410.00741

VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Proposed VideoCLIP-XL with TPCM, achieving significant improvements in long description video understanding, trained on over 2 million video-long description pairs.

Jiapeng Wang, Chengyu Wang, Kunzhe Huang et al.

2024-10-01 40
cs.RO 2410.00425

ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI

ManiSkill3 is a GPU-parallelized robotics simulator supporting 12 task domains, achieving 30,000+ FPS with 2-3x less memory, enabling fast visual RL training.

Stone Tao, Fanbo Xiang, Arth Shukla et al.

2024-10-01 47
cs.RO 2410.00157

Constraining Gaussian Process Implicit Surfaces for Robot Manipulation via Dataset Refinement

COGIS employs Gaussian Process Implicit Surfaces to online model obstacles, enhancing robot manipulation in partial observability with constraint-based dataset refinement.

Abhinav Kumar, Peter Mitrano, Dmitry Berenson

2024-10-01 47
eess.AS 2409.20007

DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

DeSTA2 uses automatic speech-text pair generation, avoiding instruction tuning, to enhance speech understanding.

Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu et al.

2024-09-30 43
cs.CV 2409.19930

EndoDepth: A Benchmark for Assessing Robustness in Endoscopic Depth Prediction

EndoDepth benchmark evaluates robustness of monocular endoscopic depth prediction using mDERS and SCARED-C dataset, revealing strengths and weaknesses of SOTA models.

Ivan Reyes-Amezcua, Ricardo Espinosa, Christian Daul et al.

2024-09-30 78
cs.CL 2409.19898

UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs

UniSumEval evaluates 9 LLMs on fine-grained, multi-dimensional summarization across 9 domains, addressing gaps in existing benchmarks.

Yuho Lee, Taewon Yun, Jason Cai et al.

2024-09-30 44
cs.IR 2409.19824

Counterfactual Evaluation of Ads Ranking Models through Domain Adaptation

Proposes a domain-adapted reward model for offline ad ranking evaluation, outperforming IPS and non-generalized models.

Mohamed A. Radwan, Himaghna Bhattacharjee, Quinn Lanners et al.

2024-09-30 74
cs.RO 2409.19778

Lessons Learned from Developing a Human-Centered Guide Dog Robot for Mobility Assistance

Human-centered quadruped guide dog robot utilizing visual foundation models and RL control, achieving 95% obstacle avoidance success and 2-hour battery life.

Hochul Hwang, Ken Suzuki, Nicholas A Giudice et al.

2024-09-30 35
Prev 1 ... 331 332 333 334 335 336 337 ... 573 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home