Event-based Motion Deblurring via Multi-Temporal Granularity Fusion
MTGNet enhances event camera deblurring via multi-temporal granularity fusion, outperforming existing methods.
Xiaopeng Lin, Hongwei Ren, Yulong Huang et al.
MTGNet enhances event camera deblurring via multi-temporal granularity fusion, outperforming existing methods.
Xiaopeng Lin, Hongwei Ren, Yulong Huang et al.
This study evaluates GPT-4 mini's in-context learning for music entity recognition, revealing significant impact of entity exposure on performance and robustness.
Simon Hachmeier, Robert Jäschke
PEMC integrates ML predictors with Monte Carlo, reducing variance by 30-55% while maintaining unbiasedness and efficiency.
Fengpei Li, Haoxian Chen, Jiahe Lin et al.
Proposed a task-oriented dialogue system for Wolof using Rasa, cross-lingual transfer, and in-house machine translation, achieving intent classification F1 of 0.995.
Derguene Mbaye, Moussa Diallo
Entropy-regularized process reward model (ER-PRM) with KL regularization improves mathematical reasoning by guiding multi-step inference.
Hanning Zhang, Pengcheng Wang, Shizhe Diao et al.
METIS jointly optimizes query scheduling and configuration adaptation, reducing RAG latency by 1.64-2.54× while maintaining high quality.
Siddhant Ray, Rui Pan, Zhuohan Gu et al.
DEFAME employs a six-stage dynamic multimodal evidence retrieval framework, surpassing traditional text-only methods with significant accuracy gains.
Tobias Braun, Mark Rothermel, Marcus Rohrbach et al.
Diffusion-based 3D scene reconstruction from a single RGB image, achieving 12.04% AP3D and 13.43% F-Score improvements.
Manuel Dahnert, Angela Dai, Norman Müller et al.
Proposes SARA shield, using reachability analysis to classify contact types and ensure robot kinetic energy stays below injury thresholds in human environments.
Jakob Thumm, Julian Balletshofer, Leonardo Maglanoc et al.
Proposes STILL-2 framework combining imitation, exploration, and self-improvement to develop industry-level slow-thinking reasoning systems.
Yingqian Min, Zhipeng Chen, Jinhao Jiang et al.
WaLLoC combines wavelet transforms and autoencoders for efficient compression, outperforming state-of-the-art models in rate-distortion and downstream tasks.
Dan Jacobellis, Neeraja J. Yadwadkar
Score and Distribution Matching transforms diffusion policies into single-step generators, achieving 6× inference speedup with state-of-the-art action quality.
Bofang Jia, Pengxiang Ding, Can Cui et al.
phi-4 is a 14B parameter model leveraging synthetic data, surpassing GPT-4 in STEM QA, with innovative training strategies.
Marah Abdin, Jyoti Aneja, Harkirat Behl et al.
FlowEdit edits SD3/FLUX without inversion and reaches SOTA.
Vladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas et al.
Multi-GraspLLM generates multi-hand semantic-guided grasp poses using multimodal LLM, significantly improving accuracy.
Haosheng Li, Weixin Mao, Weipeng Deng et al.
FLIP employs flow-based generative planning with pixel-level flow and video models, achieving 100% success in long-horizon manipulation tasks.
Chongkai Gao, Haozhuo Zhang, Zhixuan Xu et al.
CogNav employs LLM-based cognitive modeling with a heterogeneous map, boosting ObjectNav success by at least 14%.
Yihan Cao, Jiazhao Zhang, Zhinan Yu et al.
Transforms slow bidirectional video diffusion into fast autoregressive models, achieving 9.4 FPS with high quality via distillation and ODE initialization.
Tianwei Yin, Qiang Zhang, Richard Zhang et al.
Introducing Predictive Motion Priors (PMP) based on CVAE for decoupled upper-limb prediction and robust lower-limb control, significantly improving humanoid robot manipulation.
Chenhao Lu, Xuxin Cheng, Jialong Li et al.
Introduces 3DSRBench, a benchmark with 2772 annotated QA pairs, to evaluate large multimodal models' 3D spatial reasoning, revealing current limitations especially in uncommon viewpoints.
Wufei Ma, Haoyu Chen, Guofeng Zhang et al.