Logics-Parsing-Omni Technical Report
Omni Parsing framework standardizes multimodal data parsing via a unified taxonomy and progressive parsing paradigm.
Xin An, Jingyi Cai, Xiangyang Chen et al.
Omni Parsing framework standardizes multimodal data parsing via a unified taxonomy and progressive parsing paradigm.
Xin An, Jingyi Cai, Xiangyang Chen et al.
Proposed a Variational Latent Equilibrium (VLE) method to approximate BPTT in a biologically plausible manner, enhancing complex spatiotemporal pattern learning.
Simon Brandt, Paul Haider, Walter Senn et al.
StateFactory uses hierarchical object-attribute structures for zero-shot reward prediction, reducing EPIC distance by 60%.
Yijun Shen, Delong Chen, Xianming Hu et al.
OddGridBench reveals MLLMs' deficiency in visual discrepancy detection, proposing OddGrid-GRPO to enhance performance.
Tengjin Weng, Wenhao Jiang, Jingyi Wang et al.
RoomTour3D-IGR learns implicit geometry from web videos, improving NaviLLM by over 6% across four VLN benchmarks.
Mingfei Han, Haihong Hao, Liang Ma et al.
TTC layer enhances LLM reasoning via LQR planning, achieving 27.8% improvement on MATH-500.
Peihao Wang, Shan Yang, Xijun Wang et al.
ZeroWBC leverages egocentric videos and vision-language models to generate and execute whole-body humanoid behaviors without teleoperation, achieving diverse scene-aware actions.
Haoran Yang, Jiacheng Bao, Yucheng Xin et al.
WS-Net addresses weak signal collapse via state-space and weak signal attention fusion, achieving up to 55% RMSE and 63% SAD reductions.
Zekun Long, Ali Zia, Guanyiman Fu et al.
MEMO enhances multi-agent LLM game performance by memory-augmented context optimization, boosting win rate from 25.1% to 49.5% with reduced variance.
Yunfei Xie, Kevin Wang, Bobby Cheng et al.
Proposes a duality-based weakly polynomial algorithm for parametric submodular minimization with complexity O(n² log(nM∥d∥₁)+n³ log(nM∥d∥₁)+SFM)
Swati Gupta, Alec Zhu
SecAgent enhances mobile GUI agent efficiency with semantic context and a Chinese dataset, matching 7B-8B model performance.
Yiping Xie, Song Chen, Jingxuan Xing et al.
Introduces adaptive looping and memory banks in Transformers, achieving 22% math BPB improvement and recovering commonsense performance.
Markus Frey, Behzad Shomali, Ali Hamza Bashir et al.
HDR-NSFF introduces 4D spatio-temporal neural scene flow fields for high-quality dynamic HDR scene reconstruction, outperforming 2D fusion methods.
Shin Dong-Yeon, Kim Jun-Seong, Kwon Byung-Ki et al.
MM-TS method enhances contrastive learning with long-tail data through dynamic temperature and margin schedules.
Siarhei Sheludzko, Dhimitrios Duka, Bernt Schiele et al.
Ramsa corpus provides baseline ASR and TTS for Emirati Arabic; Whisper-large-v3-turbo excels.
Rania Al-Sabbagh
QualiTeacher enhances image restoration by conditioning on pseudo-label quality, significantly improving model generalization.
Fengyang Xiao, Jingjia Feng, Peng Hu et al.
HILA framework with Dual-Loop Policy Optimization enables adaptive human–agent collaboration, outperforming state-of-the-art multi-agent systems by 10%+ on reasoning benchmarks.
Wei Yang, Defu Cao, Jiacheng Pang et al.
Proposes DMRAL framework combining relation graphs and sub-question guided reasoning, boosting large-scale multi-table numerical QA by 24% retrieval and 55% accuracy.
Feng Luo, Hai Lan, Hui Luo et al.
Proposed Masked Motion Diffusion Model (MMDM) with Kinematic Attention Aggregation for robust motion reconstruction.
Junkun Jiang, Jie Chen, Ho Yin Au et al.
UniUncer enhances driving accuracy by integrating dynamic-static uncertainty, reducing trajectory error by 7%.
Yu Gao, Jijun Wang, Zongzheng Zhang et al.