The Third Pillar of Causal Analysis? A Measurement Perspective on Causal Representations
Proposes T-MEX score within measurement model framework to evaluate causal representations in CRL.
Dingling Yao, Shimeng Huang, Riccardo Cadei et al.
Proposes T-MEX score within measurement model framework to evaluate causal representations in CRL.
Dingling Yao, Shimeng Huang, Riccardo Cadei et al.
FSDrive employs visual spatio-temporal Chain-of-Thought (CoT) to unify future scene prediction and trajectory planning, improving accuracy and safety in autonomous driving.
Shuang Zeng, Xinyuan Chang, Mengwei Xie et al.
OphNet-3D dataset enables dynamic 3D hand-instrument reconstruction in ophthalmic surgery, reducing MPJPE to 2.3mm and improving interaction metrics by 23%.
Ming Hu, Zhengdi Yu, Feilong Tang et al.
QwenLong-L1 employs progressive context scaling and RL algorithms (GRPO, DAPO) to enhance long-text reasoning, outperforming existing models.
Fanqi Wan, Weizhou Shen, Shengyi Liao et al.
Proposes Agent Distillation, transferring LLM agent behaviors into small models with retrieval and code tools, achieving performance comparable to larger models.
Minki Kang, Jongwon Jeong, Seanie Lee et al.
Proposes AMRIV, an adaptive IV estimator achieving semiparametric efficiency under noncompliance.
Miruna Oprescu, Brian M Cho, Nathan Kallus
Direct3D-S2 uses Spatial Sparse Attention for efficient 3D generation, achieving significant speedups.
Shuang Wu, Youtian Lin, Feihu Zhang et al.
Introduced Value-Guided Search (VGS) using a 1.5B value model trained on 2.5M reasoning traces for efficient long-context reasoning.
Kaiwen Wang, Jin Peng Zhou, Jonathan Chang et al.
Layerwise Compression-Expression phenomenon reveals task info capture and generation in ICL.
Jiachen Jiang, Yuxin Dong, Jinxin Zhou et al.
CrossLMM employs dual cross-attention to compress long video sequences, maintaining performance with fewer tokens.
Shilin Yan, Jiaming Han, Joey Tsai et al.
RIPT-VLA fine-tunes pretrained VLA models via reinforcement learning, boosting success rate to 97.5% with only one demonstration.
Shuhan Tan, Kairan Dou, Yue Zhao et al.
Proposes TON framework with Thought Dropout and GRPO for selective reasoning, reducing 90% inference length while maintaining accuracy.
Jiaqi Wang, Kevin Qinghong Lin, James Cheng et al.
Capped Squared Loss learns contextual value distributions with O(dξ²cmax⁴/ε⁸δ²) samples.
Anna Heuser, Thomas Kesselheim
DeepRec enhances recommendation by multi-turn interactions between LLMs and TRMs, significantly improving performance.
Bowen Zheng, Xiaolei Wang, Enze Liu et al.
KRIS-Bench employs a cognitive-inspired taxonomy to evaluate models' reasoning, covering 22 tasks across 7 dimensions with a focus on knowledge plausibility.
Yongliang Wu, Zonghui Li, Xinting Hu et al.
Proposes OneDC, a one-step diffusion image codec with semantic distillation, reducing bitrate by 39% and decoding time by 20×.
Naifu Xue, Zhaoyang Jia, Jiahao Li et al.
ToDi adaptively combines FKL and RKL per token via probability ratio, significantly improving distillation accuracy.
Seongryong Jung, Suwan Yoon, DongGeon Kim et al.
Think-RM models internal reasoning to enable long-horizon inference, outperforming BT RM and scaled GenRM by 8% on RM-Bench.
Ilgee Hong, Changlong Yu, Liang Qiu et al.
IRONIC framework achieves state-of-the-art zero-shot sarcasm detection using multi-modal coherence relations.
Aashish Anantha Ramakrishnan, Aadarsh Anantha Ramakrishnan, Dongwon Lee
Using lightweight fine-tuning to probe concept erasure reversibility, revealing existing methods only achieve superficial suppression.
Ping Liu, Chi Zhang