Training deep physical neural networks with local physical information bottleneck
The Physical Information Bottleneck (PIB) framework enhances energy efficiency and speed in deep physical neural networks.
Hao Wang, Ziao Wang, Xiangpeng Liang et al.
The Physical Information Bottleneck (PIB) framework enhances energy efficiency and speed in deep physical neural networks.
Hao Wang, Ziao Wang, Xiangpeng Liang et al.
Proposes two training-free trajectory smoothing methods (Look-Ahead and Look-Back) to improve image generation stability and quality in flow-based models.
Yan Luo, Henry Huang, Todd Y. Zhou et al.
UI-Venus-1.5 achieves 77.6% success rate in AndroidWorld navigation via model merging and online reinforcement learning.
Venus Team, Changlong Gao, Zhangxuan Gu et al.
MotionCrafter employs a 4D VAE to jointly reconstruct dense geometry and scene flow from monocular videos, achieving 38.64% geometry and 25% motion improvements.
Ruijie Zhu, Jiahao Lu, Wenbo Hu et al.
AnomSeer combines ExpCoT and TimerPO to enhance fine-grained reasoning and detection accuracy in TSAD, outperforming baselines with 58.8% F1.
Junru Zhang, Lang Feng, Haoran Shi et al.
VideoVeritas combines perception pretraining and fact-based reasoning, using the MintVid dataset to detect deepfake videos with 85.7% accuracy.
Hao Tan, Jun Lan, Senyuan Shi et al.
Instance-disentangled attention improves multi-region editing in flow matching models, reducing attribute leakage and enhancing local control.
Carmine Zaccagnino, Fabio Quattrini, Enis Simsar et al.
LEFT fuses tri-view features with analysis-synthesis cycle consistency for unsupervised time series anomaly detection, achieving over 3% ROC and 6% PR improvements.
Dezheng Wang, Tong Chen, Guansong Pang et al.
Inspiration Seeds leverages CLIP sparse autoencoders for unsupervised visual feature decomposition, enabling diverse non-literal image combinations without text prompts.
Kfir Goldberg, Elad Richardson, Yael Vinker
OneLive introduces a dynamically unified generative recommendation framework with real-time content encoding and time-aware attention, significantly improving live-streaming performance.
Shen Wang, Yusheng Huang, Ruochen Yang et al.
MINT uses multi-scale spectral decomposition to disentangle behavior intent from execution details, improving transfer and generalization in imitation learning.
Renming Huang, Chendong Zeng, Wenjing Tang et al.
RLTR introduces cross-model transfer rewards to enhance reasoning robustness and efficiency, achieving +3.6% in Maj@64 on MATH-500 and 2.5× training step reduction.
Hyunseok Lee, Soheil Abbasloo, Jihoon Tack et al.
NarraScore generates soundtracks for long videos using affective control, enhancing narrative consistency.
Yufan Wen, Zhaocheng Liu, YeGuo Hua et al.
LycheeMemory enables long-context reasoning via compressed memory, reducing GPU usage by 2x and speeding inference by 6x.
Zhuoen Chen, Dongfang Li, Meishan Zhang et al.
E-VAds-R1 achieves 109.2% improvement in commercial intent reasoning using MG-GRPO reward design.
Xianjie Liu, Yiman Hu, Liang Wu et al.
Introduced SWE-ContextBench, a benchmark for evaluating context understanding and reuse in coding models, improving accuracy and efficiency significantly.
Jiayuan Zhu, Junde Wu, Minhao Hu et al.
Document reconstruction enables scalable long-context RLVR, improving RULER and LongBench v2 performance.
Yao Xiao, Lei Wang, Yue Deng et al.
SkillRL enhances performance by 15.3% in tasks like ALFWorld via recursive skill-augmented reinforcement learning.
Peng Xia, Jianwen Chen, Hanyang Wang et al.
CrossTraffic constrains LLM execution with a knowledge graph, achieving MAE<0.50 and F1=1.0 for invalid-input detection.
Rei Tamaru, Bin Ran
Combines Prompt engineering, Marked Words, SVM, and JSD to detect biases in LLM-generated recommendations.
Ke Xu, Shera Potka, Alex Thomo