Verifiable abstention makes AI leak diagnosis accountable in urban water distribution networks
Verifiable abstention makes AI leak diagnosis accountable, with 214 correct actions out of 550 events.
Tianwei Mu, Yue Wang, Mingzhe Yuan et al.
Verifiable abstention makes AI leak diagnosis accountable, with 214 correct actions out of 550 events.
Tianwei Mu, Yue Wang, Mingzhe Yuan et al.
This study systematically evaluates the fragility of memory-based self-improving agents, revealing high variance and task order sensitivity, and proposes information enrichment to mitigate performance degradation.
Qinyuan Ye, Yu Li, Yada Pruksachatkun et al.
Using latent variable graded response models, this study measures users' willingness to deploy and accept agent-mediated communication in online dating, revealing a significant asymmetry with deployment willingness three times higher.
Daria Leshchikova, Valentina V. Kuskova, Dmitry Zaytsev et al.
HLSR integrates real-time edge speeds with short-term forecasts, using dual-threshold congestion detection and driver personalization, significantly improving congestion mitigation efficiency.
Xiao Wang, Shun Ren Yang, Hui Nien Hung
Baobab compiles SROIQ ontologies into differentiable circuits, enabling robust neuro-symbolic learning with complex logic features.
Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho et al.
SRT, based on SuperMemo-2, dynamically schedules sample review, improving old knowledge retention by 5-37% while maintaining new learning.
Alankar Atreya, Devesh Batra, Yoages Kumar Mantri et al.
This study analyzes how LLM agents design AI methods, revealing heavy reliance on recombining human-designed algorithmic spaces, with limited innovation.
Yikang Yang, Zhengxin Yang, Luzhou Peng et al.
ASI-Bench evaluates AI's autonomous scientific exploration; scores drop from 50.91 to 26.62 as guidance decreases, showing reliance on human input.
Junwei Zhou, Zhen Sun, Binyu Li et al.
DiSCO reduces unsafe content in image generation via adversarial prompt optimization, achieving a 37.7% ASR reduction.
Tong Zhang, Motasem Alfarra, Carlos Hinojosa et al.
Introduces Internal Compliance Score (ICS), a training-free activation readout that reveals rule blindness in compliance detectors, validated across multiple benchmarks and models.
Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu et al.
This paper introduces TTP-D, combining MILP, metaheuristics, and DRL to optimize truck-drone collection, boosting profit and efficiency.
Kabir Murjani, Abhay Sobhanan
ADMITOR enhances optimization modeling precision with label-free certification, achieving candidate precision of 0.927.
Junbo Jacob Lian, Huiling Chen, Hanzhang Qin et al.
Hierarchical byte-level model with multi-byte prediction (MBP) using variable-length windows and LCA masks achieves optimal speed-performance trade-off.
Abraham Toluwase Owodunni, Chibuzor Okocha, Christan Grant et al.
FoT uses early voting and rollout pruning based on hesitation signals to cut 28.8% attention FLOPs while maintaining SC@32 accuracy.
Chanhee Park, Sungbin Han, Jeongho Yoon et al.
S2-MoE employs routing-aware adaptive speculative expansion and expert reuse gating to accelerate MoE inference on edge devices, achieving up to 5.3× speedup.
Haochen Huang, Shengxuan Qiu, Meng Li
Proposes Influence Calibration (ICSD) to address trust-utility mismatch, improving task success rates by 3-5% on benchmarks.
Qizhen Lan, Xi Xiao, Xiangchen Guan et al.
Proposes a task-relevant state transfer framework using predictive sufficiency, with fixed-length bit bounds for session handover.
Masahiro Kato, Taka Kato
Wyvern employs a multi-agent framework combining retrieval, generation, and claim verification, achieving 87% improvement in figure informativeness.
Beatrice Alessandra Motetti, Emilien Guandalino, Daniele Jahier Pagliari et al.
Introduces DHD protocol to measure trajectory value of multi-agent messages, revealing that wrong messages can be helpful; over 40% of such messages aid reasoning in benchmarks.
Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang et al.
ScienceFlow uses recoverable executable states and ESTRA for long-horizon autonomous research, boosting search efficiency.
Mingming Zhao, Jiqian Dong, Kangping Xu et al.