A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding
A-MAR framework enhances multimodal art retrieval explanation quality through structured reasoning plans.
Shuai Wang, Hongyi Zhu, Jia-Hong Huang et al.
A-MAR framework enhances multimodal art retrieval explanation quality through structured reasoning plans.
Shuai Wang, Hongyi Zhu, Jia-Hong Huang et al.
SafetyALFRED evaluates safety planning in multimodal LLMs in kitchen settings, finding good hazard recognition but low risk mitigation success.
Josue Torres-Fonseca, Naihao Deng, Yinpei Dai et al.
Large language models exhibit normative conformity, revealing underlying mechanisms.
Mikako Bito, Keita Nishimoto, Kimitaka Asatani et al.
MathNet provides a global multimodal benchmark for mathematical reasoning and retrieval, covering 30,676 Olympiad-level problems from 47 countries.
Shaden Alshammari, Kevin Wen, Abrar Zainal et al.
BLF system achieves state-of-the-art binary forecasting performance on ForecastBench using sequential Bayesian updating of linguistic beliefs.
Kevin Murphy
ClawEnvKit automates environment generation for claw-like agents, reducing costs by 13,800x.
Xirui Li, Ming Li, Derry Xu et al.
Proposes EvoOR-Agent, a co-evolution framework optimizing reasoning paths and architectures, achieving 15% performance gains on heterogeneous benchmarks.
Jiahao Huang, Peilan Xu, Xiaoya Nan et al.
SkillFlow benchmark demonstrates 8.43% improvement in task success via lifelong skill evolution, using a dual-agent iterative framework.
Ziao Zhang, Kou Shi, Shiting Huang et al.
HalluClear employs classification, three-stage evaluation, and closed-loop reasoning to reduce GUI hallucinations effectively.
Chao Jin, Wenkui Yang, Hao Sun et al.
Proposed DeepInsightTheorem framework enhances informal theorem proving by identifying core techniques, significantly outperforming baselines.
Yunhe Li, Hao Shi, Bowen Deng et al.
Using the CompCQ framework, this study analyzes LLM-generated competency questions across domains, revealing generation characteristics.
Reham Alharbi, Valentina Tamma, Terry R. Payne et al.
The study shows language models exhibit strong spatial transfer in shortest path problems but fail in length scaling due to recursive instability.
Yao Tong, Jiayuan Ye, Anastasia Borovykh et al.
Diagnosing LLM judge reliability using transitivity analysis and conformal prediction sets, revealing 33%-67% documents with at least one 3-cycle.
Manan Gupta, Dhruv Kumar
Study reveals LLMs and VLMs struggle with viewpoint rotation understanding without vision, proposes VRUBench dataset, and improves performance via selective fine-tuning.
Zhen Yang, Ping Jian, Zhongbin Guo et al.
Introduces IRS framework, enhancing multimodal humor understanding with incongruity-resolution supervision, 72B model approaches expert level on NYCC.
Hatice Merve Vural, Doga Kukul, Ege Erdem Ozlu et al.
Policy-Guided Hybrid Simulation (PGHS) achieves 8.80% group simulation error on Meituan, improving over baselines by 45.8% and 40.9%.
Ziyang Chen, Renbing Chen, Daowei Li et al.
OpenMobile synthesizes environment-aware instructions and trajectories with policy switching, achieving 51.7% success on AndroidWorld.
Kanzhi Cheng, Zehao Li, Zheng Ma et al.
AIBuildAI employs hierarchical multi-agent architecture to automate AI model building from task description, achieving 63.1% medal rate, surpassing all baselines.
Ruiyi Zhang, Peijia Qin, Qi Cao et al.
RiskWebWorld offers 1,513 tasks showcasing challenges for GUI agents in e-commerce risk management, with top models achieving 49.1% success.
Renqi Chen, Zeyin Tao, Jianming Guo et al.
Proposes Fun-TSG, a function-driven multivariate time series generator with variable-level anomaly labels, enabling controllable, transparent synthetic data creation.
Pierre Lotte, André Péninou, Olivier Teste