GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
GUI Knowledge Bench reveals VLMs' knowledge gap in GUI tasks, evaluating 292 applications.
Chenrui Shi, Zedong Yu, Zhi Gao et al.
GUI Knowledge Bench reveals VLMs' knowledge gap in GUI tasks, evaluating 292 applications.
Chenrui Shi, Zedong Yu, Zhi Gao et al.
SplitFlow employs flow decomposition and aggregation for inversion-free text-guided image editing, enhancing semantic fidelity and attribute disentanglement.
Sung-Hoon Yoon, Minghan Li, Gaspard Beaudouin et al.
This paper introduces FreeArt3D, a training-free framework leveraging pre-trained static 3D diffusion models (e.g., Trellis) to generate high-fidelity articulated objects by extending SDS into 3D-4D space, optimizing geometry, texture, and articulation parameters from sparse images.
Chuhao Chen, Isabella Liu, Xinyue Wei et al.
Synthetic data reveals that current MIL models underperform the Bayesian optimal in capturing spatial dependencies, even with 10,000 samples.
Ethan Harvey, Dennis Johan Loevlie, Michael C. Hughes
This work proves that a Transformer with nonlinear MLP is asymptotically equivalent to a polynomial predictor, highlighting data quality and mixing effects on ICL.
Samet Demir, Zafer Dogan
Toolathlon benchmarks 32 real-world apps with 604 tools; top model Claude-4.5-Sonnet achieves only 38.6% success in complex long-horizon tasks.
Junlong Li, Wenshuo Zhao, Jian Zhao et al.
ScaleCall shows that ETR minimizes latency, while listwise TRR improves ambiguity resolution; instruction boosts Code N@10 from 14.26% to 32.38%.
Richard Osuagwu, Thomas Cook, Maraim Masoud et al.
MDP-Agent enhances AI search by transforming documents into LLM-ready inputs, achieving 35% accuracy improvement.
Hongjin Qian, Zheng Liu
SeeingEye employs agentic information flow with structured intermediate representations to enable multimodal reasoning in text-only LLMs, outperforming larger end-to-end models.
Weijia Zhang, Zijia Liu, Haoru Li et al.
Study how Chain-of-Thought affects Transformer learning using a three-parameter logistic curve to quantify accuracy.
Zihan Pengmei, Costas Mavromatis, Zhengyuan Shen et al.
Leveraging Gemma 3, MetricX-25 and GemSpanEval enhance translation evaluation and error detection with multi-task learning and generative span output.
Juraj Juraska, Tobias Domhan, Mara Finkelstein et al.
WebLeaper introduces a tree-structured reasoning framework with multi-source data synthesis, significantly improving WebAgent's search efficiency and effectiveness.
Zhengwei Tao, Haiyang Shen, Baixuan Li et al.
R3 framework uses reinforcement learning to optimize retrieval in RAG, improving performance by 5.2%.
Jiawei Zhou, Lei Chen
APTBench evaluates base LLMs' agentic potential during pretraining via real-world trajectories, focusing on planning, action, and atomic skills.
Jiarui Qin, Yunjia Xi, Junjie Huang et al.
MGA optimizes GUI tasks using an observe-first and memory-enhanced approach, improving efficiency.
Weihua Cheng, Junming Liu, Yifei Sun et al.
Pie system enables flexible, efficient LLM serving via APIs and inferlets, boosting throughput by 1.3x-3.4x.
In Gim, Zhiyao Ma, Seung-seob Lee et al.
Constructed TicToc dataset to evaluate LLMs' temporal awareness; models show less than 65% alignment with human preferences under timestamp info.
Yize Cheng, Arshia Soltani Moakhar, Chenrui Fan et al.
Combining SHAP and causal discovery algorithms to enhance fault detection accuracy and interpretability in industrial processes.
Pedro Cortes dos Santos, Matheus Becali Rocha, Renato A Krohling
Proposes a convex optimization framework for identifying switched network systems, jointly recovering node dynamics and graph topology from sampled data without prior labels.
Kaito Iwasaki, Anthony Bloch, Maani Ghaffari
Proposes TIRE framework combining tracking, inpainting, and resplat for personalized 3D/4D generation, significantly improving identity preservation.
Shuhong Zheng, Ashkan Mirzaei, Igor Gilitschenski