GUICourse: From General Vision Language Models to Versatile GUI Agents
GUICourse transforms VLMs into versatile GUI agents using GUIEnv, GUIAct, and GUIChat datasets, enhancing OCR capabilities.
Wentong Chen, Junbo Cui, Jinyi Hu et al.
GUICourse transforms VLMs into versatile GUI agents using GUIEnv, GUIAct, and GUIChat datasets, enhancing OCR capabilities.
Wentong Chen, Junbo Cui, Jinyi Hu et al.
QTIP employs trellis-coded quantization (TCQ) with hardware-efficient codes to enable ultra-high-dimensional weight quantization, boosting model compression and inference speed.
Albert Tseng, Qingyao Sun, David Hou et al.
This survey links RandNLA’s three sketching paradigms to RMT-based Gaussianization; Gaussian sketches achieve condition number at most 6 for s≥2d.
Michał Dereziński, Michael W. Mahoney
miniCodeProps benchmarks neural theorem provers' ability to automatically verify simple to complex program properties in Lean 4, revealing current limitations.
Evan Lohn, Sean Welleck
Proposes near-field localization with large arrays, leveraging spherical wave models for joint distance and angle estimation, outperforming traditional far-field methods.
Zhaolin Wang, Parisa Ramezani, Yuanwei Liu et al.
Introduced SVPO, a novel algorithm using MCTS for step-level preference optimization in mathematical reasoning, achieving state-of-the-art performance.
Guoxin Chen, Minpeng Liao, Chengxi Li et al.
Identifies leading whitespace in subword vocabularies causes probability inconsistencies; proposes trailing whitespace decoding to fix, affecting cognitive and NLP tasks.
Byung-Doh Oh, William Schuler
Proposes BlockPruner, a training-free block-level pruning method for Transformer MHA and MLP, evaluated via perplexity, outperforming state-of-the-art baselines.
Longguang Zhong, Fanqi Wan, Ruijun Chen et al.
CoMM dataset employs multi-perspective filtering with Llama3 and CLIP, significantly improving content coherence and alignment, boosting multimodal model performance.
Wei Chen, Lin Li, Yongqi Yang et al.
BALANCE filters neighbors by local similarity and beats 8 baselines on 5 datasets under poisoning attacks.
Minghong Fang, Zifan Zhang, Hairi et al.
L4GM is a fast 4D Gaussian-based model that generates animated 3D objects from a single-view video in one second, outperforming traditional methods in speed and quality.
Jiawei Ren, Kevin Xie, Ashkan Mirzaei et al.
DigiRL combines offline and online RL to fine-tune a 1.3B VLM, boosting success rate from 17.7% to 67.2% in Android device control tasks.
Hao Bai, Yifei Zhou, Mert Cemri et al.
Glyph-ByT5-v2 leverages large-scale multilingual datasets and preference learning to achieve high-accuracy multilingual visual text rendering with enhanced aesthetics.
Zeyu Liu, Weicong Liang, Yiming Zhao et al.
Introduces Med-HallMark benchmark and MediHall Score for hierarchical, multi-task hallucination detection in medical LVLMs.
Jiawei Chen, Dingkang Yang, Tong Wu et al.
CarLLaVA uses LLaVA vision encoder and LLaMA backbone, achieving state-of-the-art closed-loop driving with only camera input, outperforming previous methods by 458%.
Katrin Renz, Long Chen, Ana-Maria Marcu et al.
D-NPC introduces dynamic neural point clouds for monocular scene synthesis, achieving real-time high-quality novel views.
Moritz Kappel, Florian Hahlbohm, Timon Scholz et al.
MV-Mol employs multi-modal fusion with Q-Former, using two-stage pretraining to enhance molecular property prediction accuracy by 1.24%.
Yizhen Luo, Kai Yang, Massimo Hong et al.
Introduces ImageNet3D, a large-scale dataset with 200 categories, 86,000+ instances annotated with 6D pose, position, and natural language descriptions, advancing general-purpose 3D understanding.
Wufei Ma, Guanning Zeng, Guofeng Zhang et al.
HOT3D dataset offers multi-view images and annotations for 3D hand and object tracking, boosting research.
Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon et al.
Analyzed 60,000 fine-tuned diffusion models' weight space; introduced weights2weights subspace for sampling, editing, and inversion tasks.
Amil Dravid, Yossi Gandelsman, Kuan-Chieh Wang et al.