CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
CogCoM uses Chain-of-Manipulations reasoning for precise visual reasoning, achieving state-of-the-art across 9 benchmarks with 17B parameters.
Ji Qi, Ming Ding, Weihan Wang et al.
CogCoM uses Chain-of-Manipulations reasoning for precise visual reasoning, achieving state-of-the-art across 9 benchmarks with 17B parameters.
Ji Qi, Ming Ding, Weihan Wang et al.
MobileVLM V2 employs novel architecture and high-quality data, with 1.7B parameters, outperforming comparable 3B models in accuracy and speed.
Xiangxiang Chu, Limeng Qiao, Xinyu Zhang et al.
Using linear probes to distinguish epistemic from aleatoric uncertainty in LLMs, achieving over 0.9 AUC in predictions.
Gustaf Ahdritz, Tian Qin, Nikhil Vyas et al.
Projected Diffusion Model (PDM) integrates iterative projection into reverse diffusion to enforce arbitrary constraints, achieving high-fidelity, constraint-compliant synthesis in complex domains.
Jacob K Christopher, Stephen Baek, Ferdinando Fioretto
DeepSeekMath 7B通过120B数学相关tokens和GRPO算法,显著提升开源模型数学推理能力,接近GPT-4表现。
Zhihong Shao, Peiyi Wang, Qihao Zhu et al.
Decoupled camera and object motion control in text-to-video via spatial and temporal cross-attention, without extra optimization.
Shiyuan Yang, Liang Hou, Haibin Huang et al.
Video-LaVIT achieves unified video-language pre-training with decoupled visual-motional tokenization, excelling in 13 multimodal benchmarks.
Yang Jin, Zhicheng Sun, Kun Xu et al.
Study shows LLMs can perform table-based fact verification with prompt engineering and instruction tuning.
Hanwen Zhang, Qingyi Si, Peng Fu et al.
PoCo leverages diffusion models for probabilistic multi-modal, multi-task policy composition, boosting generalization in robotics.
Lirui Wang, Jialiang Zhao, Yilun Du et al.
DeCoF leverages frame consistency to detect AI-generated videos, achieving 92% accuracy on unseen models and demonstrating strong generalization.
Long Ma, Zhiyuan Yan, Qinglang Guo et al.
BetterV fine-tunes LLMs with discriminative guidance for controlled Verilog generation, outperforming GPT-4 on VerilogEval.
Zehua Pei, Hui-Ling Zhen, Mingxuan Yuan et al.
PiCO uses consistency optimization for unsupervised peer review among LLMs, achieving rankings closer to human preferences.
Kun-Peng Ning, Shuo Yang, Yu-Yang Liu et al.
Using Dutch image descriptions and eye-tracking data, this study quantifies human visuo-linguistic signal variation and evaluates pretrained models' ability to capture it.
Ece Takmaz, Sandro Pezzelle, Raquel Fernández
ReEvo integrates LLMs with reflective evolution to enhance heuristic search for NP-hard problems, achieving state-of-the-art results.
Haoran Ye, Jiarui Wang, Zhiguang Cao et al.
LitLLM toolkit uses Retrieval Augmented Generation to significantly reduce literature review time.
Shubham Agarwal, Gaurav Sahu, Abhay Puri et al.
Proposes compositional generative models, combining smaller models for data-efficient learning and better generalization.
Yilun Du, Leslie Kaelbling
Proposes CodeAct, a framework enabling LLMs to generate executable Python code for actions, significantly improving success rates (up to 20%) in complex multi-tool tasks via integrated Python interpreter and multi-turn interactions.
Xingyao Wang, Yangyi Chen, Lifan Yuan et al.
Proposes a modular tree-based framework for c-approximate window search, achieving up to 75× speedup with Vamana on standard datasets.
Joshua Engels, Benjamin Landrum, Shangdi Yu et al.
Proposes Formal-LLM framework using automaton-guided plan generation, achieving over 50% performance boost and ensuring plan validity.
Zelong Li, Wenyue Hua, Hao Wang et al.
Learn planning-based reasoning via trajectory collection and process reward synthesis, 7B model surpasses GPT-3.5-Turbo.
Fangkai Jiao, Chengwei Qin, Zhengyuan Liu et al.