cs.CV 2503.10589

Long Context Tuning for Video Generation

Long Context Tuning (LCT) extends pre-trained video diffusion models' context window, enabling scene-level multi-shot generation with high consistency.

Yuwei Guo, Ceyuan Yang, Ziyan Yang et al.

2025-03-14 28
cs.SE 2503.09089

LocAgent: Graph-Guided LLM Agents for Code Localization

LocAgent employs graph-based representation and multi-hop reasoning, achieving 92.7% accuracy in code localization with 86% cost reduction using fine-tuned Qwen-2.5-Coder-Instruct-32B.

Zhaoling Chen, Xiangru Tang, Gangda Deng et al.

2025-03-12 106 citations 88
cs.CV 2503.08507

Referring to Any Person

This paper introduces RexSeek, a model combining multimodal large language models and object detection, achieving superior multi-person referring understanding on the HumanRef dataset.

Qing Jiang, Lin Wu, Zhaoyang Zeng et al.

2025-03-11 24 citations 28