DynSUP: Dynamic Gaussian Splatting from An Unposed Image Pair
DynSUP uses two unposed images to achieve dynamic Gaussian splatting, significantly enhancing dynamic scene synthesis.
Weihang Li, Weirong Chen, Shenhan Qian et al.
DynSUP uses two unposed images to achieve dynamic Gaussian splatting, significantly enhancing dynamic scene synthesis.
Weihang Li, Weirong Chen, Shenhan Qian et al.
HyperForget uses hypernetworks for machine unlearning, maintaining high accuracy on retain sets.
Jose Miguel Lara Rangel, Stefan Schoepf, Jack Foster et al.
Proposes Video-3D LLM, integrating 3D position encoding into video representations, achieving SOTA on 5 3D scene benchmarks with 58.1% [email protected] on ScanRefer.
Duo Zheng, Shijia Huang, Liwei Wang
SSACL+ICLCR integrates data-lake tables; F1 +4.2% and accuracy +18.9%.
Daomin Ji, Hui Luo, Zhifeng Bao et al.
TQA-Bench evaluates LLMs in multi-table QA with context lengths from 8K to 64K tokens.
Zipeng Qiu, Chenyue Li, You Peng et al.
AMO sampler enhances text rendering in diffusion models by adaptive overshooting, improving accuracy by 35.9% without extra training.
Xixi Hu, Keyang Xu, Bo Liu et al.
Trajectory attention explicitly models pixel trajectories, boosting fine-grained camera motion control with improved long-range consistency.
Zeqi Xiao, Wenqi Ouyang, Yifan Zhou et al.
SoLM generates structured objects conforming to complex schemas via self-supervised denoising, cost-efficiently.
Amir Tavanaei, Kee Kiat Koo, Hayreddin Ceker et al.
FonTS employs a two-stage DiT pipeline with parameter-efficient fine-tuning and style adapters to achieve precise word-level typography and style control, significantly improving text rendering quality.
Wenda Shi, Yiren Song, Dengming Zhang et al.
AC3D leverages spectral analysis and model optimization to enable precise 3D camera control in video diffusion transformers, improving visual quality by 10% and training speed by 15%.
Sherwin Bahmani, Ivan Skorokhodov, Guocheng Qian et al.
Introduces VDMini, accelerating video diffusion models by pruning and consistency loss, achieving 2.5x speedup.
Yiming Wu, Zhenghao Chen, Huan Wang et al.
G3Flow integrates foundation models with 3D generative and pose tracking to enhance robotic manipulation success rates by over 20% in complex tasks.
Tianxing Chen, Yao Mu, Zhixuan Liang et al.
ChatRex enhances multimodal LLM perception with decoupled design, achieving 72.8% recall on COCO dataset.
Qing Jiang, Gen Luo, Yuqin Yang et al.
Proposes Safety-as-Policy combining virtual scenario generation and reflection to enhance robot safety, achieving 52.63% safety rate in synthetic data and 75% in real-world tests.
Minheng Ni, Lei Zhang, Zihan Chen et al.
LLM-powered GUI agents automate tasks via natural language and visual processing, enhancing human-computer interaction efficiency.
Chaoyun Zhang, Shilin He, Jiaxu Qian et al.
Proposed nonlinear Ramp ADC in memristor arrays enables efficient RNN activation approximation, boosting energy and area efficiency.
Junyi Yang, Ruibin Mao, Mingrui Jiang et al.
Type-R employs post-processing to automatically correct typos in text-to-image outputs, significantly improving text accuracy without degrading image quality.
Wataru Shimoda, Naoto Inoue, Daichi Haraguchi et al.
CHOICE benchmark systematically evaluates 23 remote sensing tasks for large vision-language models, revealing strengths and gaps in perception and reasoning, with 10,507 problems across 50 cities.
Xiao An, Jiaxing Sun, Zihan Gui et al.
MiniKV achieves 86% KV cache compression with 98.5% accuracy via 2-bit layer-discriminative KV cache.
Akshat Sharma, Hangliang Ding, Jianping Li et al.
Proposed a multimodal VQA framework integrating external knowledge and reasoning, achieving 78.5% accuracy on VQA 2.0.
Jiayi Kuang, Jingyou Xie, Haohao Luo et al.