Kling-Omni Technical Report
Kling-Omni is a unified multimodal video synthesis framework integrating instruction understanding, editing, and reasoning, achieving high fidelity and efficiency.
Kling Team, Jialu Chen, Yuanzheng Ci et al.
Kling-Omni is a unified multimodal video synthesis framework integrating instruction understanding, editing, and reasoning, achieving high fidelity and efficiency.
Kling Team, Jialu Chen, Yuanzheng Ci et al.
Proposes TODSynth framework with MM-DiT and CRFM, significantly improving synthetic data quality for remote sensing semantic segmentation.
Yunkai Yang, Yudong Zhang, Kunquan Zhang et al.
N3D-VLM integrates native 3D perception with spatial reasoning, achieving state-of-the-art 3D grounding and inference.
Yuxin Wang, Lei Ke, Boqiang Zhang et al.
VoP method leverages web-scale knowledge, boosting ML models' success rate in city navigation from 20-30% to over 60%.
Dwip Dalal, Utkarsh Mishra, Narendra Ahuja et al.
Proposes 'Off-The-Grid' architecture for sub-pixel Gaussian primitive detection, significantly improving 3D scene reconstruction quality and efficiency.
Arthur Moreau, Richard Shaw, Michal Nazarczuk et al.
Introduces O-Voxel and a 4B-parameter flow model for high-fidelity 3D generation with complex topology and detailed appearance, achieving fast inference and superior quality.
Jianfeng Xiang, Xiaoxue Chen, Sicheng Xu et al.
ART employs a category-agnostic transformer to reconstruct complete 3D articulated objects from sparse multi-state RGB images, predicting geometry, texture, and articulation parameters.
Zizhang Li, Cheng Zhang, Zhengqin Li et al.
LINA predicts prompt-specific interventions to improve causal consistency and out-of-distribution instruction following in diffusion models, outperforming SOTA.
Shu Yu, Chaochao Lu
Introduced Next Scene Prediction task using Qwen-VL and LTX, achieving 73% causal consistency.
Xinjie Li, Zhimin Chen, Rui Zhao et al.
CoRe3D introduces a unified 3D reasoning framework combining semantic Chain-of-Thought and spatial geometric reasoning, significantly improving 3D understanding and generation metrics.
Tianjiao Yu, Xinzhuo Li, Yifan Shen et al.
FysicsWorld offers a full-modality benchmark supporting bidirectional I/O across image, video, audio, and text, covering 16 tasks.
Yue Jiang, Dingkang Yang, Minghao Han et al.
Proposes AdaptiveDetector, combining YOLOv11 and VLM with adaptive thresholds and GRPO, boosting zero-shot polyp recall by 14-22% under challenging conditions.
Shengkai Xu, Hsiang Lun Kao, Tianxiang Xu et al.
WeDetect achieves fast open-vocabulary object detection via retrieval, achieving SOTA across 15 benchmarks with high inference efficiency.
Shenghao Fu, Yukun Su, Fengyun Rao et al.
UFVideo unifies multi-scale video understanding, integrating global, pixel, and temporal info, outperforming GPT-4 with 7.3% improvement across benchmarks.
Hewen Pan, Cong Wei, Dashuang Liang et al.
Proposes Group Diffusion, leveraging cross-sample attention to improve image generation, achieving up to 32.2% FID reduction.
Sicheng Mo, Thao Nguyen, Richard Zhang et al.
Proposed SEPL method enhances noisy domain generalization via feature probing and prediction ensemble.
Wang Lu, Jindong Wang
Tri-Bench tests VLM spatial reasoning under camera tilt and object interference, with ~69% average accuracy.
Amit Bendkhale
InfiniteVL synergizes linear and sparse attention for efficient unlimited-input vision-language models, achieving 1.7x decoding speedup.
Hongyuan Tao, Bencheng Liao, Shaoyu Chen et al.
Fast-ARDiff accelerates AR+diffusion generation using entropy-informed methods, achieving 4.3x speedup on ImageNet.
Zhen Zou, Xiaoxiao Ma, Jie Huang et al.
First comprehensive survey on body and face motion generation, covering datasets, evaluation metrics, and techniques like GANs and DDPMs for multimodal tasks.
Lownish Rai Sookha, Nikhil Pakhale, Mudasir Ganaie et al.