An Evaluation of DUSt3R/MASt3R/VGGT 3D Reconstruction on Photogrammetric Aerial Blocks
DUSt3R, MASt3R, and VGGT achieve up to 50% point cloud completeness improvement on sparse aerial image sets.
Xinyi Wu, Steven Landgraf, Markus Ulrich et al.
DUSt3R, MASt3R, and VGGT achieve up to 50% point cloud completeness improvement on sparse aerial image sets.
Xinyi Wu, Steven Landgraf, Markus Ulrich et al.
PositionIC achieves high-fidelity image customization with spatial precision and identity consistency using BMPDS and visibility-aware attention.
Junjie Hu, Tianyang Han, Kai Ma et al.
MorphoNAS grows complex neural networks through morphogenetic self-organization, achieving low-complexity 6-7 neuron solutions in CartPole tasks.
Mykola Glybovets, Sergii Medvid
QuestA introduces partial solutions during RL training, boosting 1.5B models' math reasoning by over 10% on benchmarks.
Jiazheng Li, Hongzhou Lin, Hong Lu et al.
Leveraging asynchronous cross-border market prices via ensemble models improves day-ahead electricity price forecasts in Belgium by 22% and Sweden by 9%.
Maria Margarida Mascarenhas, Jilles De Blauwe, Mikael Amelin et al.
GraspGen, a diffusion transformer with on-generator discriminator, achieves 94.7% AUC in 6-DOF grasping, outperforming prior methods.
Adithyavairavan Murali, Balakumar Sundaralingam, Yu-Wei Chao et al.
DiffRhythm+ generates high-quality full-length songs with multimodal style control and preference optimization.
Huakang Chen, Yuepeng Jiang, Guobin Ma et al.
PhysX-3D introduces a physics-grounded 3D asset generation framework using PhysXNet and PhysXGen, covering five key physical properties.
Ziang Cao, Zhaoxi Chen, Liang Pan et al.
Infherno employs multi-step LLM agents with external tools to convert unstructured clinical notes into FHIR-compliant structured resources.
Johann Frei, Nils Feldhus, Lisa Raithel et al.
Dark-EvGS combines event camera data with 3D Gaussian Splatting for multi-view bright frame synthesis in low-light scenes.
Jingqian Wu, Peiqi Duan, Zongqiang Wang et al.
SimDiffRec uses semantic similarity-guided diffusion to enhance sequential recommendation performance, outperforming baselines.
Jinkyeong Choi, Yejin Noh, Donghyeon Park
P-COQS employs binary search for DP quantile estimation, ensuring privacy and coverage in conformal prediction.
Ogonnaya M. Romanus, Roberto Molinari
PoliAnalyzer uses NLP for personalized privacy policy analysis, achieving F1 scores of 90-100%.
Rui Zhao, Vladyslav Melnychuk, Jun Zhao et al.
Task-Oriented Human Grasp Synthesis via Context- and Task-Aware Diffusers significantly improves grasp quality.
An-Lun Liu, Yu-Wei Chao, Yi-Ting Chen
RCG framework integrates real-world crash semantics into adversarial scenario generation, boosting safety and realism in autonomous driving testing.
Benjamin Stoler, Juliet Yang, Jonathan Francis et al.
MoR introduces dynamic recursive depths within shared Transformer layers, reducing validation perplexity by up to 15% and improving few-shot accuracy, while enhancing efficiency.
Sangmin Bae, Yujin Kim, Reza Bayat et al.
FOCAL leverages foundation models at inference to canonicalize images, boosting robustness against complex transformations.
Utkarsh Singhal, Ryan Feng, Stella X. Yu et al.
MoVieS employs pixel-aligned Gaussian primitives with explicit motion supervision for 4D scene reconstruction in 1 second from monocular videos.
Chenguo Lin, Yuchen Lin, Panwang Pan et al.
IGD integrates multimodal understanding and parameter prediction with diffusion models, enabling editable multi-scenario graphic design from natural language instructions.
Yadong Qu, Shancheng Fang, Yuxin Wang et al.
This paper introduces a posterior sampling-based expected improvement method, significantly reducing cumulative regret in Bayesian optimization.
Shion Takeno, Yu Inatsu, Masayuki Karasuyama et al.