ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
ArtiScene leverages image intermediary, enabling training-free, text-driven 3D scene generation with high aesthetic quality.
Zeqi Gu, Yin Cui, Zhaoshuo Li et al.
ArtiScene leverages image intermediary, enabling training-free, text-driven 3D scene generation with high aesthetic quality.
Zeqi Gu, Yin Cui, Zhaoshuo Li et al.
C3PO introduces a central path-inspired modification to PPO, improving constraint satisfaction and reward performance in constrained RL.
Nikola Milosevic, Johannes Müller, Nico Scherf
Introduced a framework for nonlinearly-constrained local Bayesian optimization, reducing function evaluations.
André L. Marchildon, David W. Zingg
AnnaAgent uses multi-agent collaboration and tertiary memory to realistically simulate seeker emotional evolution and multi-session memory in psychological counseling.
Ming Wang, Peidong Wang, Lin Wu et al.
Multimodal GenAI with autoregressive LLMs enhances human motion understanding and generation quality.
Muhammad Islam, Tao Huang, Euijoon Ahn et al.
PointODE leverages Neural ODE for lightweight point cloud feature extraction, with only 0.58M parameters, enabling FPGA acceleration for edge devices.
Keisuke Sugiura, Mizuki Yasuda, Hiroki Matsutani
Introduces TSGD-M, a sampling momentum method that scales textual gradients effectively within limited context windows, improving prompt optimization performance.
Zixin Ding, Junyuan Hong, Zhan Shi et al.
RoboMoRe significantly enhances robot design efficiency via joint optimization of morphology and reward.
Jiawei Fang, Yuxuan Sun, Chengtian Ma et al.
ComposeRAG employs a modular architecture with question decomposition, retrieval decision, and verification, achieving up to 15% accuracy improvement in multi-hop QA.
Ruofan Wu, Youngwon Lee, Fan Shu et al.
PhySense: Principle-based physics reasoning benchmark reveals LLMs' reasoning flaws in physics problems.
Yinggan Xu, Yue Liu, Zhiqiang Gao et al.
TW-GRPO framework enhances video reasoning with focused thinking, achieving 50.4% accuracy on CLEVRER.
Jisheng Dang, Jingze Wu, Teng Wang et al.
Voice conversion improves cross-domain robustness for Arabic dialect ID, achieving 34.1% accuracy increase.
Badr M. Abdullah, Matthew Baas, Bernd Möbius et al.
Proposes RT-X Net, a cross-attention transformer that fuses RGB and thermal images, achieving superior low-light enhancement with PSNR of 27.75dB.
Raman Jha, Adithya Lenka, Mani Ramanagopal et al.
Proposes SCRIPT encoding based on Unicode script/category, with constrained BPE merging, achieving robust multilingual pretokens and reducing partial UTF-8 sequences.
Sander Land, Catherine Arnett
Proposes VG LLM, leveraging a pre-trained 3D geometry encoder to learn spatial priors from videos, significantly improving 3D scene understanding and reasoning.
Duo Zheng, Shijia Huang, Yanyang Li et al.
Proposes an Active Inference-based distributed AI framework for autonomous service management in the Computing Continuum, achieving over 90% SLO fulfillment.
Victor Casamayor Pujol, Boris Sedlak, Tommaso Salvatori et al.
BinConv uses Cumulative Binary Encoding to model ordinal info, significantly improving time series forecasting accuracy.
Andrei Chernov, Vitaliy Pozdnyakov, Ilya Makarov
Proposed a cross-level attribution method revealing a 'mid-activation, late-amplification' pattern, improving layer efficiency by 37% in sparse MoE models.
Junzhuo Li, Bo Wang, Xiuze Zhou et al.
Object-Centric Concept Bottlenecks (OCB) integrates pretrained object detection and concept discovery to enhance performance and interpretability in complex visual tasks, achieving 68.84% accuracy on COCOLogic.
David Steinmann, Wolfgang Stammer, Antonia Wüst et al.
AReaL employs a fully asynchronous RL system, achieving up to 2.77× training speedup while maintaining or improving model performance, by decoupling generation and training.
Wei Fu, Jiaxuan Gao, Xujie Shen et al.