Flexible Motion Generation from Language and Style References
FlexMoGen framework generates high-quality human motions from language and style references.
Kai Weixian Lan, Bodie Criswell, Briana Fedkiw et al.
FlexMoGen framework generates high-quality human motions from language and style references.
Kai Weixian Lan, Bodie Criswell, Briana Fedkiw et al.
LLM digital twins reduce human measurement via statistical substitutability, but behavioral fidelity is insufficient.
Steven Wang, Kyle Hunt, Shaojie Tang et al.
Kalman Delta Networks improve linear attention by modeling uncertainty, enhancing perplexity and accuracy in 750M and 1.3B parameter models.
Ngoc Bui, Tinglin Huang, Rex Ying
Online Draft Co-Training with Zigzag Ring Attention and TapChannel achieves up to 1.88× end-to-end RL speedup.
Zili Wang, Zhaopeng Qiu, Yuekai Zhang et al.
CantoneseLLM v2 combines CPT, DPO, and RLVR, reaching 73.16 on HKCanto-Eval with Cantonese Traditional-Chinese reasoning.
Tsz Chung Cheng, Chung Shing Cheng, Chaak Ming Lau et al.
Proposes DH-BPE combining minimum-token segmentation exposure with hierarchical BPE for fixed vocab optimization.
Kenny Shao
ContextFlow uses conditional flow matching for fine-tuning-free robot adaptation, reaching 73.5% average LIBERO success and surpassing ICRT by 35 points.
Jian Ding, Xianjie Dai, Roei Herzig et al.
MyBuddy significantly improves performance on the BuddyVQA benchmark through multimodal chain-of-thought reasoning.
Hangyu Qin, Junbin Xiao, Shenglang Zhang et al.
GAN-Blot: Generates 46K synthetic WB images, enhancing control over protein-band structure and visual style.
Hao-Chiang Shao, Fong-Yi Lin, Te-An Chien et al.
Introduced a new theoretical framework for Masked Pretraining (MPT) to address dimensional collapse.
Qi Zhang, Runyu Zhou, Yifei Wang et al.
A training-free large-scale scene mesh generation framework combining local continuity and global appearance consistency, improving geometric and visual fidelity.
SangEun Lee, Wonseok Chae, Hoyoung Yoo et al.
Proposes MovieGrid, a spatial grid-based post-training method, increasing multi-shot video generation by 6.05× in a 1616-frame video.
Jiawei Mao, Haoqin Tu, Hardy Chen et al.
A scale–accuracy response framework finds that stronger ImageNet-1K classifiers tolerate smaller inputs: Pearson r=-0.890.
Anish Monsley Kirupakaran
EvoSafeHarness optimizes model- and domain-specific safety mechanisms, reducing attack success rate from 45.6% to 10.0%.
Nanxi Li, Yingzi Ma, Yulong Cao et al.
Proposed a method for emotion recognition using joint probability matrix learning to improve accuracy.
Tingyi Lin, Wen-Ren Yang, Kuanwei Chen
UniMate uses a topology-aware diffusion transformer to animate diverse skeletons, outperforming existing methods.
Linzhan Mou, Jiahui Lei, Zhiyang Dou et al.
WearableQA benchmark evaluates AI's health reasoning on real wearable data, with performance ranging from 19.6% to 72.9%.
Ji Soo Lee, Xilun Chen, Pierce Chuang et al.
Diffusion TV offers a tangible experience of diffusion models using a modified CRT TV, letting audiences simulate the denoising process.
Sihwa Park
RegionFed achieves personalized query understanding with gradient conflict analysis, reaching 92.27% accuracy.
Quoc H. Nguyen, Ali Lafzi, Abhijeet Phatak et al.
ROBORMBENCH reveals reward instability in vision-language models under semantically equivalent instructions, featuring 2,390 trajectories and 21,673 paraphrases.
Wonje Jeung, Sangyeon Yoon, Hyesoo Hong et al.