FireRed-Image-Edit-1.0 Technical Report
FireRed-Image-Edit employs a diffusion transformer with 1.6B samples, achieving state-of-the-art instruction-based image editing.
Super Intelligence Team, Changhao Qiao, Chao Hui et al.
FireRed-Image-Edit employs a diffusion transformer with 1.6B samples, achieving state-of-the-art instruction-based image editing.
Super Intelligence Team, Changhao Qiao, Chao Hui et al.
This paper revisits single-minus gluon tree amplitudes, deriving a piecewise-constant closed-form expression valid in half-collinear and complex momentum configurations, satisfying key physical constraints.
Alfredo Guevara, Alexandru Lupsasca, David Skinner et al.
DreamID-Omni framework achieves controllable human-centric audio-video generation, surpassing existing commercial models.
Xu Guo, Fulong Ye, Qichao Sun et al.
SafeNeuron identifies safety neurons via ES and SAS metrics, freezes them, and fine-tunes remaining parameters, boosting robustness against neuron pruning with minimal utility loss.
Zhaoxin Wang, Jiaming Liang, Fengbin Zhu et al.
FAIL employs adversarial imitation learning with pathwise and policy gradient algorithms, improving image generation with only 13,000 samples, surpassing traditional preference methods.
Yeyao Ma, Chen Li, Xiaosong Zhang et al.
Proposes G-OPD framework using reward extrapolation to surpass teacher performance, validated on math reasoning and code tasks.
Wenkai Yang, Weijie Liu, Ruobing Xie et al.
AssetFormer uses autoregressive Transformer to generate modular 3D assets, enhancing UGC content creation quality.
Lingting Zhu, Shengju Qian, Haidi Fan et al.
VLAW improves vision-language-action models via iterative online interaction, achieving a 39.2% success rate increase.
Yanjiang Guo, Tony Lee, Lucy Xiaoyang Shi et al.
Proposes Composition-RL, automatically combining prompts to improve reasoning, achieving up to 21.4% pass@1 gains on 30B models.
Xin Xu, Clive Bai, Kai Yang et al.
Introduces Min-kNN distance for black-box detection of RLVR training data via structural convergence, achieving 70% AUC.
Hongbo Zhang, Yue Yang, Jianhao Yan et al.
ABot-N0 employs a hierarchical ‘Brain-Action’ architecture, unifying 5 navigation tasks with 16.9M trajectories, achieving SOTA performance.
Zedong Chu, Shichao Xie, Xiaolong Wu et al.
LASER combines SeqVault with STA/GSTA for low-latency ultra-long CTR modeling, delivering +2.36% ADVV and +2.08% revenue online.
Tianhe Lin, Ziwei Xiong, Baoyuan Ou et al.
NRT model enhances reasoning by self-generating reasoning paths without external verifiers, significantly improving complex reasoning tasks.
Yuanfu Wang, Zhixuan Liu, Xiangtian Li et al.
Krause Attention reduces attention sinks by promoting local synchronization, improving performance and reducing complexity.
Jingkun Liu, Yisong Yue, Max Welling et al.
AuroraRL leverages sparse parameter updates with delta checkpoints, reducing bandwidth by 79× and boosting throughput 1.3-9.5× in decentralized RL training.
Chaoyi Ruan, Geng Luo, Xinyi Wan et al.
Proposed retrieval-aware distillation retains only 2% of attention heads, recovering 95% performance.
Aviv Bick, Eric P. Xing, Albert Gu
Integrating lightweight learned depth priors into VINS-Mono reduces ATE by up to 28.3%, improving robustness in low-texture environments.
Arda Alniak, Sinan Kalkan, Mustafa Mert Ankarali et al.
Proposes Wasserstein-based design features to predict query quality and guide adaptive acquisition in Bayesian Optimization.
Antonio Candelieri, Francesco Archetti
pplx-embed combines diffusion pretraining with contrastive curricula, reaching 69.66 MTEB and 81.96 ConTEB nDCG@10.
Sedigheh Eslami, Maksim Gaiduk, Markus Krimmel et al.
Introduces GameDevBench, a benchmark for AI in complex game development, with a success rate of 53.8%.
Wayne Chi, Yixiong Fang, Arnav Yayavaram et al.