InFoBench: Evaluating Instruction Following Ability in Large Language Models
Introduces DRFR metric and InFoBench benchmark, significantly improving LLM instruction-following evaluation reliability.
Yiwei Qin, Kaiqiang Song, Yebowen Hu et al.
Introduces DRFR metric and InFoBench benchmark, significantly improving LLM instruction-following evaluation reliability.
Yiwei Qin, Kaiqiang Song, Yebowen Hu et al.
Activation Beacon introduces progressive activation compression, enabling efficient long-text processing with up to 8x compression, doubling inference speed.
Peitian Zhang, Zheng Liu, Shitao Xiao et al.
GRAM achieves 73.68% ANLS on multi-page DocVQA, enhancing long-sequence processing efficiency.
Tsachi Blau, Sharon Fogel, Roi Ronen et al.
DeepSeek LLM scales open-source models; 67B model outperforms LLaMA-2 70B in code, math, reasoning.
DeepSeek-AI, :, Xiao Bi et al.
Proposes a duality hypothesis between LLMs and Tulving's memory theory, exploring consciousness as an emergent ability.
Jitang Li, Jinzheng Li
TinyLlama trains a 1.1B Llama-2-style model on up to 3T tokens, outperforming similarly sized open baselines.
Peiyuan Zhang, Guangtao Zeng, Tianduo Wang et al.
CRVPINN accelerates robust variational PINNs via point collocation and LU decomposition, solving multiple PDEs efficiently.
Marcin Łoś, Tomasz Służalec, Paweł Maczuga et al.
GUESS leverages multi-scale skeleton abstraction and cascaded latent diffusion for text-driven human motion synthesis.
Xuehao Gao, Yang Yang, Zhenyu Xie et al.
CG-STVG improves video grounding accuracy with context guidance, setting new records on HCSTVG and VidSTG.
Xin Gu, Heng Fan, Yan Huang et al.
VideoStudio combines LLM and diffusion models to generate multi-scene videos with content consistency, outperforming SOTA with detailed script and reference image guidance.
Fuchen Long, Zhaofan Qiu, Ting Yao et al.
Extends causal models with impossible worlds to formalize mathematical explanations, addressing the issue that all mathematical facts are true in all models.
Joseph Y. Halpern
Emulates insect path integration on BrainScaleS-2 with spike-based short-term memory, achieving high-precision autonomous navigation.
Korbinian Schreiber, Timo Wunderlich, Philipp Spilger et al.
Darwin3 employs a novel ISA supporting large-scale SNNs with on-chip learning, achieving 2.35 million neurons and 28.3x code density improvement.
De Ma, Xiaofei Jin, Shichun Sun et al.
MosaicBERT integrates FlashAttention, ALiBi, GLU, and other techniques, achieving 79.6 GLUE score in 1.13 hours, vastly improving training efficiency.
Jacob Portes, Alex Trott, Sam Havens et al.
LLoVi combines short-term visual captions with large models for long video QA, achieving 50.3% accuracy.
Ce Zhang, Taixi Lu, Md Mohaiminul Islam et al.
Segment3D leverages SAM-generated pseudo labels and a two-stage training scheme to achieve high-quality, label-free 3D scene segmentation, outperforming supervised baselines.
Rui Huang, Songyou Peng, Ayca Takmaz et al.
Q-Align trains LMMs with discrete text levels, achieving SOTA in image/video quality assessment.
Haoning Wu, Zicheng Zhang, Weixia Zhang et al.
DL3DV-10K is a large-scale real-world scene dataset with 65 POI categories, 5120万 frames, supporting advanced 3D vision tasks.
Lu Ling, Yichen Sheng, Zhi Tu et al.
Comprehensive survey of 62 code LLMs, 15 pre-training objectives, and 112 tasks, highlighting their architectures, applications, and future directions.
Quanjun Zhang, Chunrong Fang, Yang Xie et al.
This study compares seven energy-based learning algorithms on deep convolutional Hopfield networks, revealing negative perturbations outperform positive ones, with Centered EP achieving state-of-the-art results.
Benjamin Scellier, Maxence Ernoult, Jack Kendall et al.