cs.LG 2412.11006

Entropy-Regularized Process Reward Model

Entropy-regularized process reward model (ER-PRM) with KL regularization improves mathematical reasoning by guiding multi-step inference.

Hanning Zhang, Pengcheng Wang, Shizhe Diao et al.

2024-12-15 28 citations 31
eess.IV 2412.09405

Learned Compression for Compressed Learning

WaLLoC combines wavelet transforms and autoencoders for efficient compression, outperforming state-of-the-art models in rate-distortion and downstream tasks.

Dan Jacobellis, Neeraja J. Yadwadkar

2024-12-13 43
cs.CL 2412.08905

Phi-4 Technical Report

phi-4 is a 14B parameter model leveraging synthetic data, surpassing GPT-4 in STEM QA, with innovative training strategies.

Marah Abdin, Jyoti Aneja, Harkirat Behl et al.

2024-12-12 34
cs.CV 2412.07825

3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark

Introduces 3DSRBench, a benchmark with 2772 annotated QA pairs, to evaluate large multimodal models' 3D spatial reasoning, revealing current limitations especially in uncommon viewpoints.

Wufei Ma, Haoyu Chen, Guofeng Zhang et al.

2024-12-11 32