SALSA-CLRS: A Sparse and Scalable Benchmark for Algorithmic Reasoning
SALSA-CLRS extends CLRS benchmark, enhancing scalability and sparsity in algorithmic reasoning.
Julian Minder, Florian Grötschla, Joël Mathys et al.
SALSA-CLRS extends CLRS benchmark, enhancing scalability and sparsity in algorithmic reasoning.
Julian Minder, Florian Grötschla, Joël Mathys et al.
Introducing Quasi-Monte Carlo (QMC) methods to efficiently approximate 3D sliced Wasserstein distances with theoretical guarantees.
Khai Nguyen, Nicola Bariletto, Nhat Ho
MINT evaluates LLMs' multi-turn tool use and feedback leveraging, with performance gains of 2-17% across 20 models, revealing training impacts and evaluation gaps.
Xingyao Wang, Zihan Wang, Jiateng Liu et al.
RotateIt employs multimodal perception with vision and touch, trained in simulation, to enable multi-axis in-hand object rotation without fine-tuning in real-world.
Haozhi Qi, Brent Yi, Sudharshan Suresh et al.
Survey of multimodal foundation models evolving from specialized to general-purpose, covering vision and language capabilities.
Chunyuan Li, Zhe Gan, Zhengyuan Yang et al.
Safety Chip adds LTL safety to LLM robots, reaching 100% safety in simulation.
Ziyi Yang, Shreyas S. Raman, Ankit Shah et al.
SafeShift uses counterfactual probing to identify safety-critical scenarios, reducing collision rates by 10% in autonomous driving trajectory prediction.
Benjamin Stoler, Ingrid Navarro, Meghdeep Jana et al.
Proposes a GNN-based multi-modal biological network framework, enhancing network inference and personalized medicine.
Marinka Zitnik, Michelle M. Li, Aydin Wells et al.
Adding 3% safety examples to LLaMA significantly improves safety without reducing capability.
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio et al.
Framework for LLM-based agents enhances potential for general intelligence.
Zhiheng Xi, Wenxiang Chen, Xin Guo et al.
Proposes convergence analysis for online vector-valued kernel regression, achieving order-optimal error bounds under noise and smoothness assumptions.
Michael Griebel, Peter Oswald
GARA algorithm abstracts goal space via reachability analysis for efficient learning and transfer.
Mehdi Zadem, Sergio Mover, Sao Mai Nguyen
C-Pack introduces C-MTEB, C-MTP, and BGE, significantly advancing Chinese text embedding performance.
Shitao Xiao, Zheng Liu, Peitian Zhang et al.
Analyzing training trajectories reveals that Syntactic Attention Structure (SAS) causes phase transitions, crucial for grammar acquisition in MLMs.
Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho et al.
SafetyBench evaluates LLM safety; GPT-4 excels.
Zhexin Zhang, Leqi Lei, Lindong Wu et al.
Proposes a novel differentiable JPEG method with full modeling of key operations, achieving an average PSNR improvement of 3.47dB over previous approaches.
Christoph Reich, Biplob Debnath, Deep Patel et al.
Proposes CURB-SG, a multi-agent LiDAR-based urban 3D scene graph for enhanced environment understanding in autonomous driving.
Elias Greve, Martin Büchner, Niclas Vödisch et al.
InstaFlow achieves high-quality one-step text-to-image generation using Rectified Flow, with an FID of 22.4.
Xingchao Liu, Xiwen Zhang, Jianzhu Ma et al.
Proposed co-learning of synaptic delays, weights, and neuronal adaptation in SNN, achieving state-of-the-art speech recognition accuracy with fewer parameters.
Lucas Deckers, Laurens Van Damme, Ing Jyh Tsang et al.
Score-based grasping primitive GraspGF combined with residual RL achieves 56.5% success on diverse objects, outperforming baselines with high efficiency.
Tianhao Wu, Mingdong Wu, Jiyao Zhang et al.