MidiTok Visualizer: a tool for visualization and analysis of tokenized MIDI symbolic music
MidiTok Visualizer simplifies symbolic music research by visualizing tokenized MIDI data.
Michał Wiszenko, Kacper Stefański, Piotr Malesa et al.
MidiTok Visualizer simplifies symbolic music research by visualizing tokenized MIDI data.
Michał Wiszenko, Kacper Stefański, Piotr Malesa et al.
MusicFlow uses cascaded flow matching for text-guided music generation, achieving 2-5x smaller model size and 5x fewer iterations.
K R Prajwal, Bowen Shi, Matthew Lee et al.
SWE-Search combines MCTS with self-improvement, achieving 23% performance gains in software tasks.
Antonis Antoniades, Albert Örwall, Kexun Zhang et al.
OGBench provides a comprehensive benchmark with 8 environments and 85 datasets to evaluate offline goal-conditioned RL algorithms across multiple capabilities.
Seohong Park, Kevin Frans, Benjamin Eysenbach et al.
MAPO integrates momentum into natural language gradient descent, reducing convergence time by 77.9% and boosting peak F1 score by 5.28%.
Anthony Cui, Pranav Nandyalam, Andrew Rufail et al.
LRDS leverages prior mode locations to improve multi-modal distribution sampling efficiency.
Maxence Noble, Louis Grenioux, Marylou Gabrié et al.
Using sparse autoencoders (SAEs) for interpretable knowledge unlearning in language models; limited effectiveness with current techniques.
Eoin Farrell, Yeu-Tong Lau, Arthur Conmy
MoGe predicts 3D point clouds from single images using affine-invariant representations, enhanced by global and multi-scale local supervision, achieving state-of-the-art accuracy.
Ruicheng Wang, Sicheng Xu, Cassie Dai et al.
OSCAR uses state-aware reasoning and re-planning to control OS, enhancing user productivity.
Xiaoqiang Wang, Bang Liu
Experts in MoE models excel in memorization tasks, outperforming dense models in these scenarios.
Samy Jelassi, Clara Mohri, David Brandfonbrener et al.
MazeNet is a deep learning method for solving OARSMT with 100% accuracy.
Gabriel Díaz Ramos, Toros Arikan, Richard G. Baraniuk
Introduced Concept-Guided Conditional Diffusion and Prototype Networks to enhance model interpretability.
Alba Carballo-Castro, Sonia Laguna, Moritz Vandenhirtz et al.
This study reveals that robot imitation learning generalization follows a power-law with environment and object diversity, with 4 hours of data collection achieving ~90% success in new settings.
Fanqi Lin, Yingdong Hu, Pingyue Sheng et al.
AgentStore improves OSWorld benchmark performance to 23.85% using MetaAgent and AgentToken strategy.
Chengyou Jia, Minnan Luo, Zhuohang Dang et al.
Proposed Structure Language Model (SLM) combines discrete VAE and conditional language modeling, achieving 20-100x faster protein conformation sampling.
Jiarui Lu, Xiaoyin Chen, Stephen Zhewen Lu et al.
AVHBench is a benchmark for evaluating cross-modal hallucinations in audio-visual LLMs, revealing their limited understanding of complex relationships.
Kim Sung-Bin, Oh Hyun-Bin, JungMok Lee et al.
LAMPS predicts API call strategies to optimize scheduling, reducing latency by 27%-85% and TTFT by 4%-96%.
Rana Shahout, Cong Liang, Shiji Xin et al.
TranSPORTmer employs a unified transformer framework with Set Attention Blocks to perform trajectory forecasting, imputation, inference, and state classification, outperforming SOTA on sports datasets.
Guillem Capellera, Luis Ferraz, Antonio Rubio et al.
Introduces a multi-token prediction model using tensor decomposition to enhance sampling efficiency while maintaining accuracy.
Artem Basharin, Andrei Chertkov, Ivan Oseledets
Proposes similarity-adjusted surprisal, leveraging Ricotta and Szeidl’s diversity index, to enhance reading time prediction beyond standard surprisal.
Clara Meister, Mario Giulianelli, Tiago Pimentel