OoO-Spec: Out-of-Order Semantic Speculation for Fast Tool Calling
OoO-Spec employs asynchronous semantic speculation, achieving up to 5.34× speedup in tool calls with a single trained sidecar model.
Zhiheng Zhang, Mujie Xu, Feiyu Sun et al.
OoO-Spec employs asynchronous semantic speculation, achieving up to 5.34× speedup in tool calls with a single trained sidecar model.
Zhiheng Zhang, Mujie Xu, Feiyu Sun et al.
AdaMTP uses entropy-based boundary detection and dynamic masking to improve multi-token prediction accuracy and speed.
Ziqiang Cui, Han Shi, Bowei He et al.
Introduced COMPINT suite to address session constraint loss in context compaction, achieving 90% retention.
Zhiqi Wang, Yichi Zhang, Dongwon Lee et al.
Data Turnstile framework generates high-quality function-calling data, boosting Qwen3-0.6B accuracy to 75.9%.
Goutham Ramakrishnan, Megha Sharma
DuPLeR employs dual-path structural reasoning combined with multimodal LLM priors to improve few-shot knowledge graph completion, achieving significant performance gains.
Jinlan Liu, Zhiying Tu, Yongchao Xing et al.
Metis integrates native memory into foundation models, using a parametric memory state and self-supervised training to enhance long-term reasoning.
Zeyu Zhang, Ziliang Guo, Yihang Sun et al.
This paper systematically analyzes lossy verification in speculative decoding, classifying into truncation and collaborative methods, revealing their mechanisms, pitfalls, and control principles.
Tianyu Wang, Yuxuan Zhou, Wenbin Wang et al.
Proposes a human-in-the-loop workflow using SciSummNet and GPT-4o-mini for scientific summary simplification, with expert editing to preserve terminology and claims.
Kyuri Im, Michael Färber
Proposes an instruction-free alignment approach for audio-language models, training only a lightweight projector, achieving performance comparable to heavily fine-tuned models.
Xuanru Zhou, Yiwen Shao, Jiahong Li et al.
GEMCo uses human-crafted proxy conversations validated against real data, enabling ethical sharing of sensitive counseling dialogues with minimal distribution gap.
Philipp Steigerwald, Eric Rudolph, Mara Stieler et al.
Study enhances legal QA performance via context-injected fine-tuning, showing significant gains for Qwen3.5 at 0.8B parameters.
Moniruzzaman Mahadi, Abrar Mohammed Tanzim Alam, Sayma Siddika Monalisa et al.
Study finds Chinese debates show higher semantic repetition than English, introducing 'Prior-Argument Similarity' metric.
Huiqian Lai
IKS-Instruct creates a 24,795-pair multilingual dataset covering 41 Indian Knowledge System techniques, significantly enhancing model cultural depth and pedagogical ability.
Shwetha Singaravelu, Gayathri Muruganantham, Lakshmi Rajendran et al.
Proposes BHARATI, a morphology-aware tokenizer with subword fertility analysis, reducing sequence length by 90% on classical Indian language corpora.
Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran et al.
Proposes Co-E system with bidirectional graph-text memory for training-free multi-hop QA, outperforming baselines.
Hieu Man, Thien Huu Nguyen
This paper derives power-law scaling laws for from-scratch vision-language pretraining, revealing distinct behaviors for language and multimodal objectives under fixed compute.
Haoyuan Wu, Aoqi Wu, Hai Wang et al.
Unified moral-value dataset for instruction tuning enhances AI's ethical alignment, improving value-behavior consistency.
Zhaohui Zeng, Florian Mai
REFACT enhances long-text reasoning by adaptive fact restatement, reducing tokens and improving evidence density.
Zhensheng Jin, Xin Dai, Zhenghao Liu et al.
Proposes two reliability-aware distillation methods, CHAD and EWAD+CPDP, outperforming standard KD with +0.0219 ROUGE-L on low-resource summarization.
Dipto Sumit, Ankan Kumar Roy Srizon, Sadia Khair Rodela et al.
Proposes Sentence Splitter, a T5-based model that automatically identifies factual sentence boundaries, improving knowledge graph completion and QA by 4-6%.
Ahmad Pouramini, Mahsa Afsharizadeh