KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant
KISS Sorcar combines five layered agents and Git isolation, reaching 62.2% on Terminal Bench 2.0.
Koushik Sen
KISS Sorcar combines five layered agents and Git isolation, reaching 62.2% on Terminal Bench 2.0.
Koushik Sen
MuSS creates a large-scale cinematic dataset with cross-shot matching, improving multi-shot narrative coherence and subject consistency.
Haojie Zhang, Di Wu, Bingyan Liu et al.
Kolmogorov-Arnold Networks achieve universality with a single non-affine function.
Vugar Ismailov
DeepSeek-V4 employs hybrid CSA and HCA attention, supports million-token contexts, reducing inference FLOPs to 27%, enabling ultra-long sequence processing.
DeepSeek-AI, Anyi Xu, Bangcai Lin et al.
Learn&Drop accelerates CNN training by layer dropping, reducing ResNet-152 forward propagation FLOPs by 83.74%.
Giorgio Cruciata, Luca Cruciata, Liliana Lo Presti et al.
Proposes node-wise beam search (NBS) with frontier-based expected gain and RRAG graph for efficient active perception path planning, outperforming SOTA by 20%.
Kaixian Qu, Han Wang, Victor Klemm et al.
MoSS framework integrates tactile and torque signals into VLAs, boosting success rates to 49.0% in contact-rich tasks.
Jimin Lee, Huiwon Jang, Myungkyu Koo et al.
Study shows architecture choice crucial for symbolic regression target recovery using EML operator.
Chakshu Gupta
Proposed a Collocation-based Robust Physics-Informed Neural Network (CRVPINN) for simulating pollution propagation under thermal inversion conditions on Spitsbergen.
Leszek Siwik, Maciej Sikora, Natalia Leszczyńska et al.
RINSE evaluates demonstration quality via trajectory smoothness, achieving 16% higher success on RoboMimic with one-sixth data.
Soham Kulkarni, Raayan Dhar, Yuchen Cui
AutoPyVerifier uses DAG search to automatically learn compact Python verifiers, improving target F1 by up to 55 points.
Pouya Pezeshkpour, Estevam Hruschka
Budget-efficient scaling law fitting via active experiment selection achieves full dataset performance using only 10% of the budget.
Sijie Li, Shanda Li, Haowei Lin et al.
RecoverFormer employs causal Transformer with latent mode and contact prediction, achieving 100% success in humanoid recovery, zero-shot transfer, and robustness under unseen disturbances.
Zihui Liu
Study reveals representational harms in LLM narratives against Global Majority nationalities using a QA model on 500,000 stories.
Ilana Nguyen, Harini Suresh, Thema Monroe-White et al.
Proved the undecidability of the plan existence problem even with modal depth at most 1 and no postconditions.
Antonis Achilleos
Using BantuMorph v7, a neural model recovers historical lexical structures in Bantu languages from modern data, confirming 90.9% noun candidates align with Proto-Bantu forms.
Hillary Mutisya, John Mugane
GCImOpt learns efficient goal-conditioned policies by imitating optimal trajectories, significantly improving control task success rates and efficiency.
Jon Goikoetxea, Jesús F. Palacián
Zero-shot morphological discovery in low-resource Bantu languages via cross-lingual transfer and unsupervised clustering.
Hillary Mutisya, John Mugane
Aligning Dense Retrievers with LLM Utility via Distillation, UAE improves Recall@1 by 30.59% on QASPER benchmark.
Rajinder Sandhu, Di Mu, Cheng Chang et al.
Scaling large point-in-time language models (4B params, 1T tokens) narrows performance gap with unrestricted models, ensuring temporal validity and economic relevance.
Bryan Kelly, Semyon Malamud, Johannes Schwab et al.