GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CL 2411.09972

Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems

Using large language models as user agents to evaluate task-oriented dialogue systems, enhancing diversity and task completion rates.

Taaha Kazi, Ruiliang Lyu, Sizhe Zhou et al.

2024-11-15 54
cs.CV 2411.08380

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

EgoVid-5M provides 5M curated clips, while EgoDreamer jointly controls egocentric generation with text and kinematics.

Xiaofeng Wang, Kang Zhao, Feng Liu et al.

2024-11-13 33
cs.CV 2411.08033

GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation

GaussianAnything employs point cloud structured latent space with diffusion models for multi-modal 3D generation, enabling high-quality, editable outputs.

Yushi Lan, Shangchen Zhou, Zhaoyang Lyu et al.

2024-11-13 34
cs.CL 2411.07763

Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows

Spider 2.0 framework evaluates language models on enterprise text-to-SQL workflows, solving only 21.3% of tasks.

Fangyu Lei, Jixuan Chen, Yuxiao Ye et al.

2024-11-12 29
cs.CL 2411.07237

Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries

Contextualized evaluations synthesize context to assess language model responses to underspecified queries, significantly altering evaluation conclusions.

Chaitanya Malaviya, Joseph Chee Chang, Dan Roth et al.

2024-11-12 52
cs.CL 2411.07175

Continual Memorization of Factoids in Language Models

Introduces continual memorization framework; REMIX data mixing significantly reduces language model forgetting.

Howard Chen, Jiayi Geng, Adithya Bhaskar et al.

2024-11-12 36
cs.NE 2411.07057

Randomized Forward Mode Gradient for Spiking Neural Networks in Scientific Machine Learning

RFG replaces backpropagation with randomized weight perturbations, reaching 0.0347 Poisson error and an estimated 66% lower training cost.

Ruyin Wan, Qian Zhang, George Em Karniadakis

2024-11-11 33
stat.ML 2411.06140

Deep Nonparametric Conditional Independence Tests for Images

DNCIT combines learned image embeddings with nonparametric CITs and validates MRI–behavior null findings in UK Biobank data.

Marco Simnacher, Xiangnan Xu, Hani Park et al.

2024-11-09 34
cs.CV 2411.05706

Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models

Proposes Image2Text2Image framework using diffusion models for label-free image caption evaluation, achieving high correlation with human judgment.

Jia-Hong Huang, Hongyi Zhu, Yixian Shen et al.

2024-11-09 44
cs.CV 2411.05900

Enhancing Cardiovascular Disease Prediction through Multi-Modal Self-Supervised Learning

Multi-modal self-supervised learning enhances CVD prediction, improving balanced accuracy by 7.6% on limited data using ECG, CMR, and clinical info.

Francesco Girlanda, Olga Demler, Bjoern Menze et al.

2024-11-09 47
cs.CV 2411.05222

Don't Look Twice: Faster Video Transformers with Run-Length Tokenization

Introduces Run-Length Tokenization (RLT), accelerating video Transformer training by 30% with only a 0.1% accuracy drop.

Rohan Choudhury, Guanglei Zhu, Sihan Liu et al.

2024-11-08 0
cs.CL 2411.05000

Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks?

This study evaluates 17 leading LLMs on their ability to follow multiple information threads through near-million-token contexts, revealing performance degradation but strong thread safety.

Jonathan Roberts, Kai Han, Samuel Albanie

2024-11-08 46
cs.CV 2411.04989

SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation

SG-I2V enables zero-shot trajectory control in image-to-video generation using pre-trained diffusion models, achieving high-quality, controllable videos without fine-tuning.

Koichi Namekata, Sherwin Bahmani, Ziyi Wu et al.

2024-11-08 31
cs.RO 2411.04983

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

DINO-WM leverages pre-trained visual patch features for zero-shot planning, outperforming state-of-the-art models with 56% LPIPS improvement and 45% success rate increase.

Gaoyue Zhou, Hengkai Pan, Yann LeCun et al.

2024-11-08 351 citations 45
cs.AI 2411.04872

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

FrontierMath benchmark tests advanced mathematical reasoning; AI solves less than 2%, exposing a large gap with human experts.

Elliot Glazer, Ege Erdil, Tamay Besiroglu et al.

2024-11-08 40
cs.IR 2411.04602

Self-Calibrated Listwise Reranking with Large Language Models

Proposes SCaLR, a self-calibrated listwise reranking framework using explicit relevance scores, improving efficiency and global comparability in large candidate sets.

Ruiyang Ren, Yuhao Wang, Kun Zhou et al.

2024-11-07 38
cs.CL 2411.04368

Measuring short-form factuality in large language models

SimpleQA benchmark evaluates language models on short factual questions, challenging and easy to grade.

Jason Wei, Nguyen Karina, Hyung Won Chung et al.

2024-11-07 21
cs.LG 2411.03538

Long Context RAG Performance of Large Language Models

Long-context RAG enhances LLM performance, but only a few models maintain accuracy above 64k tokens.

Quinn Leng, Jacob Portes, Sam Havens et al.

2024-11-06 23
cs.RO 2411.03287

The Future of Intelligent Healthcare: A Systematic Analysis and Discussion on the Integration and Impact of Robots Using Large Language Models for Healthcare

Robots using large language models can enhance healthcare efficiency, addressing aging and workforce shortages.

Souren Pashangpour, Goldie Nejat

2024-11-06 1
cs.CY 2412.05282

International Scientific Report on the Safety of Advanced AI (Interim Report)

This report analyzes rapid AI capability growth via scaling and interpretability limits, emphasizing resource constraints and safety challenges.

Yoshua Bengio, Sören Mindermann, Daniel Privitera et al.

2024-11-05 31
Prev 1 ... 323 324 325 326 327 328 329 ... 573 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home