Simultaneous identification of models and parameters of scientific simulators
SBMI method infers joint probability distributions of model components and parameters using neural networks.
Cornelius Schröder, Jakob H. Macke
SBMI method infers joint probability distributions of model components and parameters using neural networks.
Cornelius Schröder, Jakob H. Macke
Introduces JEEBench, a challenging dataset from Indian IIT JEE exams, revealing GPT-4's performance below 40% on complex problems.
Daman Arora, Himanshu Gaurav Singh, Mausam
Proposes RoHT framework with hierarchical question decomposition tree, achieving 29.7% EM improvement on KQA Pro, surpassing SOTA.
Jiajie Zhang, Shulin Cao, Tingjia Zhang et al.
This study systematically compares distillation objectives (prediction, hidden states, attention) and teacher layer initialization, revealing attention transfer's robustness and low-layer initialization's benefits, especially on GLUE tasks.
Xinpeng Wang, Leonie Weissweiler, Hinrich Schütze et al.
Large language models excel in sentiment analysis but struggle with complex tasks.
Wenxuan Zhang, Yue Deng, Bing Liu et al.
Formalized jailbreak taxonomy; tested 3700 prompts across GPT, BLOOM, OPT, FLAN-T5-XXL; success rates >78%; detection via GPT-4 accuracy 92%.
Abhinav Rao, Sachin Vashistha, Atharva Naik et al.
Introduced MetricEval framework using measurement theory to evaluate NLG metric reliability and validity.
Ziang Xiao, Susu Zhang, Vivian Lai et al.
Proposes DAP, a post-hoc method to correct popularity bias in GCN recommenders, significantly improving tail item recommendations.
Jiajia Chen, Jiancan Wu, Jiawei Chen et al.
Proposes AutoCompressors, an unsupervised method compressing long texts into summary vectors, improving perplexity and in-context learning efficiency.
Alexis Chevalier, Alexander Wettig, Anirudh Ajith et al.
Contrastive decoding (CAD) improves factual consistency across models by adjusting output probabilities based on context, especially effective in knowledge conflict scenarios.
Weijia Shi, Xiaochuang Han, Mike Lewis et al.
BLIP-Diffusion combines pre-trained subject representations with diffusion models for zero-shot and fast fine-tuning controllable text-to-image generation.
Dongxu Li, Junnan Li, Steven C. H. Hoi
Proposed ALCE benchmark for automatic evaluation of LLMs generating citation-supported texts, covering fluency, correctness, and citation quality.
Tianyu Gao, Howard Yen, Jiatong Yu et al.
VIPER leverages pretrained video prediction models as reward signals, enabling complex behavior learning without task-specific rewards.
Alejandro Escontrela, Ademi Adeniji, Wilson Yan et al.
Proposes Diffusion Hyperfeatures to fuse multi-scale, multi-timestep features from diffusion models, improving semantic keypoint correspondence with 85.3% accuracy on SPair-71k.
Grace Luo, Lisa Dunlap, Dong Huk Park et al.
Proposes LLM-based dynamic model selection combining CoT and PAL, achieving 96.8% on GSM8K, surpassing state-of-the-art.
James Xu Zhao, Yuxi Xie, Kenji Kawaguchi et al.
CREATOR enables LLMs to generate tools via documentation and code, disentangling abstract creation from concrete reasoning, improving performance on math and tabular tasks.
Cheng Qian, Chi Han, Yi R. Fung et al.
QLoRA combines 4-bit quantization, LoRA, and paging to finetune 65B models on a single GPU, achieving near ChatGPT performance.
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman et al.
SciMON integrates knowledge retrieval and iterative optimization, significantly enhancing LLMs' ability to generate novel scientific ideas grounded in literature.
Qingyun Wang, Doug Downey, Heng Ji et al.
FActScore decomposes long texts into atomic facts, computes support ratios, and enhances factual evaluation granularity and automation.
Sewon Min, Kalpesh Krishna, Xinxi Lyu et al.
Scaling a 1.5M multi-turn instruction dataset UltraChat and fine-tuning LLaMA-13B yields UltraLLaMA, outperforming Vicuna with a 9.02/10 score.
Ning Ding, Yulin Chen, Bokai Xu et al.