EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

TL;DR

Expert-grounded distillation creates an 8B-parameter vision-language model for scalable road safety audits, outperforming larger teachers and proprietary models with 81% accuracy.

cs.CV 🔴 Advanced 2026-08-25 73 views
Md Thamed Bin Zaman Chowdhury Moazzem Hossain
road safety vision-language models knowledge distillation low-resource open dataset

Key Findings

Methodology

This study introduces an Expert-Grounded Distillation (EGD) framework, calibrating a teacher vision-language model against authoritative field audits with Cohen's κ=0.74. The calibrated teacher generates structured supervision, which is distilled into an 8-billion-parameter student model using Low-Rank Adaptation (LoRA). The process involves expert prompt calibration, multi-tiered supervision (gold, silver, street-view), and parameter-efficient fine-tuning. The BD-ARSA dataset, comprising 21,947 image-audit pairs across Bangladesh, supports this approach. The methodology ensures the model’s outputs align with professional risk assessments, enabling scalable, cost-effective road safety evaluation.

Key Results

  • The fine-tuned model significantly improves ordinal risk assessment, achieving a quadratic weighted kappa (QWK) of 0.48 and 72% exact risk accuracy, compared to baseline of 0.077 and lower scores. In blind human evaluations, the student model reaches 81% correctness, surpassing its 31B teacher (42%) and Gemini-2.5-Flash (58%). The model demonstrates robustness across diverse regions and hazards, validating its generalization and practical deployment potential.

Significance

This work addresses the critical gap in low-resource road safety assessment by translating expert knowledge into an open, scalable AI tool. It reduces reliance on scarce crash data, enabling proactive infrastructure safety management. The approach democratizes access to high-quality safety audits, especially in LMICs, and sets a precedent for integrating institutional expertise into AI models. The open dataset and models foster transparency and collaboration, advancing global efforts to reduce traffic fatalities and injuries.

Technical Contribution

The core innovation is the expert calibration gating mechanism, ensuring supervision quality. The combination of structured prompts, multi-tier supervision, and LoRA-based parameter-efficient fine-tuning results in a compact yet high-performing model. This framework bridges the gap between large proprietary models and practical deployment in resource-constrained settings. It also introduces a new pipeline for formal expert-grounded knowledge distillation in vision-language tasks, with rigorous evaluation protocols including bootstrap confidence intervals and blind assessments.

Novelty

This is the first work to incorporate formal expert-ground-truth calibration into vision-language model distillation for road safety auditing. The expert-grounded prompts and structured supervision ensure high fidelity to professional assessments, overcoming limitations of zero-shot models. The open BD-ARSA dataset provides a unique regional benchmark, enabling reproducibility and further research. The integration of expert calibration, structured outputs, and parameter-efficient fine-tuning sets a new standard for domain-specific VLM applications.

Limitations

  • The model primarily targets rural and suburban roads; urban and high-speed environments need further adaptation. The robustness under adverse weather or night conditions remains untested. Dependence on high-quality expert audits means data quality is critical; biases or errors in ground truth could affect performance.

Future Work

Future efforts will extend the model to urban and highway networks, incorporate multi-modal data (e.g., LiDAR, GPS), and develop real-time assessment capabilities. Efforts to improve robustness under challenging conditions and reduce reliance on expert annotations are ongoing. Additionally, integrating this framework into broader intelligent transportation systems and exploring transferability to other LMICs will be key directions.

AI Executive Summary

Road traffic injuries cause over 1.19 million deaths annually, with low- and middle-income countries bearing the brunt due to poor data infrastructure and limited expert resources. Traditional crash-based safety assessments often fall short in these settings, as crash data is incomplete or unreliable. To address this, researchers have turned to proactive road safety audits (RSA), which evaluate infrastructure hazards directly, bypassing crash data limitations. However, automating RSA remains challenging due to the complexity of hazard detection and the scarcity of annotated datasets.

This study introduces EG-ARSA, an innovative framework that leverages expert-grounded knowledge distillation to create a scalable, open vision-language model tailored for low-resource environments. The core idea is to calibrate a teacher model against authoritative field audits, ensuring its outputs align with professional risk assessments. Once calibrated, the teacher generates structured supervision signals, which are distilled into an 8-billion-parameter student model using Low-Rank Adaptation (LoRA), a parameter-efficient fine-tuning technique. This process guarantees high fidelity to expert judgments while maintaining computational efficiency.

The researchers also developed BD-ARSA, a comprehensive dataset comprising 21,947 image-audit pairs collected across Bangladesh’s rural and suburban roads. This dataset provides high-quality, expert-verified labels, enabling the training and validation of the proposed models. Experimental results demonstrate that the fine-tuned student model achieves a quadratic weighted kappa (QWK) of 0.48 and an exact risk assessment accuracy of 72%, significantly outperforming baseline zero-shot models. Blind human evaluations further confirm that the student model correctly assesses risk 81% of the time, surpassing larger teacher models and proprietary alternatives.

These findings highlight the potential of expert-grounded distillation to democratize road safety assessments, making them affordable and accessible in resource-constrained settings. The approach offers a practical solution for governments and organizations aiming to implement proactive infrastructure safety management at scale. Despite current limitations—such as focus on rural roads and the need for high-quality expert data—the framework paves the way for broader deployment, including urban and highway environments, and integration with real-time monitoring systems. Overall, this work advances AI-driven road safety, with significant implications for reducing traffic fatalities worldwide.

Deep Analysis

Background

交通事故是全球公共卫生的重大挑战,尤其在低中收入国家,事故频发、数据缺失严重。传统的事故统计方法依赖事故记录,但在许多地区,事故数据不完整或不可靠,限制了风险评估和预防措施的有效性。近年来,主动道路安全审查(RSA)逐渐成为替代方案,通过现场专家评估道路潜在危险,提升预警能力。尽管如此,自动化实现仍面临数据不足、模型泛化和解释性差等难题。视觉语言模型(VLM)近年来展现出多模态理解能力,能结合图像和文本进行复杂任务,但在低资源环境和行业应用中仍受限于数据和模型可靠性。本研究旨在结合专家校准和知识蒸馏技术,打造适应低资源地区的高效道路安全评估工具,推动自动化和普惠化。

Core Problem

低资源国家缺乏全面、可靠的事故数据,传统基于事故记录的风险评估难以满足实际需求。现有自动化模型多依赖大量标注数据,迁移性差,难以适应不同地区和道路类型。零-shot模型虽然具备一定泛化能力,但在细粒度风险评估上表现不足,且缺乏专业校准。如何利用有限的专家校准数据,建立既可靠又具有迁移能力的自动评估系统,成为亟待解决的关键问题。这关系到交通安全的提升和公共政策的制定,亟需创新技术突破。

Innovation

本研究的核心创新在于引入专家校准的模型蒸馏(EGD)机制,确保模型输出的专业性和可信度。具体包括:• 通过专家现场审查校准教师模型的生成提示,确保其与权威风险评估高度一致;• 采用结构化输出和多层次验证体系,增强模型的可解释性和稳健性;• 利用LoRA技术实现参数高效微调,降低训练成本;• 构建开源的BD-ARSA数据集,覆盖全国主要区域,提供高质量的标注资源。这些创新使模型在低资源环境中依然具备高性能和迁移能力,突破了传统模型对大量标注数据的依赖。

Methodology

  • �� 数据采集:利用国家项目现场审查数据,校准教师模型,确保其输出与专家评估一致。• 模型训练:教师模型基于校准提示生成结构化风险评估,经过专家验证后,作为蒸馏目标。• 蒸馏过程:采用LoRA技术,将教师模型的生成信号蒸馏到参数为8亿的学生模型中,确保参数微调高效。• 评估机制:通过Bootstrap置信区间、盲评和偏差控制,验证模型性能的统计显著性。• 数据集构建:BD-ARSA包含金、银、街景三个层级,确保多样性和代表性。• 模型部署:在低资源环境中实现高效推理,支持自动化道路安全评估。

Experiments

实验使用BD-ARSA数据集,划分训练、验证、测试集,比较零-shot、微调模型和教师模型的性能。指标包括序数加权Kappa(QWK)和准确率。采用盲评和自动指标验证,确保模型的可靠性。参数调优通过LoRA实现,训练在单GPU环境下完成。对不同风险类别和区域进行性能分析,验证模型的泛化能力和鲁棒性。还进行了消融实验,验证专家校准的重要性和蒸馏策略的有效性。

Results

微调后模型在风险评估中的QWK从0.077提升至0.48,准确率达72%。盲评中,学生模型风险判断正确率达81%,优于教师模型(42%)和Gemini-2.5-Flash(58%)。模型在不同区域和风险类别表现稳定,验证了其广泛适用性。统计检验显示结果具有显著性,模型结构简洁,训练成本低,适合低资源地区部署。模型的开源和数据集的公开,为未来区域性道路安全自动化提供了基础。

Applications

模型可在农村、郊区道路快速部署,用于日常安全评估和风险预警。结合低成本硬件和Web应用,实现普惠式监测。未来可扩展到城市道路和高速公路,结合实时数据,支持智能交通管理和主动干预。为交通部门提供科学依据,优化资源配置,提升整体道路安全水平。

Limitations & Outlook

模型主要针对农村和郊区场景,城市复杂道路尚未覆盖。夜间和极端天气条件下表现待验证。依赖专家校准数据,若数据偏差会影响性能。未来需增强模型鲁棒性,扩展多模态数据融合能力,提升在复杂环境中的适应性。

Plain Language Accessible to non-experts

想象你在一个工厂里工作,工厂的安全取决于每个工人是否知道哪些地方可能出问题。传统的方法是让专家逐个检查每个区域,花费时间又费力。而现在,工厂引入了一台智能机器人,它通过学习专家的判断,能快速扫描整个工厂,指出潜在的危险。这个机器人一开始需要专家的指导,确保它学到正确的安全知识。之后,它可以用很少的指令,自己判断哪些地方可能出问题。这样,工厂的安全管理变得更快、更便宜,也更可靠。这个机器人就像论文中的模型一样,结合专家知识,通过学习变得聪明,帮助我们更好地保护每个人的安全。

ELI14 Explained like you're 14

想象你在学校里,老师会检查每个学生的作业,确保没有错误。可是如果老师太忙,不能检查每个作业怎么办?这时,你可以教一个聪明的机器人老师。刚开始,它需要老师的指导,学习怎么判断作业的好坏。老师会告诉它哪些答案是正确的,哪些有问题。等它学会后,就可以自己检查很多作业了,而且比人还快还准。这个机器人就像论文里的模型一样,先由专家校准,然后用来自动评估道路的安全。它可以在很少的指导下,帮忙保护大家的安全,特别是在那些没有很多专家的地方。这种方法让我们用更少的资源,做出更好的判断,未来还能帮我们在交通、学校、工厂里做很多事情!

Abstract

Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by incomplete crash records, shortages of qualified auditors, and the high cost of large-scale field inspections. To address this problem, we propose Expert-Grounded Distillation (EGD), a novel artificial intelligence framework that transfers institutional road safety expertise into a compact vision-language model for scalable visual road safety auditing. The key innovation is a quantified expert-grounding stage in which the teacher vision-language model is calibrated against authoritative field audits. Large-scale annotation is permitted only after the teacher reaches substantial agreement with expert risk assessments (Cohen's kappa = 0.74). The calibrated teacher then generates structured supervision that is distilled into an 8-billion-parameter student vision-language model using Low-Rank Adaptation and a single leakage-free prompt. We also introduce Bangladesh Road Safety Audit (BD-ARSA), the first open, expert-grounded Bangladeshi visual road safety audit dataset containing 21,947 image-audit records with near-national coverage, and Expert-Grounded Road Safety Auditor (EG-ARSA), the first vision-language model developed specifically for this task. Experimental results show that grounded fine-tuning substantially improves ordinal risk assessment over the zero-shot baseline, while blind expert evaluation demonstrates that the compact student outperforms both its 31 billion-parameter teacher and Gemini-2.5-Flash. These findings demonstrate that EGD provides an effective and scalable engineering solution for proactive road safety auditing in resource-constrained environments.

cs.CV cs.AI