BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams
BLUEX v2 evaluates 21 LLMs on Brazilian second-phase exams, using multimodal, open-ended questions, with performance gap of 4.92 points.
Key Findings
Methodology
The study employs a rubric-based evaluation protocol where LLMs are assessed as judges against structured criteria derived from official answers. The dataset includes 395 questions with 919 subquestions, processed through stages of data extraction, OCR, image captioning with Gemini 3.1, subject classification via Sabiá-4, and rubric generation. Models are queried via API, with prompts including question text, context-aware image captions, and previous answers for sequential subquestions. The judge (Sabiá-4) outputs binary judgments on criteria, aggregated into scores. Human validation shows substantial agreement (κ ≈ 0.69), confirming reliability.
Key Results
- The top model, Gemini 3.1 Pro, scores 9.10/10, while the weakest, LLaMA-3.2-11B Vision, scores 4.18/10, with a performance gap of 4.92 points, demonstrating strong discriminative power.
- Mathematical Reasoning (avg. 7.52) and Image Understanding (avg. 7.79) are the most challenging capabilities, with image questions scoring 0.54 points lower than text-only questions on average.
- The LLM judge achieves 89.5% agreement with human raters, validating the automated protocol's effectiveness for academic content evaluation.
Significance
This work addresses the critical gap in evaluating Portuguese-language LLMs on complex, multimodal, open-ended tasks, providing a standardized benchmark across multiple subjects and years. It advances the understanding of models' capabilities in deep reasoning and multimodal comprehension, essential for real-world applications like education, automated grading, and AI-assisted learning. The methodology offers a scalable, objective alternative to manual grading, fostering broader adoption in low-resource language contexts.
Technical Contribution
The paper introduces a novel pipeline combining automated data extraction, OCR, context-aware image captioning, and structured rubric generation, enabling scalable, fine-grained evaluation of LLMs. The use of LLMs as judges, validated against human raters, is a significant innovation, ensuring reliability and consistency. The structured rubric approach allows detailed diagnostics of model strengths and weaknesses across multiple disciplines and modalities, setting a new standard for multilingual, multimodal AI evaluation.
Novelty
This is the first benchmark to incorporate second-phase Brazilian university entrance exam questions with multimodal content into an open-ended, rubric-based evaluation framework. It uniquely combines multi-year, multi-subject data with automated scoring validated against human judgments, pushing beyond prior multiple-choice or recognition-based assessments. The integration of multimodal content and structured rubric generation represents a major step forward in low-resource language AI evaluation.
Limitations
- The evaluation relies on captioned descriptions rather than raw images, which may limit visual reasoning fidelity. Future work should incorporate direct image processing.
- Scope is limited to specific subjects and question types; broader coverage is needed for generalization.
- API costs and rate limits restrict large-scale deployment, impacting scalability. Further optimization is required for real-time applications.
Future Work
Future directions include integrating multimodal models capable of processing raw images, expanding subject and question diversity, and refining rubric generation for finer-grained scoring. Enhancing the interpretability of model responses and extending validation across more languages and domains will further solidify this benchmark as a standard for multilingual, multimodal AI evaluation.
AI Executive Summary
BLUEX v2 represents a significant advancement in evaluating large language models within the context of Brazilian second-phase university entrance exams. By leveraging a multi-year, multi-subject dataset that includes complex, open-ended questions with multimodal content, the benchmark provides a comprehensive platform for assessing deep reasoning, language understanding, and multimodal comprehension. The evaluation pipeline integrates automated data extraction, OCR, context-aware image captioning with Gemini 3.1, and subject classification via Sabiá-4, culminating in a structured rubric-based scoring system. This system, validated against human judgments, ensures reliable and objective assessment of model performance.
The experimental results reveal a performance gap of 4.92 points among 21 state-of-the-art models, with the best reaching 9.10 and the weakest 4.18. Notably, mathematical reasoning and image understanding emerge as the most challenging capabilities, highlighting areas for future improvement. The study confirms that multimodal content introduces an average penalty of 0.54 points, emphasizing the importance of visual reasoning in real academic settings.
This work has broad implications for AI research and education, offering a scalable, objective, and multilingual benchmark that addresses the limitations of traditional multiple-choice assessments. It paves the way for more nuanced, multi-disciplinary evaluation frameworks, fostering the development of models capable of complex reasoning across languages and modalities. Despite current limitations such as reliance on captioned descriptions and API costs, the platform sets a new standard for low-resource language AI evaluation, with promising avenues for future enhancement and application.
Deep Analysis
Background
近年来,深度学习模型如GPT、BERT和T5极大推动了自然语言处理的发展,但大多集中在英语环境。葡萄牙语作为第五大母语,缺乏系统的深度评测工具。早期的BLUEX基于多项选择题,解决了数据匮乏问题,但未能涵盖深度推理和生成能力。随着模型能力提升,评估需求也从识别题转向开放式、多模态、多学科的复杂任务,推动了多模态理解和深度推理的研究。当前,缺乏针对葡语环境的多模态、深度评测平台,限制了模型在实际应用中的能力验证。
Core Problem
现有评测多偏重识别和选择,难以反映模型在复杂推理和生成能力上的真实水平。葡语环境中缺少高质量、多模态、开放式的评估数据,限制了模型能力的全面衡量。尤其在教育场景中,深度理解和多模态内容的评估需求不断增长,亟需建立标准化、多维度的评测体系,推动模型在实际应用中的性能提升。
Innovation
本研究提出了BLUEX v2,结合多模态内容和结构化评分,首次引入Brazilian第二阶段高考题,涵盖9个学科、4个年份,建立了多学科、多模态、多模型的评估平台。创新点包括:• 利用OCR和Gemini 3.1生成上下文感知的图像字幕,确保多模态内容的自动处理;• 通过Sabiá-4自动生成结构化评分标准,实现细粒度、客观的评估;• 跨学科、多模型、多年份的评估框架,验证了LLM作为判评者的可行性。这些创新极大丰富了低资源语言环境中的AI评测手段。
Methodology
- �� 数据采集:从官方PDF提取题目、答案,结合OCR识别图像内容。• 数据清洗:多轮人工校验,去除重复和无关题目。• 图像描述:用Gemini 3.1生成上下文感知字幕,描述图像与题目的关系。• 学科分类:利用Sabiá-4模型自动标注题目所属学科。• 评分标准:由Sabiá-4基于官方答案自动生成结构化评分准则,拆解答案为多个可检测的二元项。• 评估流程:模型通过API接口回答题目,结合题干、字幕和前序答案,逐步生成回答。• 判评机制:Sabiá-4依据评分标准逐项打分,输出二值判断,最后合成得分。• 结果分析:模型得分转换为0-10分制,统计整体表现,验证模型在不同学科和内容复杂度下的能力。
Experiments
采用2022-2025年Brazilian第二阶段高考题,评估21个模型,包括前沿API模型和开源模型。指标为0-10分,验证模型在多学科、多模态内容中的表现差异。通过人机一致性验证,确保自动评分的可靠性。分析不同能力维度(如数学推理、图像理解)对模型性能的影响,进行子类别分析,验证模型在实际学术场景中的适用性。还通过对比不同题型和难度的子集,评估模型的泛化能力。
Results
模型表现差异显著,最高Gemini 3.1 Pro得分9.10,最低LLaMA-3.2-11B Vision得分4.18,差距4.92分。数学推理和图像理解为最难能力,平均得分分别为7.52和7.79,图像题平均比纯文本低0.54分。判评器与人类评审的Kappa值达0.69,验证了自动评估的可靠性。模型在STEM科目表现较弱,验证了符号推理和视觉理解的挑战。多模态内容带来的平均性能惩罚为0.54分,强调多模态理解的重要性。
Applications
该平台可用于教育评估、模型能力诊断和多模态AI研发。帮助高校和企业快速评估模型在复杂学科和多模态内容中的表现,推动AI在自动评分、智能辅导和个性化学习中的应用。未来可结合实际教学场景,开发智能批改和个性化辅导系统,提升教学效率和公平性。
Limitations & Outlook
目前评估依赖字幕描述,未直接处理原始图像,可能影响视觉推理的真实性。题目范围有限,未来需扩展到更多学科和题型。API调用成本高,限制大规模评测的可扩展性。模型判评的准确性虽高,但仍存在偏差,需持续优化判评机制。
Plain Language Accessible to non-experts
想象你在一家大工厂工作,工厂里有许多不同的机器,每台机器负责不同的任务。有些机器只会识别简单的指令,有些机器还能自己思考和做出复杂决定。科学家们希望让这些机器像人一样理解复杂的问题,比如解答难题或理解图片内容。于是他们设计了一个特别的考试,就像学校的考试一样,里面有很多不同类型的问题,包括文字和图片。然后,他们用另一台“聪明的机器”来检查这些机器的答案,判断它们是不是答得对。这个过程就像老师批改作业一样。通过不断让机器参加考试、被检查,科学家们可以知道这些机器到底有多聪明,未来还能用在学校、医院、工厂,让我们的生活变得更方便。这就像让机器人变得更聪明,帮我们解决各种难题一样。
ELI14 Explained like you're 14
想象你在参加一个超级难的考试,里面不仅有写作文的问题,还有很多图片,比如地图、图表和化学公式。你的任务是用自己的话回答问题,还要理解图片里的内容。科学家们也在做类似的事情,他们用一种超级聪明的电脑模型来帮忙回答这些复杂的问题。这个模型就像一个非常厉害的朋友,能理解文字和图片,还能给出详细的答案。为了让这个“朋友”变得更聪明,科学家们让它参加了很多考试,看看它答得怎么样。他们还让另一台模型像老师一样,检查它的答案是否正确。结果显示,这些模型在一些科目,比如数学和图片理解方面还不够强,但在其他方面表现不错。未来,这样的技术可以用在自动批改作业、智能辅导甚至帮助老师设计更好的考试,让学习变得更有趣、更高效。是不是很酷?
Abstract
Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended, discursive tasks that demand deeper reasoning and generation capabilities. While the original BLUEX benchmark addressed the scarcity of Portuguese evaluation datasets through multiple-choice questions from Brazilian university entrance exams, it did not cover the more challenging second-phase examinations, which require free-form written responses. In this work, we introduce BLUEX v2, a benchmark derived from the second-phase entrance exams of Brazil's two leading universities: UNICAMP (Comvest) and USP (Fuvest), spanning exam years 2022--2025. Our dataset comprises 395 questions unfolding into 919 graded subquestions, with 55.7% of questions containing associated images (represented as context-aware captions during inference to enable evaluation across both vision-capable and text-only models). Each question is annotated with subject area, official reference answers, LLM-generated rubric criteria, and six cognitive capability tags. We evaluate 21 state-of-the-art LLMs using an LLM-as-a-judge protocol. Results reveal a 4.92-point performance spread across models (4.18-9.10 on a 0-10 scale), with Mathematical Reasoning and Image Understanding emerging as the hardest capability dimensions. The evaluation code, model outputs, and dataset are publicly available at https://github.com/TropicAI-Research/BLUEXv2 and on Hugging Face at https://huggingface.co/datasets/Tropic-AI/BLUEX-v2.