TextMineX: Data, Evaluation Framework and Ontology-guided LLM Pipeline for Humanitarian Mine Action
TextMineX enhances humanitarian mine action text knowledge extraction accuracy by 44.2% using ontology-guided LLM pipeline.
Key Findings
Methodology
TextMineX introduces an ontology-guided large language model pipeline for extracting knowledge from humanitarian mine action (HMA) domain texts. This method structures HMA reports into (subject, relation, object) triples and evaluates extraction accuracy using a bias-aware framework combining human-annotated triples and an LLM-as-Judge protocol. The dataset from the Cambodian Mine Action Centre ensures real-world relevance.
Key Results
- Experiments show ontology-aligned prompts improve extraction accuracy by 44.2%, reduce hallucinations by 22.5%, and enhance format adherence by 20.9%.
- A bias-aware evaluation framework reduces position bias in reference-free scoring.
- Experiments across different LLM models validate the effectiveness of ontology-aligned prompts.
Significance
The study automates the extraction and organization of critical knowledge in the HMA domain, addressing long-standing issues of information transferability between agencies. By improving extraction accuracy and reducing hallucinations, TextMineX offers new technical means for knowledge management in humanitarian mine action.
Technical Contribution
TextMineX significantly improves knowledge extraction accuracy through ontology-guided prompt strategies and is the first to apply LLMs for knowledge extraction in the HMA domain. The method combines layout-aware document chunking and multi-perspective evaluation, offering new possibilities for domain-specific IE applications.
Novelty
TextMineX is the first system to apply ontology-guided LLM pipelines for knowledge extraction in the HMA domain. It significantly improves extraction accuracy and practicality through context-level reasoning and multi-ontology aggregation compared to existing methods.
Limitations
- The method may have limitations in handling complex contexts, especially involving multi-sentence reasoning.
- Further validation is needed for applicability in other domains.
- Dependence on ontology may limit its application in unstructured data.
Future Work
Future work could include extending to other humanitarian domains, further optimizing ontology alignment strategies, and exploring methods to reduce ontology dependence.
AI Executive Summary
Humanitarian Mine Action (HMA) addresses the challenge of detecting and removing landmines from conflict regions. Despite the publication of life-saving operational knowledge by HMA agencies, much of this information is buried in unstructured reports, limiting the transferability of information between agencies. To address this issue, researchers propose TextMineX, the first dataset, evaluation framework, and ontology-guided large language model (LLM) pipeline for knowledge extraction from text in the HMA domain. TextMineX structures HMA reports into (subject, relation, object) triples, creating domain-specific knowledge. To ensure real-world relevance, the dataset from the Cambodian Mine Action Centre was utilized. Additionally, a bias-aware evaluation framework was introduced, combining human-annotated triples and an LLM-as-Judge protocol to mitigate position bias in reference-free scoring. Experimental results show that ontology-aligned prompts improve extraction accuracy by 44.2%, reduce hallucinations by 22.5%, and enhance format adherence by 20.9% compared to baseline models. The research team publicly releases the dataset and code to facilitate further research and application. The introduction of TextMineX is not only a technical advancement but also a humanitarian imperative, making life-saving insights more accessible, actionable, and transferable across many HMA agencies.
Deep Analysis
Background
Humanitarian Mine Action (HMA) aims to detect and remove landmines from conflict regions to return land to civilian use. Despite the publication of life-saving operational knowledge by HMA agencies, much of this information is buried in unstructured reports, limiting the transferability of information between agencies. To address this challenge, researchers propose TextMineX, the first dataset, evaluation framework, and ontology-guided large language model (LLM) pipeline for knowledge extraction from text in the HMA domain.
Core Problem
The knowledge in the HMA domain is largely in unstructured text, making it difficult to share and utilize information across different agencies. Existing LLMs perform poorly in domain-specific information extraction, lacking the ability to handle complex contexts.
Innovation
TextMineX significantly improves knowledge extraction accuracy through ontology-guided prompt strategies. The method combines layout-aware document chunking and multi-perspective evaluation, offering new possibilities for domain-specific IE applications. Through context-level reasoning and multi-ontology aggregation, TextMineX significantly improves extraction accuracy and practicality.
Methodology
- �� Utilized the dataset from the Cambodian Mine Action Centre for real-world application validation.
- �� Introduced a bias-aware evaluation framework combining human-annotated triples and an LLM-as-Judge protocol.
- �� Improved extraction accuracy through ontology-aligned prompt strategies.
- �� Employed layout-aware document chunking and multi-perspective evaluation.
Experiments
Experiments were conducted across different LLM models, including GPT-4o and Llama3-70B. The dataset from the Cambodian Mine Action Centre was used for validation, evaluating extraction accuracy, hallucination rate, and format adherence. Results show significant improvement in extraction accuracy with ontology-aligned prompt strategies.
Results
Experimental results show ontology-aligned prompts improve extraction accuracy by 44.2%, reduce hallucinations by 22.5%, and enhance format adherence by 20.9% compared to baseline models. The bias-aware evaluation framework reduces position bias in reference-free scoring.
Applications
TextMineX can be directly applied to information extraction and knowledge management in the humanitarian mine action domain, helping different agencies share and utilize critical operational knowledge more effectively.
Limitations & Outlook
The method may have limitations in handling complex contexts, especially involving multi-sentence reasoning. Further validation is needed for applicability in other domains. Dependence on ontology may limit its application in unstructured data.
Plain Language Accessible to non-experts
Imagine you're in a huge library with many books about how to clear landmines, but none of the books have an index. TextMineX is like a super librarian that quickly scans these books, identifies the key points of each, and organizes them into an easy-to-understand index. This way, other librarians can quickly find the information they need without flipping through pages. This process is like preparing all ingredients and tools in a kitchen before cooking, so you can make delicious dishes faster.
ELI14 Explained like you're 14
Imagine you're playing a game about clearing landmines. The game has many levels, each with different challenges. TextMineX is like a super helper that quickly finds the key information for each level, like where the mines are and how to clear them. This way, you can pass the levels faster without guessing. It's like having a smart friend in school who helps you find the important parts in your textbook, so you can finish your homework faster!
Glossary
Ontology
An ontology is a formal representation of concepts and their relationships within a domain.
In TextMineX, ontologies guide the knowledge extraction process.
LLM (Large Language Model)
A large language model is an AI model capable of understanding and generating natural language text.
TextMineX uses LLMs for knowledge extraction.
HMA (Humanitarian Mine Action)
HMA refers to the activities of detecting and removing landmines in conflict regions.
TextMineX focuses on extracting knowledge from HMA domain texts.
Triple
A triple is structured data consisting of a subject, relation, and object.
TextMineX structures HMA reports into triples.
Hallucination
Hallucination refers to information generated by a model that is inconsistent with the input.
TextMineX reduces hallucinations through ontology-aligned prompt strategies.
Open Questions Unanswered questions from this research
- 1 How to improve knowledge extraction accuracy without relying on ontologies?
- 2 How to extend TextMineX to other humanitarian domains?
Applications
Immediate Applications
Humanitarian Mine Action
TextMineX can help mine action agencies share and utilize critical operational knowledge more effectively.
Long-term Vision
Cross-domain Knowledge Management
TextMineX's technology can be extended to other fields requiring knowledge extraction and management, such as healthcare and finance.
Abstract
Humanitarian Mine Action (HMA) addresses the challenge of detecting and removing landmines from conflict regions. Much of the life-saving operational knowledge produced by HMA agencies is buried in unstructured reports, limiting the transferability of information between agencies. To address this issue, we propose TextMineX: the first dataset, evaluation framework and ontology-guided large language model (LLM) pipeline for knowledge extraction from text in the HMA domain. TextMineX structures HMA reports into (subject, relation, object)-triples, thus creating domain-specific knowledge. To ensure real-world relevance, we utilized the dataset from our collaborator Cambodian Mine Action Centre (CMAC). We further introduce a bias-aware evaluation framework that combines human-annotated triples with an LLM-as-Judge protocol to mitigate position bias in reference-free scoring. Our experiments show that ontology-aligned prompts improve extraction accuracy by up to 44.2%, reduce hallucinations by 22.5%, and enhance format adherence by 20.9% compared to baseline models. We publicly release the dataset and code.