Enemray: Toward Capable Language Models for Hassaniya

TL;DR

Enemray enhances Hassaniya language capability with stability-plasticity objective, retaining multilingual reasoning and safety.

cs.CL 🔴 Advanced 2026-09-14 5 views
Cheikh Ahmed
language model Hassaniya multilingual stability plasticity

Key Findings

Methodology

Enemray employs layer-selective continual pretraining and supervised post-training, integrating Hassaniya-specific parameter updates into instruction-tuned parameter space. Using the Gemma 4 E4B architecture, the model learns language updates on Hassaniya text and develops conversational, cultural, literary, and cross-lingual behavior through supervised learning.

Key Results

  • Enemray achieves the best English-Hassaniya translation among compared models and the highest score in Mauritanian translation error detection, while retaining capabilities in mathematical reasoning, knowledge, code generation, and function calling.
  • The model excels in natural Hassaniya text and multi-task dialogues, surpassing existing resources.
  • Layer-selective continual pretraining updates language in early and late layers, reducing trainable parameters.

Significance

Enemray significantly elevates the status of Hassaniya in NLP, addressing the gap in low-resource language representation. It enhances Hassaniya capabilities while retaining multilingual reasoning and safety, advancing multilingual models' application in low-resource languages.

Technical Contribution

Enemray introduces a novel parameter-space composition method through layer-selective continual pretraining and supervised post-training, significantly enhancing low-resource language modeling capabilities.

Novelty

Enemray is the first to combine Hassaniya language capability with multilingual reasoning through layer-selective continual pretraining, pioneering a new direction in low-resource language modeling.

Limitations

  • The model's performance in extreme low-resource scenarios needs improvement, especially in complex dialogue contexts.
  • Limited support for dialects beyond Hassaniya, requiring further expansion.

Future Work

Future work includes expanding to other low-resource languages, optimizing performance in complex dialogue scenarios, and exploring more cross-lingual applications.

AI Executive Summary

Enemray is a model focused on the Hassaniya language, aiming to enhance its application in multilingual environments. Existing solutions often face data scarcity and model capability limitations when dealing with low-resource languages. Enemray successfully combines Hassaniya language capability with multilingual reasoning through layer-selective continual pretraining and supervised post-training.

The model employs the Gemma 4 E4B architecture, achieving significant language capability improvements through language updates on natural Hassaniya text. Experimental results show that Enemray excels in English-Hassaniya translation and Mauritanian translation error detection, outperforming existing open and proprietary models.

While Enemray makes significant strides in Hassaniya language modeling, its performance in extreme low-resource scenarios needs improvement. Future research directions include expanding to other low-resource languages, optimizing performance in complex dialogue scenarios, and exploring more cross-lingual applications.

Deep Analysis

Background

Hassaniya is a Western Arabic variety primarily spoken in Mauritania, with a rich oral and literary tradition. However, it is underrepresented in computational resources, with existing resources scattered across sentiment data, parallel corpora, and cultural instructions.

Core Problem

The application of Hassaniya in NLP is limited by data scarcity and model capability constraints. Existing multilingual models often face stability-plasticity issues when handling low-resource languages, struggling to retain multilingual reasoning capabilities.

Innovation

Enemray combines Hassaniya language capability with multilingual reasoning through layer-selective continual pretraining and supervised post-training. This method updates language in early and late layers, reducing trainable parameters, and develops conversational, cultural, literary, and cross-lingual behavior through supervised learning.

Methodology

  • �� Use Gemma 4 E4B architecture for layer-selective continual pretraining.
  • �� Perform language updates on natural Hassaniya text.
  • �� Transfer language updates to instruction-tuned model.
  • �� Develop conversational, cultural, literary, and cross-lingual behavior through supervised learning.

Experiments

Experiments used natural Hassaniya text and multi-task dialogue datasets, evaluating model performance in English-Hassaniya translation and Mauritanian translation error detection. Key hyperparameters included learning rate and batch size.

Results

Enemray excels in English-Hassaniya translation and Mauritanian translation error detection, outperforming existing open and proprietary models while retaining capabilities in mathematical reasoning, knowledge, code generation, and function calling.

Applications

Enemray can enhance Hassaniya language application in multilingual environments, suitable for translation, dialogue systems, and cultural knowledge dissemination.

Limitations & Outlook

The model's performance in extreme low-resource scenarios needs improvement, especially in complex dialogue contexts. Additionally, limited support for dialects beyond Hassaniya requires further expansion.

Plain Language Accessible to non-experts

Imagine you're learning a new language, like Hassaniya. Enemray is like a super language teacher that not only teaches you how to speak the language but also how to engage in complex dialogues and translations. It continuously learns and updates its knowledge to ensure you don't make mistakes when using the language. Like a smart assistant, it helps you use Hassaniya in different scenarios, whether in everyday conversations or professional translations.

ELI14 Explained like you're 14

Hey there! Did you know? Enemray is like a super cool language robot that helps you learn a language called Hassaniya. This language is popular in Mauritania, but there aren't many resources online. Enemray keeps learning and updating itself, getting smarter and smarter. It can help you translate and even teach you how to chat and write in this language. Isn't that awesome?

Glossary

Continual Pretraining

A method of additional training on a specific language to enhance the model's language capability.

Used for language updates on Hassaniya text.

Instruction Tuning

Adjusting model parameters through supervised learning to improve performance on specific tasks.

Used to develop conversational, cultural, and cross-lingual behavior.

Parameter-Space Composition

A method of transferring language updates into instruction-tuned models.

Used to transfer Hassaniya language capability into multilingual models.

Gemma 4 E4B

A multimodal model architecture supporting text, image, and audio inputs.

The foundational architecture for Enemray.

Low-Resource Language

A language underrepresented in computational resources.

Hassaniya is a low-resource language.

Open Questions Unanswered questions from this research

  • 1 How to improve model performance in extreme low-resource scenarios?
  • 2 How to expand model support to other dialects?

Applications

Immediate Applications

Translation Services

Enemray can enhance Hassaniya language translation capabilities, suitable for businesses and organizations requiring multilingual support.

Long-term Vision

Cultural Dissemination

Enhancing Hassaniya language capabilities to promote cultural exchange and knowledge dissemination, advancing multilingual societies.

Abstract

We introduce Enemray, a Hassaniya-centric language model that enables general-purpose interaction in Hassaniya. Enemray is trained around a stability--plasticity objective: acquire strong Hassaniya linguistic and cultural competence while preserving the general reasoning, multilingual, instruction-following, and safety behaviors of a capable instruction-tuned model. The development pipeline separates language acquisition from behavioral specialization. A separately assembled continual-pretraining corpus provides broad exposure to natural Hassaniya and Mauritanian text; layer-selective continual pretraining learns a compact language-specific parameter update; that update is transferred into the instruction-tuned parameter space; and supervised post-training develops conversational, cultural, literary, task-oriented, and cross-lingual behavior. The supervised corpus integrates selected public Hassaniya and Mauritanian resources with a substantially larger body of newly collected, reconstructed, curated, and constructed instruction data, while policy-generated replay provides a retention signal from the reference model's own behavior distribution. The resulting collection is substantially larger and broader in purpose than existing Hassaniya text resources. In evaluation, Enemray achieves the strongest English to Hassaniya translation among the compared open and proprietary models and the highest overall score on Mauritanian translation error detection, while retaining most of the general capabilities of its instruction-tuned base model on mathematical reasoning, knowledge, code generation, and function calling. This report describes the motivation, data construction, model design, training methodology, and evaluation of Enemray.

cs.CL cs.AI