Falcon-Emirati-7B: An Arabic LLM Specialized for Emirati Dialect and Culture
Falcon Emirati 7B: An Arabic LLM Specialized for Emirati Dialect and Culture Arabic includes multiple varieties that differ in vocabulary, grammar, rhythm, and cultural context....
By AI Engineering Team
Falcon-Emirati-7B: An Arabic LLM Specialized for Emirati Dialect and Culture
Arabic includes multiple varieties that differ in vocabulary, grammar, rhythm, and cultural context. Modern Standard Arabic (MSA) is common in news and formal writing, while everyday conversations in the UAE often take place in Emirati Arabic, a Gulf dialect with its own expressions and conventions.
Emirati poetry, particularly nabati poetry, as well as proverbs and short anecdotes, often communicates meanings that cannot be captured through literal translation. A model trained primarily on MSA may understand the individual words in an Emirati sentence while missing its intended meaning.
Falcon-Emirati-7B was developed to address this gap. Built on Falcon-H1-Arabic, it is designed to understand and generate Emirati Arabic while accounting for vocabulary, tone, grammar, and cultural context.
Built on Falcon-H1-Arabic
Falcon-Emirati-7B was not trained from scratch. It builds on Falcon-H1-Arabic, an Arabic model family based on the Falcon-H1 hybrid architecture. Each block runs State Space Models, specifically Mamba, and Transformer attention in parallel, then fuses their outputs before the block's projection.
This design combines Mamba's linear-time efficiency for long sequences with attention's ability to model long-range dependencies. These capabilities are relevant to Arabic, a morphologically rich language. The Falcon-H1-Arabic family includes 3B, 7B, and 34B parameter variants, with context windows of up to 128K and 256K tokens. Its training data included MSA, Gulf, Levantine, Egyptian, and Maghrebi Arabic, together with English and multilingual data.
That foundation provided broad Arabic understanding, long-context capabilities, and existing exposure to dialectal Arabic. Falcon-Emirati-7B further adapts the model to Emirati vocabulary, grammar, and cultural knowledge.
The 7B variant was selected as a balance between capability and cost. The 34B model might provide additional quality, but its training and serving requirements are higher. The 3B model offers less capacity for the linguistic and cultural detail required by dialect specialization.
Why Dialect Adaptation Is Difficult
Adapting a general Arabic model to Emirati Arabic presents several challenges:
- Emirati Arabic is primarily spoken. It appears less frequently in online writing than MSA and some other Gulf and Levantine dialects, limiting the amount of available text.
- Meaning is often non-literal. Idioms, proverbs, and poetic references depend on shared cultural knowledge rather than surface vocabulary alone.
- There is no established adaptation recipe. The appropriate amount of dialectal data, the balance between dialect and MSA, and the relative value of continued pretraining, supervised fine-tuning, and preference optimization are not standardized.
The development process therefore involved testing different data mixtures, training stages, and supervision strategies. Automatic scores were combined with native-speaker assessments to identify changes that improved the model in practice.
Data Sources
The development team created an Emirati-focused data pipeline based on three complementary sources.
1. Authentic Emirati-Dialect Web Data
The dataset included curated content from Emirati websites and forums written natively in the dialect rather than translated or transliterated from MSA. This material provided examples of everyday phrasing, colloquial expressions, and the natural interaction between Emirati Arabic and MSA in online communication.
2. MSA Material About Emirati Culture and Identity
MSA-language articles and references about Emirati culture, heritage, language, history, customs, values, and social norms were added alongside dialectal text. This material was intended to provide cultural knowledge, even when the model was not using dialectal language in its response.
3. Synthetic Data Guided by Glossaries and Style Rules
Authentic dialectal text did not cover all the topics needed by a conversational model. Synthetic Emirati data was therefore generated to fill gaps. The generation process used glossaries, dictionaries, and rules focused on Emirati vocabulary and grammar rather than allowing a model to produce unrestricted Gulf-style Arabic.
These constraints were intended to reduce output that was grammatically valid but sounded unnatural to Emirati speakers.
Finding an Adaptation Strategy
Because no standard MSA-to-dialect training recipe exists, the development process included ablation studies examining:
- How much dialectal data to add
- Which training stage should receive the data
- How to balance authentic and synthetic examples
- How to avoid overfitting to synthetic patterns
- How much MSA cultural material was needed for cultural grounding
Automatic evaluation was paired with review by native speakers. This was important because metrics alone do not fully measure naturalness, tone, or cultural appropriateness.
Evaluation Methodology
Progress was measured through manual review and automatic evaluation.
Native-Speaker Review
Native Emirati speakers assessed generated answers for correctness, naturalness, tone, and cultural appropriateness. These properties are difficult to capture with a single benchmark score.
Alyah Benchmark
Quantitative evaluation used Alyah, a benchmark for Emirati-dialect capabilities in Arabic language models. Alyah contains 1,173 manually collected, native-authored multiple-choice samples covering topics such as:
- Everyday greetings and etiquette
- Figurative language
- Heritage knowledge
- Emirati poetry
- Language and dialect
The benchmark focuses on areas where dialect and cultural knowledge are especially important.
Results on Alyah
Falcon-Emirati-7B achieved 84.83% on Alyah. In the reported comparison, it scored higher than the other Arabic and multilingual models evaluated, including models with substantially more parameters.
The comparison excludes Falcon-H1-Arabic family models because Falcon-Emirati-7B is built on that family.
The results indicate that model size alone does not guarantee dialect competence. Arabic-focused models generally performed better than broadly multilingual models, but broad Arabic coverage was not sufficient by itself. Targeted training was still needed for poetry, heritage knowledge, and language-and-dialect questions.
The Alyah results also show that strong general-purpose models can lose capability on culturally embedded dialectal content. Increasing model size does not automatically resolve that gap without dialect-specific data and evaluation.
Open-Ended Generation and Dialect Fidelity
Multiple-choice accuracy measures whether a model can select the correct answer from several choices. It does not show whether the model will produce Emirati Arabic when responding freely.
To measure this capability, open-ended generation was evaluated on the same 1,173 Alyah questions. Gemini 3.7 Flash served as the judge, comparing:
- Falcon-Emirati-7B
- ALLaM-7B-Instruct-preview
- gemma-3-27b-it
- Jais-2-8B-Chat
- Fanar-2-27B-Instruct
Each answer was assessed separately for factual correctness and whether it used Emirati Arabic rather than MSA. The evaluation reported both partial-credit and pass/fail results, along with abstention rates.
Falcon-Emirati-7B led on correctness. Its larger distinction appeared in dialect fidelity, where it scored 0.52 partial credit, compared with 0.05 for ALLaM, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat, and effectively 0.00 for Fanar-2-27B-Instruct.
The competing models often identified the correct answer but returned it in MSA, even when prompted in Emirati Arabic. Falcon-Emirati-7B was the only model in the comparison that consistently responded in the requested dialect.
Fanar-2-27B-Instruct also had the highest abstention rate, declining to answer 26.2% of the time. The other models abstained in fewer than 5% of cases. Fanar-2-27B-Instruct had a partial-credit correctness score of 0.27, the lowest among the five models.
The dialect-fidelity pattern remained consistent across Alyah categories, from greetings and daily expressions to poetry. Competing models performed relatively better in greetings, where Emirati Arabic and MSA overlap more substantially.
Pairwise Category Comparisons
A separate pairwise evaluation compared Falcon-Emirati-7B with Jais-2-8B-Chat, ALLaM-7B-Instruct-preview, and Fanar-2-27B-Instruct. Gemini 3.7 Flash saw two answers for each Alyah question without knowing which model produced them, then selected the better answer.
Falcon-Emirati-7B won most categories against all three competitors, with the largest differences appearing in categories requiring dialectal or cultural fluency.
Against Jais-2-8B-Chat, the largest differences were:
- Poetry & Creative Expression: 0.69 vs. 0.31
- Language & Dialect: 0.62 vs. 0.38
Against ALLaM-7B-Instruct-preview:
- Poetry & Creative Expression: 0.66 vs. 0.34
- Language & Dialect: 0.58 vs. 0.42
Against Fanar-2-27B-Instruct, Falcon-Emirati-7B won every category. The widest margins were:
- Poetry & Creative Expression: 0.88 vs. 0.12
- Religious & Social Sensitivity: 0.80 vs. 0.20
Greetings & Daily Expressions was the closest category. Falcon-Emirati-7B lost narrowly to Jais-2-8B-Chat, 0.46 vs. 0.54, and tied ALLaM-7B-Instruct-preview at 0.50. It beat Fanar-2-27B-Instruct in the category, 0.70 vs. 0.30.
The results suggest that generic Arabic models can sound relatively natural in greetings because Emirati Arabic and MSA share more vocabulary and phrasing in that area. Falcon-Emirati-7B's advantage was more pronounced in poetry, figurative language, and heritage-related content.
Emirati Cultural Understanding
The model was also evaluated on the UAE portion of ArabCulture-Dialogue, a benchmark for cultural understanding in Arabic. In its multiple-choice task, models select the most culturally appropriate reply from three options.
The evaluation used 283 UAE scenarios in both Emirati Arabic and MSA, with different amounts of location information. The reported overall accuracy was:
- Falcon-Emirati-7B: 85.57%
- ALLaM-7B: 83.39%
- Jais-2-8B: 73.79%
- Fanar-2-27B: 71.50%
Falcon-Emirati-7B achieved the highest score among the four evaluated models on this task.
Limitations and Responsible Use
Falcon-Emirati-7B, like other language models, can reflect biases in its training data and produce incorrect answers. Errors may be more likely with rare expressions, highly localized references, or cases with limited training examples.
Dialectal and cultural judgments can also be subjective, and native speakers may disagree about the most appropriate response. The model should therefore be evaluated for the intended use case before being used for sensitive, official, or high-stakes applications.
Citation
@misc{falcon_emirati_7b_2026,
title = {Falcon-Emirati-7B: A Dialect-Specialized Arabic LLM for the Emirati Dialect},
author = {Shaikha Alsuwaidi, Omar Alkaabi, Maitha Alhammadi, Hamza Alobeidli, Ahmed Alzubaidi, Mohammed Alyafeai, Leen AlQadi, Basma Boussaha, Hakim Hacid},
organization = {Technology Innovation Institute},
year = {2026}
}