Lifestyle

Technology Innovation Institute releases Falcon-Emirati-7B: an Arabic AI model focused on the Emirati dialect that, by its own tests, answers in dialect more often than rivals

Technology Innovation Institute (TII) released Falcon-Emirati-7B on October 6, 2026, an AI language model built to understand and use the spoken dialect of the UAE (United Arab Emirates). TII says other Arabic models often get the content right but reply in written Standard Arabic. This article summarizes the training method, evaluation figures and usage limits TII published; all data are TII's own claims.

About 9 min read

Technology Innovation Institute releases Falcon-Emirati-7B: an Arabic AI model focused on the Emirati dialect that, by its own tests, answers in dialect more often than rivals
Image: Mokaair (Original editorial artwork)

What happened

According to a post published by the Technology Innovation Institute (TII) team on the Hugging Face blog on October 6, 2026, the institute has launched Falcon-Emirati-7B. TII says it is a dialect-specialized model built on its Arabic model Falcon-H1-Arabic, with the goal of understanding and writing the Emirati dialect the way a native speaker would, including its vocabulary, tone and the cultural context behind it.

TII explains its reasons for the release: Modern Standard Arabic (MSA for short) is the written language used in news and textbooks, but everyday conversation, humor, negotiation and storytelling in the UAE mostly take place in the Emirati dialect. TII argues that a model that only knows MSA may translate an Emirati sentence word for word and still miss what it really means. TII also says Falcon-Emirati-7B is already available on its chat platform.

How the model was built

TII says the base model, Falcon-H1-Arabic, uses the Falcon-H1 hybrid architecture, which runs two text-processing techniques in parallel inside each block: state space models (Mamba) and Transformer attention. The family comes in three sizes by parameter count, 3B, 7B and 34B (B stands for billion; the larger the number, the larger the model), with context windows (the amount of text a model can process at once) of up to 128K and 256K tokens (a token is the unit a model splits text into). TII says it chose the 7B version because it strikes the best balance between quality and the cost of training and inference (the model actually running and producing answers): the cost of 34B is not worthwhile for a dialect-specialized chat model, while 3B does not have enough headroom to carry the cultural and linguistic depth required.

TII points to three difficulties in teaching a model a dialect: the Emirati dialect is mostly spoken, so there is far less written text online than for MSA; the meaning is often not the literal one; and there is no established training recipe. According to TII, the training data came from three sources:

  • Content from Emirati websites and forums written natively in the dialect, rather than translated or transliterated from MSA.
  • Material written in MSA about Emirati culture, traditions and language, used to give the model background on the relevant topics.
  • Synthetic data, meaning Emirati dialect text generated by another model; TII says generation was constrained by vocabulary lists, dictionaries and strict style rules, and was used to fill out coverage of everyday topics.

TII says the training process used both human evaluation by native Emirati speakers and automatic scoring, because automatic metrics alone are not enough to measure naturalness, tone and cultural fit.

The evaluation results TII published

The main benchmark TII used (a standard test for comparing the abilities of different models) is Alyah. TII explains that this is a multiple-choice test released by its team and the community specifically to assess Emirati dialect ability. It has 1,173 questions, collected by hand by native Emirati speakers, covering everyday greetings and etiquette, figurative language, traditional knowledge and Emirati poetry. TII reports that Falcon-Emirati-7B scored 84.83% on Alyah, and says it leads all the other Arabic and multilingual models it was compared against, including several that are many times larger; the comparison excluded the Falcon-H1-Arabic family.

Multiple-choice questions only show whether a model can recognize the correct answer, so TII also ran an open-ended generation evaluation: models wrote out their own answers to the same 1,173 questions, and another AI model, Gemini 3.7 Flash, acted as judge. The judge scored two things separately: whether the content was correct, and whether the answer was actually written in Emirati dialect rather than MSA (referred to below as "dialect fidelity"). TII says other models often know the correct answer but still default to answering in MSA even when asked in Emirati dialect. TII also tested four models on 283 scenarios from the UAE portion of the cultural understanding test ArabCulture-Dialogue, a task that requires choosing the most culturally appropriate reply from three options.

Source: TII's own evaluation results published on the Hugging Face blog, not verified by a third party.
ModelDialect fidelity (partial credit, i.e. the judge scores by degree)Accuracy on the UAE portion of ArabCulture-DialogueTII's other notes
Falcon-Emirati-7B0.5285.57%Alyah multiple-choice score of 84.83%
ALLaM-7B-Instruct-preview0.0583.39%In pairwise comparison, tied with Falcon-Emirati-7B at 0.50 in the greetings category
gemma-3-27b-it0.03TII did not include it in this testTook part only in the open-ended generation evaluation
Jais-2-8B-Chat0.0273.79%Narrowly beat Falcon-Emirati-7B in the greetings category, 0.54 to 0.46
Fanar-2-27B-InstructEffectively 0.0071.50%Refused to answer in 26.2% of cases; correctness of 0.27 was the lowest of the five

TII also ran pairwise comparisons: answers from Falcon-Emirati-7B and one other model were given side by side to Gemini 3.7 Flash, which picked the better one without being told which answer came from which model, and win rates were then tallied by category. TII says Falcon-Emirati-7B beat the three rivals in most categories. In the "poetry and creative expression" category, for example, it was 0.69 to 0.31 against Jais-2-8B-Chat, 0.66 to 0.34 against ALLaM-7B-Instruct-preview, and 0.88 to 0.12 against Fanar-2-27B-Instruct. In the "greetings and everyday expressions" category, TII says Falcon-Emirati-7B was 0.46 to 0.54 against Jais-2-8B-Chat (behind), tied at 0.50 with ALLaM-7B-Instruct-preview, but 0.70 to 0.30 against Fanar-2-27B-Instruct. TII explains that this is the category where the Emirati dialect and MSA overlap most, so general Arabic models find it easier to sound natural here.

What it means for general readers

The conclusion TII draws from these results is that model size alone does not deliver dialect ability; dialect ability has to be trained deliberately, with data and evaluation built for that dialect. For ordinary users, this is a reminder that a chatbot "understanding a language" and "being able to respond the way local people actually speak" are two different things. If TII's claims hold, a model may get the content right and still reply in a bookish, formal register rather than the spoken dialect the person asking used.

This release is aimed at the Emirati dialect. The difficulties TII describes (less written data for spoken dialects, meaning that depends on cultural context) come from its own analysis. When assessing any language model's localization ability, readers can pay attention to who designed the evaluation, who did the scoring, and whether native speakers were involved.

Frequently asked questions

What is Falcon-Emirati-7B?

According to TII, it is a dialect-specialized AI language model built on the 7B version of Falcon-H1-Arabic, with the goal of understanding and writing the Emirati dialect, including its vocabulary, tone and cultural context.

How does it differ from general Arabic models?

TII says other models often know the correct answer but still default to answering in Modern Standard Arabic even when asked in Emirati dialect. In TII's evaluation where models wrote out their own answers, Falcon-Emirati-7B scored 0.52 on whether it answered in Emirati dialect, while the other four models tested ranged from about 0.00 to 0.05. This is the result of TII's own evaluation.

Can these evaluation results be trusted?

So far there are only the figures TII itself has published, and the post includes no independent verification. According to TII, the Alyah test was released by its team and the community, and part of the evaluation was judged by another AI model, Gemini 3.7 Flash. TII also mentions that native Emirati speakers carried out human review during training, but acknowledges that judgments about dialect and culture are subjective.

Can the general public use it now?

TII says Falcon-Emirati-7B is already available on its chat platform. TII's post does not say whether the model itself is available for download, nor does it mention license terms.

What are the model's known limitations?

TII notes that the model may reflect biases in its training data and may make mistakes with rare expressions, highly localized references or where data is scarce. TII recommends evaluating it for the specific context before sensitive, official or high-risk use.

Why 7B rather than a larger model?

TII says 7B strikes the best balance between quality and the cost of training and running the model. 34B might improve quality somewhat further, but the cost is not worthwhile for a dialect-specialized chat model; 3B does not have enough headroom to carry the cultural and linguistic depth required.

Browse the latest news in this topic

Sources

Lifestyle