Skip to main content
Back to Blog
AI/MLInnovation
4 October 20264 min readUpdated 7 October 2026

Falcon ASR: Arabic and Multilingual Speech Recognition

Falcon ASR: Arabic and Multilingual Speech Recognition Published October 7, 2026 The Technology Innovation Institute (TII) in Abu Dhabi has introduced Falcon ASR, a 1.6 billion...

By AI Engineering Team

Falcon ASR: Arabic and Multilingual Speech Recognition

Published October 7, 2026

The Technology Innovation Institute (TII) in Abu Dhabi has introduced Falcon-ASR, a 1.6 billion-parameter speech recognition model focused particularly on Arabic and the Emirati dialect. The model also supports English, French, Spanish, and Portuguese.

In TII's evaluation, Falcon-ASR achieved an average word error rate (WER) of 20.92% across six Arabic test sets. The best published result in the leaderboard snapshot used for comparison was 23.17%. On an internal Emirati evaluation, Falcon-ASR recorded the lowest word and character error rates among the systems compared.

The model also provides word-level timestamps, associating each transcribed word with its position in the audio.

Recognizing spoken Arabic

Arabic speech varies according to region, speaker, and recording conditions. A model that performs well on a formal news broadcast may still have difficulty with an Emirati conversation or audio recorded over a telephone line. Dialectal Arabic also has fewer transcribed resources than Modern Standard Arabic (MSA), making training and evaluation more challenging.

Falcon-ASR was trained on Emirati Arabic, MSA, other Gulf and Arabic dialects, and English. The training goal was to transcribe everyday speech, including dialectal forms and switches between languages.

Arabic benchmark results

The Open Universal Arabic ASR Leaderboard, maintained by the ELM Research Center, ranks systems using the equal-weight average WER across six test sets. It also reports character error rate (CER). Lower values indicate better performance. Falcon-ASR was evaluated using this protocol.

ModelParametersAvg WER (%)Avg CER (%)
Falcon-ASR1.6B20.928.79
Audar-ASR-V1-Turbo2.35B23.179.23
Cohere Transcribe Arabic (07-2026)2.0B25.8711.80
omniASR LLM 7B7.0B28.3212.52

WER means word error rate, while CER means character error rate. Lower values indicate better results.

TII evaluated Falcon-ASR on the same six benchmarks using the leaderboard's pinned manifests. The competitor figures were the published leaderboard averages checked on September 30, 2026. Falcon-ASR's average WER was 2.25 percentage points lower than the best published result in that snapshot.

Evaluation on Emirati speech

Public evaluation data already includes Emirati speech. The Casablanca dataset has a UAE subset. TII supplemented that coverage with an internal evaluation of additional Emirati and Gulf speech, using held-out recordings and human-validated transcripts.

Falcon-ASR achieved a WER of 22.73% and a CER of 10.19% in this internal evaluation:

ModelParametersWER (%)CER (%)
Falcon-ASR1.6B22.7310.19
Qwen3-Omni-30B-A3B-Instruct30.0B (3.0B active)26.8012.72
Audar-ASR-V1-Turbo2.35B27.8913.75
Cohere Transcribe Arabic (07-2026)2.0B31.0518.07
Qwen3-ASR-1.7B-hf2.0B31.5213.35
Audar-ASR-V1-Flash0.78B32.8715.36

Falcon-ASR had the lowest WER and CER among the systems compared in this evaluation. Its WER was 4.07 percentage points lower than Qwen3-Omni, which produced the next-best result.

Training for different recording conditions

Training data included background noise, overlapping speech, music, room reverberation, and telephony effects. It also included variations in speaking speed and pitch. The same treatment was applied to Emirati recordings to expose the model to conditions found in meetings, calls, and other everyday recordings.

English and other languages

Falcon-ASR transcribes English using the same model weights. On the seven public English test sets used by the Hugging Face Open ASR Leaderboard, it achieved a mean WER of 5.74%.

Test setWER (%)
LibriSpeech clean1.75
LibriSpeech other4.21
SPGISpeech2.02
VoxPopuli3.87
GigaSpeech8.15
AMI8.33
Earnings-2211.86

The model also supports French, Spanish, and Portuguese. All five languages use the same weights, without requiring a language flag. The output is a transcript in the language spoken.

Model foundation

Falcon-ASR builds on the Falcon3-Audio work. The architecture and training approach for Falcon3-Audio are described in the paper Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data.

Demonstration and planned access

A Hugging Face demonstration space provides an interface for testing Falcon-ASR and examining its transcription output. API access and native applications are planned.