Article

Falcon-OCR-Arabic Targets Receipts, Forms and Tables

TII’s Arabic OCR specialist warrants a document-specific evaluation. Read the reported scores, access gaps and checks for receipts, forms and tables.

Editorial illustration for Falcon-OCR-Arabic Targets Receipts, Forms and Tables: a document represents the research briefing. Not documentary evidence.

On October 6, 2026, the Technology Innovation Institute (TII) announced Falcon-OCR-Arabic, an Arabic adaptation of its 270M-parameter OCR model. For teams digitizing receipts, forms and tables, it offers a new candidate for extracting text from document images. The immediate decision is whether to evaluate it on your document mix; production access needs a separate check.

TII says it adapted the existing architecture using supervised finetuning on real and synthetic Arabic documents, followed by reinforcement learning. Its reported results support a business-document shortlist, not a universal replacement for general vision-language models.

Read the scores as transcription evidence

TII evaluated 17 models on its own 11,974-sample Arabic dataset spanning 15 categories, using the OmniDocBench evaluation framework. The figures below are TII-reported measurements, not independently reproduced results or scores on the standard OmniDocBench dataset.

Model

Arabic text score

Table TEDS

Falcon-OCR-Arabic

81.87%

59.95%

Gemini 3.5 Flash

84.34%

51.30%

Unadapted Falcon OCR

55.39%

24.83%

The text score is based on normalized edit distance per page, including diacritics. Table TEDS measures similarity in structure and content; a missing table scores zero. These metric definitions matter: 81.87% does not mean that proportion of invoices is error-free, and 59.95% is not the share of tables extracted perfectly.

For a business workflow, an altered invoice total or identifier may matter more than many harmless spacing differences. Treat transcription fidelity, exact critical-field correctness and table-cell alignment as separate acceptance criteria.

Choose by document type, not the headline average

TII reports category leads on official documents, administrative forms, receipts and invoices. Dense layouts tell a different story: newspaper text scores are 55.4% for Falcon-OCR-Arabic versus 72.2% for Gemini 3.5 Flash. The category breakdown therefore supports keeping a general-model baseline in any mixed-document evaluation.

Calculated workload check: books account for 4,296 ÷ 11,974 = 35.88% of the benchmark; receipts and invoices together account for (246 + 122) ÷ 11,974 = 3.07%. These shares use TII’s published sample counts, rounded to two decimals. Assuming the stated per-page averaging uses those counts, the overall mean gives book pages far more influence than receipts and invoices.

An accounts-payable team should weight its local result by its actual incoming documents while still reporting each category separately. Small reported category leads are a reason to test, not a guarantee: the receipt and invoice subsets contain only 246 and 122 samples, respectively, and the announcement supplies no uncertainty intervals.

Arabic access is not the base-model download

The announcement points to the Falcon Arabic OCR playground. Its landing page responded during the October 6 checks, but inference, login requirements and document limits were not tested.

The public Falcon-OCR model card identifies the base model, lists Apache-2.0 and documents local loading and Docker/vLLM serving. Neither that card nor the announcement establishes that these weights are the Arabic adaptation. The checked TII public model inventory did not identify a separately named Arabic OCR checkpoint.

Arabic-specific downloadable weights, licensing terms, a documented production API and pricing remain unverified in the inspected sources. That does not rule out private access. Before integration, obtain the exact supported model or service, permitted use, data-retention terms, input/output limits, rate limits and price. Do not infer these from the base model or build against an undocumented playground backend.

A comparison that can settle the decision

The following is a proposed evaluation, not a test performed for this article. Confirm permitted access first and begin with authorized, non-sensitive documents.

  1. Build a held-out Arabic document set. Separate receipts, invoices, forms, books and dense pages. Include the scan quality you actually receive, handwriting, mixed Arabic/Latin lines, diacritics and both Western and Eastern Arabic numerals. Have competent Arabic reviewers verify reference transcriptions and fields.

  2. Compare the same images. Run the current OCR pipeline, an approved general vision-language model and the specialist on identical held-out inputs. Record model versions, prompts, rendering resolution, preprocessing and output limits. Tune on a separate development set so the final comparison stays untouched.

  3. Score what downstream systems need. Check exact amounts, dates and identifiers; missing or invented text; right-to-left reading order; and whether each table value remains attached to the correct row and column. State Unicode and diacritic normalization rules, and retain raw-text fidelity scores so normalization cannot hide consequential errors.

  4. Set acceptance criteria before reviewing outputs. Choose thresholds from the consequences of errors, not the release’s average score. Count failed and escalated documents. Retain source-page references so a reviewer can check uncertain fields before they update business records.

  5. Measure the whole workflow. Include rendering, extraction, validation, retries and human correction in latency and cost. Measure typical and slow-case latency at intended concurrency. A 270M-parameter count alone does not establish cheaper accepted documents.

Proposed cost metric: total processing and review cost across all attempted documents ÷ documents meeting your predefined acceptance criteria. Include failed attempts and retries without double counting. If no documents pass, report the cost and zero accepted documents rather than a ratio.

Keep extraction quality separate from any later language-generation task. If the next step is drafting Emirati replies from verified records, RohitAI’s Falcon-Emirati guide addresses that different model-selection question.

Evaluate Falcon-OCR-Arabic when Arabic business-document extraction is a material workload. Adopt it only for document categories where local critical-field checks and operating requirements pass. Retain the baseline elsewhere, and defer production integration until the Arabic-specific access terms are clear.

Methodology: AI-assisted reporting and analysis of TII’s announcement, model documentation, public model inventory and related RohitAI coverage, checked October 6, 2026. Calculations use published counts. No model inference, deployment, Arabic-speaker trial or measured cost comparison was performed for this article.