On October 8, 2026, LightOn announced LightOnOCR-3, an Apache-2.0 model family for teams extracting information from scanned documents and building retrieval-augmented generation (RAG) systems. The 0.8B, 1B and 4B checkpoints are publicly downloadable.
LightOnOCR-3 combines transcription, labeled page regions, image descriptions and chart tables in one downloadable model family. That gives builders a way to connect extracted content to its position on a page. Whether it can replace an existing OCR and layout pipeline depends on the documents, the errors that matter and the cost of reviewing them.
Choose a checkpoint by document type
The model cards identify an architectural split: 0.8B and 4B use Qwen3.5, while 1B retains the LightOnOCR-2 architecture. LightOn recommends 4B for most tasks. For a new evaluation, compare 0.8B and 4B with your current parser; include 1B when compatibility with an existing LightOnOCR-2 deployment matters.
The table uses LightOn’s versioned ParseBench results and French-document benchmark results, not independent measurements. French handwriting scores are converted from fractions to a 0–100 scale. The final column is our suggested evaluation role.
Checkpoint | Architecture | ParseBench overall | French handwriting | Why include it |
|---|---|---|---|---|
Qwen3.5 | 74.61 | 33.3 | Compact baseline with the full output interface | |
LightOnOCR-2 family | 71.40 | 37.0 | Existing serving-stack compatibility | |
Qwen3.5 | 75.14 | 46.1 | Compare difficult pages and handwriting |
Calculated from those rows, 0.8B exceeds 1B by 3.21 ParseBench points. Moving from 0.8B to 4B adds only 0.53 overall points there, but 12.8 points on French handwriting. That is a reason to include representative handwriting in a pilot, not proof that the larger model is worth its operating cost or that either gap is statistically significant.
These scores are not percentages of error-free documents. The benchmark notes also warn that the French dataset revision contains 2,318 tests rather than the published leaderboard’s 2,335. Keep the scorer and dataset revision attached to results; do not combine rankings from different setups.
The published olmOCR summaries differ: the launch article gives olmOCR overall scores of 86.3, 85.5 and 84.5 for 4B, 0.8B and 1B. The pinned repository CSV records 86.1, 85.4 and 84.5 after its PP3 formatting postprocessing. The sources do not explain the difference. Use a named evaluation configuration, rather than averaging the two.
Start with the pinned local client
The following setup is adapted from the maintainer’s client instructions and vLLM’s serving options. It assumes Git, uv and a working vLLM-compatible GPU environment. It pins the repository and 0.8B checkpoint checked for this article and binds the server to loopback. These commands have not been executed as a deployment test.
git clone https://github.com/lightonai/LightOnOCR.git
cd LightOnOCR
git checkout 36755d461be079737860a5f03ae0c803501269e9
uv sync --extra vllm
uv run vllm serve lightonai/LightOnOCR-3-0.8B \
--revision 4a953edfc77f0e435532c503dd69dd74663d44c3 \
--host 127.0.0.1 --port 8010 \
--limit-mm-per-prompt '{"image": 1}' \
--max-model-len 18000 \
--default-chat-template-kwargs '{"enable_thinking": false}' \
--mm-processor-cache-gb 0 --no-enable-prefix-cachingOnce the server is ready, open another terminal in the same repository and run this against a PDF you are authorized to process:
uv run lightonocr paper.pdf --mode groundingThe serving dependency set pins vLLM 0.30.0 and Transformers 5.16.1. Maintainers report looping with vLLM 0.27.1 and a 1B compatibility problem with Transformers 5.17 on vLLM 0.30.0. Those are their advisories, not failures reproduced for this guide.
Record the checkpoint revision, dependency lock, device, sampling settings and rendered page size with each run. The 18,000-token setting above is a configured context limit, not a tested guarantee that every page fits. Detect truncated responses, repeated output and missing regions before accepting a page.
The client saves page images, Markdown and parsed grounding JSON under out/<name>/. Its normal output includes a math-delimiter repair; for an untouched model response through the Python client, use ocr.ocr(page, fix_dollars=False). Keep the raw response as well as any normalized representation. Client output documentation.
Use the two supported modes, then preserve provenance
The 0.8B model card documents two inputs, not an open-ended instruction-following interface:
plain: send one page image without a text instruction for transcription.grounding: send the image with only the textgroundingfor labeled content blocks and bounding boxes.
The card says other instructions fall outside the training distribution. Asking it directly for invoice-field JSON or a document summary is therefore not the documented workflow. Treat those as separate downstream steps with their own validation. Disable thinking for the Qwen3.5-based variants, as their cards specify.
The grounded output specification represents boxes on a 0–1000 scale and uses a + suffix for continuation blocks. To map a box to the rendered image, calculate x_pixels = x × image_width / 1000 and y_pixels = y × image_height / 1000. Validate coordinate order and bounds, preserve page transforms when mapping back to PDF coordinates, and retain continuation markers when reconstructing reading order.
Chart tables need a different check from transcription. LightOn’s training description says chart targets copy printed values but estimate unprinted values from axes. A bounding box identifies the region; it does not establish numerical accuracy. Preserve the chart crop, verify series names, units and axis scales, and do not treat every table cell as a number literally printed in the source.
For RAG ingestion, attach document ID, page number, region coordinates, access permissions and extraction revision to each chunk. Keep literal text distinguishable from generated image descriptions and chart reconstructions. Parse model-produced HTML tables as data, not trusted active markup; validate cross-page table joins separately.
Extraction still leaves indexing, retrieval and answer verification to the application. The EmbeddingGemma 2 local retrieval guide covers the complementary embedding layer. That is an architectural connection, not evidence of a tested LightOnOCR integration.
Evaluate accepted pages, not just throughput
Client defaults and benchmark settings differ. The client profiles use a 2,048-pixel longest edge for 0.8B/4B and 1,540 for 1B. LightOn’s speed analysis reports that 0.8B/4B reach their best olmOCR scores at 400 DPI with a five-megapixel cap, while reducing resolution improves throughput. The speed tests used one H100 per model. Do not pair high-resolution accuracy with low-resolution peak speed in a production forecast.
A useful pilot would keep development pages separate from a held-out evaluation set and compare candidates on the same source documents. The following is a proposed evaluation, not work performed for this article:
Stratify by language, scan quality, handwriting and layout. Include charts, forms, tiny text, multi-column pages and tables that cross page boundaries.
Have reviewers establish reference transcriptions, critical identifiers and amounts, table row/column relationships, chart series and page regions. Define acceptable numerical tolerances before scoring.
Count omissions, invented content, reading-order errors, malformed boxes and failed or truncated responses. Measure critical fields separately from aggregate text quality.
Measure supported answers and correct source citations in the downstream RAG system separately. Passing extraction checks does not establish answer quality.
French challenge-set results do not establish performance on Arabic business documents. For that workload, the Falcon-OCR-Arabic evaluation guide offers a related language-specific checklist.
For a cost comparison, use the same workload and accounting period on both sides. A proposed decision metric is:
Cost per accepted page =
(compute or API charges + rendering + storage + operations + review costs)
/ pages meeting the predefined acceptance criteriaInclude failed attempts and retries in those costs. If no pages pass, report that outcome rather than dividing by zero. Compare latency at the required concurrency and utilization, not just maximum throughput. This framework can show whether local extraction is worthwhile without assuming downloadable weights make operation free.
No LightOnOCR-3 hosted-service rate, SLA, minimum VRAM guarantee or measured local-versus-hosted break-even was established for this guide. Local inference also does not make the whole document workflow local: check storage, telemetry, embeddings and answer generation before making that claim.
The next step is a limited comparison with your current parser. Promote a checkpoint only when its accepted outputs, review burden and operating costs meet your requirements; retain the existing path for document categories it does not handle reliably.
Methodology: AI-assisted reporting and analysis based on LightOn’s announcement, model cards, repository code and published benchmark artifacts, checked on October 9, 2026. All reported benchmark results are LightOn’s. No inference, deployment, cost measurement or client test was performed for this article.
