Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs
Jina AI, part of Elastic, has released jina-ocr-v1, an end-to-end visual document parser.It takes PDFs, scans, tables, charts or invoices and returns clean Markdown in 1 pass.The model has 3.4B total parameters, with about 570M decoder parameters active per token.
A speculative decoding head ships inside the checkpoint.Jina AI built it to serve on low-budget GPUs such as the NVIDIA L4.The technical report lists 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench.Is it deployable?Yes, for research and non-commercial use.The open weights are about 6.
8 GB in BF16 and run on Transformers or vLLM.The CC BY-NC 4.0 license means commercial use requires contacting Jina AI.What is jina-ocr-v1?The model post-trains DeepSeek-OCR and keeps its 2 efficiency components.
DeepEncoder has about 380M parameters and chains SAM, a 16x convolutional compressor and CLIP-L.It turns a 1024×1024 page view from 4,096 patches into 256 visual tokens.A dynamic-resolution mode adds up to 9 local tiles at 100 tokens each.That caps a page at 1,156 visual tokens.
The decoder is DeepSeek-3B-MoE with 12 layers, 64 routed experts and 2 shared experts.Top-6 routing activates about 570M parameters per token.The position limit is 32,768.Output is Markdown, with tables in HTML and formulas in LaTeX.
How FastMTP Speculative Decoding Works OCR output is near-deterministic and locally structured.That makes it a good fit for speculative decoding.Jina AI adds a FastMTP head: 1 dense draft block applied recursively for K=3 steps.Draft parameters stay constant as depth grows.
The decoder then verifies the drafts greedily.It accepts the longest prefix that matches its own choices and commits 1 more token itself.If all 3 drafts match, that extra token is a bonus.The committed text always equals plain greedy decoding, so the speedup is lossless.At K=3 the model commits 2.
73 tokens per step on average.(function(){var f=document.getElementById("mtp-jina-ocr-v1");window.addEventListener("message",function(e){if(f&&e.source===f.contentWindow&&e.data&&e.data.jxH){f.style.height=e.data.
jxH+"px";}});})(); Post-Training With Dense Verifiable Rewards Post-training combines instruction alignment, robustness fine-tuning on degraded pages, and GRPO.Every reward term is deterministic code scored against a reference transcription.
The terms cover content, formulas, tables, structural validity, unit tests, repetition and format.The terms are multiplied, and each one is graded, so partly correct pages earn partial credit.Structural, unit-test and format terms are floored at 0.2, and the table term at 0.1.
The repetition term has no floor, because loops can inflate the content score.On natural pages, the formula and table rewards apply to few samples.Jina AI therefore built JinaOCRSynth, synthetic pages packed with both, each carrying olmOCR-Bench-style unit tests.
An agent also merges candidate checkpoints under a fixed evaluation budget.The draft head is trained last, against the frozen final verifier.Benchmarks and Throughput ModelParams as listed in the paperOmniDocBench v1.6olmOCR-Benchjina-ocr-v13B/570M91.1483.4DeepSeek-OCR3B/570Mnot listed76.
0DeepSeek-OCR-23B/570M90.25not listedPaddleOCR-VL-1.60.9B96.34not listedchandra-ocr-24Bnot listed85.8Qwen3-VL-235B235B/22B89.78not listed For MoE models, params show decoder total and active counts.The whole jina-ocr-v1 model is about 3.4B.The model does not lead on accuracy.PaddleOCR-VL-1.
6 and HunyuanOCR-1.5 (94.74) score higher on OmniDocBench.chandra-ocr-2 and dots.mocr (83.9) score higher on olmOCR-Bench.Post-training does add 7.4 points over the DeepSeek-OCR backbone on olmOCR-Bench.Throughput is the main result.On 1 A100 40 GB at concurrency 32, jina-ocr-v1 parses 2.
57 pages per second.That is the highest of 14 systems Jina AI measured, against 1.22 for olmOCR-2 and 0.38 for chandra-ocr-2.It emits 1,085 output tokens per page.Jina AI says that is the shortest output among systems scoring above 83.On an NVIDIA L4 at batch size 1, eager decoding rises from 42.
7 to 83.1 tokens per second.That is a 1.95x speedup at a 57.6% acceptance rate.With CUDA graphs the baseline is already 158.3 tokens per second.There, K=1 works best at 185.6 tokens per second, a 1.17x gain.How to Run It The quickest route is Jina Reader.Send a URL to r.jina.
ai with the header X-Respond-With: jina-ocr-v1.Reader fetches the page or PDF, runs the model and returns Markdown.An X-Page header transcribes 1 page of a longer document.Jina AI also hosts an OpenAI-compatible endpoint at https://api.jina.ai/v1/chat/completions.
A hosted demo is available for quick tests.For self-hosting, weights and custom code ship in 1 repository and load with trustremotecode=True.FastMTP requires vLLM 0.21 or later and a one-time register() call.The Transformers path runs the MoE decoder alone and ignores the draft weights.
Key Takeaways 3.4B total parameters, about 570M active per token, built on DeepSeek-OCR.FastMTP drafts 3 tokens per step, and greedy verification keeps decoding lossless.Scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench.Reaches 2.
57 pages per second on 1 A100, the highest of 14 measured systems.Available on Hugging Face and through a Jina Reader header today.Check out the Paper, Model weights, Release post, Model page and Announcement.All credit goes to the researcher of this project.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs appeared first on MarkTechPost.
Related
相關文章

無問芯穹與華環電子簽署戰略合作,共同探索國產異構算力AI基礎設施新方向
無問芯穹與華環電子簽署戰略合作協議,雙方將結合各自在AI軟體平台、網路通信與硬體研發的優勢,共同探索國產異構算力基礎設施的協同方案。此次合作聚焦於智算中心解決方案及「Token工廠」新模式,目標是推動計算、網路與AI原生基礎設施深度融合,為AI規模化應用提供高效穩定的支撐。

優步全球範圍裁員 10%,被裁員工稱 AI 已大舉滲透日常工作
作者:清源 責編:清源 評論: 9 月 18 日消息,據《商業內幕》今天(18 日)晚間報道,在優步(Uber),AI 已經滲透到員工工作的許多環節,從回答 Slack 裡的內部問題,到替乘客行程中聯繫客服時收到的消息撰寫回復。6 名近期遭裁員的員工透露,過去幾個月,AI 在工作中的使用範圍明顯擴大,其中一些人甚至會通過提示詞讓 AI 完成相當一部分任務。

智譜 ZCode 被質疑“偷傳代碼”:官方回應稱問題已修復,將開源代碼庫、引入第三方審查
作者:清源 責編:清源 評論: 感謝網友 咩咩洋 的線索投遞!9 月 18 日消息,針對社區中有關代碼庫數據上傳的討論,智譜旗下編程產品 ZCode 今天(18 日)通過智譜官方群組向受影響用戶致歉,併發布回應稱已第一時間完成自查,相關問題目前已經修復。

月之暗面遞表之後,Kimi 的成色要被驗算三遍
舒澤品牌手記2026.09.18 18:16 · 來自浙江全文4982字00:00 / 14:05Anthropic 的 30 萬次指控,會成為招股書的第幾頁?文 | 舒澤品牌手記9月17日,月之暗面發佈了一套金融行業解決方案。按官方披露,中信建投、中金公司、易方達等數十家金融機構已經在用 Kimi 處理投研建模、風險排查和盡調材料——研究人員把管理層報表、審計報告和盡調文件交給 Kimi,拿回一份可以繼續調整假設的 Excel 模型。同一天,深圳商報記者就港股上市進展、股東架構調整等事項向月之暗面發去採訪函。

Calibre上手 AI 互動寫作:電子書管理器搖身變成"文字冒險遊戲引擎"
這個遊戲默認藏而不發,不會跟著 Calibre 啟動就冒出來。用戶得主動在"首選項 — 工具欄和菜單"裡把它請到主工具欄,才算真正激活。它的玩法很清晰:由 AI 在後臺搭起並掌管一個虛構世界,用戶通過不斷輸入文字來推著故事往前走,等於把"讀電子書"這件事,翻轉成了"和 AI 一起寫故事"。

Anthropic 用 AI 造 AI:Claude 已主導公司四分之一研發工作
截至今年 8 月,Claude 已主導公司 26% 的 AI 研發任務,而在 2 月這一佔比還不足 1%,增長速度十分驚人。六級自動化框架:超 90% 達到人機協作,但沒有“完全自主”Anthropic 採用六級自動化框架評估 AI 參與程度,其中 AL4 級別定義為“AI 主導”:人類只給出高層目標指令,由 AI 完成任務的絕大部分流程,人負責監督審核。