{"id":106782,"date":"2026-07-25T19:40:53","date_gmt":"2026-07-25T19:40:53","guid":{"rendered":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/"},"modified":"2026-07-25T19:40:53","modified_gmt":"2026-07-25T19:40:53","slug":"datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown","status":"publish","type":"post","link":"https:\/\/youzum.net\/it\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/","title":{"rendered":"Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown"},"content":{"rendered":"<p class=\"wp-block-paragraph\"><strong><a href=\"https:\/\/pxllnk.co\/c1pzvpb\" target=\"_blank\" rel=\"noreferrer noopener\">Datalab has released Marker 2<\/a><\/strong>, a full rewrite of its open source document conversion pipeline. Marker converts PDF, image, PPTX, DOCX, XLSX, HTML, and EPUB files into markdown, JSON, HTML, or chunks. The Datalab team rebuilt it around three components shipped over the preceding months: <strong>Surya OCR 2<\/strong>, a <strong>20M-param fast layout model<\/strong>, and a rebuilt <strong>pdftext<\/strong> that is 3\u00d7 faster than the previous one.<\/p>\n<p class=\"wp-block-paragraph\">The main result comes from <a href=\"https:\/\/github.com\/allenai\/olmocr\/tree\/main\/olmocr\/bench\">olmOCR-bench<\/a>, a third-party benchmark from Allen AI. <a href=\"https:\/\/pxllnk.co\/c1pzvpb\" target=\"_blank\" rel=\"noreferrer noopener\">Marker 2\u2019s <\/a>balanced mode scores 76.0% overall and 83.5% on born-digital PDFs. It sustains 2.9 pages per second on a single B200 GPU. That is over 5\u00d7 the throughput of MinerU\u2019s pipeline backend, which scores 72.7% at 0.54 pages per second. Docling scores 50.3% at 2.1 pages per second on the same harness.<\/p>\n<h2 class=\"wp-block-heading\"><strong>What\u2019s New in Marker 2<\/strong><\/h2>\n<p class=\"wp-block-paragraph\"><strong><a href=\"https:\/\/pxllnk.co\/c1pzvpb\" target=\"_blank\" rel=\"noreferrer noopener\">Marker 2 <\/a>exposes three conversion paths instead of one:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li>balanced \u2014 the Surya VLM handles layout, and the whole page is re-OCR\u2019d whenever embedded text is bad. Highest quality, best on GPU. <strong>76.0%<\/strong> olmOCR-bench.<\/li>\n<li>fast \u2014 a lightweight rf-detr\/onnx layout detector plus pdftext, with minimal, surgical VLM use. <strong>66.6%<\/strong>, and far cheaper.<\/li>\n<li>\u2013disable_ocr \u2014 pure text-layer extraction, no VLM calls at all. Runs entirely on CPU. <strong>43.6%<\/strong>, 23.7 pg\/s.<\/li>\n<\/ul>\n<p class=\"wp-block-paragraph\">Mode is now <strong>device-aware by default<\/strong>: balanced on GPU, fast on CPU\/MPS, overridable with \u2013mode. Full CPU support is the second structural change. fast \u2013disable_ocr needs no GPU and no inference server, and the 20M layout model still reads columns, tables and headers on CPU.<\/p>\n<p class=\"wp-block-paragraph\">The third change is architectural, and it is the one that produces the throughput numbers. Many thin CPU workers share a single Surya inference server. The parent process budgets VLM concurrency across them, so throughput scales with server capacity rather than per-process VRAM. Datalab reports that balanced mode sustains ~2.9 pg\/s against a ~0.3 pg\/s single-stream rate on the same hardware.<\/p>\n<p class=\"wp-block-paragraph\">Breaking changes are worth flagging before an upgrade. Python 3.10+ is now required. Packaging moved from Poetry to <strong>uv<\/strong>, with hatchling as the build backend, though pip install marker-pdf is unchanged. The structured-extraction converter and extractors were removed; Datalab points users to the hosted API or a \u2013use_llm workflow instead.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Comparison<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">The scoring benchmark is <strong>olmOCR-bench<\/strong> from Ai2: 1,403 PDFs with roughly 8,400 pass\/fail unit tests covering math rendering, table structure, reading order, headers and footers, and old scans. The overall score is the macro-average across the 8 categories, computed with the official olmOCR-bench checker. Throughput is sustained <strong>concurrent<\/strong> pg\/s on one B200 host, not single-stream latency.<\/p>\n<p class=\"wp-block-paragraph\">A note on provenance. olmOCR-bench is a third-party benchmark from Ai2, but every score and throughput figure below comes from Datalab\u2019s own runs. All of them are reproducible through the open harness in the Marker repository, which ships competitor runners for MinerU, Docling and LiteParse alongside Marker\u2019s own.<\/p>\n<p class=\"wp-block-paragraph\">These numbers also reflect one benchmark\u2019s document mix measured on a single hardware setup, so results on your own documents may differ. Teams evaluating these systems should run the harness against their own corpus, which is the only way to know how the four rank on the documents they actually process.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img fetchpriority=\"high\" decoding=\"async\" width=\"2048\" height=\"1087\" data-attachment-id=\"81307\" data-permalink=\"https:\/\/www.marktechpost.com\/2026\/07\/24\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/image-535\/\" data-orig-file=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/07\/image-4.png\" data-orig-size=\"2048,1087\" data-comments-opened=\"0\" data-image-meta='{\"aperture\":\"0\",\"credit\":\"\",\"camera\":\"\",\"caption\":\"\",\"created_timestamp\":\"0\",\"copyright\":\"\",\"focal_length\":\"0\",\"iso\":\"0\",\"shutter_speed\":\"0\",\"title\":\"\",\"orientation\":\"0\",\"alt\":\"\"}' data-image-title=\"image\" data-image-description=\"\" data-image-caption=\"\" data-large-file=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/07\/image-4-1024x544.png\" src=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/07\/image-4.png\" alt=\"\" class=\"wp-image-81307\" \/><\/figure>\n<\/div>\n<h3 class=\"wp-block-heading\"><strong>Marker 2 vs MinerU<\/strong><\/h3>\n<p class=\"wp-block-paragraph\">MinerU\u2019s pipeline backend is the closest architectural match. Both read the PDF text layer and OCR selectively. On overall score, Marker balanced leads 76.0 to 72.7. On born-digital documents the two are effectively tied: 83.5 against 83.3.<\/p>\n<p class=\"wp-block-paragraph\">The separation is throughput. Marker balanced sustains 2.9 pg\/s against MinerU\u2019s 0.54 pg\/s, a 5.4\u00d7 gap at a higher score. Marker fast sustains 7.4 pg\/s, roughly 13.7\u00d7 MinerU\u2019s pipeline rate, but scores 6.1 points below MinerU to do it.<\/p>\n<p class=\"wp-block-paragraph\">MinerU also ships a <strong>VLM backend<\/strong>, which Datalab states scores higher than its pipeline backend. That backend is a full-page-VLM approach and is not in this table. AI teams evaluating MinerU should benchmark that path separately.<\/p>\n<h3 class=\"wp-block-heading\"><strong>Marker 2 vs Docling<\/strong><\/h3>\n<p class=\"wp-block-paragraph\">Docling is the widest margin among the GPU pipelines. Marker balanced leads 76.0 to 50.3 overall and 83.5 to 64.0 on born-digital, while also running faster: 2.9 pg\/s against 2.1 pg\/s. Datalab notes Docling was run on its default pipeline, which uses the text layer for born-digital pages and OCR for image regions.<\/p>\n<p class=\"wp-block-paragraph\">Docling\u2019s counterweight is governance and format breadth, not accuracy. The codebase is <strong>MIT-licensed<\/strong>, it originated at IBM Research, and it is hosted as a project in the <strong>LF AI &amp; Data Foundation<\/strong>. Its input list also extends past documents into audio and email formats.<\/p>\n<h3 class=\"wp-block-heading\"><strong>Marker 2 vs LiteParse<\/strong><\/h3>\n<p class=\"wp-block-paragraph\">LiteParse, from the LlamaIndex team, is a Rust document parser. It does not compete on the same axis. On CPU it scores 22.4 overall and 20.4 with OCR off, against Marker\u2019s CPU-only 43.6.<\/p>\n<p class=\"wp-block-paragraph\">But LiteParse with OCR disabled reports <strong>1721 pg\/s<\/strong> \u2014 roughly 73\u00d7 Marker\u2019s CPU mode, which is the tradeoff. Marker\u2019s fast \u2013disable_ocr runs a 20M layout model on CPU and still recovers structure, which is why it more than doubles a plain text dump\u2019s score. LiteParse has no layout model and collapses on anything non-linear.<\/p>\n<h3 class=\"wp-block-heading\"><strong>Marker 2 vs the full-page VLM tier<\/strong><\/h3>\n<p class=\"wp-block-paragraph\">The Datalab team emphasizes that Marker is designed as a pipeline rather than a VLM, clarifying that these are distinct tools. In this evaluation, their hosted <strong>Chandra 2<\/strong> scores 85.8, while Gemini Flash 3.5 via API scores 76.4. Datalab\u2019s Chandra repository also positions Ai2\u2019s olmOCR 2 at 82.4 and dots.ocr 1.5 at 83.9 within a separate table. For scans, math-heavy pages, and achieving top accuracy, the VLM tier remains superior to all listed pipelines.<\/p>\n<p class=\"wp-block-paragraph\">Marker\u2019s balanced mode narrows the performance gap to just 0.4 points behind Gemini Flash 3.5 overall, and it even outperforms it on born-digital documents by a margin of 83.5 to 79.1\u2014without requiring a per-page API call.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img decoding=\"async\" width=\"2048\" height=\"1087\" data-attachment-id=\"81306\" data-permalink=\"https:\/\/www.marktechpost.com\/2026\/07\/24\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/image-534\/\" data-orig-file=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/07\/image-4.png\" data-orig-size=\"2048,1087\" data-comments-opened=\"0\" data-image-meta='{\"aperture\":\"0\",\"credit\":\"\",\"camera\":\"\",\"caption\":\"\",\"created_timestamp\":\"0\",\"copyright\":\"\",\"focal_length\":\"0\",\"iso\":\"0\",\"shutter_speed\":\"0\",\"title\":\"\",\"orientation\":\"0\",\"alt\":\"\"}' data-image-title=\"image\" data-image-description=\"\" data-image-caption=\"\" data-large-file=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/07\/image-4-1024x544.png\" src=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/07\/image-4.png\" alt=\"\" class=\"wp-image-81306\" \/><\/figure>\n<\/div>\n<h2 class=\"wp-block-heading\"><strong>Per-category behavior<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">The mode you pick changes the failure profile, not just the score. Each row is one olmOCR-bench category, scored across all three modes. Math is the sharp edge: fast mode reads equations from the PDF text layer instead of VLM-OCRing them, so arXiv math falls from 83.9 to 23.4, and \u2013disable_ocr scores 0.0 there by design. Outside the two math categories, old scans is the weakest split in every mode, topping out at 43.2.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large is-resized\"><img decoding=\"async\" width=\"1024\" height=\"698\" data-attachment-id=\"81309\" data-permalink=\"https:\/\/www.marktechpost.com\/2026\/07\/24\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/image-537\/\" data-orig-file=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/07\/image-6.png\" data-orig-size=\"1374,936\" data-comments-opened=\"0\" data-image-meta='{\"aperture\":\"0\",\"credit\":\"\",\"camera\":\"\",\"caption\":\"\",\"created_timestamp\":\"0\",\"copyright\":\"\",\"focal_length\":\"0\",\"iso\":\"0\",\"shutter_speed\":\"0\",\"title\":\"\",\"orientation\":\"0\",\"alt\":\"\"}' data-image-title=\"image\" data-image-description=\"\" data-image-caption=\"\" data-large-file=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/07\/image-6-1024x698.png\" src=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/07\/image-6-1024x698.png\" alt=\"\" class=\"wp-image-81309\" \/><\/figure>\n<\/div>\n<h2 class=\"wp-block-heading\"><strong>Licensing<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">This is where the four systems diverge most for commercial teams:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Marker<\/strong>: code is Apache 2.0. Model weights use a modified AI Pubs OpenRAIL-M license \u2014 free for research, personal use, and startups under $5M funding\/revenue. Beyond that, commercial use of the weights requires a paid license.<\/li>\n<li><strong>MinerU<\/strong>: now under the MinerU Open Source License, based on Apache 2.0 with added conditions. A separate commercial license is required above 100M MAU or $20M monthly revenue, and online services built on it must disclose that fact.<\/li>\n<li><strong>Docling<\/strong>: MIT, with model licenses tracked separately in their original packages.<\/li>\n<li><strong>LiteParse<\/strong>: open source, from run-llama, with LlamaParse positioned as the paid cloud path for hard documents.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\"><strong>Use Case- Comparison<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Score alone does not pick the tool. Corpus type, hardware, licensing band and output format decide it. Try the interactive picker below to filter ten deployment scenarios by constraint and by tool, and see which parser fits your use case.<\/p>\n<div>\n<\/div>\n<p class=\"wp-block-paragraph\">\n<h2 class=\"wp-block-heading\"><strong>Key Takeaways<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li>Marker 2 balanced scores 76.0% on olmOCR-bench at 2.9 pg\/s \u2014 over 5\u00d7 MinerU\u2019s pipeline throughput.<\/li>\n<li>It beats Docling on both axes at once: 76.0% against 50.3%, and 2.9 pg\/s against 2.1 pg\/s.<\/li>\n<li>LiteParse trades structure for speed \u2014 1721 pg\/s with OCR off, but 20.4% against Marker\u2019s 43.6% on CPU.<\/li>\n<li>Fast mode with \u2013disable_ocr runs entirely on CPU, no inference server, at 23.7 pg\/s.<\/li>\n<li>Licensing splits the field: Docling is MIT, MinerU stays free to $20M monthly revenue, and Marker\u2019s weights need a paid license above $5M.<\/li>\n<li>All benchmark and throughput numbers ship with a reproducible benchmarks\/ harness.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\"><strong>Interactive Dynamic Explainer<\/strong><\/h2>\n<div>\n<\/div>\n<\/p><p class=\"wp-block-paragraph\">\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n<\/p><p class=\"wp-block-paragraph\">\n<\/p><p class=\"wp-block-paragraph\"><strong>Links:<\/strong><a href=\"https:\/\/pxllnk.co\/c1pzvpb\" target=\"_blank\" rel=\"noreferrer noopener\"> GitHub repo<\/a> |<a href=\"https:\/\/github.com\/datalab-to\/marker\/releases\/tag\/v2.0.0\"> Release notes<\/a> |<a href=\"https:\/\/www.datalab.to\/blog\/marker-2\"> Blog post<\/a> |<a href=\"https:\/\/x.com\/VikParuchuri\/status\/2079545884681830784\"> Announcement tweet<\/a><\/p>\n<p class=\"wp-block-paragraph\">\n<\/p><p>The post <a href=\"https:\/\/www.marktechpost.com\/2026\/07\/24\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/\">Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown<\/a> appeared first on <a href=\"https:\/\/www.marktechpost.com\/\">MarkTechPost<\/a>.<\/p>","protected":false},"excerpt":{"rendered":"<p>Datalab has released Marker 2, a full rewrite of its open source document conversion pipeline. Marker converts PDF, image, PPTX, DOCX, XLSX, HTML, and EPUB files into markdown, JSON, HTML, or chunks. The Datalab team rebuilt it around three components shipped over the preceding months: Surya OCR 2, a 20M-param fast layout model, and a rebuilt pdftext that is 3\u00d7 faster than the previous one. The main result comes from olmOCR-bench, a third-party benchmark from Allen AI. Marker 2\u2019s balanced mode scores 76.0% overall and 83.5% on born-digital PDFs. It sustains 2.9 pages per second on a single B200 GPU. That is over 5\u00d7 the throughput of MinerU\u2019s pipeline backend, which scores 72.7% at 0.54 pages per second. Docling scores 50.3% at 2.1 pages per second on the same harness. What\u2019s New in Marker 2 Marker 2 exposes three conversion paths instead of one: balanced \u2014 the Surya VLM handles layout, and the whole page is re-OCR\u2019d whenever embedded text is bad. Highest quality, best on GPU. 76.0% olmOCR-bench. fast \u2014 a lightweight rf-detr\/onnx layout detector plus pdftext, with minimal, surgical VLM use. 66.6%, and far cheaper. \u2013disable_ocr \u2014 pure text-layer extraction, no VLM calls at all. Runs entirely on CPU. 43.6%, 23.7 pg\/s. Mode is now device-aware by default: balanced on GPU, fast on CPU\/MPS, overridable with \u2013mode. Full CPU support is the second structural change. fast \u2013disable_ocr needs no GPU and no inference server, and the 20M layout model still reads columns, tables and headers on CPU. The third change is architectural, and it is the one that produces the throughput numbers. Many thin CPU workers share a single Surya inference server. The parent process budgets VLM concurrency across them, so throughput scales with server capacity rather than per-process VRAM. Datalab reports that balanced mode sustains ~2.9 pg\/s against a ~0.3 pg\/s single-stream rate on the same hardware. Breaking changes are worth flagging before an upgrade. Python 3.10+ is now required. Packaging moved from Poetry to uv, with hatchling as the build backend, though pip install marker-pdf is unchanged. The structured-extraction converter and extractors were removed; Datalab points users to the hosted API or a \u2013use_llm workflow instead. Comparison The scoring benchmark is olmOCR-bench from Ai2: 1,403 PDFs with roughly 8,400 pass\/fail unit tests covering math rendering, table structure, reading order, headers and footers, and old scans. The overall score is the macro-average across the 8 categories, computed with the official olmOCR-bench checker. Throughput is sustained concurrent pg\/s on one B200 host, not single-stream latency. A note on provenance. olmOCR-bench is a third-party benchmark from Ai2, but every score and throughput figure below comes from Datalab\u2019s own runs. All of them are reproducible through the open harness in the Marker repository, which ships competitor runners for MinerU, Docling and LiteParse alongside Marker\u2019s own. These numbers also reflect one benchmark\u2019s document mix measured on a single hardware setup, so results on your own documents may differ. Teams evaluating these systems should run the harness against their own corpus, which is the only way to know how the four rank on the documents they actually process. Marker 2 vs MinerU MinerU\u2019s pipeline backend is the closest architectural match. Both read the PDF text layer and OCR selectively. On overall score, Marker balanced leads 76.0 to 72.7. On born-digital documents the two are effectively tied: 83.5 against 83.3. The separation is throughput. Marker balanced sustains 2.9 pg\/s against MinerU\u2019s 0.54 pg\/s, a 5.4\u00d7 gap at a higher score. Marker fast sustains 7.4 pg\/s, roughly 13.7\u00d7 MinerU\u2019s pipeline rate, but scores 6.1 points below MinerU to do it. MinerU also ships a VLM backend, which Datalab states scores higher than its pipeline backend. That backend is a full-page-VLM approach and is not in this table. AI teams evaluating MinerU should benchmark that path separately. Marker 2 vs Docling Docling is the widest margin among the GPU pipelines. Marker balanced leads 76.0 to 50.3 overall and 83.5 to 64.0 on born-digital, while also running faster: 2.9 pg\/s against 2.1 pg\/s. Datalab notes Docling was run on its default pipeline, which uses the text layer for born-digital pages and OCR for image regions. Docling\u2019s counterweight is governance and format breadth, not accuracy. The codebase is MIT-licensed, it originated at IBM Research, and it is hosted as a project in the LF AI &amp; Data Foundation. Its input list also extends past documents into audio and email formats. Marker 2 vs LiteParse LiteParse, from the LlamaIndex team, is a Rust document parser. It does not compete on the same axis. On CPU it scores 22.4 overall and 20.4 with OCR off, against Marker\u2019s CPU-only 43.6. But LiteParse with OCR disabled reports 1721 pg\/s \u2014 roughly 73\u00d7 Marker\u2019s CPU mode, which is the tradeoff. Marker\u2019s fast \u2013disable_ocr runs a 20M layout model on CPU and still recovers structure, which is why it more than doubles a plain text dump\u2019s score. LiteParse has no layout model and collapses on anything non-linear. Marker 2 vs the full-page VLM tier The Datalab team emphasizes that Marker is designed as a pipeline rather than a VLM, clarifying that these are distinct tools. In this evaluation, their hosted Chandra 2 scores 85.8, while Gemini Flash 3.5 via API scores 76.4. Datalab\u2019s Chandra repository also positions Ai2\u2019s olmOCR 2 at 82.4 and dots.ocr 1.5 at 83.9 within a separate table. For scans, math-heavy pages, and achieving top accuracy, the VLM tier remains superior to all listed pipelines. Marker\u2019s balanced mode narrows the performance gap to just 0.4 points behind Gemini Flash 3.5 overall, and it even outperforms it on born-digital documents by a margin of 83.5 to 79.1\u2014without requiring a per-page API call. Per-category behavior The mode you pick changes the failure profile, not just the score. Each row is one olmOCR-bench category, scored across all three modes. Math is the sharp edge: fast mode reads equations from the PDF text layer instead of VLM-OCRing them, so arXiv math falls from 83.9 to 23.4, and \u2013disable_ocr scores 0.0 there by design.<\/p>","protected":false},"author":2,"featured_media":106783,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"pmpro_default_level":"","site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"_pvb_checkbox_block_on_post":false,"footnotes":""},"categories":[52,5,7,1],"tags":[],"class_list":["post-106782","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-club","category-committee","category-news","category-uncategorized","pmpro-has-access"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.3 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown - YouZum<\/title>\n<meta name=\"description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/youzum.net\/it\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/\" \/>\n<meta property=\"og:locale\" content=\"it_IT\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown - YouZum\" \/>\n<meta property=\"og:description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta property=\"og:url\" content=\"https:\/\/youzum.net\/it\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/\" \/>\n<meta property=\"og:site_name\" content=\"YouZum\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/DroneAssociationTH\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-25T19:40:53+00:00\" \/>\n<meta name=\"author\" content=\"admin NU\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Scritto da\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin NU\" \/>\n\t<meta name=\"twitter:label2\" content=\"Tempo di lettura stimato\" \/>\n\t<meta name=\"twitter:data2\" content=\"6 minuti\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/\"},\"author\":{\"name\":\"admin NU\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\"},\"headline\":\"Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown\",\"datePublished\":\"2026-07-25T19:40:53+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/\"},\"wordCount\":1274,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"image\":{\"@id\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#primaryimage\"},\"thumbnailUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp\",\"articleSection\":[\"AI\",\"Committee\",\"News\",\"Uncategorized\"],\"inLanguage\":\"it-IT\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/\",\"url\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/\",\"name\":\"Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown - YouZum\",\"isPartOf\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#primaryimage\"},\"image\":{\"@id\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#primaryimage\"},\"thumbnailUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp\",\"datePublished\":\"2026-07-25T19:40:53+00:00\",\"description\":\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\",\"breadcrumb\":{\"@id\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#breadcrumb\"},\"inLanguage\":\"it-IT\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"it-IT\",\"@id\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#primaryimage\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp\",\"width\":2048,\"height\":1087},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/youzum.net\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/yousum.gpucore.co\/#website\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"name\":\"YouSum\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/yousum.gpucore.co\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"it-IT\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\",\"name\":\"Drone Association Thailand\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"it-IT\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"width\":300,\"height\":300,\"caption\":\"Drone Association Thailand\"},\"image\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/DroneAssociationTH\/\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\",\"name\":\"admin NU\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"it-IT\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"caption\":\"admin NU\"},\"url\":\"https:\/\/youzum.net\/it\/members\/adminnu\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown - YouZum","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/youzum.net\/it\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/","og_locale":"it_IT","og_type":"article","og_title":"Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown - YouZum","og_description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","og_url":"https:\/\/youzum.net\/it\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/","og_site_name":"YouZum","article_publisher":"https:\/\/www.facebook.com\/DroneAssociationTH\/","article_published_time":"2026-07-25T19:40:53+00:00","author":"admin NU","twitter_card":"summary_large_image","twitter_misc":{"Scritto da":"admin NU","Tempo di lettura stimato":"6 minuti"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#article","isPartOf":{"@id":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/"},"author":{"name":"admin NU","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c"},"headline":"Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown","datePublished":"2026-07-25T19:40:53+00:00","mainEntityOfPage":{"@id":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/"},"wordCount":1274,"commentCount":0,"publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"image":{"@id":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#primaryimage"},"thumbnailUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp","articleSection":["AI","Committee","News","Uncategorized"],"inLanguage":"it-IT","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/","url":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/","name":"Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown - YouZum","isPartOf":{"@id":"https:\/\/yousum.gpucore.co\/#website"},"primaryImageOfPage":{"@id":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#primaryimage"},"image":{"@id":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#primaryimage"},"thumbnailUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp","datePublished":"2026-07-25T19:40:53+00:00","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","breadcrumb":{"@id":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#breadcrumb"},"inLanguage":"it-IT","potentialAction":[{"@type":"ReadAction","target":["https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/"]}]},{"@type":"ImageObject","inLanguage":"it-IT","@id":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#primaryimage","url":"https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp","width":2048,"height":1087},{"@type":"BreadcrumbList","@id":"https:\/\/youzum.net\/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/youzum.net\/"},{"@type":"ListItem","position":2,"name":"Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown"}]},{"@type":"WebSite","@id":"https:\/\/yousum.gpucore.co\/#website","url":"https:\/\/yousum.gpucore.co\/","name":"YouSum","description":"","publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/yousum.gpucore.co\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"it-IT"},{"@type":"Organization","@id":"https:\/\/yousum.gpucore.co\/#organization","name":"Drone Association Thailand","url":"https:\/\/yousum.gpucore.co\/","logo":{"@type":"ImageObject","inLanguage":"it-IT","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","width":300,"height":300,"caption":"Drone Association Thailand"},"image":{"@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/DroneAssociationTH\/"]},{"@type":"Person","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c","name":"admin NU","image":{"@type":"ImageObject","inLanguage":"it-IT","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","caption":"admin NU"},"url":"https:\/\/youzum.net\/it\/members\/adminnu\/"}]}},"rttpg_featured_image_url":{"full":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp",2048,1087,false],"landscape":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp",2048,1087,false],"portraits":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp",2048,1087,false],"thumbnail":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q-150x150.webp",150,150,true],"medium":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q-300x159.webp",300,159,true],"large":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q-1024x544.webp",1024,544,true],"1536x1536":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q-1536x815.webp",1536,815,true],"2048x2048":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q.webp",2048,1087,false],"trp-custom-language-flag":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q-18x10.webp",18,10,true],"woocommerce_thumbnail":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q-300x300.webp",300,300,true],"woocommerce_single":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q-600x318.webp",600,318,true],"woocommerce_gallery_thumbnail":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/07\/image-4-Vm2z0Q-100x100.webp",100,100,true]},"rttpg_author":{"display_name":"admin NU","author_link":"https:\/\/youzum.net\/it\/members\/adminnu\/"},"rttpg_comment":0,"rttpg_category":"<a href=\"https:\/\/youzum.net\/it\/category\/ai-club\/\" rel=\"category tag\">AI<\/a> <a href=\"https:\/\/youzum.net\/it\/category\/committee\/\" rel=\"category tag\">Committee<\/a> <a href=\"https:\/\/youzum.net\/it\/category\/news\/\" rel=\"category tag\">News<\/a> <a href=\"https:\/\/youzum.net\/it\/category\/uncategorized\/\" rel=\"category tag\">Uncategorized<\/a>","rttpg_excerpt":"Datalab has released Marker 2, a full rewrite of its open source document conversion pipeline. Marker converts PDF, image, PPTX, DOCX, XLSX, HTML, and EPUB files into markdown, JSON, HTML, or chunks. The Datalab team rebuilt it around three components shipped over the preceding months: Surya OCR 2, a 20M-param fast layout model, and a&hellip;","_links":{"self":[{"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/posts\/106782","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/comments?post=106782"}],"version-history":[{"count":0,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/posts\/106782\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/media\/106783"}],"wp:attachment":[{"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/media?parent=106782"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/categories?post=106782"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/tags?post=106782"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}