DITA migration specialists — we reply within one business day [email protected] [email protected]
Insights

AI + deterministic conversion: how we get quality and cost-efficiency in the same migration

By the DocentraX team · August 2026 · 7 min read

If you're planning a DITA migration in 2026, you've probably run a few of your pages through an AI chatbot and been impressed. You've probably also gotten a quote from a traditional conversion vendor and wondered why it costs so much. Here's the honest answer to the question behind both experiments: neither pure AI nor pure rules can carry a production migration alone — and how you combine them decides both the quality of your DITA and the size of your invoice.

What AI conversion gets right — and where it breaks down

Modern AI models are genuinely remarkable at reading documents. Give one a scanned page and it will read the text, guess the headings, and produce something that looks like structured output. For a 10-page sample, the result can be striking.

Production migrations are not 10-page samples. They are 5,000, 50,000, sometimes 500,000 pages — and at that scale, four problems appear that no demo shows you:

What deterministic conversion gets right — and where it stops

Deterministic conversion is the opposite temperament: rules, mappings and transformations engineered for your specific content. Rules are boring in the best possible way:

But rules have a hard limit: they need something to hold on to. A FrameMaker book, a Flare project, an HTML export — these carry recoverable structure, and rules convert them beautifully. A scanned page from a 1980s maintenance manual carries no structure at all. There is nothing for a rule to read — only pixels and, after OCR, undifferentiated text. This is exactly where pure-rules vendors give up, and exactly where AI shines.

The DocentraX pipeline: rules first, AI where it earns its place

Deterministic where it matters. AI where it helps. Humans where it counts.

Our conversion pipeline runs both approaches in a deliberate order:

  1. Analysis first. We study your corpus and build conversion rules and mappings for your documents — your heading conventions, your table styles, your warning formats. This is where migration quality is actually decided.
  2. Rules do the heavy lifting. Everything with recognizable structure is converted deterministically: repeatable, validated, and with zero AI cost on those pages — which, even in messy corpora, is most of them.
  3. AI handles what rules can't reach. On scanned and damaged pages it transcribes the words OCR mangled — from exactly the fragments a read-only quality pass flagged — proofreads within strict bounds that let it fix recognition damage but never rephrase, and titles rebuilt figures. Structure recovery and topic typing stay deterministic, and AI output flows into the deterministic pipeline, not around it, so the final DITA is still validated and consistent.
  4. Every AI decision is reviewed and remembered. AI suggestions are cached, auditable and human-reviewed. Approved decisions join your customer-specific knowledge base — so the AI share of the work falls with every project, and so does the cost.
  5. Humans close the loop. Our review team and yours iterate together. Every accepted fix becomes a rule. When your content imports into your CCMS — valid, consistent, complete — the migration is done.

What this means for you, in practice

Your concernHow the hybrid pipeline answers it
"Will the DITA be consistent across 20,000 topics?" Rules produce identical structures for identical patterns — consistency is a property of the pipeline, not a hope.
"Will it import into our CCMS?" DTD validation and link resolution are built into the pipeline; validation reports ship with every delivery.
"What about our scanned manuals?" OCR, rule-based figure and heading repair, and tightly bounded AI transcription feed the deterministic pipeline — the hard cases get the extra intelligence without giving up the guarantees.
"Why isn't this priced like a pure-AI service?" It's priced better: AI is spent only on the pages that need it, and re-runs during review cost compute, not tokens. Your budget buys quality, not redundant inference.
"What happens on the next project?" Your terms, metadata and approved decisions are remembered per customer — every subsequent migration starts smarter and cheaper.

Frequently asked questions

Is AI-converted DITA valid enough for a CCMS import?

Raw AI output routinely fails DTD validation, which is why a CCMS rejects it. In our pipeline every deliverable is validated against the DITA DTDs before delivery, so what you import is what the grammar accepts — proven on every run, not presumed.

What does AI actually do in the DocentraX pipeline?

Three narrow jobs: transcribing words OCR could not read, from exactly the page fragments a read-only quality pass flagged; proofreading within strict bounds that allow it to fix recognition damage but never to rephrase; and titling figures rebuilt from the source. Structure recovery and topic typing are deterministic rules — and every AI decision is cached, so a re-run makes no new AI calls.

Why is a hybrid pipeline cheaper than a pure-AI service?

AI is spent only on the fragments that genuinely need it, and every decision it makes is remembered — so review iterations cost compute, not tokens, and the AI share of the work falls from project to project as your approved decisions accumulate.

Related reading

The bottom line

AI made document conversion demos easy. It did not make document migrations easy — those are still won by engineering discipline: rules that guarantee consistency, AI applied precisely where it adds value, and humans accountable for the result. That combination is what we've built DocentraX around, and it's why we can make you a simple offer instead of a sales pitch:

Send us 25 pages. The messier, the better.

We'll convert them to DITA free of charge — through the real pipeline, not a demo — and review the output with you. Then we'll discuss pricing one-to-one.

Request your free sample conversion