Ask a documentation team what is blocking their move to DITA and you will usually hear a number: pages, books, gigabytes. Ask a second question — which part of it worries you — and the answer changes shape entirely.
Most teams are sitting on two piles.
The first pile converts. FrameMaker books, MadCap Flare projects, RoboHelp output, Word files with real styles applied. The structure is there in the source: headings are headings, procedures are numbered lists, tables are tables. Any competent conversion vendor will quote for this pile, and most will do a reasonable job.
The second pile is why the project never starts. Scanned page images from before anyone thought about reuse. PDFs from the 1990s where the text is technically present but carries no styling information at all. Print-era manuals held in a format that was only ever meant to be printed. Mixed archives nobody has opened in a decade.
That second pile is rarely the larger one. It is almost always the one that decides whether the migration happens.
A migration needs a number before it can get a budget. To produce a number, somebody has to look at the content and say how much of it is machine-recoverable. For the first pile that is straightforward. For the second it is genuinely hard — and the honest answer from most vendors is a wide range with a lot of manual re-keying inside it.
Wide ranges do not get approved. So the migration is deferred, the content stays where it is for another year, and the people who understood it leave.
Traditional conversion tooling is built on rules that read structure out of the source and map it to DITA elements. This works well, and it is genuinely the right approach — when there is structure to read.
On a scanned page there is none. There is no heading element to detect, because there is no element at all: there is a picture of a heading. On an unstyled legacy PDF the situation is subtler and often worse — text is present, so tools will happily ingest it, but nothing distinguishes a heading from a caption from a body paragraph except size, position and human judgment.
Rules cannot recover what was never encoded. That is not a failure of the approach; it is the boundary of it.
The obvious answer in 2026 is to hand the whole thing to a large language model. It is a reasonable instinct, and on a single page the result is often impressive. In production, three problems appear.
The two approaches fail in opposite directions, which is exactly what makes the combination work.
Deterministic conversion rules — written for the specific patterns in your content, not generic ones — do the bulk of the work. They produce identical output on every run. They validate against the DITA DTDs. They cost nothing to re-run. And when review finds a defect, the fix goes into the rule, which means that defect is gone permanently rather than gone from one file.
AI is spent only where rules genuinely cannot decide: transcribing the words OCR mangled, from exactly the page fragments a read-only quality pass flagged; proofreading within tight bounds that let it fix recognition damage but never rephrase; titling the figures rebuilt from the scan; and — where you switch it on — enriching topics with terms and metadata. Everything structural stays deterministic: a chapter heading the scan lost is restored by cross-checking the book's own table of contents, and whether a topic is a concept, a task or a reference is decided by rule, the same way on every run. The AI decisions that remain are reviewed, cached and auditable — and because they are cached, the second run does not pay for them again.
The result is that the cost curve bends the right way. The first pass is the expensive one. Every pass after it is close to free, which is what allows review to actually converge instead of being rationed.
To check whether this holds on the genuinely worst input, we used a service manual printed in 1925 — scanned as page images, no usable text layer, no styling information, and a page design that predates every convention modern tooling assumes.
Through the pipeline it came out as well-formed DITA topics under a real bookmap with per-chapter submaps, with no placeholder or dummy topics standing in for content that failed to convert. OCR damage was repaired. Figures that the scan had shredded across page boundaries were reassembled by re-rendering them from the source itself.
AI touched only the fragments the quality pass flagged, and every decision it made was cached — so the passes that followed made no AI calls at all. The cost curve stayed bent the right way even on the worst input we could find.
One last thing worth insisting on, whoever you work with. A migration is not finished when files are delivered. It is finished when the content is imported, valid and accepted in your CCMS — Heretto, Paligo, IXIASOFT, Tridion Docs, or a plain DITA repository.
Everything in between — DTD validation, resolved cross-references and conrefs, preserved tables and images, a working ditamap with keys and relationship tables rather than a folder of loose topics — is what separates content your team can maintain for twenty years from content they will quietly rebuild by hand.
If you have a pile that vendors have been declining, that is the one we would like to see. Send us up to 25 pages and we will convert them to DITA free of charge, through the same pipeline we use in production. You get the topics, the ditamap, the validation report, and an honest assessment of what a full migration takes — including if the answer is that this one is harder than it looks.
Only after you have seen the output on your own content do we talk about price.
We'll convert them to DITA free of charge — through the real pipeline, not a demo — and review the output with you. Then we'll discuss pricing one-to-one.
Request your free sample conversion