DITA migration specialists — we reply within one business day [email protected] [email protected]
Planning guide

DITA Conversion Cost: What Actually Drives the Number

By the DocentraX team · August 2026 · 7 min read

There is no honest per-page rate card for DITA conversion cost. What exists instead is a short list of variables that decide the number — and almost every one of them is something you can measure in your own content before you speak to a single supplier. This guide sets out what really drives DITA migration cost, what reduces each driver, and how to get a real figure for your own corpus rather than an average of everybody else's.

Why a DITA conversion cost cannot be quoted from a page count

Ask three suppliers "how much does DITA conversion cost per page?" and you will get three confident numbers, all of them guesses. A page is not a unit of work. Ten pages of a tidy MadCap Flare project arrive in DITA almost untouched; ten pages of a scanned service manual have to be rebuilt almost line by line. Page count tells you the size of a job and almost nothing about its difficulty — and difficulty is what you pay for.

What follows is the anatomy of the number: eight variables that genuinely move DITA conversion cost, and the lever that reduces each. Use it to interrogate a quote, ours or anyone else's. If you are still scoping the work itself, start with the DITA conversion overview and come back here to price it.

The eight real drivers of DITA conversion cost

1. Source format and how much recoverable structure it carries

The largest single variable, and it is not about how old the content is. It is about how much of the author's intent the source still holds in a form software can act on. Sources fall into three tiers:

  1. Structured projects. Flare, RoboHelp, FrameMaker, Sphinx, DocBook, Markdown. The project already knows its own reading order and the topic hierarchy your readers navigate, and it usually holds variables, snippets and conditions as live content rather than as typed-out text. We read such a project the way its own authoring tool reads it, so the hierarchy you publish today is the hierarchy you get in DITA — not a guess assembled from folder names. These are the cheapest corpora to move; Flare to DITA conversion shows what that looks like in practice.
  2. Semi-structured exports. Word with real heading styles, InDesign exports, Drupal exports, plain HTML sites. There is no published navigation to read, so reading order follows the order the files are stored in and topic nesting is inferred from heading levels. Recoverable — but inference is a judgement, and judgements need review. Word to DITA conversion sits here whenever authors used styles properly.
  3. Rendered or unstructured output. PDF, and any Word file where someone typed bold 14-point text instead of applying a heading style. Nothing semantic survives; headings, lists and tables have to be rebuilt from what the page merely looks like, and scanned pages need OCR before anything else can begin. PDF to DITA conversion is the most expensive route into DITA and always will be.

The practical implication: find the upstream source before you buy a conversion. If the PDF was produced from FrameMaker or Word and that file still exists on somebody's drive, converting the original removes an entire cost tier.

2. Corpus size — and, more decisively, corpus consistency

Size behaves as you expect: a mapping decision is made once and then applied everywhere it fits, so unit cost falls as volume rises. Consistency is the variable buyers underestimate. A very large corpus produced by one team from one template is a cheaper migration than a far smaller pile assembled from four acquisitions, three authoring tools and a decade of drift. Decisions amortise across content that is genuinely alike; they do not amortise across content that merely looks alike. Before requesting a quote, count your content families, not only your pages.

3. How many distinct styles and templates exist

Every conversion rests on a style-to-element map: a "Caution" paragraph style becomes a note of type caution, a "Step" style becomes a task step, a "Code" style becomes a code block. Building that map is skilled work and it scales with the number of distinct styles, not with the number of documents using them. Corpora where authors invented styles freely cost more to map than disciplined corpora, however large those disciplined corpora are.

4. Conditions, variables, snippets and reuse

Single-sourcing machinery is where cheap conversions quietly fail. Product names held as variables, snippets nested inside other snippets, glossary terms, conditional text publishing several audience variants from one file — all of it has to arrive on the DITA side as the equivalent construct, not as flattened text that has forgotten why it was ever reusable. That means variables resolved or re-expressed as keys, snippets landing in place or becoming conref targets, and your conditions mapped onto your own ditaval and profiling scheme. The more audiences, products and regions you publish against, the more this drives DITA migration cost — and the more it costs you later if a supplier skips it.

5. Image and table load

Tables are the expensive objects in technical content: CALS conversion has to survive spans, nested lists inside cells and column widths that carried meaning. Images cost differently — every reference must still resolve after files move, captions must become real figure titles, and embedded screenshots must be lifted out and named. A page of dense specification tables is not comparable to a page of prose, and a quote that treats them alike is averaging over a risk it has not measured.

6. Depth of tailoring to your own DITA rules

There are two honest products here. A generic conversion gives valid, well-formed DITA with structure inferred, generic topic types and machine-made names — a correct starting point. A customer-specific conversion also maps to your specialization and element rules, your style semantics, your conditional-text scheme, your reuse strategy, your metadata model, your naming conventions and your folder layout, and validates against your DTD. The first is cheaper to buy and more expensive to own, because your team spends the weeks after delivery hand-fixing element choices, conditions and metadata. Choose deliberately.

7. The number of review iterations

Every migration discovers things in review: an undocumented style, a legacy note convention, a table that should have been a list. Iterations are not the question — what each one costs is. Where conversion is hand work, an iteration means doing the hand work again. Where the conversion is governed by a rule you can change, an iteration means changing that rule and running the corpus again, which is why a correction can be applied everywhere it belongs rather than only where a reviewer happened to notice it.

8. Target, CCMS and delivery requirements

The finish line moves the price. Delivering one composite DITA document is not the same as delivering topic-per-file output with a bookmap, a generated key map, submaps and an asset layout your CCMS accepts on import — see composite DITA to split topics for the difference. Add required metadata, index terms, keys, or a validation report that must pass against your schema, and you have added scope.

Cost drivers and the levers that reduce them

Cost driverWhat reduces it
Unstructured source (PDF, unstyled Word)Hunt for the upstream original; OCR once, well; convert the source project rather than its printed output
Style sprawlRationalise styles before conversion, or accept coarser mapping for the ones nobody uses
Corpus inconsistencySplit into content families and convert each on its own terms instead of forcing one compromise map
Conditions, variables, snippetsDocument the profiling scheme up front; decide the target key and conref strategy before conversion, not after
Table and image complexitySample the worst tables early, so decisions are made against reality rather than against the easy pages
Depth of tailoringTailor where it changes authoring behaviour — notes, tasks, conditions, metadata — and accept generic mapping elsewhere
Review iterationsCorrections applied across the whole corpus at once instead of hand rework; one consolidated review round instead of trickled feedback
CCMS acceptance riskValidate against your own DTD before delivery with a DITA completeness check, rather than after an import fails
Paying twice for duplicate contentNormalise and de-duplicate during migration so reuse is created, not deferred — see Content Convergence

How the pricing model changes what you actually pay

Three models dominate DITA migration services, and they distribute risk very differently. The model matters as much as the rate, because it decides who absorbs the surprises your archive is holding.

The DocentraX economics, stated plainly

Our conversions are governed by mapping decisions that carry from project to project, rather than by hand work rebuilt from nothing each time. That produces a particular cost shape, and it is worth comparing against however your other candidates work; the full argument for it is made in AI + deterministic conversion.

How to get a real DITA conversion cost for your corpus

There is no published price list here, and any figure quoted before someone has looked at your files is theatre. The honest route is a sample: send twenty-five representative pages, and we convert them against your own requirements before anybody signs anything.

  1. Choose badly, on purpose. Send the dense tables, the condition-heavy topics, the chapter nobody wants to touch — not your cleanest pages.
  2. Include the machinery. If conditions, variables or snippets exist, include files that use them, plus the profiling scheme they belong to.
  3. Send your target rules. Your DTD or specialization, naming conventions and folder expectations, so the sample is measured against the standard the real project will face.
  4. Judge the output, then the price. Open it in your editor, validate it against your schema, try importing it into your CCMS. Twenty-five real pages will price the remaining thousands far better than any rate card.

Frequently asked questions

How much does DITA conversion cost per page?

There is no defensible per-page rate, because a page is not a unit of work. The same page count can represent wildly different amounts of effort depending on whether the source is a structured Flare or DocBook project or a scanned PDF, how many distinct styles exist, and how much conditional text and reuse has to be carried across. Any supplier quoting a rate before seeing your files is averaging over risk they have not measured — ask them to convert a sample of your worst content instead.

What is the biggest driver of DITA migration cost?

Source format, specifically how much recoverable structure it still carries. A structured project already knows its own reading order and holds variables, snippets and conditions as live content, so the real topic hierarchy can be rebuilt rather than guessed. Printed output such as PDF carries only the look of the page, so headings, lists and tables have to be reconstructed, and scanned pages need OCR first. Corpus consistency runs a close second: many content families cost more than many pages.

Is PDF to DITA conversion more expensive than Flare to DITA?

Yes, and structurally so. A Flare project openly carries its table of contents, variables, snippets, glossary and conditions, so the work is mapping rather than reconstruction. A PDF carries only marks on a page: structure has to be rebuilt, and scanned material needs OCR and repair before conversion can begin at all. If the PDF was generated from a FrameMaker or Word original that still exists somewhere, converting that original instead is the single largest saving available to you.

Does AI make DITA conversion cheaper?

Only when it is used narrowly. Rules are cheaper, faster and repeatable for everything they can reach — structure, styles, tables, links, conditions, reuse — and they give identical output on every run. AI earns its cost on material rules cannot reach, such as damaged text on scanned pages. Watch the billing model as closely as the technology: if the automated work is metered every time it runs, every review iteration costs you again — a model that turns your reviewers' thoroughness into a billable event.

How do I budget for a DITA migration project?

Budget for four things, not one: the conversion itself, the tailoring to your specialization and metadata model, the review iterations your team will run, and the acceptance work in your CCMS. Inventory your content families, distinct styles and profiling dimensions first, since those drive DITA migration cost more than page count does. Then convert a representative sample against your own DTD and use what it reveals to size the rest of the estate.

Do review iterations increase DITA conversion pricing?

That depends entirely on how the supplier works. Where conversion is hand work, each iteration is fresh hand work and is billed accordingly. Where the conversion is governed by rules you can change, an iteration means changing a decision and running the corpus again, so corrections propagate everywhere they belong rather than only where a reviewer noticed them. Ask explicitly how corrections are billed before signing: per-iteration billing quietly discourages the review depth a migration depends on.

Related conversions

Send us 25 pages. The messier, the better.

We'll convert them to DITA free of charge — through the real pipeline, not a demo — and review the output with you. Then we'll discuss pricing one-to-one.

Request your free sample conversion