There is no honest per-page rate card for DITA conversion cost. What exists instead is a short list of variables that decide the number — and almost every one of them is something you can measure in your own content before you speak to a single supplier. This guide sets out what really drives DITA migration cost, what reduces each driver, and how to get a real figure for your own corpus rather than an average of everybody else's.
Ask three suppliers "how much does DITA conversion cost per page?" and you will get three confident numbers, all of them guesses. A page is not a unit of work. Ten pages of a tidy MadCap Flare project arrive in DITA almost untouched; ten pages of a scanned service manual have to be rebuilt almost line by line. Page count tells you the size of a job and almost nothing about its difficulty — and difficulty is what you pay for.
What follows is the anatomy of the number: eight variables that genuinely move DITA conversion cost, and the lever that reduces each. Use it to interrogate a quote, ours or anyone else's. If you are still scoping the work itself, start with the DITA conversion overview and come back here to price it.
The largest single variable, and it is not about how old the content is. It is about how much of the author's intent the source still holds in a form software can act on. Sources fall into three tiers:
The practical implication: find the upstream source before you buy a conversion. If the PDF was produced from FrameMaker or Word and that file still exists on somebody's drive, converting the original removes an entire cost tier.
Size behaves as you expect: a mapping decision is made once and then applied everywhere it fits, so unit cost falls as volume rises. Consistency is the variable buyers underestimate. A very large corpus produced by one team from one template is a cheaper migration than a far smaller pile assembled from four acquisitions, three authoring tools and a decade of drift. Decisions amortise across content that is genuinely alike; they do not amortise across content that merely looks alike. Before requesting a quote, count your content families, not only your pages.
Every conversion rests on a style-to-element map: a "Caution" paragraph style becomes a note of type caution, a "Step" style becomes a task step, a "Code" style becomes a code block. Building that map is skilled work and it scales with the number of distinct styles, not with the number of documents using them. Corpora where authors invented styles freely cost more to map than disciplined corpora, however large those disciplined corpora are.
Single-sourcing machinery is where cheap conversions quietly fail. Product names held as variables, snippets nested inside other snippets, glossary terms, conditional text publishing several audience variants from one file — all of it has to arrive on the DITA side as the equivalent construct, not as flattened text that has forgotten why it was ever reusable. That means variables resolved or re-expressed as keys, snippets landing in place or becoming conref targets, and your conditions mapped onto your own ditaval and profiling scheme. The more audiences, products and regions you publish against, the more this drives DITA migration cost — and the more it costs you later if a supplier skips it.
Tables are the expensive objects in technical content: CALS conversion has to survive spans, nested lists inside cells and column widths that carried meaning. Images cost differently — every reference must still resolve after files move, captions must become real figure titles, and embedded screenshots must be lifted out and named. A page of dense specification tables is not comparable to a page of prose, and a quote that treats them alike is averaging over a risk it has not measured.
There are two honest products here. A generic conversion gives valid, well-formed DITA with structure inferred, generic topic types and machine-made names — a correct starting point. A customer-specific conversion also maps to your specialization and element rules, your style semantics, your conditional-text scheme, your reuse strategy, your metadata model, your naming conventions and your folder layout, and validates against your DTD. The first is cheaper to buy and more expensive to own, because your team spends the weeks after delivery hand-fixing element choices, conditions and metadata. Choose deliberately.
Every migration discovers things in review: an undocumented style, a legacy note convention, a table that should have been a list. Iterations are not the question — what each one costs is. Where conversion is hand work, an iteration means doing the hand work again. Where the conversion is governed by a rule you can change, an iteration means changing that rule and running the corpus again, which is why a correction can be applied everywhere it belongs rather than only where a reviewer happened to notice it.
The finish line moves the price. Delivering one composite DITA document is not the same as delivering topic-per-file output with a bookmap, a generated key map, submaps and an asset layout your CCMS accepts on import — see composite DITA to split topics for the difference. Add required metadata, index terms, keys, or a validation report that must pass against your schema, and you have added scope.
| Cost driver | What reduces it |
|---|---|
| Unstructured source (PDF, unstyled Word) | Hunt for the upstream original; OCR once, well; convert the source project rather than its printed output |
| Style sprawl | Rationalise styles before conversion, or accept coarser mapping for the ones nobody uses |
| Corpus inconsistency | Split into content families and convert each on its own terms instead of forcing one compromise map |
| Conditions, variables, snippets | Document the profiling scheme up front; decide the target key and conref strategy before conversion, not after |
| Table and image complexity | Sample the worst tables early, so decisions are made against reality rather than against the easy pages |
| Depth of tailoring | Tailor where it changes authoring behaviour — notes, tasks, conditions, metadata — and accept generic mapping elsewhere |
| Review iterations | Corrections applied across the whole corpus at once instead of hand rework; one consolidated review round instead of trickled feedback |
| CCMS acceptance risk | Validate against your own DTD before delivery with a DITA completeness check, rather than after an import fails |
| Paying twice for duplicate content | Normalise and de-duplicate during migration so reuse is created, not deferred — see Content Convergence |
Three models dominate DITA migration services, and they distribute risk very differently. The model matters as much as the rate, because it decides who absorbs the surprises your archive is holding.
Our conversions are governed by mapping decisions that carry from project to project, rather than by hand work rebuilt from nothing each time. That produces a particular cost shape, and it is worth comparing against however your other candidates work; the full argument for it is made in AI + deterministic conversion.
There is no published price list here, and any figure quoted before someone has looked at your files is theatre. The honest route is a sample: send twenty-five representative pages, and we convert them against your own requirements before anybody signs anything.
There is no defensible per-page rate, because a page is not a unit of work. The same page count can represent wildly different amounts of effort depending on whether the source is a structured Flare or DocBook project or a scanned PDF, how many distinct styles exist, and how much conditional text and reuse has to be carried across. Any supplier quoting a rate before seeing your files is averaging over risk they have not measured — ask them to convert a sample of your worst content instead.
Source format, specifically how much recoverable structure it still carries. A structured project already knows its own reading order and holds variables, snippets and conditions as live content, so the real topic hierarchy can be rebuilt rather than guessed. Printed output such as PDF carries only the look of the page, so headings, lists and tables have to be reconstructed, and scanned pages need OCR first. Corpus consistency runs a close second: many content families cost more than many pages.
Yes, and structurally so. A Flare project openly carries its table of contents, variables, snippets, glossary and conditions, so the work is mapping rather than reconstruction. A PDF carries only marks on a page: structure has to be rebuilt, and scanned material needs OCR and repair before conversion can begin at all. If the PDF was generated from a FrameMaker or Word original that still exists somewhere, converting that original instead is the single largest saving available to you.
Only when it is used narrowly. Rules are cheaper, faster and repeatable for everything they can reach — structure, styles, tables, links, conditions, reuse — and they give identical output on every run. AI earns its cost on material rules cannot reach, such as damaged text on scanned pages. Watch the billing model as closely as the technology: if the automated work is metered every time it runs, every review iteration costs you again — a model that turns your reviewers' thoroughness into a billable event.
Budget for four things, not one: the conversion itself, the tailoring to your specialization and metadata model, the review iterations your team will run, and the acceptance work in your CCMS. Inventory your content families, distinct styles and profiling dimensions first, since those drive DITA migration cost more than page count does. Then convert a representative sample against your own DTD and use what it reveals to size the rest of the estate.
That depends entirely on how the supplier works. Where conversion is hand work, each iteration is fresh hand work and is billed accordingly. Where the conversion is governed by rules you can change, an iteration means changing a decision and running the corpus again, so corrections propagate everywhere they belong rather than only where a reviewer noticed them. Ask explicitly how corrections are billed before signing: per-iteration billing quietly discourages the review depth a migration depends on.
We'll convert them to DITA free of charge — through the real pipeline, not a demo — and review the output with you. Then we'll discuss pricing one-to-one.
Request your free sample conversion