DITA conversion turns the content you already own — trapped in Flare, Word, PDF, FrameMaker, HTML and a dozen other formats — into clean, valid, standards-based DITA you can single-source, reuse and publish anywhere. DocentraX runs every conversion, migration and transformation through one disciplined pipeline: we read each source the way the tool that produced it reads it, rebuild what the content means rather than copy how it looked, and validate the result against your own schema before anything ships. The result is not DITA in name only — it is content engineered to drop straight into your CCMS.
The value of your documentation was never in the tool that produced it. It lives in the words, the structure and the reuse logic your writers built over years — and in most organisations that value is locked inside a proprietary project, a folder of Word files, or a published site nobody can edit at the source. You cannot reuse a safety warning across ten manuals, you cannot publish the same topic to a help portal and a PDF from one source, and every correction means editing the same sentence in six places.
DITA is the OASIS open standard for topic-based technical content, and its whole reason to exist is single-sourcing, reuse through conref and keyref, conditional publishing with ditaval profiling, and multi-channel output through the DITA Open Toolkit. The classic barrier to entry has always been the migration itself: getting years of legacy content into clean, valid XML. That barrier is what DocentraX removes. Whether you are running a full legacy documentation migration or converting a single source format, the goal is the same — turn trapped content into a reusable content asset.
Every DocentraX conversion — from any source, into DITA or out of it — runs through one disciplined, five-stage lifecycle. It is a hybrid method: a repeatable, auditable and private core that behaves exactly the same way on the thousandth document as on the first, with AI applied only in the narrow places where it genuinely earns its place, and every AI judgment reused rather than paid for twice — the full argument for that hybrid is in AI + deterministic conversion.
We open each source the way the tool that produced it would open it. A Flare project is read as a Flare project, a RoboHelp help system as a help system, a FrameMaker book as a book, a Sphinx site as the site your readers browse today. That matters for one very practical reason: wherever the source publishes real navigation, the reading order you get in DITA is the navigation you actually designed. And where the source genuinely carries none — an InDesign or Drupal export, a folder of loose HTML — reading order follows the order the files are stored in, and we say so and flag it for review rather than dressing inference up as design. We then inventory every topic, asset, style, variable, condition and link, so you know the true size and shape of the estate before a single element is mapped.
We map the source's structure and styles into clean, valid DITA, rebuilding real semantics rather than copying appearance: headings become a topic hierarchy, numbered procedures become tasks, tables become CALS, and notes, footnotes, cross-references and images are reconstructed as proper DITA elements. Because the mapping is expressed as rules rather than hand-cut one-offs, tailoring it to your content is configuration, not a fresh custom build each time.
Your variables resolve to real values, and every reused snippet lands in place — including the ones nested inside other snippets, which is exactly where cheaper conversions quietly drop content. Conditions and profiling carry across, and your metadata, keys, IDs and naming conventions are applied. Optionally, Knowledge Fabric writes discovered keywords straight into topic prologs as metadata, and Content Convergence collapses duplicate topics and repeated blocks into a governed reuse library so your DITA behaves like a real single source of truth.
Before anything ships, the Completeness Check runs at repository scale against your own DTD — well-formedness, DTD conformance, broken links, xrefs, conref, keyref, image references, duplicate ids and keys, and circular map chains — together with the checks a grammar cannot make: table grids proved arithmetically sound, spans and morerows included, and a sweep that reports every file your maps never reference. "Done" stops being a judgment call and becomes a report you clear.
Output arrives in exactly the structure your platform expects: topic-per-file or composite, a map or bookmap, your folder layout, assets foldered, every reference intact. Content conservation is a zero-tolerance rule throughout — structure may degrade gracefully when it cannot be determined confidently, but text is never dropped.
A generic conversion gives you valid, well-formed DITA with inferred structure — a correct starting point, but one that uses generic infotypes, generated IDs and the source's own style names. It is real DITA; it is just not your DITA yet. A customer-specific conversion adds the layer that makes the output drop straight into your CCMS with no rework.
| Aspect | Generic run | Customer-specific run |
|---|---|---|
| Style handling | Source style names carried through | Your styles mapped to semantic elements (a "Caution" style becomes a note of type caution) |
| Infotypes | Generic topics | Your concept / task / reference specialization |
| Conditions | Carried as generic profiling attributes where the source records them | Converted to your ditaval profiling scheme |
| Reuse | Text inlined once | Mapped to your conref / keyref / keys strategy |
| IDs and files | Generated | Your naming and folder conventions |
| Validation | Standard DITA grammar | Against your own DTD or specialization |
The difference is weeks of work. Skip the tailoring and your team hand-fixes element mapping, rebuilds conditions and re-establishes reuse before the content is usable. Agree it once, as configuration, and every future batch inherits it — which is why we treat customer-specific mapping as the natural destination for any serious migration rather than an upsell at the end of one.
DocentraX converts the whole landscape of legacy formats. Each source has its own converter, built around how that format actually organises content and what its authors typically rely on:
| Source format | What survives the move |
|---|---|
| MadCap Flare to DITA conversion | We read your Flare project the way Flare itself does, so the topic hierarchy your readers navigate today is the hierarchy you get in DITA. Variables resolve to real values, reused snippets land in place including those nested inside other snippets, your glossary comes across, and topics outside the table of contents that your content links to are kept rather than quietly lost. |
| RoboHelp to DITA conversion | Your help system's contents file defines the reading order, topic nesting is rebuilt from your heading structure, and your authoring styles become real DITA semantics instead of leftover formatting. |
| FrameMaker to DITA conversion | Published FrameMaker help is read as the book it came from, so book, chapter and section hierarchy survives end to end. |
| Structured FrameMaker to DITA | The chapter sequence of your structured book drives the map, so a long, deeply nested manual arrives in DITA in the order it was written to be read. |
| InDesign to DITA conversion | Design-led pages are recovered as content: headings become real topic structure, so layout files finally become reusable topics. An InDesign export carries no table of contents, so reading order follows the stored file order — inference we flag for your review, not a hierarchy we pretend to find. |
| Drupal to DITA conversion | Your exported pages become clean, editable topics, so web content joins the same single source as your manuals. A Drupal export carries no navigation — that lives in the site's database, not the export — so reading order follows the export's folder structure and is reviewed with you rather than presented as designed. |
| Sphinx to DITA conversion | Your existing navigation defines the reading order, so the structure your readers know survives, and your site's landing page is carried across as front matter. |
| HTML to DITA conversion | The general-purpose route for any HTML that did not come from a recognised authoring tool: heading levels build the topic nesting. |
| Word to DITA conversion | One structured topic tree per document, driven by your heading styles. Word's numbering is rebuilt as genuine lists, tables become CALS, footnotes and images are carried across, and your house styles map to real semantic elements. |
| PDF to DITA conversion | Even flat, final-form PDFs are reconstructed into editable, structured content before conversion, with dedicated repair passes for scanned books. Image-only scans need OCR before conversion can begin — plan for that step up front. |
| Markdown to DITA conversion | Your site's own navigation sets the reading order. Tables become CALS, fenced code becomes code blocks, admonitions become real notes, and front matter becomes topic metadata. |
| DocBook to DITA conversion | Typed topics rather than a generic wrapper, with exact asset resolution where your export records the real filename of each graphic — and a logged warning wherever an image reference cannot be recovered, so nothing goes missing silently. |
Most teams clear the migration barrier three painful ways: a services firm doing manual migration (slow, expensive and inconsistent between writers), a one-off script that handles the easy 80% and breaks on snippets and conditions, or a desktop utility that ignores project structure and just walks the published HTML. All three leave you with generic output and weeks of cleanup.
DocentraX is engineered as a repeatable system instead of a one-off: every source read on its own terms, reading order taken from your real navigation, customer-specific mapping that is configured rather than rebuilt, a content-conservation guarantee, repository-scale validation, reuse and normalization, and the Knowledge Fabric intelligence layer on top. It is an end-to-end system, not a point tool — which is why you can run a thousand documents the same way you run one, and re-run any of them at the cost of compute rather than a fresh manual project.
A generic conversion gives you valid DITA. A customer-specific DocentraX conversion gives you DITA that drops into your CCMS and behaves like a single source of truth.
Conversion is the beginning, not the end. Once your content is clean DITA you can reshape it, publish it to other formats, and mine it for knowledge. Explore how we transform DITA into and out of other formats — DocBook, Markdown, PowerPoint, and between composite and topic shapes — and how a managed DITA migration project takes you from a pile of legacy files to production-ready content in your platform, with nothing lost along the way.
DITA conversion is the process of turning content authored in another format — such as MadCap Flare, Word, FrameMaker, HTML or PDF — into DITA, the OASIS open standard for topic-based technical content. Done well, it rebuilds real semantics: headings become a topic hierarchy, procedures become tasks, tables become CALS, and reuse is expressed with conref and keyref. The output is valid, single-sourceable content rather than a superficial format swap.
DocentraX converts MadCap Flare, Adobe RoboHelp, FrameMaker (published help and structured), Adobe InDesign exports, Drupal, Sphinx, plain HTML, Word/DOCX, PDF, Markdown and DocBook. Each has a dedicated converter built around how that format really organises content. Where the source publishes navigation — a Flare or RoboHelp project, a Sphinx site, a structured book — reading order and hierarchy come straight from it; where it does not, as with InDesign and Drupal exports or PDF, structure is rebuilt from headings and stored file order and flagged for your review.
A generic conversion produces valid, well-formed DITA with inferred structure, generic infotypes, generated IDs and the source's own style names — a correct starting point. A customer-specific conversion additionally maps your authoring styles to the right semantic elements, converts conditions to your ditaval profiling scheme, aligns reuse with your conref and keyref strategy, applies your metadata and naming conventions, and validates against your own DTD, so the output drops into your CCMS with no rework.
No. Content conservation is a zero-tolerance rule across every stage of the pipeline. Structure may degrade gracefully when it genuinely cannot be determined — an unplaceable heading is filed under a neighbouring topic and flagged for cleanup — but visible text is never discarded, and each stage is auditable so any structural degradation is reported rather than discovered later.
The core of every conversion is deterministic and rule-based: repeatable, auditable and private — content from authoring-tool projects is processed entirely under our control, not passed to outside services. The one exception is PDF, where rebuilding a final-form document into editable structure relies on a specialist cloud conversion service; we confirm that with you before anything is uploaded. AI itself is used only where it genuinely earns its place — repairing words that OCR mangled in a scanned manual, for example — and it is opt-in, sends only the flagged fragments, and reuses every judgment so re-runs never pay for the same call twice.
Every conversion output passes through the Completeness Check, which runs at repository scale against your own DTD: well-formedness, DTD conformance, broken links, cross-references, conref and keyref resolution, image references, duplicate ids and keys, and circular map chains — plus the checks no grammar can make: CALS table grids proved arithmetically sound, with spans, morerows and column definitions all consistent, and a sweep that reports every file your maps never reference. It produces an actionable report so completeness is proven before the content reaches your platform.
We'll convert them to DITA free of charge — through the real pipeline, not a demo — and review the output with you. Then we'll discuss pricing one-to-one.
Request your free sample conversion