DITA migration specialists — we reply within one business day [email protected] [email protected]
DITA migration & conversion services

Your legacy content, migrated to clean, valid DITA.

DocentraX converts structured and unstructured legacy content into topic-based, CCMS-ready DITA — including the hard cases other vendors turn away: scanned PDFs, decades-old PDFs with no styling, and mixed legacy archives.

FrameMaker · MadCap Flare · RoboHelp · MS Word · InDesign · HTML & Drupal · Markdown · DocBook · legacy XML · scanned & unstyled PDF — delivered as valid DITA with a working ditamap, imported and accepted in your CCMS.

Quality does not come from throwing documents at an AI. It comes from an AI-enabled conversion pipeline: deterministic rules built for your content do the heavy lifting — repeatable, validatable, cost-effective — and AI is applied precisely where it adds value: recognizing structure in unstructured pages and enriching your content with metadata.

What we migrate

Every source. One destination: DITA.

We focus on one thing and do it exceptionally well: migrating customer legacy content — structured or unstructured — into clean, topic-based DITA your team can maintain for the next twenty years.

Structured & semi-structured sources

Authoring tools and help systems with recoverable structure. Deterministic converters map every heading, procedure, table, note and cross-reference to the right DITA element — identically, every run.

The hard cases — scanned & legacy PDF

Where most conversion vendors stop, we start. Scanned page images, very old PDFs with no styling information, print-era manuals, mixed archives — we rebuild the structure itself: OCR and AI-assisted layout recognition, conversion rules developed for your specific documents, and human review on every deliverable.

Valid, not just converted

Every topic validates against the DITA DTDs. Links, cross-references and tables are resolved and preserved — with validation reports included in the delivery.

Architecture, not just files

You receive a working ditamap with keys, relationship tables and a reuse strategy — ready to import into any DITA CCMS or publishing pipeline.

Accountable to the end

The migration is finished when your content is imported, valid and accepted in your CCMS — Heretto, Paligo, IXIASOFT, Tridion Docs or any DITA-compliant system — not when files land in a folder.

In-depth guides: DITA conversion · DITA migration · DITA transformation · browse all conversion guides →

Our approach

Deterministic where it matters. AI where it helps.

Pure-AI conversion is impressive on a demo and unsustainable in production: every page costs tokens, every run can differ, and nobody can certify the output. Pure rules break down on unstructured sources. We engineered the combination.

🔧 Deterministic core

  • Conversion rules developed for your content patterns
  • Identical output on every run — fixes are permanent
  • DTD validation, link and conref resolution built in
  • Zero AI cost on the pages rules handle — that is most of them

✨ AI where rules can't reach

  • Structure recognition on scanned and unstyled pages
  • Topic-type detection: concept, task, reference
  • Terms, keywords and semantic metadata enrichment
  • Every AI decision reviewed, cached and auditable

Analyze

We study your corpus, identify the patterns, and build the conversion rules and mappings for your documents specifically.

Convert

The AI-enabled pipeline runs: deterministic converters first, AI on the genuinely ambiguous remainder — quality with cost control.

Review

Our review team and yours iterate together. Every fix becomes a permanent rule — the same issue never returns.

Deliver

Valid DITA, working ditamap, validation report — imported and accepted in your CCMS. That's when we call it done.

Read the full article: How we combine AI and deterministic conversion →

From any source to intelligent DITA

One pipeline, every legacy format

Transparent engagement

See it on your content first. 25 pages, free.

A migration is a trust purchase — so we earn the trust before we ask for the budget. No generic rate card, no commitment: judge us on your own documents.

  1. 1

    Send up to 25 pages

    Any format — a scanned manual, an old PDF, FrameMaker books, Word files. The messier, the better.

  2. 2

    We convert them to DITA — free

    Through the same pipeline we use in production: rules, AI assist where needed, validation.

  3. 3

    Review the output together

    You see exactly what your DITA will look like: topics, ditamap, metadata, validation report.

  4. 4

    Pricing, one-to-one

    If you like what you see, we shape a pricing model around your volume, formats and timeline — together.

The platform behind the service

A real conversion platform, not a folder of scripts

Every migration runs on the DocentraX conversion portal — the same platform your team can use for ongoing conversions after the migration.

Modular conversion pipelines

Every converter is a versioned module with full job logs, validation reports and repeatable results — multi-step pipelines chain conversions in one job.

Knowledge that compounds

Terms, acronyms, product names and metadata are learned per customer and reused on every subsequent project — the AI share of the work falls as your knowledge base grows, and so does the cost.

Enterprise-grade isolation

Multi-tenant with strict per-organization isolation, role-based access, audit logging and cloud-native scaling — your content never mixes with anyone else's.

After the migration

Documentation that reports on itself

Because your content passes through our knowledge pipeline, it arrives smarter than it left: the platform can tell you which topics are duplicated, which glossary terms conflict, which procedures are inconsistent across manuals, and which documents are outdated — across your whole corpus.

News & insights

From our team

All news & articles →

Who we work with

Wherever legacy content meets high stakes

Technical publications Manufacturing manuals Aerospace documentation Medical device documentation Government publications Educational content Product documentation Knowledge bases Training material
Leadership

The expertise behind every migration

Lief Erickson

Lief Erickson

Co-founder and Principal Strategist

With more than 25 years in the content industry, Lief is the founder of Intuitive Stack, a content strategy consultancy for businesses with outdated technical documentation practices. He holds a master’s degree in Content Strategy from FH Joanneum in Austria, where he also teaches information architecture. A technical writer and information architect by background, his expertise spans taxonomies, metadata, search optimization, and ContentOps, and he speaks frequently on information architecture, metadata, search, and AI-driven content strategy. His focus is helping organizations modernize their technical documentation and content operations so they can focus on their next innovation.

FAQ

Questions we hear from every migration team

Which source formats can you convert to DITA?

Structured sources: FrameMaker, MadCap Flare, RoboHelp, MS Word, InDesign, HTML and Drupal exports, Markdown, DocBook and legacy XML. Unstructured sources: PDFs of any age — including scanned PDFs and very old PDFs with no styling information — plus mixed legacy archives. If your format isn't listed, ask: the pipeline is modular by design.

Can you convert scanned or image-only PDFs to DITA?

Yes — scanned and unstyled legacy PDFs are our specialty. We combine OCR and AI-assisted structure recognition with deterministic conversion rules developed for your specific content, then a human review loop, so the delivered DITA is valid, consistent and CCMS-ready.

How do you keep AI-assisted conversion reliable — and affordable?

Deterministic rules do the bulk of the work, so results are repeatable and most pages cost no AI tokens at all. AI is applied only where rules cannot decide: recognizing structure in unstructured pages, resolving ambiguity, enriching metadata. Every output is validated against the DITA DTDs and reviewed before delivery. That combination is what keeps quality high and the price cost-effective — read the full article.

How much does DITA conversion cost?

Start with the free sample: send up to 25 pages and we convert them at no charge. Once you've seen the output quality on your own content, we discuss a pricing model one-to-one that fits your volume, formats and timeline. No generic per-page rate card, no surprises.

Which CCMS platforms do you deliver into?

Any DITA-compliant CCMS — including Heretto, Paligo, IXIASOFT, Tridion Docs and Oxygen Content Fusion — or a plain DITA repository. We stay accountable until your content is imported, valid and accepted in your target system.

Is the delivered DITA actually valid?

Yes. Every topic is validated against the DITA DTDs, links and cross-references are resolved, tables and images preserved, and you receive a working ditamap architecture — not a folder of loose files. Validation reports are part of the delivery.

Ready to see your content in DITA?

Tell us what you're sitting on — FrameMaker books, a shelf of scanned manuals, a wiki export — and we'll tell you honestly what the migration takes. The first 25 pages are on us.

Talk to us

Tell us about your content challenge — we usually reply within one business day.

[email protected]
[email protected]