DocentraX converts structured and unstructured legacy content into topic-based, CCMS-ready DITA — including the hard cases other vendors turn away: scanned PDFs, decades-old PDFs with no styling, and mixed legacy archives.
FrameMaker · MadCap Flare · RoboHelp · MS Word · InDesign · HTML & Drupal · Markdown · DocBook · legacy XML · scanned & unstyled PDF — delivered as valid DITA with a working ditamap, imported and accepted in your CCMS.
Quality does not come from throwing documents at an AI. It comes from an AI-enabled conversion pipeline: deterministic rules built for your content do the heavy lifting — repeatable, validatable, cost-effective — and AI is applied precisely where it adds value: recognizing structure in unstructured pages and enriching your content with metadata.
We focus on one thing and do it exceptionally well: migrating customer legacy content — structured or unstructured — into clean, topic-based DITA your team can maintain for the next twenty years.
Authoring tools and help systems with recoverable structure. Deterministic converters map every heading, procedure, table, note and cross-reference to the right DITA element — identically, every run.
Where most conversion vendors stop, we start. Scanned page images, very old PDFs with no styling information, print-era manuals, mixed archives — we rebuild the structure itself: OCR and AI-assisted layout recognition, conversion rules developed for your specific documents, and human review on every deliverable.
Every topic validates against the DITA DTDs. Links, cross-references and tables are resolved and preserved — with validation reports included in the delivery.
You receive a working ditamap with keys, relationship tables and a reuse strategy — ready to import into any DITA CCMS or publishing pipeline.
The migration is finished when your content is imported, valid and accepted in your CCMS — Heretto, Paligo, IXIASOFT, Tridion Docs or any DITA-compliant system — not when files land in a folder.
In-depth guides: DITA conversion · DITA migration · DITA transformation · browse all conversion guides →
Pure-AI conversion is impressive on a demo and unsustainable in production: every page costs tokens, every run can differ, and nobody can certify the output. Pure rules break down on unstructured sources. We engineered the combination.
We study your corpus, identify the patterns, and build the conversion rules and mappings for your documents specifically.
The AI-enabled pipeline runs: deterministic converters first, AI on the genuinely ambiguous remainder — quality with cost control.
Our review team and yours iterate together. Every fix becomes a permanent rule — the same issue never returns.
Valid DITA, working ditamap, validation report — imported and accepted in your CCMS. That's when we call it done.
Read the full article: How we combine AI and deterministic conversion →
The pipeline doesn't just move your content — it structures it: topics are recognized and split, the ditamap is architected for reuse, and every topic arrives validated, with terms, keywords and semantic metadata already in place.
A migration is a trust purchase — so we earn the trust before we ask for the budget. No generic rate card, no commitment: judge us on your own documents.
Any format — a scanned manual, an old PDF, FrameMaker books, Word files. The messier, the better.
Through the same pipeline we use in production: rules, AI assist where needed, validation.
You see exactly what your DITA will look like: topics, ditamap, metadata, validation report.
If you like what you see, we shape a pricing model around your volume, formats and timeline — together.
Every migration runs on the DocentraX conversion portal — the same platform your team can use for ongoing conversions after the migration.
Every converter is a versioned module with full job logs, validation reports and repeatable results — multi-step pipelines chain conversions in one job.
Terms, acronyms, product names and metadata are learned per customer and reused on every subsequent project — the AI share of the work falls as your knowledge base grows, and so does the cost.
Multi-tenant with strict per-organization isolation, role-based access, audit logging and cloud-native scaling — your content never mixes with anyone else's.
Because your content passes through our knowledge pipeline, it arrives smarter than it left: the platform can tell you which topics are duplicated, which glossary terms conflict, which procedures are inconsistent across manuals, and which documents are outdated — across your whole corpus.
With more than 25 years in the content industry, Lief is the founder of Intuitive Stack, a content strategy consultancy for businesses with outdated technical documentation practices. He holds a master’s degree in Content Strategy from FH Joanneum in Austria, where he also teaches information architecture. A technical writer and information architect by background, his expertise spans taxonomies, metadata, search optimization, and ContentOps, and he speaks frequently on information architecture, metadata, search, and AI-driven content strategy. His focus is helping organizations modernize their technical documentation and content operations so they can focus on their next innovation.
Structured sources: FrameMaker, MadCap Flare, RoboHelp, MS Word, InDesign, HTML and Drupal exports, Markdown, DocBook and legacy XML. Unstructured sources: PDFs of any age — including scanned PDFs and very old PDFs with no styling information — plus mixed legacy archives. If your format isn't listed, ask: the pipeline is modular by design.
Yes — scanned and unstyled legacy PDFs are our specialty. We combine OCR and AI-assisted structure recognition with deterministic conversion rules developed for your specific content, then a human review loop, so the delivered DITA is valid, consistent and CCMS-ready.
Deterministic rules do the bulk of the work, so results are repeatable and most pages cost no AI tokens at all. AI is applied only where rules cannot decide: recognizing structure in unstructured pages, resolving ambiguity, enriching metadata. Every output is validated against the DITA DTDs and reviewed before delivery. That combination is what keeps quality high and the price cost-effective — read the full article.
Start with the free sample: send up to 25 pages and we convert them at no charge. Once you've seen the output quality on your own content, we discuss a pricing model one-to-one that fits your volume, formats and timeline. No generic per-page rate card, no surprises.
Any DITA-compliant CCMS — including Heretto, Paligo, IXIASOFT, Tridion Docs and Oxygen Content Fusion — or a plain DITA repository. We stay accountable until your content is imported, valid and accepted in your target system.
Yes. Every topic is validated against the DITA DTDs, links and cross-references are resolved, tables and images preserved, and you receive a working ditamap architecture — not a folder of loose files. Validation reports are part of the delivery.
Tell us what you're sitting on — FrameMaker books, a shelf of scanned manuals, a wiki export — and we'll tell you honestly what the migration takes. The first 25 pages are on us.
Tell us about your content challenge — we usually reply within one business day.