Not all legacy content comes from a tidy tool project — sometimes it is just HTML, and that is exactly what our general-purpose HTML to DITA conversion is built for. DocentraX covers conversion, migration and transformation in one pipeline, turning unstructured, hand-written or exported HTML into clean, valid DITA by reading the topic structure your own headings already imply, so even orphaned content becomes single-sourced, reusable and multi-channel.
Unstructured HTML is content with no future beyond a browser window. You cannot reuse it, single-source it, publish it to a new channel or manage it in a CCMS, because there is nothing structured there to manage. It sits in a folder, valuable but inert, quietly going stale. And it is almost always the content nobody wants to touch — which is exactly why it never gets migrated, and exactly why it keeps costing you.
Doing it by hand is the least attractive migration of all: no project to lean on, inconsistent markup written by several authors in several eras, and no navigation to recover the reading order from. A deterministic HTML to DITA conversion that derives structure from the content itself is what makes this material movable at all, and what turns it from a permanent backlog item into a realistic line in your DITA migration services plan.
The content nobody wants to migrate is often the content you can least afford to lose. This is how it moves.
This is the path for HTML that did not come from a documentation tool at all: hand-written pages, an export from a system that no longer exists, a folder inherited with an acquisition. There is no table of contents to lean on, so reading order comes from the way your content is actually organised — we walk your folders in the order they are laid out, following each branch to the bottom before moving on, so the sequence a reader would take through the directory is the sequence they get through the DITA map.
Topic structure comes from your own headings — and, when your pages have none, from the styling that stood in for them. A page whose headings step down from one level to the next describes a hierarchy whether or not anyone ever wrote it down, and that hierarchy becomes a real DITA topic tree. The hard case — the one generic converters fail — is HTML whose hierarchy lives only in presentation: every "heading" is just a paragraph set larger, bolder, capitalised or centred. Our structure inference reads exactly that evidence, together with enumeration patterns like chapter and part numbering, and rebuilds a genuine heading outline from it; documents that already have some headings are only gap-filled, existing headings are never rewritten, and how deep the inferred outline may go is tunable, so a shallow set of pages does not come back over-nested and a deep manual does not come back flat. From there tables become CALS tables that survive translation and print, and notes, cross-references and images are rebuilt as proper elements rather than markup lookalikes; where a customer-specific profile identifies your procedure styling, those sections arrive as DITA tasks with real steps. Inconsistent formatting is normalised: the decorative markup, stray spacing and one-off styling that accumulate in hand-written pages are resolved into consistent semantics instead of being carried forward. Content conservation is absolute — structure degrades gracefully where it cannot be inferred, but no text is ever dropped.
Every HTML to DITA migration runs the same five-stage lifecycle.
Establish what you actually have: an inventory of pages, headings, images, recurring styles and links, plus a reading order derived from how the files are organised.
Turn heading levels — real or inferred from styling — into a DITA topic hierarchy, with tables to CALS and notes, cross-references and images rebuilt as real elements; a customer-specific profile maps your procedure styling to tasks with real steps.
Apply your metadata, keys, IDs and naming conventions, and optionally run Knowledge Fabric to write keywords and index terms into topic prologs, so content this neglected becomes searchable again.
Run the DITA completeness check before delivery: DTD conformance against the DITA grammars or your own specialization, resolved links, conrefs and keyrefs, verified image references, duplicate ids and keys caught, circular map chains detected.
Hand over one clean composite DITA document per title — or, when your platform wants topic-per-file, run the splitting step for individual typed topics with a bookmap and key map — with assets and references intact.
A generic run gives you valid DITA with inferred structure, generic topic types and generated IDs — correct, but shaped only by what the raw markup revealed. It can also produce a conversion report listing every recurring style no mapping covers: the concrete punch list a customer-specific profile is built from. A customer-specific conversion aligns it to your model. With a source this raw the gap between the two is at its widest, and closing it is where the value sits:
Because this tailoring is configuration rather than bespoke code written from scratch each time, it applies to the next batch as readily as the first. Without it the work does not vanish — it lands on your writers, who turn inferred structure into your structure by hand, one file at a time.
| Your concern | How we answer it |
|---|---|
| There is no project or navigation to work from. | Structure is inferred from your own headings — or, where there are none, from the font sizes, emphasis, capitalisation and numbering that stood in for them — and reading order from the way your files are organised. |
| The markup is inconsistent from one file to the next. | Recurring styles and patterns are normalised into consistent DITA semantics rather than copied across. |
| This is the content nobody wants to touch. | The conversion is deterministic and repeatable, so nobody hand-migrates page by page. |
| We cannot afford to lose anything from these old files. | Content conservation is zero-tolerance: structure may degrade gracefully, text is never dropped. |
If your HTML actually came from a documentation tool, the tool-specific conversion understands that tool's output — its navigation where the export carries one, its export format and templates where it does not — and produces a better result than any general-purpose path can. Use InDesign to DITA conversion for InDesign output, Sphinx to DITA conversion for Sphinx documentation sites, or Drupal to DITA conversion for Drupal exports. This general-purpose route exists precisely for the HTML that fits none of them — and in most estates that turns out to be a surprising amount of it.
DITA gives you single-sourcing, conref and keyref reuse, conditional publishing and multi-channel output through the DITA Open Toolkit — once content is migrated cleanly. Unstructured HTML is the hardest source of all, because there is no structure to recover, only structure to infer. Teams normally re-key it, abandon it, or run a scraper that produces one flat topic per page and calls that a migration. DocentraX infers a real topic hierarchy from your headings, derives reading order from how your content is actually organised, and applies the same content-conservation guarantee and repository-scale validation as every other source we handle — one DITA conversion pipeline, not a point tool bolted on the side.
Your most neglected content becomes standards-based DITA: single-sourced, reusable through conref and keyref, conditionally publishable, and deliverable to any channel through DITA-OT. The folder of forgotten HTML turns into a governed content asset like everything else — validated against your own schema, addressable, and finally earning the space it takes up.
The general-purpose HTML to DITA conversion handles any HTML that did not come from a documentation tool. Reading order is derived from the way your files are organised, topic nesting is inferred from the heading levels in the pages themselves, and DITA semantics are rebuilt properly — CALS tables, notes, cross-references and figures, plus tasks where a customer-specific profile identifies your procedure styling — before the output is validated against your own DTD.
Yes. Loose HTML has no table of contents, so reading order comes from how your content is organised in folders and topic nesting is derived from your heading levels — or, for pages with no heading markup at all, from an outline inferred from the styling and numbering that implied one. How deep that hierarchy goes is tuned to your content, so the result matches how your material is genuinely structured rather than following an arbitrary default.
Use the general-purpose HTML to DITA conversion when your HTML is not from a documentation tool. If it came from InDesign, Sphinx, Drupal, Flare, RoboHelp or FrameMaker, the source-specific conversion understands that tool's own output — following its navigation where the export carries one, and recognising its export format and templates where it does not — so start from the matching page instead.
No. Content conservation is a zero-tolerance rule: where structure cannot be confidently inferred it degrades gracefully, but visible text is never dropped. The completeness check verifies links, references and image references before anything is delivered.
It accepts .htm, .html and .xhtml pages, or a single .zip of your HTML together with the images it references, and returns one clean composite DITA document per title — or, through the splitting step, individual typed topics with a bookmap and key map in the structure your platform expects.
We'll convert them to DITA free of charge — through the real pipeline, not a demo — and review the output with you. Then we'll discuss pricing one-to-one.
Request your free sample conversion