Knowledge Fabric is DocentraX's content enrichment and metadata automation layer — it discovers the terms, acronyms, products and relationships buried in your DITA and writes real keywords straight into topic prologs. Once your content is converted, migrated or transformed into clean DITA, Knowledge Fabric turns a repository of documents into an evolving enterprise knowledge asset that gets sharper every time a human touches it.
Once your legacy content is clean, single-sourced DITA, a bigger opportunity opens up. Every topic you converted is full of knowledge that is invisible to your tools: the product names, the acronyms, the component relationships, the terms that mean something very specific in your domain. A converted repository is a searchable pile of files — it can tell you which topic contains a string, but it cannot tell you that a product's internal code name, its marketing name and the abbreviation your engineers use are all the same thing, that one component belongs to three products, or that two topics describe the same procedure in different words.
The cost of leaving that gap open is quiet but constant: terminology that drifts across a corpus, a glossary maintained by hand and always slightly out of date, duplicated topics nobody knew were duplicates, and every departing expert taking undocumented knowledge with them. Knowledge Fabric captures that knowledge as structured, versioned assets that belong to your organization — so it accumulates instead of evaporating, and your metadata is generated from the content itself rather than from a stretched team's spare time.
Every human decision becomes permanent organizational knowledge — the corpus learns instead of forgetting.
Discovery comes first, and it is deterministic: the same content yields the same terms, acronyms, products, components and relationships every time, and none of your content is ever shipped to an AI to be read — discovery is analysis, not prompting. What it finds is held as real, versioned assets that belong to your organization alone — a glossary, an acronym list, a product and component catalog, a knowledge graph, and search that understands meaning rather than matching characters. Every human decision layered on top of that — an approved term, a corrected definition, two topics confirmed as duplicates — is remembered, so a judgment your reviewer makes once is never asked of them again. AI sits as an optional last layer for the genuinely ambiguous calls only — shown distilled candidates, never your repository — never as the whole engine, and never as the thing standing between your content and your metadata.
Those capabilities drive a concrete set of outcomes. You can enrich topics by writing discovered keywords directly into their DITA prologs; export the glossary as CSV or as DITA glossentry topics; search your corpus by meaning and surface genuinely related topics; detect duplicates; and follow a term's whole history — when it first became a recognized term, how its definition has evolved, and which topics each change touches. It reads DITA, Markdown, HTML, XML and plain text, and it works best over the deduplicated masters that Content Convergence produces — reasoning over one canonical copy of each thing rather than a dozen scattered near-copies. That is what makes enrichment metadata automation rather than metadata data-entry: the DITA keywords are discovered from your content and written back into it as real metadata.
| Your concern | How Knowledge Fabric answers it |
|---|---|
| Will our content be sent to an AI? | Not unless you switch the optional AI layer on — and even then it sees only distilled candidate terms, never your topics. Its answers are kept, so the same question is never asked twice. |
| Our glossary is always out of date. | Terminology is discovered from the content itself and governed centrally, and the CSV and DITA glossentry exports stay in step with the corpus instead of drifting away from it. |
| We already have a vocabulary. | A customer-specific configuration seeds Knowledge Fabric with your existing terms, catalog and acronym expansions, so it speaks your language on the first run instead of proposing one you already have. |
| Adding metadata by hand never gets done. | Discovered keywords are written into topic prologs automatically, as real DITA metadata inside the content rather than a list somebody maintains beside it. |
| How does it improve over time? | Every human correction is kept permanently, and the share of decisions that needs AI at all is measured run by run — it demonstrably falls as the knowledge base grows, so accuracy rises while cost goes down, and the improvement belongs to your organization. |
| We re-ingest after every release. | Re-ingest is incremental: unchanged topics are recognized and skipped, so a full corpus pass after a small change takes seconds, and concepts already decided never re-enter the pipeline. |
A generic run discovers terms, acronyms and entities from your content and gives you a starting glossary and a meaning-aware index — immediately useful, but built from scratch. A customer-specific configuration seeds Knowledge Fabric with your existing vocabulary, your product and component catalog, your acronym expansions and your preferred canonical terms, and tunes the judgment calls to your domain — so what it produces speaks your organization's language on the very first run.
Take acronyms in a medical-device corpus. A generic run will correctly flag "UDI" as a frequent acronym and offer an expansion inferred from context. A customer-specific run already knows UDI means Unique Device Identifier in your regulatory sense, knows which products it applies to, places it in the knowledge graph beside your labeling components, and writes the approved keyword into every relevant topic's prolog — then remembers any correction your reviewer makes, so it is never made twice. Without that tailoring, your team spends its review effort reconciling machine-discovered terms against the vocabulary it already had.
Terminology and metadata are the parts of a DITA program teams most want and least manage to sustain. The usual approach is a glossary spreadsheet somebody owns until they leave, index terms added by hand when there is time, and taxonomy work that stalls the moment the initial project ends. There is no feedback loop, so quality decays. Knowledge Fabric answers that decay directly: discovery is deterministic, so it is repeatable — and private, because your content is never sent to an AI; AI is an optional layer over the hard cases, not the whole engine, and the share of the work it does is measured and shrinks; and human judgment is captured rather than spent, so the corpus gets sharper every time somebody touches it. Because it works over the normalized, deduplicated content from Content Convergence and plugs into the same pipeline as every DocentraX converter and each DITA conversion we run, it is an intelligence layer over your whole content operation, not a disconnected taxonomy tool.
With Knowledge Fabric in place, your converted content stops being a static deliverable and becomes a living asset. Terminology is consistent because it is discovered and governed centrally. Glossaries and prolog keywords maintain themselves and flow back into the topics as real DITA metadata. Writers find related topics and catch duplicates through search that understands meaning, instead of relying on who happens to remember what. Institutional knowledge accumulates in a versioned store that belongs to your organization instead of walking out of the door with the people who held it. This is the payoff of the whole pipeline: legacy content converted, cleaned, normalized and validated — and then enriched into an enterprise knowledge layer that makes every downstream use, from single-sourced publishing to onboarding to search, measurably smarter.
Content enrichment is the process of discovering the terms, acronyms, products and relationships in your content and writing them back as real metadata — keywords in topic prologs, plus a governed glossary. Knowledge Fabric automates this across the whole corpus rather than leaving it to manual tagging.
Yes. Discovered keywords are written directly into each topic's prolog as real DITA metadata — merged with what is already there, never duplicated — so the content carries its own metadata instead of depending on hand tagging that never quite gets done.
Not by default. Discovery is deterministic and sends nothing to an AI. An optional AI layer sees only distilled candidate terms — never your repository — for the genuinely ambiguous cases, and its answers are kept, so the same question is never asked twice.
It discovers terminology from the content itself and governs it centrally as a versioned glossary, acronym list and product and component catalog. Every human correction is kept permanently, so the vocabulary becomes more accurate over time instead of decaying between projects.
Yes. The glossary exports as CSV or as DITA glossentry topics, and the same knowledge drives search by meaning, related topics, duplicate detection and term history across your corpus.
We'll convert them to DITA free of charge — through the real pipeline, not a demo — and review the output with you. Then we'll discuss pricing one-to-one.
Request your free sample conversion