Tens of thousands of pages, migrated from DITA to AsciiDoc

A telecommunications technology company needed its full documentation library moved out of DITA and into a docs-as-code toolchain — without a multi-year project or a drop in quality.

pages migrated
Tens of thousands
single source format
DITA → AsciiDoc
human-reviewed output
100%

Scale that ruled out a manual approach

The library had grown over many years in DITA: deep topic hierarchies, conditional profiling, thousands of cross-references and a long tail of inconsistent markup. Manual conversion at that volume was not viable, and a plain script would have left too many broken links and malformed tables for writers to fix by hand.

Scripts and AI for the volume, writers for the judgement

Tools used

DITA AsciiDoc Python conversion scripts LLM-assisted cleanup Git CI publishing
  1. 01

    Analysed the DITA source to classify content patterns and estimate where automated conversion would hold and where it would break.

  2. 02

    Built a scripted conversion pipeline from DITA to AsciiDoc, preserving structure, conditional content and cross-references.

  3. 03

    Used LLM-assisted cleanup for the residual problems scripts could not resolve: irregular tables, inline markup drift and inconsistent terminology.

  4. 04

    Routed every converted page through writer review with a defined checklist before it entered the new Git repository.

  5. 05

    Set up an AsciiDoc build and CI publishing pipeline, and trained the in-house team on the new workflow.

Delivered, and handed over

The complete library now lives in plain-text AsciiDoc under version control, publishes automatically and can be edited by engineers and writers alike. The combination of scripted conversion, AI-assisted cleanup and human review delivered the migration in a fraction of the time a manual approach would have taken.

Next story: 40+ training modules in two weeks →

Have a similar project? Let's talk.

Contact us