Pipeline Overview
GitHub Actions can watch your documentation repository and automatically convert any newly committed PDFs, Word docs, PowerPoints, or HTML files into Markdown. The converted output lands in the repo (or an artifacts bucket) so your AI pipelines always consume clean text.
- Trigger: `pull_request` or scheduled workflow.
- Conversion: Call the MDConvert API with the changed files.
- Validation: Run linting, front matter checks, and token count reports.
- Publishing: Commit Markdown back or upload to S3/GCS for downstream jobs.
Sample Workflow File
Drop this workflow into .github/workflows/markdown-pipeline.yml:
name: Markdown Conversion Pipeline
on:
pull_request:
paths:
- 'source-docs/**'
workflow_dispatch: {}
jobs:
convert-to-markdown:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Set up Node
uses: actions/setup-node@v4
with:
node-version: '20'
- name: Install dependencies
run: npm install --global @mdconvert/cli
- name: Convert documents to Markdown
env:
MDCONVERT_API_KEY: ${{ secrets.MDCONVERT_API_KEY }}
run: |
mdconvert batch --input ./source-docs --output ./markdown --format markdown --metadata yes
- name: Lint Markdown
run: npx markdownlint ./markdown/**/*.md
- name: Upload artifacts
uses: actions/upload-artifact@v4
with:
name: markdown-output
path: ./markdown
Customize the CLI command to include OCR, table preservation, or persona-specific templates. If you prefer raw API calls, swap the CLI step for a `curl` script.
Hardening the Workflow
- Secrets Management: Store API keys in GitHub Secrets; limit scope to conversion operations.
- Diff Review: Commit Markdown output in a separate branch so reviewers can inspect diffs before merging.
- Pre-merge Gates: Require Markdown lint, PII scan, and minimum coverage of metadata fields.
- Notifications: Use `actions/github-script` to comment on pull requests with token savings and conversion summary.
- Scheduled Refresh: Add a weekly cron trigger to reconvert everything so knowledge stays current.
Pair this with a deployment step that pushes Markdown to your vector database refresh job or static knowledge site. When docs update, your AI systems follow automatically.