This directory contains tutorials focused on Docling, IBM's open-source toolkit for parsing and converting documents from a wide range of formats.
Docling is an IBM open-source library for parsing documents and exporting them to preferred formats. It supports:
- Input formats: PDF, DOCX, PPTX, XLSX, Images, HTML, AsciiDoc, Markdown
- Output formats: Markdown, JSON
- OCR support: Optical character recognition for scanned documents
Common use cases include processing medical records, banking documents, travel documents, and any scenario where unstructured data needs to be converted into a machine-readable structured format.
Each tutorial in this directory includes its own setup and installation instructions. Please refer to the individual tutorial files for specific requirements.
Common requirements:
- Python 3.10 - 3.13
- IBM watsonx.ai account
- Install dependencies (see Installation above)
- Navigate to this directory:
cd tutorials/15-docling - Open and run the tutorials
Use Docling with Python to convert unstructured data contained in a group of scanned files into a structured format.
- Author: Erika Russi
- Topics: Document parsing, unstructured-to-structured conversion, OCR, Docling export formats
- Prerequisites: RAG dependencies (
requirements-rag.txt) - Estimated time: 30–40 minutes
The following Docling tutorials are located in the 01-rag-and-retrieval category:
Build a document-based question answering system using Docling and IBM Granite 3.1.
- Topics: Document parsing, Q&A systems, Docling + Ollama integration, large context windows
- Prerequisites: RAG dependencies + Ollama
Advanced RAG with reasoning capabilities using DeepSeek and Docling on watsonx.
- Topics: Reasoning over documents, advanced retrieval, Docling integration, watsonx
- Prerequisites: RAG dependencies + Ollama
- Structured data: Information organized into a fixed field within a record or file (SQL databases, JSON, XML, CSV, Excel spreadsheets). Ready for efficient data processing, analysis, and management.
- Unstructured data: Information that does not conform to a predefined data model or schema. Typically text-heavy (emails, social media posts, customer reviews) or non-text formats (audio recordings, video files, images). Makes up ~90% of enterprise information.
- Analysis automation: Run real-time queries and analytics
- Machine interpretability: Algorithms can process structured data directly
- Integration: Feed into databases, pipelines, and downstream AI systems
- Document Loading: Ingest files (PDF, DOCX, images, etc.)
- Parsing: Extract text, tables, and layout information
- OCR (optional): Recognize text in scanned images
- Export: Output as Markdown or JSON for downstream use
- Healthcare: Scan and structure medical records and lab reports
- Finance: Process banking documents, invoices, and statements
- Legal: Extract structured data from contracts and filings
- Logistics: Parse shipping documents and travel records
- RAG pipelines: Pre-process documents before embedding and retrieval
- Docling GitHub Repository
- IBM Think: Unstructured Data
- IBM Think: Structured vs. Unstructured Data
- IBM Think: OCR
- Main Repository README
- RAG and Retrieval Tutorials
Found an issue or want to add a new Docling tutorial? Please open an issue or pull request in the GitHub repository.
See the LICENSE file in the repository root for license information.