Skip to content
NicholasRyanPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Lord of the Rings — Local RAG

Ask questions about the Lord of the Rings screenplays. Entirely on your own machine.

Retrieval-augmented generation over all three film scripts, powered by qwen2.5:14b through Ollama and ChromaDB. No API keys, no cloud inference, no data leaving your laptop.

tests Python Ollama Embeddings Vector store Runs offline License


Why this exists

Most RAG tutorials wire up an OpenAI key, chunk a PDF every 500 characters, and call it done. This project is the opposite of that on both counts:

  • Nothing leaves your machine. Generation and embeddings both run locally through Ollama. The only network request the app ever makes is downloading the scripts once.
  • Chunking respects the document. Screenplays have obvious structural seams — scene headings — and splitting on those instead of blind character windows is what makes the citations meaningful. See Design decisions.

It's a small, readable codebase: seven modules, no LangChain, no framework indirection. If you want to understand what a RAG pipeline actually does, you can read the whole thing in one sitting.


Demo

Note

This block is illustrative, not a recorded session. The shape is accurate — streamed answer, then ranked sources with similarity scores and script position — but the wording is an example. Your answers will differ; local models are not deterministic across runs.

$ uv run lotr-rag chat

  One index to find them, one model to answer them.
  LOTR screenplay RAG -- qwen2.5:14b + nomic-embed-text via Ollama

  725 chunks indexed. /help for commands.

you > How is Boromir's death staged in the script?

rag > The screenplay stages it as a slow collapse rather than a single blow. Boromir is
struck repeatedly while defending the hobbits, and the action lines emphasise that he keeps
fighting between hits [The Fellowship of the Ring - EXT. AMON HEN - DAY]. Aragorn arrives
only after the fatal wound, and the scene shifts registers from battle to a quiet exchange
between the two men before Boromir dies [The Fellowship of the Ring - EXT. AMON HEN - DAY].

  sources: The Fellowship of the Ring / EXT. AMON HEN - DAY; The Fellowship of the Ring / ...

you > /sources
  [1] The Fellowship of the Ring - EXT. AMON HEN - DAY  (score 0.784, 91% in)
  [2] The Fellowship of the Ring - EXT. AMON HEN - DAY  (score 0.761, 92% in)
  [3] The Two Towers - EXT. THE RIVER - DAY             (score 0.634, 3% in)

you > /film two-towers
Restricted to The Two Towers.

Quickstart

Requirements: Python 3.11+, uv, and Ollama. qwen2.5:14b is roughly a 9 GB download and wants 16 GB+ of RAM to be comfortable.

# 1. Pull the models (once)
ollama pull qwen2.5:14b
ollama pull nomic-embed-text
ollama serve                  # skip if Ollama is already running

# 2. Install dependencies
git clone https://github.com/NicholasRyan/lotr-rag.git
cd lotr-rag
uv sync

# 3. Download, chunk and embed the screenplays (~2-5 min on Apple silicon)
uv run lotr-rag ingest

# 4. Ask it things
uv run lotr-rag chat

Check your setup at any point with uv run lotr-rag status — it reports which models Ollama has pulled, which scripts are cached, and how many chunks are indexed per film.

Running on a smaller machine

qwen2.5:14b is the default, not a requirement. Every model is environment-configurable:

LOTR_CHAT_MODEL=qwen2.5:7b uv run lotr-rag chat     # ~4.7 GB instead of ~9 GB
LOTR_CHAT_MODEL=llama3.1:8b uv run lotr-rag chat    # different family entirely

Swapping the chat model is free — it doesn't touch the index. Swapping the embedding model requires uv run lotr-rag ingest --rebuild, because the stored vectors and your query vectors have to come from the same model.


Usage

Command What it does
lotr-rag ingest Download, chunk and embed all three screenplays
lotr-rag ingest --refresh Re-download the scripts even if cached
lotr-rag ingest --rebuild Clear the collection and re-embed from scratch
lotr-rag chat Interactive chat loop with conversation memory
lotr-rag ask "<question>" One-shot question, prints answer and sources, exits
lotr-rag ask "<q>" --film two-towers Restrict retrieval to a single film
lotr-rag status Models, cached scripts, index state
lotr-rag reset Delete the vector store (cached scripts are kept)

Inside chat:

Command Effect
/sources Full ranked sources for the last answer, with scores
/film <name> Restrict retrieval — fellowship, two-towers, return-of-the-king, or all
/k <n> Change how many chunks are retrieved per question
/clear Forget the conversation history
/help, /quit Self-explanatory

Films accept aliases, so --film fotr, --film ttt and --film rotk all work.


How it works

  IMSDb pages
      │
      │  scrape.py      fetch with a browser UA, pull <td class="scrtext">,
      ▼                 normalise whitespace, cache to data/raw/*.txt
  plain text
      │
      │  chunk.py       split on sluglines (EXT. WEATHERTOP - NIGHT),
      ▼                 then window long scenes with overlap
  Chunk(text, film, scene, scene_index, position)
      │
      │  store.py       nomic-embed-text ──▶ ChromaDB (cosine, persistent)
      ▼
  data/chroma/  ~725 chunks
      │
      ▼
  question ──▶ embed ──▶ top-k chunks ──▶ rag.py prompt ──▶ qwen2.5:14b ──▶ answer + citations

Each chunk carries its film, scene heading, scene index, and position — how far through the script it falls, which is what produces the 91% in marker in the sources list.


Design decisions

The parts worth arguing about, and why they landed where they did.

Chunking follows scene boundaries, not character counts

A fixed 500-character window will happily cut through the middle of a line of dialogue and staple it to an unrelated stage direction. For a screenplay that's a waste of obvious structure: sluglines like EXT. WEATHERTOP - NIGHT mark exactly where one unit of action ends and the next begins.

So chunk.py splits on sluglines first, then windows any scene too long to fit in one chunk. The practical payoff is that a retrieved chunk almost always sits inside a single scene, which makes a citation like [The Two Towers - EXT. FANGORN FOREST - DAY] a real locator rather than decoration. It also gives the embedding model a coherent unit of meaning instead of a fragment spanning two locations.

The slugline regex handles INT., EXT., INT./EXT., I/E, optional scene numbers, and a fallback pattern for all-caps location lines ending in a time of day, which some transcripts use instead.

Query and document embeddings must be identical

This one bit during development and is worth writing down, because it's the failure mode that doesn't announce itself.

ChromaDB 1.x has two entry points on an embedding function: __call__ when indexing documents, and embed_query when searching. The split exists so models trained with asymmetric task prefixes can use them. nomic-embed-text supports exactly such prefixes (search_document: / search_query:), so it's tempting to add the query prefix now that there's a hook for it.

Don't — not unless you rebuild the index with the matching document prefix. Prefixing only one side produces no error at all, just quietly worse retrieval. Both paths here route through the same call, and there's a test asserting they return identical vectors for identical input.

Ingestion is idempotent

Chunk IDs are deterministic (fellowship:00042) and writes are upserts, so re-running ingest refreshes in place rather than silently doubling your index. Downloaded scripts are cached to data/raw/, so only the very first run touches the network — iterating on chunking parameters costs you embedding time, not bandwidth.

Failures should be actionable

Every error path tries to tell you what to actually do: a missing model prints the exact ollama pull command, an unreachable server prints the host it tried and suggests ollama serve, an empty index points you at ingest, and a scrape failure names the file to drop in by hand. The scraper also refuses suspiciously short downloads rather than cheerfully embedding an error page.

The prompt is deliberately strict

The system prompt tells the model to answer only from retrieved context, to say so when the context doesn't cover the question, to cite film and scene per claim, and to distinguish films from books. That last one matters more than you'd expect: a 14B model has plenty of latent Tolkien knowledge and will happily blend book canon into an answer about the screenplay unless you push back on it.


Configuration

Everything is environment-configurable, with the defaults in config.py:

Variable Default Purpose
OLLAMA_HOST http://localhost:11434 Ollama server address
LOTR_CHAT_MODEL qwen2.5:14b Generation model
LOTR_EMBED_MODEL nomic-embed-text Embedding model (rebuild index if changed)
LOTR_CHUNK_SIZE 1400 Chunk size in characters
LOTR_CHUNK_OVERLAP 200 Overlap between windows of a long scene
LOTR_TOP_K 8 Chunks retrieved per question
LOTR_NUM_CTX 8192 Model context window
LOTR_TEMPERATURE 0.2 Low, because this is a factual retrieval task
LOTR_HISTORY_TURNS 4 Conversation turns kept in chat
LOTR_DATA_DIR ./data Where scripts and the index live
LOTR_TIMEOUT 300 Request timeout, seconds (first 14B load is slow)

Pointing it at a different corpus

The pipeline isn't LOTR-specific — only SOURCES in config.py is. Add a Source entry with a key, title, year and url, and it'll be scraped, chunked and indexed like the rest. Leave url blank and drop a data/raw/<key>.txt file in yourself, and the scraper is skipped entirely — which is the easiest way to run this over any plain-text corpus you already have. Non-screenplay text still works; you just lose the scene-heading structure and fall back to plain overlapping windows.


Project layout

src/lord_of_the_rings_movie/
  config.py          paths, models, chunk sizes, source definitions — all env-overridable
  ollama_client.py   /api/embed + streaming /api/chat, with actionable error messages
  scrape.py          IMSDb fetch, <pre> extraction, whitespace normalisation
  chunk.py           slugline detection, scene splitting, overlap windowing
  store.py           ChromaDB collection + the Ollama embedding function
  rag.py             context assembly, prompt construction, grounded answer loop
  cli.py             ingest / chat / ask / status / reset
tests/
  test_pipeline.py   offline tests — no Ollama, no network, no index required

Roughly 900 lines total, dependencies limited to httpx, chromadb and beautifulsoup4.


Tests

uv run pytest              # everything
uv run pytest -v           # per-test names
uv run pytest tests/test_store.py

60 tests, and every one runs offline — no Ollama server, no network, no prebuilt index. That's possible because the embedding client is injected rather than constructed internally, so a deterministic fake replaces the single external dependency wholesale. CI runs the same suite on Python 3.11 and 3.12.

Module Covered by What it pins down
chunk.py test_pipeline.py Scene splitting, chunk metadata, ID uniqueness, long-scene windowing
scrape.py test_pipeline.py HTML extraction, the longest-<pre> fallback, whitespace normalisation
store.py test_store.py Embedding-function contract, upsert idempotency, film filtering, reset
rag.py test_rag.py Context assembly, citation labels, history truncation, empty-retrieval fallback
cli.py test_cli.py All five subcommands, alias resolution, the pre-flight guard paths

The test worth reading is test_query_and_document_vectors_are_identical. Chroma calls __call__ to index and embed_query to search; making those two paths disagree produces no error at all, just quietly worse retrieval. It is the one regression in this codebase that a normal test suite would never catch, so it gets an explicit assertion.

Troubleshooting

Symptom Cause and fix
Could not reach Ollama at ... Server isn't running. ollama serve, or set OLLAMA_HOST.
Missing Ollama model(s) The error prints the exact ollama pull command.
The vector store is empty Run uv run lotr-rag ingest.
AttributeError: ... has no attribute 'embed_query' An older copy of store.py against ChromaDB 1.x. Fixed in current code — see Design decisions.
Downloaded only N characters IMSDb's layout changed. Save that script as text to data/raw/<key>.txt and re-run ingest.
Answers ignore the retrieved context Lower LOTR_TEMPERATURE, or raise LOTR_TOP_K if retrieval is returning too little to work with.
First question takes forever Ollama is loading 9 GB of weights. Subsequent questions are much faster.

About the screenplays

No screenplay text is included in this repository, and none should be committed to it.

The scripts are third-party material hosted on IMSDb. ingest downloads them to data/raw/ on your own machine for personal study, and data/ is gitignored — both the cached scripts and the vector index stay local. If you fork this, keep it that way: ship the pipeline, not the corpus.


Limitations and future work

  • Retrieval is pure dense similarity. Hybrid search (BM25 + embeddings) would handle exact name lookups better — dense retrieval is comparatively weak at "find the scene where X is said verbatim".
  • No reranker. A cross-encoder pass over the top-k would improve precision at the cost of another model in the loop.
  • Conversation history isn't used to rewrite queries. Follow-ups like "what about the other one?" retrieve against the literal question, so they underperform. Query rewriting from history is the obvious next step.
  • Scene detection is regex-based and depends on the transcript being reasonably formatted.
  • No evaluation harness. A small labelled question set with retrieval hit-rate would make chunking changes measurable rather than vibes-based.

License

MIT — © 2026 Nicholas Ryan.

The licence covers this code. It does not cover the screenplays, which are not mine to relicense and are not in the repository.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages