Ask questions about the Lord of the Rings screenplays. Entirely on your own machine.
Retrieval-augmented generation over all three film scripts, powered by
qwen2.5:14b through Ollama and ChromaDB.
No API keys, no cloud inference, no data leaving your laptop.
Most RAG tutorials wire up an OpenAI key, chunk a PDF every 500 characters, and call it done. This project is the opposite of that on both counts:
- Nothing leaves your machine. Generation and embeddings both run locally through Ollama. The only network request the app ever makes is downloading the scripts once.
- Chunking respects the document. Screenplays have obvious structural seams — scene headings — and splitting on those instead of blind character windows is what makes the citations meaningful. See Design decisions.
It's a small, readable codebase: seven modules, no LangChain, no framework indirection. If you want to understand what a RAG pipeline actually does, you can read the whole thing in one sitting.
Note
This block is illustrative, not a recorded session. The shape is accurate — streamed answer, then ranked sources with similarity scores and script position — but the wording is an example. Your answers will differ; local models are not deterministic across runs.
$ uv run lotr-rag chat
One index to find them, one model to answer them.
LOTR screenplay RAG -- qwen2.5:14b + nomic-embed-text via Ollama
725 chunks indexed. /help for commands.
you > How is Boromir's death staged in the script?
rag > The screenplay stages it as a slow collapse rather than a single blow. Boromir is
struck repeatedly while defending the hobbits, and the action lines emphasise that he keeps
fighting between hits [The Fellowship of the Ring - EXT. AMON HEN - DAY]. Aragorn arrives
only after the fatal wound, and the scene shifts registers from battle to a quiet exchange
between the two men before Boromir dies [The Fellowship of the Ring - EXT. AMON HEN - DAY].
sources: The Fellowship of the Ring / EXT. AMON HEN - DAY; The Fellowship of the Ring / ...
you > /sources
[1] The Fellowship of the Ring - EXT. AMON HEN - DAY (score 0.784, 91% in)
[2] The Fellowship of the Ring - EXT. AMON HEN - DAY (score 0.761, 92% in)
[3] The Two Towers - EXT. THE RIVER - DAY (score 0.634, 3% in)
you > /film two-towers
Restricted to The Two Towers.Requirements: Python 3.11+, uv, and Ollama.
qwen2.5:14b is roughly a 9 GB download and wants 16 GB+ of RAM to be comfortable.
# 1. Pull the models (once)
ollama pull qwen2.5:14b
ollama pull nomic-embed-text
ollama serve # skip if Ollama is already running
# 2. Install dependencies
git clone https://github.com/NicholasRyan/lotr-rag.git
cd lotr-rag
uv sync
# 3. Download, chunk and embed the screenplays (~2-5 min on Apple silicon)
uv run lotr-rag ingest
# 4. Ask it things
uv run lotr-rag chatCheck your setup at any point with uv run lotr-rag status — it reports which models Ollama
has pulled, which scripts are cached, and how many chunks are indexed per film.
qwen2.5:14b is the default, not a requirement. Every model is environment-configurable:
LOTR_CHAT_MODEL=qwen2.5:7b uv run lotr-rag chat # ~4.7 GB instead of ~9 GB
LOTR_CHAT_MODEL=llama3.1:8b uv run lotr-rag chat # different family entirelySwapping the chat model is free — it doesn't touch the index. Swapping the embedding
model requires uv run lotr-rag ingest --rebuild, because the stored vectors and your query
vectors have to come from the same model.
| Command | What it does |
|---|---|
lotr-rag ingest |
Download, chunk and embed all three screenplays |
lotr-rag ingest --refresh |
Re-download the scripts even if cached |
lotr-rag ingest --rebuild |
Clear the collection and re-embed from scratch |
lotr-rag chat |
Interactive chat loop with conversation memory |
lotr-rag ask "<question>" |
One-shot question, prints answer and sources, exits |
lotr-rag ask "<q>" --film two-towers |
Restrict retrieval to a single film |
lotr-rag status |
Models, cached scripts, index state |
lotr-rag reset |
Delete the vector store (cached scripts are kept) |
Inside chat:
| Command | Effect |
|---|---|
/sources |
Full ranked sources for the last answer, with scores |
/film <name> |
Restrict retrieval — fellowship, two-towers, return-of-the-king, or all |
/k <n> |
Change how many chunks are retrieved per question |
/clear |
Forget the conversation history |
/help, /quit |
Self-explanatory |
Films accept aliases, so --film fotr, --film ttt and --film rotk all work.
IMSDb pages
│
│ scrape.py fetch with a browser UA, pull <td class="scrtext">,
▼ normalise whitespace, cache to data/raw/*.txt
plain text
│
│ chunk.py split on sluglines (EXT. WEATHERTOP - NIGHT),
▼ then window long scenes with overlap
Chunk(text, film, scene, scene_index, position)
│
│ store.py nomic-embed-text ──▶ ChromaDB (cosine, persistent)
▼
data/chroma/ ~725 chunks
│
▼
question ──▶ embed ──▶ top-k chunks ──▶ rag.py prompt ──▶ qwen2.5:14b ──▶ answer + citations
Each chunk carries its film, scene heading, scene index, and position — how far through the
script it falls, which is what produces the 91% in marker in the sources list.
The parts worth arguing about, and why they landed where they did.
A fixed 500-character window will happily cut through the middle of a line of dialogue and
staple it to an unrelated stage direction. For a screenplay that's a waste of obvious
structure: sluglines like EXT. WEATHERTOP - NIGHT mark exactly where one unit of action ends
and the next begins.
So chunk.py splits on sluglines first, then windows any scene too long to fit in one chunk.
The practical payoff is that a retrieved chunk almost always sits inside a single scene, which
makes a citation like [The Two Towers - EXT. FANGORN FOREST - DAY] a real locator rather
than decoration. It also gives the embedding model a coherent unit of meaning instead of a
fragment spanning two locations.
The slugline regex handles INT., EXT., INT./EXT., I/E, optional scene numbers, and a
fallback pattern for all-caps location lines ending in a time of day, which some transcripts
use instead.
This one bit during development and is worth writing down, because it's the failure mode that doesn't announce itself.
ChromaDB 1.x has two entry points on an embedding function: __call__ when indexing documents,
and embed_query when searching. The split exists so models trained with asymmetric task
prefixes can use them. nomic-embed-text supports exactly such prefixes
(search_document: / search_query:), so it's tempting to add the query prefix now that
there's a hook for it.
Don't — not unless you rebuild the index with the matching document prefix. Prefixing only one side produces no error at all, just quietly worse retrieval. Both paths here route through the same call, and there's a test asserting they return identical vectors for identical input.
Chunk IDs are deterministic (fellowship:00042) and writes are upserts, so re-running ingest
refreshes in place rather than silently doubling your index. Downloaded scripts are cached to
data/raw/, so only the very first run touches the network — iterating on chunking parameters
costs you embedding time, not bandwidth.
Every error path tries to tell you what to actually do: a missing model prints the exact
ollama pull command, an unreachable server prints the host it tried and suggests
ollama serve, an empty index points you at ingest, and a scrape failure names the file to
drop in by hand. The scraper also refuses suspiciously short downloads rather than cheerfully
embedding an error page.
The system prompt tells the model to answer only from retrieved context, to say so when the context doesn't cover the question, to cite film and scene per claim, and to distinguish films from books. That last one matters more than you'd expect: a 14B model has plenty of latent Tolkien knowledge and will happily blend book canon into an answer about the screenplay unless you push back on it.
Everything is environment-configurable, with the defaults in config.py:
| Variable | Default | Purpose |
|---|---|---|
OLLAMA_HOST |
http://localhost:11434 |
Ollama server address |
LOTR_CHAT_MODEL |
qwen2.5:14b |
Generation model |
LOTR_EMBED_MODEL |
nomic-embed-text |
Embedding model (rebuild index if changed) |
LOTR_CHUNK_SIZE |
1400 |
Chunk size in characters |
LOTR_CHUNK_OVERLAP |
200 |
Overlap between windows of a long scene |
LOTR_TOP_K |
8 |
Chunks retrieved per question |
LOTR_NUM_CTX |
8192 |
Model context window |
LOTR_TEMPERATURE |
0.2 |
Low, because this is a factual retrieval task |
LOTR_HISTORY_TURNS |
4 |
Conversation turns kept in chat |
LOTR_DATA_DIR |
./data |
Where scripts and the index live |
LOTR_TIMEOUT |
300 |
Request timeout, seconds (first 14B load is slow) |
The pipeline isn't LOTR-specific — only SOURCES in config.py is. Add a Source entry with
a key, title, year and url, and it'll be scraped, chunked and indexed like the rest.
Leave url blank and drop a data/raw/<key>.txt file in yourself, and the scraper is skipped
entirely — which is the easiest way to run this over any plain-text corpus you already have.
Non-screenplay text still works; you just lose the scene-heading structure and fall back to
plain overlapping windows.
src/lord_of_the_rings_movie/
config.py paths, models, chunk sizes, source definitions — all env-overridable
ollama_client.py /api/embed + streaming /api/chat, with actionable error messages
scrape.py IMSDb fetch, <pre> extraction, whitespace normalisation
chunk.py slugline detection, scene splitting, overlap windowing
store.py ChromaDB collection + the Ollama embedding function
rag.py context assembly, prompt construction, grounded answer loop
cli.py ingest / chat / ask / status / reset
tests/
test_pipeline.py offline tests — no Ollama, no network, no index required
Roughly 900 lines total, dependencies limited to httpx, chromadb and beautifulsoup4.
uv run pytest # everything
uv run pytest -v # per-test names
uv run pytest tests/test_store.py60 tests, and every one runs offline — no Ollama server, no network, no prebuilt index. That's possible because the embedding client is injected rather than constructed internally, so a deterministic fake replaces the single external dependency wholesale. CI runs the same suite on Python 3.11 and 3.12.
| Module | Covered by | What it pins down |
|---|---|---|
chunk.py |
test_pipeline.py |
Scene splitting, chunk metadata, ID uniqueness, long-scene windowing |
scrape.py |
test_pipeline.py |
HTML extraction, the longest-<pre> fallback, whitespace normalisation |
store.py |
test_store.py |
Embedding-function contract, upsert idempotency, film filtering, reset |
rag.py |
test_rag.py |
Context assembly, citation labels, history truncation, empty-retrieval fallback |
cli.py |
test_cli.py |
All five subcommands, alias resolution, the pre-flight guard paths |
The test worth reading is test_query_and_document_vectors_are_identical. Chroma calls
__call__ to index and embed_query to search; making those two paths disagree produces
no error at all, just quietly worse retrieval. It is the one regression in this codebase
that a normal test suite would never catch, so it gets an explicit assertion.
| Symptom | Cause and fix |
|---|---|
Could not reach Ollama at ... |
Server isn't running. ollama serve, or set OLLAMA_HOST. |
Missing Ollama model(s) |
The error prints the exact ollama pull command. |
The vector store is empty |
Run uv run lotr-rag ingest. |
AttributeError: ... has no attribute 'embed_query' |
An older copy of store.py against ChromaDB 1.x. Fixed in current code — see Design decisions. |
Downloaded only N characters |
IMSDb's layout changed. Save that script as text to data/raw/<key>.txt and re-run ingest. |
| Answers ignore the retrieved context | Lower LOTR_TEMPERATURE, or raise LOTR_TOP_K if retrieval is returning too little to work with. |
| First question takes forever | Ollama is loading 9 GB of weights. Subsequent questions are much faster. |
No screenplay text is included in this repository, and none should be committed to it.
The scripts are third-party material hosted on IMSDb. ingest downloads
them to data/raw/ on your own machine for personal study, and data/ is gitignored — both
the cached scripts and the vector index stay local. If you fork this, keep it that way: ship
the pipeline, not the corpus.
- Retrieval is pure dense similarity. Hybrid search (BM25 + embeddings) would handle exact name lookups better — dense retrieval is comparatively weak at "find the scene where X is said verbatim".
- No reranker. A cross-encoder pass over the top-k would improve precision at the cost of another model in the loop.
- Conversation history isn't used to rewrite queries. Follow-ups like "what about the other one?" retrieve against the literal question, so they underperform. Query rewriting from history is the obvious next step.
- Scene detection is regex-based and depends on the transcript being reasonably formatted.
- No evaluation harness. A small labelled question set with retrieval hit-rate would make chunking changes measurable rather than vibes-based.
MIT — © 2026 Nicholas Ryan.
The licence covers this code. It does not cover the screenplays, which are not mine to relicense and are not in the repository.