Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
72 changes: 68 additions & 4 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,8 @@ curl -i -H "Authorization: Bearer $CRON_SECRET" localhost:3000/api/prompts
| -------------------- | ---------------------------------------------------------- |
| `app/api/*/route.ts` | HTTP surface: prompts, results, cron, webhook, mcp |
| `lib/runner.ts` | Scheduling: which prompts are due, submit, sweep |
| `lib/extract.ts` | Pure derivation: links and brand matches from a response |
| `lib/refresh.ts` | Writes the derived tables, a batch per tick |
| `lib/cloro.ts` | The only place that calls the cloro API |
| `lib/engines.ts` | Engine slugs, task types, per-engine payload shape |
| `lib/webhooks.ts` | Callback URL, token derivation, signature checks |
Expand Down Expand Up @@ -123,9 +125,71 @@ submit twice. Keep that update atomic.
- Errors return `{ "error": { "message": ... } }`. Use `withErrors`.
- Prefer adding to an existing `lib/` module over creating a new one.

## Derived tables

`result_sources` and `result_brand_mentions` are computed from
`results.response` and can always be thrown away and rebuilt. Three rules
hold them together:

**Store the misses, not only the hits.** `result_brand_mentions` gets a row
for every completed result and every enabled brand, mentioned or not.
Share of voice needs a denominator, and a table of hits alone cannot show
the difference between "never named" and "never asked".

**A brand edit reopens the whole history.** Adding a brand, renaming one,
or changing its aliases or domains changes what the extractor would have
produced for answers that already arrived. Every write path that touches
those fields calls `markAllForReextraction()`. Skip it and the new brand's
chart begins on the day somebody remembered to add it, which reads as a
brand that appeared from nowhere. `isOwn` is exempt: it is a label the
extractor never reads.

**The refresh runs last in the tick and may stop early.** Submissions are
time-sensitive; this is not. It is bounded by a batch size and a time
budget, and the leftover work stays queued in `results.extraction_revision`
for the next tick. Bump `EXTRACTION_REVISION` when the extraction rules
change, and the whole history is re-derived on its own.

Extraction runs in the app, not in the database. Neon's free tier has no
`pg_cron`, so a materialised view would have nothing to refresh it.

`lib/brand-candidates.json` is a fourth extraction input, and the only one
that is a FILE rather than a table. Nothing can call
`markAllForReextraction()` when a file changes, so `EXTRACTION_STAMP`
folds the sorted candidate list into the value written to
`results.extraction_revision`. An edited file simply stops matching what is
stored and the next tick re-derives on its own. That column holds a
fingerprint of the extraction inputs, not a version number.

`scripts/seed.mjs` fills a local database with synthetic answers so the
Grafana panels can be built without waiting a month for real data. It
writes prompts, brands and raw results, and derives nothing: run the tick
afterwards and the app fills the derived tables through the code that runs
in production. Prompts are seeded disabled, because an enabled prompt is
due the moment it exists and the tick would submit it to the real API.

## Out of scope

Do not add a web UI, a login system, or an analysis layer that scores or
classifies answers. The product stores raw answers and lets agents
interpret them. Keep the dependency list small: this has to stay free to
run on a hobby plan.
Do not add a web UI or a login system. Keep the dependency list small:
this has to stay free to run on a hobby plan.

**Do not add anything that scores, ranks or judges an answer.** The
product stores raw answers and lets agents interpret them.

The brand extraction added in `lib/extract.ts` is the one thing near that
line, and it stays on the safe side by being mechanical: the user declares
the brands, and the code does literal case-insensitive matching and
hostname comparison. It decides _whether a name is present_, never how
good an answer is, who is winning, or which brands are worth tracking. A
sentiment score, a quality grade, a recommendation, or a built-in list of
competitors would all cross it.

<!-- BEGIN:nextjs-agent-rules -->

# This is NOT the Next.js you know

This version has breaking changes — APIs, conventions, and file structure may all differ from your training data. Read the relevant guide in `node_modules/next/dist/docs/` (resolved from this file's directory; in monorepos the `next` package may not be visible from the repo root) before writing any code. Heed deprecation notices.

This block is written and re-added by `next dev` — verify at `node_modules/next/dist/server/lib/generate-agent-files.js`. Removing it from a diff only re-creates the uncommitted change; committing it with your work keeps the tree clean.

<!-- END:nextjs-agent-rules -->
120 changes: 105 additions & 15 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,8 @@ That one sentence is a `create_prompt` call, a `run_prompt` call and a
add prompts, run them and query the answers without a human in the loop.
- **Your data, queryable** — plain Postgres, so an agent can also read it
with SQL, and the [Grafana starter](./grafana/README.md) charts it.
- **$0 to run** — fits in Vercel's free tier, database included.
- **$0 to run** — fits in Vercel's free tier, database included, until the
derived tables outgrow it (see [Data & retention](#data--retention)).
- **Fully async** — scrapes are submitted as
[cloro async tasks](https://cloro.dev/docs) and results come back by
webhook, so no serverless function ever waits on a scrape.
Expand Down Expand Up @@ -167,6 +168,11 @@ All endpoints except the webhook require
| `DELETE` | `/api/prompts/:id` | Delete a prompt and (cascade) its results |
| `POST` | `/api/prompts/:id/run` | Submit the prompt to its engines now; returns pending task ids (202) |
| `GET` | `/api/results` | Query results (filters below) |
| `GET` | `/api/brands` | List tracked brands |
| `POST` | `/api/brands` | Track a brand (`name`, `aliases[]`, `domains[]`, `isOwn`, `enabled`) |
| `GET` | `/api/brands/:id` | Get one brand |
| `PATCH` | `/api/brands/:id` | Update any subset of the brand fields |
| `DELETE` | `/api/brands/:id` | Delete a brand and (cascade) its mention rows |
| `GET` | `/api/cron` | Scheduler tick — same bearer token |
| `POST` | `/api/webhook` | cloro result callback — auth via token in the callback URL |
| `*` | `/api/mcp` | MCP endpoint (Streamable HTTP) |
Expand All @@ -185,13 +191,17 @@ The MCP endpoint is the primary way to use geo-tracker. It is not a
read-only reporting layer: an agent can create prompts, trigger runs and
pull the stored answers, which is the whole product surface.

| Tool | What the agent can do |
| --------------- | ----------------------------------------------------- |
| `list_prompts` | See what is being tracked, and when each last ran |
| `create_prompt` | Add a prompt, pick the engines, set how often it runs |
| `run_prompt` | Run one now instead of waiting for the schedule |
| `get_results` | Query runs by prompt, engine or status |
| `get_result` | Pull one full raw engine answer for analysis |
| Tool | What the agent can do |
| ---------------------- | ----------------------------------------------------- |
| `list_prompts` | See what is being tracked, and when each last ran |
| `create_prompt` | Add a prompt, pick the engines, set how often it runs |
| `run_prompt` | Run one now instead of waiting for the schedule |
| `get_results` | Query runs by prompt, engine or status |
| `get_result` | Pull one full raw engine answer for analysis |
| `list_brands` | See which brands are being looked for |
| `track_brand` | Start looking for a brand, with aliases and domains |
| `untrack_brand` | Stop looking for one, and drop its derived rows |
| `get_brand_visibility` | How often each brand was named, and cited |

Things worth asking an agent once it is connected:

Expand Down Expand Up @@ -253,12 +263,80 @@ The same tick also sweeps: pending results whose webhook was missed are
polled from the cloro API and backfilled, so nothing is lost if a webhook
delivery fails.

## Brand visibility

Tell geo-tracker which brands to look for, and every answer is flattened
into two tables you can query or chart:

```bash
curl -X POST https://<your-app>.vercel.app/api/brands \
-H "Authorization: Bearer $CRON_SECRET" \
-H "content-type: application/json" \
-d '{"name":"Acme","aliases":["Acme Corp"],"domains":["acme.io"],"isOwn":true}'
```

- `result_sources` — one row per link an engine returned, tagged by where
it came from (`source`, `citation_pill`, `organic`, `ad`, …). This is
the "which pages get cited" question.
- `result_brand_mentions` — one row per answer per brand, including the
brands that were **not** named. That is what makes share of voice
computable: a brand at 0% has rows saying so, rather than being absent.

Being **named** in the prose and being **cited** as a link are stored
separately, because they are different outcomes — an answer can recommend
you without linking you, or link you without naming you.

Adding or editing a brand re-scores every answer already stored, so a
brand you add today has full history rather than starting at zero. The
work happens in the scheduler tick, a batch at a time; the API response
tells you how many results are queued.

**A deployment with history takes a while to catch up.** The tick derives
250 results, and Vercel's Hobby plan runs one cron a day — so re-deriving
a year of answers would take months of ticks. Two ways round it, and the
tick is idempotent so either is safe:

- Point an external scheduler (GitHub Actions, cron-job.org) at
`/api/cron` every few minutes until `more` comes back `false`.
- Or call it by hand in a loop:

```bash
while curl -s -H "Authorization: Bearer $CRON_SECRET" \
https://<your-app>.vercel.app/api/cron | grep -q '"more":true'; do :; done
```

Nothing is missing while it runs. The old rows stay until each result is
rebuilt, so the dashboard shows values that are stale, never blank.

`result_search_queries` holds a third thing: the literal queries the
engines typed before retrieving anything. ChatGPT, Copilot, Grok and
Perplexity report these; the others do not.

To watch a competitor without tracking it, add its name to
`lib/brand-candidates.json`, with any alternative spellings:

```json
{ "name": "Acme", "aliases": ["Acme, Inc", "Acme Corp"] }
```

Names there that turn up in answers, and that you are not tracking,
surface as a shortlist worth adding. The file ships with fictional
placeholders to replace — geo-tracker does not guess who competes with
you, and it cannot find a brand nobody wrote down.

Nothing here scores or ranks an answer. It records whether a name is
present. What that means is the agent's call.

## Grafana

A ready-made dashboard (run volume, success rate, credits burn, failures)
lives in [`grafana/`](./grafana/README.md) — point Grafana Cloud's free
tier at your Postgres and import one JSON file.

Grafana runs outside Vercel: it is a long-running server, and Vercel hosts
serverless functions. Grafana Cloud's free tier reads your database
directly over TLS, which is all this needs.

## Local development

Contributing with a coding agent? [`AGENTS.md`](./AGENTS.md) documents the
Expand Down Expand Up @@ -286,17 +364,29 @@ so just hit `/api/cron` again after a scrape completes.

## Data & retention

Two tables: `prompts` and `results`. Each run stores one row per engine
with the full raw cloro response as `jsonb` — measured at **10–30 KB per
row**, so budget roughly 20 KB per engine per run.
Two tables hold what you asked and what came back: `prompts` and
`results`. Each run stores one row per engine with the full raw cloro
response as `jsonb` — measured at **10–30 KB per row**, so budget roughly
20 KB per engine per run.

Storage is the limit you hit first, well before anything on Vercel:
Four more are derived from those responses by the scheduler tick
(`result_sources`, `result_brand_mentions`, `result_search_queries`,
`result_candidate_mentions`). They add about **5 KB per answer**, so
budget ~25 KB per engine per run in total. That figure scales with how
many links an answer carries, not with how big its payload is, so a chatty
engine costs more here than a terse one with a large HTML blob.

Storage is the limit you hit first, well before anything on Vercel. The
figures below include the derived rows:

| Workload | Scrapes/day | Storage/month | 0.5 GB lasts |
| ------------------------------ | ----------- | ------------- | ------------ |
| 10 prompts × 3 engines × 1/day | 30 | ~18 MB | over 2 years |
| 20 prompts × 4 engines × 4/day | 320 | ~190 MB | ~3 months |
| 50 prompts × 6 engines × 8/day | 2,400 | ~1.4 GB | ~2 weeks |
| 10 prompts × 3 engines × 1/day | 30 | ~23 MB | ~1.8 years |
| 20 prompts × 4 engines × 4/day | 320 | ~240 MB | ~2 months |
| 50 prompts × 6 engines × 8/day | 2,400 | ~1.8 GB | ~8 days |

Deleting a result cascades to its derived rows, so a retention policy
needs no extra step.

The same workloads use under 1%, 1% and 7% of Vercel's free monthly
function invocations, so the compute side stays free throughout.
Expand Down
78 changes: 78 additions & 0 deletions app/api/brands/[id]/route.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,78 @@
import { eq } from "drizzle-orm";

import { isApiKeyAuthorized, unauthorized } from "@/lib/auth";
import { getDb } from "@/lib/db";
import { brands } from "@/lib/db/schema";
import {
HttpError,
isUniqueViolation,
parseOr400,
withErrors,
} from "@/lib/http";
import { affectsExtraction, markAllForReextraction } from "@/lib/refresh";
import { idSchema, updateBrandSchema } from "@/lib/validation";

export const runtime = "nodejs";

function parseId(id: string): string {
const parsed = idSchema.safeParse(id);
if (!parsed.success) throw new HttpError(400, "Invalid brand id");
return parsed.data;
}

export const GET = withErrors(async (req, { params }) => {
if (!isApiKeyAuthorized(req)) return unauthorized();
const id = parseId((await params).id);
const db = getDb();
const [row] = await db.select().from(brands).where(eq(brands.id, id));
if (!row) throw new HttpError(404, "Brand not found");
return Response.json({ brand: row });
});

export const PATCH = withErrors(async (req, { params }) => {
if (!isApiKeyAuthorized(req)) return unauthorized();
const id = parseId((await params).id);
const input = parseOr400(
updateBrandSchema,
await req.json().catch(() => null),
);

const db = getDb();
let updated;
try {
[updated] = await db
.update(brands)
.set(input)
.where(eq(brands.id, id))
.returning();
} catch (error) {
if (isUniqueViolation(error)) {
throw new HttpError(409, `A brand named "${input.name}" already exists`);
}
throw error;
}
if (!updated) throw new HttpError(404, "Brand not found");

// Only re-derive when the edit changes what the extractor reads. Renaming
// a brand or adding an alias rewrites its whole history; flipping `isOwn`
// changes a label the extractor never looks at.
const queued = affectsExtraction(input) ? await markAllForReextraction() : 0;

return Response.json({ brand: updated, queuedForExtraction: queued });
});

export const DELETE = withErrors(async (req, { params }) => {
if (!isApiKeyAuthorized(req)) return unauthorized();
const id = parseId((await params).id);
const db = getDb();
const deleted = await db
.delete(brands)
.where(eq(brands.id, id))
.returning({ id: brands.id });
if (deleted.length === 0) throw new HttpError(404, "Brand not found");

// No re-extraction: the foreign key cascades this brand's mention rows
// away, and no other brand's rows depend on it. Deleting is the one brand
// change that costs nothing.
return Response.json({ deleted: true });
});
55 changes: 55 additions & 0 deletions app/api/brands/route.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,55 @@
import { asc } from "drizzle-orm";

import { isApiKeyAuthorized, unauthorized } from "@/lib/auth";
import { getDb } from "@/lib/db";
import { brands } from "@/lib/db/schema";
import {
HttpError,
isUniqueViolation,
parseOr400,
withErrors,
} from "@/lib/http";
import { markAllForReextraction } from "@/lib/refresh";
import { createBrandSchema } from "@/lib/validation";

export const runtime = "nodejs";

export const GET = withErrors(async (req) => {
if (!isApiKeyAuthorized(req)) return unauthorized();
const db = getDb();
const rows = await db.select().from(brands).orderBy(asc(brands.name));
return Response.json({ brands: rows });
});

export const POST = withErrors(async (req) => {
if (!isApiKeyAuthorized(req)) return unauthorized();
const input = parseOr400(
createBrandSchema,
await req.json().catch(() => null),
);

const db = getDb();
let created;
try {
[created] = await db.insert(brands).values(input).returning();
} catch (error) {
// The unique index is on lower(name), so "Acme" and "acme" collide.
// That is deliberate — they are one brand — but the raw driver error
// does not say so.
if (isUniqueViolation(error)) {
throw new HttpError(409, `A brand named "${input.name}" already exists`);
}
throw error;
}

// A new brand has to be scored against answers that already arrived, not
// only the ones still to come. Without this its chart would begin on the
// day somebody remembered to add it, which reads as a brand that appeared
// out of nowhere rather than one we started watching late.
const queued = await markAllForReextraction();

return Response.json(
{ brand: created, queuedForExtraction: queued },
{ status: 201 },
);
});
Loading
Loading