diff --git a/AGENTS.md b/AGENTS.md index ac980ae..285f12d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -68,6 +68,8 @@ curl -i -H "Authorization: Bearer $CRON_SECRET" localhost:3000/api/prompts | -------------------- | ---------------------------------------------------------- | | `app/api/*/route.ts` | HTTP surface: prompts, results, cron, webhook, mcp | | `lib/runner.ts` | Scheduling: which prompts are due, submit, sweep | +| `lib/extract.ts` | Pure derivation: links and brand matches from a response | +| `lib/refresh.ts` | Writes the derived tables, a batch per tick | | `lib/cloro.ts` | The only place that calls the cloro API | | `lib/engines.ts` | Engine slugs, task types, per-engine payload shape | | `lib/webhooks.ts` | Callback URL, token derivation, signature checks | @@ -123,9 +125,71 @@ submit twice. Keep that update atomic. - Errors return `{ "error": { "message": ... } }`. Use `withErrors`. - Prefer adding to an existing `lib/` module over creating a new one. +## Derived tables + +`result_sources` and `result_brand_mentions` are computed from +`results.response` and can always be thrown away and rebuilt. Three rules +hold them together: + +**Store the misses, not only the hits.** `result_brand_mentions` gets a row +for every completed result and every enabled brand, mentioned or not. +Share of voice needs a denominator, and a table of hits alone cannot show +the difference between "never named" and "never asked". + +**A brand edit reopens the whole history.** Adding a brand, renaming one, +or changing its aliases or domains changes what the extractor would have +produced for answers that already arrived. Every write path that touches +those fields calls `markAllForReextraction()`. Skip it and the new brand's +chart begins on the day somebody remembered to add it, which reads as a +brand that appeared from nowhere. `isOwn` is exempt: it is a label the +extractor never reads. + +**The refresh runs last in the tick and may stop early.** Submissions are +time-sensitive; this is not. It is bounded by a batch size and a time +budget, and the leftover work stays queued in `results.extraction_revision` +for the next tick. Bump `EXTRACTION_REVISION` when the extraction rules +change, and the whole history is re-derived on its own. + +Extraction runs in the app, not in the database. Neon's free tier has no +`pg_cron`, so a materialised view would have nothing to refresh it. + +`lib/brand-candidates.json` is a fourth extraction input, and the only one +that is a FILE rather than a table. Nothing can call +`markAllForReextraction()` when a file changes, so `EXTRACTION_STAMP` +folds the sorted candidate list into the value written to +`results.extraction_revision`. An edited file simply stops matching what is +stored and the next tick re-derives on its own. That column holds a +fingerprint of the extraction inputs, not a version number. + +`scripts/seed.mjs` fills a local database with synthetic answers so the +Grafana panels can be built without waiting a month for real data. It +writes prompts, brands and raw results, and derives nothing: run the tick +afterwards and the app fills the derived tables through the code that runs +in production. Prompts are seeded disabled, because an enabled prompt is +due the moment it exists and the tick would submit it to the real API. + ## Out of scope -Do not add a web UI, a login system, or an analysis layer that scores or -classifies answers. The product stores raw answers and lets agents -interpret them. Keep the dependency list small: this has to stay free to -run on a hobby plan. +Do not add a web UI or a login system. Keep the dependency list small: +this has to stay free to run on a hobby plan. + +**Do not add anything that scores, ranks or judges an answer.** The +product stores raw answers and lets agents interpret them. + +The brand extraction added in `lib/extract.ts` is the one thing near that +line, and it stays on the safe side by being mechanical: the user declares +the brands, and the code does literal case-insensitive matching and +hostname comparison. It decides _whether a name is present_, never how +good an answer is, who is winning, or which brands are worth tracking. A +sentiment score, a quality grade, a recommendation, or a built-in list of +competitors would all cross it. + + + +# This is NOT the Next.js you know + +This version has breaking changes — APIs, conventions, and file structure may all differ from your training data. Read the relevant guide in `node_modules/next/dist/docs/` (resolved from this file's directory; in monorepos the `next` package may not be visible from the repo root) before writing any code. Heed deprecation notices. + +This block is written and re-added by `next dev` — verify at `node_modules/next/dist/server/lib/generate-agent-files.js`. Removing it from a diff only re-creates the uncommitted change; committing it with your work keeps the tree clean. + + diff --git a/README.md b/README.md index eff82bb..34a1315 100644 --- a/README.md +++ b/README.md @@ -34,7 +34,8 @@ That one sentence is a `create_prompt` call, a `run_prompt` call and a add prompts, run them and query the answers without a human in the loop. - **Your data, queryable** — plain Postgres, so an agent can also read it with SQL, and the [Grafana starter](./grafana/README.md) charts it. -- **$0 to run** — fits in Vercel's free tier, database included. +- **$0 to run** — fits in Vercel's free tier, database included, until the + derived tables outgrow it (see [Data & retention](#data--retention)). - **Fully async** — scrapes are submitted as [cloro async tasks](https://cloro.dev/docs) and results come back by webhook, so no serverless function ever waits on a scrape. @@ -167,6 +168,11 @@ All endpoints except the webhook require | `DELETE` | `/api/prompts/:id` | Delete a prompt and (cascade) its results | | `POST` | `/api/prompts/:id/run` | Submit the prompt to its engines now; returns pending task ids (202) | | `GET` | `/api/results` | Query results (filters below) | +| `GET` | `/api/brands` | List tracked brands | +| `POST` | `/api/brands` | Track a brand (`name`, `aliases[]`, `domains[]`, `isOwn`, `enabled`) | +| `GET` | `/api/brands/:id` | Get one brand | +| `PATCH` | `/api/brands/:id` | Update any subset of the brand fields | +| `DELETE` | `/api/brands/:id` | Delete a brand and (cascade) its mention rows | | `GET` | `/api/cron` | Scheduler tick — same bearer token | | `POST` | `/api/webhook` | cloro result callback — auth via token in the callback URL | | `*` | `/api/mcp` | MCP endpoint (Streamable HTTP) | @@ -185,13 +191,17 @@ The MCP endpoint is the primary way to use geo-tracker. It is not a read-only reporting layer: an agent can create prompts, trigger runs and pull the stored answers, which is the whole product surface. -| Tool | What the agent can do | -| --------------- | ----------------------------------------------------- | -| `list_prompts` | See what is being tracked, and when each last ran | -| `create_prompt` | Add a prompt, pick the engines, set how often it runs | -| `run_prompt` | Run one now instead of waiting for the schedule | -| `get_results` | Query runs by prompt, engine or status | -| `get_result` | Pull one full raw engine answer for analysis | +| Tool | What the agent can do | +| ---------------------- | ----------------------------------------------------- | +| `list_prompts` | See what is being tracked, and when each last ran | +| `create_prompt` | Add a prompt, pick the engines, set how often it runs | +| `run_prompt` | Run one now instead of waiting for the schedule | +| `get_results` | Query runs by prompt, engine or status | +| `get_result` | Pull one full raw engine answer for analysis | +| `list_brands` | See which brands are being looked for | +| `track_brand` | Start looking for a brand, with aliases and domains | +| `untrack_brand` | Stop looking for one, and drop its derived rows | +| `get_brand_visibility` | How often each brand was named, and cited | Things worth asking an agent once it is connected: @@ -253,12 +263,80 @@ The same tick also sweeps: pending results whose webhook was missed are polled from the cloro API and backfilled, so nothing is lost if a webhook delivery fails. +## Brand visibility + +Tell geo-tracker which brands to look for, and every answer is flattened +into two tables you can query or chart: + +```bash +curl -X POST https://.vercel.app/api/brands \ + -H "Authorization: Bearer $CRON_SECRET" \ + -H "content-type: application/json" \ + -d '{"name":"Acme","aliases":["Acme Corp"],"domains":["acme.io"],"isOwn":true}' +``` + +- `result_sources` — one row per link an engine returned, tagged by where + it came from (`source`, `citation_pill`, `organic`, `ad`, …). This is + the "which pages get cited" question. +- `result_brand_mentions` — one row per answer per brand, including the + brands that were **not** named. That is what makes share of voice + computable: a brand at 0% has rows saying so, rather than being absent. + +Being **named** in the prose and being **cited** as a link are stored +separately, because they are different outcomes — an answer can recommend +you without linking you, or link you without naming you. + +Adding or editing a brand re-scores every answer already stored, so a +brand you add today has full history rather than starting at zero. The +work happens in the scheduler tick, a batch at a time; the API response +tells you how many results are queued. + +**A deployment with history takes a while to catch up.** The tick derives +250 results, and Vercel's Hobby plan runs one cron a day — so re-deriving +a year of answers would take months of ticks. Two ways round it, and the +tick is idempotent so either is safe: + +- Point an external scheduler (GitHub Actions, cron-job.org) at + `/api/cron` every few minutes until `more` comes back `false`. +- Or call it by hand in a loop: + + ```bash + while curl -s -H "Authorization: Bearer $CRON_SECRET" \ + https://.vercel.app/api/cron | grep -q '"more":true'; do :; done + ``` + +Nothing is missing while it runs. The old rows stay until each result is +rebuilt, so the dashboard shows values that are stale, never blank. + +`result_search_queries` holds a third thing: the literal queries the +engines typed before retrieving anything. ChatGPT, Copilot, Grok and +Perplexity report these; the others do not. + +To watch a competitor without tracking it, add its name to +`lib/brand-candidates.json`, with any alternative spellings: + +```json +{ "name": "Acme", "aliases": ["Acme, Inc", "Acme Corp"] } +``` + +Names there that turn up in answers, and that you are not tracking, +surface as a shortlist worth adding. The file ships with fictional +placeholders to replace — geo-tracker does not guess who competes with +you, and it cannot find a brand nobody wrote down. + +Nothing here scores or ranks an answer. It records whether a name is +present. What that means is the agent's call. + ## Grafana A ready-made dashboard (run volume, success rate, credits burn, failures) lives in [`grafana/`](./grafana/README.md) — point Grafana Cloud's free tier at your Postgres and import one JSON file. +Grafana runs outside Vercel: it is a long-running server, and Vercel hosts +serverless functions. Grafana Cloud's free tier reads your database +directly over TLS, which is all this needs. + ## Local development Contributing with a coding agent? [`AGENTS.md`](./AGENTS.md) documents the @@ -286,17 +364,29 @@ so just hit `/api/cron` again after a scrape completes. ## Data & retention -Two tables: `prompts` and `results`. Each run stores one row per engine -with the full raw cloro response as `jsonb` — measured at **10–30 KB per -row**, so budget roughly 20 KB per engine per run. +Two tables hold what you asked and what came back: `prompts` and +`results`. Each run stores one row per engine with the full raw cloro +response as `jsonb` — measured at **10–30 KB per row**, so budget roughly +20 KB per engine per run. -Storage is the limit you hit first, well before anything on Vercel: +Four more are derived from those responses by the scheduler tick +(`result_sources`, `result_brand_mentions`, `result_search_queries`, +`result_candidate_mentions`). They add about **5 KB per answer**, so +budget ~25 KB per engine per run in total. That figure scales with how +many links an answer carries, not with how big its payload is, so a chatty +engine costs more here than a terse one with a large HTML blob. + +Storage is the limit you hit first, well before anything on Vercel. The +figures below include the derived rows: | Workload | Scrapes/day | Storage/month | 0.5 GB lasts | | ------------------------------ | ----------- | ------------- | ------------ | -| 10 prompts × 3 engines × 1/day | 30 | ~18 MB | over 2 years | -| 20 prompts × 4 engines × 4/day | 320 | ~190 MB | ~3 months | -| 50 prompts × 6 engines × 8/day | 2,400 | ~1.4 GB | ~2 weeks | +| 10 prompts × 3 engines × 1/day | 30 | ~23 MB | ~1.8 years | +| 20 prompts × 4 engines × 4/day | 320 | ~240 MB | ~2 months | +| 50 prompts × 6 engines × 8/day | 2,400 | ~1.8 GB | ~8 days | + +Deleting a result cascades to its derived rows, so a retention policy +needs no extra step. The same workloads use under 1%, 1% and 7% of Vercel's free monthly function invocations, so the compute side stays free throughout. diff --git a/app/api/brands/[id]/route.ts b/app/api/brands/[id]/route.ts new file mode 100644 index 0000000..70295d8 --- /dev/null +++ b/app/api/brands/[id]/route.ts @@ -0,0 +1,78 @@ +import { eq } from "drizzle-orm"; + +import { isApiKeyAuthorized, unauthorized } from "@/lib/auth"; +import { getDb } from "@/lib/db"; +import { brands } from "@/lib/db/schema"; +import { + HttpError, + isUniqueViolation, + parseOr400, + withErrors, +} from "@/lib/http"; +import { affectsExtraction, markAllForReextraction } from "@/lib/refresh"; +import { idSchema, updateBrandSchema } from "@/lib/validation"; + +export const runtime = "nodejs"; + +function parseId(id: string): string { + const parsed = idSchema.safeParse(id); + if (!parsed.success) throw new HttpError(400, "Invalid brand id"); + return parsed.data; +} + +export const GET = withErrors(async (req, { params }) => { + if (!isApiKeyAuthorized(req)) return unauthorized(); + const id = parseId((await params).id); + const db = getDb(); + const [row] = await db.select().from(brands).where(eq(brands.id, id)); + if (!row) throw new HttpError(404, "Brand not found"); + return Response.json({ brand: row }); +}); + +export const PATCH = withErrors(async (req, { params }) => { + if (!isApiKeyAuthorized(req)) return unauthorized(); + const id = parseId((await params).id); + const input = parseOr400( + updateBrandSchema, + await req.json().catch(() => null), + ); + + const db = getDb(); + let updated; + try { + [updated] = await db + .update(brands) + .set(input) + .where(eq(brands.id, id)) + .returning(); + } catch (error) { + if (isUniqueViolation(error)) { + throw new HttpError(409, `A brand named "${input.name}" already exists`); + } + throw error; + } + if (!updated) throw new HttpError(404, "Brand not found"); + + // Only re-derive when the edit changes what the extractor reads. Renaming + // a brand or adding an alias rewrites its whole history; flipping `isOwn` + // changes a label the extractor never looks at. + const queued = affectsExtraction(input) ? await markAllForReextraction() : 0; + + return Response.json({ brand: updated, queuedForExtraction: queued }); +}); + +export const DELETE = withErrors(async (req, { params }) => { + if (!isApiKeyAuthorized(req)) return unauthorized(); + const id = parseId((await params).id); + const db = getDb(); + const deleted = await db + .delete(brands) + .where(eq(brands.id, id)) + .returning({ id: brands.id }); + if (deleted.length === 0) throw new HttpError(404, "Brand not found"); + + // No re-extraction: the foreign key cascades this brand's mention rows + // away, and no other brand's rows depend on it. Deleting is the one brand + // change that costs nothing. + return Response.json({ deleted: true }); +}); diff --git a/app/api/brands/route.ts b/app/api/brands/route.ts new file mode 100644 index 0000000..dbae12d --- /dev/null +++ b/app/api/brands/route.ts @@ -0,0 +1,55 @@ +import { asc } from "drizzle-orm"; + +import { isApiKeyAuthorized, unauthorized } from "@/lib/auth"; +import { getDb } from "@/lib/db"; +import { brands } from "@/lib/db/schema"; +import { + HttpError, + isUniqueViolation, + parseOr400, + withErrors, +} from "@/lib/http"; +import { markAllForReextraction } from "@/lib/refresh"; +import { createBrandSchema } from "@/lib/validation"; + +export const runtime = "nodejs"; + +export const GET = withErrors(async (req) => { + if (!isApiKeyAuthorized(req)) return unauthorized(); + const db = getDb(); + const rows = await db.select().from(brands).orderBy(asc(brands.name)); + return Response.json({ brands: rows }); +}); + +export const POST = withErrors(async (req) => { + if (!isApiKeyAuthorized(req)) return unauthorized(); + const input = parseOr400( + createBrandSchema, + await req.json().catch(() => null), + ); + + const db = getDb(); + let created; + try { + [created] = await db.insert(brands).values(input).returning(); + } catch (error) { + // The unique index is on lower(name), so "Acme" and "acme" collide. + // That is deliberate — they are one brand — but the raw driver error + // does not say so. + if (isUniqueViolation(error)) { + throw new HttpError(409, `A brand named "${input.name}" already exists`); + } + throw error; + } + + // A new brand has to be scored against answers that already arrived, not + // only the ones still to come. Without this its chart would begin on the + // day somebody remembered to add it, which reads as a brand that appeared + // out of nowhere rather than one we started watching late. + const queued = await markAllForReextraction(); + + return Response.json( + { brand: created, queuedForExtraction: queued }, + { status: 201 }, + ); +}); diff --git a/app/api/mcp/route.ts b/app/api/mcp/route.ts index b17928f..757deb0 100644 --- a/app/api/mcp/route.ts +++ b/app/api/mcp/route.ts @@ -1,12 +1,15 @@ -import { and, desc, eq } from "drizzle-orm"; +import { and, asc, desc, eq, gte, sql } from "drizzle-orm"; import { createMcpHandler } from "mcp-handler"; import { z } from "zod"; import { isApiKeyAuthorized, unauthorized } from "@/lib/auth"; import { getDb } from "@/lib/db"; -import { prompts, results } from "@/lib/db/schema"; +import { brands, prompts, resultBrandMentions, results } from "@/lib/db/schema"; import { ENGINES } from "@/lib/engines"; +import { isUniqueViolation } from "@/lib/http"; +import { markAllForReextraction } from "@/lib/refresh"; import { submitPromptOnce } from "@/lib/runner"; +import { createBrandSchema } from "@/lib/validation"; export const runtime = "nodejs"; @@ -123,6 +126,121 @@ const handler = createMcpHandler( }, ); + server.registerTool( + "list_brands", + { + title: "List brands", + description: + "List the brands being looked for in answers, with their aliases and domains.", + inputSchema: z.object({}), + }, + async () => { + const rows = await getDb() + .select() + .from(brands) + .orderBy(asc(brands.name)); + return text(rows); + }, + ); + + server.registerTool( + "track_brand", + { + title: "Track a brand", + description: + "Start looking for a brand in answers. Aliases are extra spellings that count as the same brand; domains decide when a link counts as citing it. Every answer already stored is re-scored against the new brand, so its history starts full rather than empty.", + inputSchema: z.object({ + name: z.string().min(1).max(200), + aliases: z.array(z.string().min(1).max(200)).max(50).default([]), + domains: z.array(z.string().min(1).max(253)).max(50).default([]), + isOwn: z.boolean().default(false), + }), + }, + async (input) => { + const parsed = createBrandSchema.safeParse(input); + if (!parsed.success) { + return text( + parsed.error.issues + .map( + (issue) => + `${issue.path.join(".") || "input"}: ${issue.message}`, + ) + .join("; "), + ); + } + + try { + const [created] = await getDb() + .insert(brands) + .values(parsed.data) + .returning(); + const queued = await markAllForReextraction(); + return text({ brand: created, queuedForExtraction: queued }); + } catch (error) { + if (isUniqueViolation(error)) { + return text(`A brand named "${parsed.data.name}" already exists`); + } + throw error; + } + }, + ); + + server.registerTool( + "untrack_brand", + { + title: "Stop tracking a brand", + description: + "Delete a brand and every mention row derived for it. The raw answers are untouched.", + inputSchema: z.object({ brandId: z.uuid() }), + }, + async ({ brandId }) => { + const deleted = await getDb() + .delete(brands) + .where(eq(brands.id, brandId)) + .returning({ id: brands.id }); + if (deleted.length === 0) return text(`Brand ${brandId} not found`); + return text({ deleted: true }); + }, + ); + + server.registerTool( + "get_brand_visibility", + { + title: "Get brand visibility", + description: + "How often each tracked brand was named in answers, and how often its own pages were cited, over the last N days. Counts every completed run as a denominator, so a brand that was never named reports 0 rather than going missing.", + inputSchema: z.object({ + days: z.number().int().min(1).max(365).default(30), + engine: z.enum(ENGINES).optional(), + }), + }, + async ({ days, engine }) => { + const since = new Date(Date.now() - days * 24 * 60 * 60 * 1000); + const rows = await getDb() + .select({ + brand: brands.name, + isOwn: brands.isOwn, + answers: sql`count(*)::int`, + mentioned: sql`count(*) filter (where ${resultBrandMentions.mentioned})::int`, + cited: sql`count(*) filter (where ${resultBrandMentions.cited})::int`, + }) + .from(resultBrandMentions) + .innerJoin(brands, eq(brands.id, resultBrandMentions.brandId)) + .innerJoin(results, eq(results.id, resultBrandMentions.resultId)) + .where( + and( + gte(results.completedAt, since), + engine ? eq(results.engine, engine) : undefined, + ), + ) + .groupBy(brands.id, brands.name, brands.isOwn) + .orderBy( + desc(sql`count(*) filter (where ${resultBrandMentions.mentioned})`), + ); + return text(rows); + }, + ); + server.registerTool( "get_result", { diff --git a/drizzle/0001_hot_moonstone.sql b/drizzle/0001_hot_moonstone.sql new file mode 100644 index 0000000..074999a --- /dev/null +++ b/drizzle/0001_hot_moonstone.sql @@ -0,0 +1,41 @@ +CREATE TABLE "brands" ( + "id" uuid PRIMARY KEY DEFAULT gen_random_uuid() NOT NULL, + "name" text NOT NULL, + "aliases" text[] DEFAULT '{}'::text[] NOT NULL, + "domains" text[] DEFAULT '{}'::text[] NOT NULL, + "is_own" boolean DEFAULT false NOT NULL, + "enabled" boolean DEFAULT true NOT NULL, + "created_at" timestamp with time zone DEFAULT now() NOT NULL +); +--> statement-breakpoint +CREATE TABLE "result_brand_mentions" ( + "id" uuid PRIMARY KEY DEFAULT gen_random_uuid() NOT NULL, + "result_id" uuid NOT NULL, + "brand_id" uuid NOT NULL, + "mentioned" boolean NOT NULL, + "mention_count" integer DEFAULT 0 NOT NULL, + "first_position" integer, + "cited" boolean DEFAULT false NOT NULL, + "cited_source_count" integer DEFAULT 0 NOT NULL +); +--> statement-breakpoint +CREATE TABLE "result_sources" ( + "id" uuid PRIMARY KEY DEFAULT gen_random_uuid() NOT NULL, + "result_id" uuid NOT NULL, + "kind" text NOT NULL, + "position" integer, + "url" text NOT NULL, + "domain" text NOT NULL, + "label" text +); +--> statement-breakpoint +ALTER TABLE "results" ADD COLUMN "extraction_revision" integer;--> statement-breakpoint +ALTER TABLE "result_brand_mentions" ADD CONSTRAINT "result_brand_mentions_result_id_results_id_fk" FOREIGN KEY ("result_id") REFERENCES "public"."results"("id") ON DELETE cascade ON UPDATE no action;--> statement-breakpoint +ALTER TABLE "result_brand_mentions" ADD CONSTRAINT "result_brand_mentions_brand_id_brands_id_fk" FOREIGN KEY ("brand_id") REFERENCES "public"."brands"("id") ON DELETE cascade ON UPDATE no action;--> statement-breakpoint +ALTER TABLE "result_sources" ADD CONSTRAINT "result_sources_result_id_results_id_fk" FOREIGN KEY ("result_id") REFERENCES "public"."results"("id") ON DELETE cascade ON UPDATE no action;--> statement-breakpoint +CREATE UNIQUE INDEX "brands_name_idx" ON "brands" USING btree (lower("name"));--> statement-breakpoint +CREATE UNIQUE INDEX "result_brand_mentions_result_brand_idx" ON "result_brand_mentions" USING btree ("result_id","brand_id");--> statement-breakpoint +CREATE INDEX "result_brand_mentions_brand_idx" ON "result_brand_mentions" USING btree ("brand_id");--> statement-breakpoint +CREATE INDEX "result_sources_result_idx" ON "result_sources" USING btree ("result_id");--> statement-breakpoint +CREATE INDEX "result_sources_domain_idx" ON "result_sources" USING btree ("domain","kind");--> statement-breakpoint +CREATE INDEX "results_unextracted_idx" ON "results" USING btree ("completed_at") WHERE "results"."status" = 'completed'; \ No newline at end of file diff --git a/drizzle/0002_freezing_sumo.sql b/drizzle/0002_freezing_sumo.sql new file mode 100644 index 0000000..26b3a4d --- /dev/null +++ b/drizzle/0002_freezing_sumo.sql @@ -0,0 +1,21 @@ +CREATE TABLE "result_candidate_mentions" ( + "id" uuid PRIMARY KEY DEFAULT gen_random_uuid() NOT NULL, + "result_id" uuid NOT NULL, + "name" text NOT NULL, + "mention_count" integer NOT NULL +); +--> statement-breakpoint +CREATE TABLE "result_search_queries" ( + "id" uuid PRIMARY KEY DEFAULT gen_random_uuid() NOT NULL, + "result_id" uuid NOT NULL, + "kind" text NOT NULL, + "position" integer, + "query" text NOT NULL +); +--> statement-breakpoint +ALTER TABLE "result_candidate_mentions" ADD CONSTRAINT "result_candidate_mentions_result_id_results_id_fk" FOREIGN KEY ("result_id") REFERENCES "public"."results"("id") ON DELETE cascade ON UPDATE no action;--> statement-breakpoint +ALTER TABLE "result_search_queries" ADD CONSTRAINT "result_search_queries_result_id_results_id_fk" FOREIGN KEY ("result_id") REFERENCES "public"."results"("id") ON DELETE cascade ON UPDATE no action;--> statement-breakpoint +CREATE UNIQUE INDEX "result_candidate_mentions_result_name_idx" ON "result_candidate_mentions" USING btree ("result_id","name");--> statement-breakpoint +CREATE INDEX "result_candidate_mentions_name_idx" ON "result_candidate_mentions" USING btree ("name");--> statement-breakpoint +CREATE INDEX "result_search_queries_result_idx" ON "result_search_queries" USING btree ("result_id");--> statement-breakpoint +CREATE INDEX "result_search_queries_kind_idx" ON "result_search_queries" USING btree ("kind"); \ No newline at end of file diff --git a/drizzle/0003_vengeful_havok.sql b/drizzle/0003_vengeful_havok.sql new file mode 100644 index 0000000..5ed8134 --- /dev/null +++ b/drizzle/0003_vengeful_havok.sql @@ -0,0 +1,8 @@ +CREATE TABLE "extraction_state" ( + "id" boolean PRIMARY KEY DEFAULT true NOT NULL, + "stamp" integer NOT NULL, + "updated_at" timestamp with time zone DEFAULT now() NOT NULL +); +--> statement-breakpoint +DROP INDEX "results_unextracted_idx";--> statement-breakpoint +CREATE INDEX "results_unextracted_idx" ON "results" USING btree ("completed_at") WHERE "results"."status" = 'completed' AND "results"."extraction_revision" IS NULL; \ No newline at end of file diff --git a/drizzle/meta/0001_snapshot.json b/drizzle/meta/0001_snapshot.json new file mode 100644 index 0000000..d3e7ae9 --- /dev/null +++ b/drizzle/meta/0001_snapshot.json @@ -0,0 +1,585 @@ +{ + "id": "53aab6df-89ad-495f-b4a8-76d5187271f3", + "prevId": "71d66fa2-8ada-4900-a9da-f27674ff8bbc", + "version": "7", + "dialect": "postgresql", + "tables": { + "public.brands": { + "name": "brands", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "name": { + "name": "name", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "aliases": { + "name": "aliases", + "type": "text[]", + "primaryKey": false, + "notNull": true, + "default": "'{}'::text[]" + }, + "domains": { + "name": "domains", + "type": "text[]", + "primaryKey": false, + "notNull": true, + "default": "'{}'::text[]" + }, + "is_own": { + "name": "is_own", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": false + }, + "enabled": { + "name": "enabled", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": true + }, + "created_at": { + "name": "created_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": true, + "default": "now()" + } + }, + "indexes": { + "brands_name_idx": { + "name": "brands_name_idx", + "columns": [ + { + "expression": "lower(\"name\")", + "asc": true, + "isExpression": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": {}, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.prompts": { + "name": "prompts", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "name": { + "name": "name", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "prompt": { + "name": "prompt", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "engines": { + "name": "engines", + "type": "text[]", + "primaryKey": false, + "notNull": true + }, + "country": { + "name": "country", + "type": "text", + "primaryKey": false, + "notNull": true, + "default": "'US'" + }, + "runs_per_day": { + "name": "runs_per_day", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 1 + }, + "enabled": { + "name": "enabled", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": true + }, + "last_run_at": { + "name": "last_run_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": false + }, + "created_at": { + "name": "created_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": true, + "default": "now()" + } + }, + "indexes": {}, + "foreignKeys": {}, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.result_brand_mentions": { + "name": "result_brand_mentions", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "result_id": { + "name": "result_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "brand_id": { + "name": "brand_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "mentioned": { + "name": "mentioned", + "type": "boolean", + "primaryKey": false, + "notNull": true + }, + "mention_count": { + "name": "mention_count", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 0 + }, + "first_position": { + "name": "first_position", + "type": "integer", + "primaryKey": false, + "notNull": false + }, + "cited": { + "name": "cited", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": false + }, + "cited_source_count": { + "name": "cited_source_count", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 0 + } + }, + "indexes": { + "result_brand_mentions_result_brand_idx": { + "name": "result_brand_mentions_result_brand_idx", + "columns": [ + { + "expression": "result_id", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "brand_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + }, + "result_brand_mentions_brand_idx": { + "name": "result_brand_mentions_brand_idx", + "columns": [ + { + "expression": "brand_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "result_brand_mentions_result_id_results_id_fk": { + "name": "result_brand_mentions_result_id_results_id_fk", + "tableFrom": "result_brand_mentions", + "tableTo": "results", + "columnsFrom": [ + "result_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + }, + "result_brand_mentions_brand_id_brands_id_fk": { + "name": "result_brand_mentions_brand_id_brands_id_fk", + "tableFrom": "result_brand_mentions", + "tableTo": "brands", + "columnsFrom": [ + "brand_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.result_sources": { + "name": "result_sources", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "result_id": { + "name": "result_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "kind": { + "name": "kind", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "position": { + "name": "position", + "type": "integer", + "primaryKey": false, + "notNull": false + }, + "url": { + "name": "url", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "domain": { + "name": "domain", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "label": { + "name": "label", + "type": "text", + "primaryKey": false, + "notNull": false + } + }, + "indexes": { + "result_sources_result_idx": { + "name": "result_sources_result_idx", + "columns": [ + { + "expression": "result_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "result_sources_domain_idx": { + "name": "result_sources_domain_idx", + "columns": [ + { + "expression": "domain", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "kind", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "result_sources_result_id_results_id_fk": { + "name": "result_sources_result_id_results_id_fk", + "tableFrom": "result_sources", + "tableTo": "results", + "columnsFrom": [ + "result_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.results": { + "name": "results", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "prompt_id": { + "name": "prompt_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "engine": { + "name": "engine", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "task_id": { + "name": "task_id", + "type": "text", + "primaryKey": false, + "notNull": false + }, + "status": { + "name": "status", + "type": "text", + "primaryKey": false, + "notNull": true, + "default": "'pending'" + }, + "response": { + "name": "response", + "type": "jsonb", + "primaryKey": false, + "notNull": false + }, + "error": { + "name": "error", + "type": "text", + "primaryKey": false, + "notNull": false + }, + "credits_charged": { + "name": "credits_charged", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 0 + }, + "created_at": { + "name": "created_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": true, + "default": "now()" + }, + "completed_at": { + "name": "completed_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": false + }, + "extraction_revision": { + "name": "extraction_revision", + "type": "integer", + "primaryKey": false, + "notNull": false + } + }, + "indexes": { + "results_task_id_idx": { + "name": "results_task_id_idx", + "columns": [ + { + "expression": "task_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_prompt_created_idx": { + "name": "results_prompt_created_idx", + "columns": [ + { + "expression": "prompt_id", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "created_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_created_idx": { + "name": "results_created_idx", + "columns": [ + { + "expression": "created_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_pending_idx": { + "name": "results_pending_idx", + "columns": [ + { + "expression": "created_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "where": "\"results\".\"status\" = 'pending'", + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_unextracted_idx": { + "name": "results_unextracted_idx", + "columns": [ + { + "expression": "completed_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "where": "\"results\".\"status\" = 'completed'", + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "results_prompt_id_prompts_id_fk": { + "name": "results_prompt_id_prompts_id_fk", + "tableFrom": "results", + "tableTo": "prompts", + "columnsFrom": [ + "prompt_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + } + }, + "enums": {}, + "schemas": {}, + "sequences": {}, + "roles": {}, + "policies": {}, + "views": {}, + "_meta": { + "columns": {}, + "schemas": {}, + "tables": {} + } +} \ No newline at end of file diff --git a/drizzle/meta/0002_snapshot.json b/drizzle/meta/0002_snapshot.json new file mode 100644 index 0000000..e3c131d --- /dev/null +++ b/drizzle/meta/0002_snapshot.json @@ -0,0 +1,763 @@ +{ + "id": "3a88db90-917f-4b36-ae67-ca79a704dd94", + "prevId": "53aab6df-89ad-495f-b4a8-76d5187271f3", + "version": "7", + "dialect": "postgresql", + "tables": { + "public.brands": { + "name": "brands", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "name": { + "name": "name", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "aliases": { + "name": "aliases", + "type": "text[]", + "primaryKey": false, + "notNull": true, + "default": "'{}'::text[]" + }, + "domains": { + "name": "domains", + "type": "text[]", + "primaryKey": false, + "notNull": true, + "default": "'{}'::text[]" + }, + "is_own": { + "name": "is_own", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": false + }, + "enabled": { + "name": "enabled", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": true + }, + "created_at": { + "name": "created_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": true, + "default": "now()" + } + }, + "indexes": { + "brands_name_idx": { + "name": "brands_name_idx", + "columns": [ + { + "expression": "lower(\"name\")", + "asc": true, + "isExpression": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": {}, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.prompts": { + "name": "prompts", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "name": { + "name": "name", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "prompt": { + "name": "prompt", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "engines": { + "name": "engines", + "type": "text[]", + "primaryKey": false, + "notNull": true + }, + "country": { + "name": "country", + "type": "text", + "primaryKey": false, + "notNull": true, + "default": "'US'" + }, + "runs_per_day": { + "name": "runs_per_day", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 1 + }, + "enabled": { + "name": "enabled", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": true + }, + "last_run_at": { + "name": "last_run_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": false + }, + "created_at": { + "name": "created_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": true, + "default": "now()" + } + }, + "indexes": {}, + "foreignKeys": {}, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.result_brand_mentions": { + "name": "result_brand_mentions", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "result_id": { + "name": "result_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "brand_id": { + "name": "brand_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "mentioned": { + "name": "mentioned", + "type": "boolean", + "primaryKey": false, + "notNull": true + }, + "mention_count": { + "name": "mention_count", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 0 + }, + "first_position": { + "name": "first_position", + "type": "integer", + "primaryKey": false, + "notNull": false + }, + "cited": { + "name": "cited", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": false + }, + "cited_source_count": { + "name": "cited_source_count", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 0 + } + }, + "indexes": { + "result_brand_mentions_result_brand_idx": { + "name": "result_brand_mentions_result_brand_idx", + "columns": [ + { + "expression": "result_id", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "brand_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + }, + "result_brand_mentions_brand_idx": { + "name": "result_brand_mentions_brand_idx", + "columns": [ + { + "expression": "brand_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "result_brand_mentions_result_id_results_id_fk": { + "name": "result_brand_mentions_result_id_results_id_fk", + "tableFrom": "result_brand_mentions", + "tableTo": "results", + "columnsFrom": [ + "result_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + }, + "result_brand_mentions_brand_id_brands_id_fk": { + "name": "result_brand_mentions_brand_id_brands_id_fk", + "tableFrom": "result_brand_mentions", + "tableTo": "brands", + "columnsFrom": [ + "brand_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.result_candidate_mentions": { + "name": "result_candidate_mentions", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "result_id": { + "name": "result_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "name": { + "name": "name", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "mention_count": { + "name": "mention_count", + "type": "integer", + "primaryKey": false, + "notNull": true + } + }, + "indexes": { + "result_candidate_mentions_result_name_idx": { + "name": "result_candidate_mentions_result_name_idx", + "columns": [ + { + "expression": "result_id", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "name", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + }, + "result_candidate_mentions_name_idx": { + "name": "result_candidate_mentions_name_idx", + "columns": [ + { + "expression": "name", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "result_candidate_mentions_result_id_results_id_fk": { + "name": "result_candidate_mentions_result_id_results_id_fk", + "tableFrom": "result_candidate_mentions", + "tableTo": "results", + "columnsFrom": [ + "result_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.result_search_queries": { + "name": "result_search_queries", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "result_id": { + "name": "result_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "kind": { + "name": "kind", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "position": { + "name": "position", + "type": "integer", + "primaryKey": false, + "notNull": false + }, + "query": { + "name": "query", + "type": "text", + "primaryKey": false, + "notNull": true + } + }, + "indexes": { + "result_search_queries_result_idx": { + "name": "result_search_queries_result_idx", + "columns": [ + { + "expression": "result_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "result_search_queries_kind_idx": { + "name": "result_search_queries_kind_idx", + "columns": [ + { + "expression": "kind", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "result_search_queries_result_id_results_id_fk": { + "name": "result_search_queries_result_id_results_id_fk", + "tableFrom": "result_search_queries", + "tableTo": "results", + "columnsFrom": [ + "result_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.result_sources": { + "name": "result_sources", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "result_id": { + "name": "result_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "kind": { + "name": "kind", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "position": { + "name": "position", + "type": "integer", + "primaryKey": false, + "notNull": false + }, + "url": { + "name": "url", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "domain": { + "name": "domain", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "label": { + "name": "label", + "type": "text", + "primaryKey": false, + "notNull": false + } + }, + "indexes": { + "result_sources_result_idx": { + "name": "result_sources_result_idx", + "columns": [ + { + "expression": "result_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "result_sources_domain_idx": { + "name": "result_sources_domain_idx", + "columns": [ + { + "expression": "domain", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "kind", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "result_sources_result_id_results_id_fk": { + "name": "result_sources_result_id_results_id_fk", + "tableFrom": "result_sources", + "tableTo": "results", + "columnsFrom": [ + "result_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.results": { + "name": "results", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "prompt_id": { + "name": "prompt_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "engine": { + "name": "engine", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "task_id": { + "name": "task_id", + "type": "text", + "primaryKey": false, + "notNull": false + }, + "status": { + "name": "status", + "type": "text", + "primaryKey": false, + "notNull": true, + "default": "'pending'" + }, + "response": { + "name": "response", + "type": "jsonb", + "primaryKey": false, + "notNull": false + }, + "error": { + "name": "error", + "type": "text", + "primaryKey": false, + "notNull": false + }, + "credits_charged": { + "name": "credits_charged", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 0 + }, + "created_at": { + "name": "created_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": true, + "default": "now()" + }, + "completed_at": { + "name": "completed_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": false + }, + "extraction_revision": { + "name": "extraction_revision", + "type": "integer", + "primaryKey": false, + "notNull": false + } + }, + "indexes": { + "results_task_id_idx": { + "name": "results_task_id_idx", + "columns": [ + { + "expression": "task_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_prompt_created_idx": { + "name": "results_prompt_created_idx", + "columns": [ + { + "expression": "prompt_id", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "created_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_created_idx": { + "name": "results_created_idx", + "columns": [ + { + "expression": "created_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_pending_idx": { + "name": "results_pending_idx", + "columns": [ + { + "expression": "created_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "where": "\"results\".\"status\" = 'pending'", + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_unextracted_idx": { + "name": "results_unextracted_idx", + "columns": [ + { + "expression": "completed_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "where": "\"results\".\"status\" = 'completed'", + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "results_prompt_id_prompts_id_fk": { + "name": "results_prompt_id_prompts_id_fk", + "tableFrom": "results", + "tableTo": "prompts", + "columnsFrom": [ + "prompt_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + } + }, + "enums": {}, + "schemas": {}, + "sequences": {}, + "roles": {}, + "policies": {}, + "views": {}, + "_meta": { + "columns": {}, + "schemas": {}, + "tables": {} + } +} \ No newline at end of file diff --git a/drizzle/meta/0003_snapshot.json b/drizzle/meta/0003_snapshot.json new file mode 100644 index 0000000..8b6761e --- /dev/null +++ b/drizzle/meta/0003_snapshot.json @@ -0,0 +1,796 @@ +{ + "id": "7a9c3661-62ab-4fdc-9a1b-4a966f0d0c87", + "prevId": "3a88db90-917f-4b36-ae67-ca79a704dd94", + "version": "7", + "dialect": "postgresql", + "tables": { + "public.brands": { + "name": "brands", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "name": { + "name": "name", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "aliases": { + "name": "aliases", + "type": "text[]", + "primaryKey": false, + "notNull": true, + "default": "'{}'::text[]" + }, + "domains": { + "name": "domains", + "type": "text[]", + "primaryKey": false, + "notNull": true, + "default": "'{}'::text[]" + }, + "is_own": { + "name": "is_own", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": false + }, + "enabled": { + "name": "enabled", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": true + }, + "created_at": { + "name": "created_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": true, + "default": "now()" + } + }, + "indexes": { + "brands_name_idx": { + "name": "brands_name_idx", + "columns": [ + { + "expression": "lower(\"name\")", + "asc": true, + "isExpression": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": {}, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.extraction_state": { + "name": "extraction_state", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "boolean", + "primaryKey": true, + "notNull": true, + "default": true + }, + "stamp": { + "name": "stamp", + "type": "integer", + "primaryKey": false, + "notNull": true + }, + "updated_at": { + "name": "updated_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": true, + "default": "now()" + } + }, + "indexes": {}, + "foreignKeys": {}, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.prompts": { + "name": "prompts", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "name": { + "name": "name", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "prompt": { + "name": "prompt", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "engines": { + "name": "engines", + "type": "text[]", + "primaryKey": false, + "notNull": true + }, + "country": { + "name": "country", + "type": "text", + "primaryKey": false, + "notNull": true, + "default": "'US'" + }, + "runs_per_day": { + "name": "runs_per_day", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 1 + }, + "enabled": { + "name": "enabled", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": true + }, + "last_run_at": { + "name": "last_run_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": false + }, + "created_at": { + "name": "created_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": true, + "default": "now()" + } + }, + "indexes": {}, + "foreignKeys": {}, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.result_brand_mentions": { + "name": "result_brand_mentions", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "result_id": { + "name": "result_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "brand_id": { + "name": "brand_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "mentioned": { + "name": "mentioned", + "type": "boolean", + "primaryKey": false, + "notNull": true + }, + "mention_count": { + "name": "mention_count", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 0 + }, + "first_position": { + "name": "first_position", + "type": "integer", + "primaryKey": false, + "notNull": false + }, + "cited": { + "name": "cited", + "type": "boolean", + "primaryKey": false, + "notNull": true, + "default": false + }, + "cited_source_count": { + "name": "cited_source_count", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 0 + } + }, + "indexes": { + "result_brand_mentions_result_brand_idx": { + "name": "result_brand_mentions_result_brand_idx", + "columns": [ + { + "expression": "result_id", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "brand_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + }, + "result_brand_mentions_brand_idx": { + "name": "result_brand_mentions_brand_idx", + "columns": [ + { + "expression": "brand_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "result_brand_mentions_result_id_results_id_fk": { + "name": "result_brand_mentions_result_id_results_id_fk", + "tableFrom": "result_brand_mentions", + "tableTo": "results", + "columnsFrom": [ + "result_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + }, + "result_brand_mentions_brand_id_brands_id_fk": { + "name": "result_brand_mentions_brand_id_brands_id_fk", + "tableFrom": "result_brand_mentions", + "tableTo": "brands", + "columnsFrom": [ + "brand_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.result_candidate_mentions": { + "name": "result_candidate_mentions", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "result_id": { + "name": "result_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "name": { + "name": "name", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "mention_count": { + "name": "mention_count", + "type": "integer", + "primaryKey": false, + "notNull": true + } + }, + "indexes": { + "result_candidate_mentions_result_name_idx": { + "name": "result_candidate_mentions_result_name_idx", + "columns": [ + { + "expression": "result_id", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "name", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + }, + "result_candidate_mentions_name_idx": { + "name": "result_candidate_mentions_name_idx", + "columns": [ + { + "expression": "name", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "result_candidate_mentions_result_id_results_id_fk": { + "name": "result_candidate_mentions_result_id_results_id_fk", + "tableFrom": "result_candidate_mentions", + "tableTo": "results", + "columnsFrom": [ + "result_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.result_search_queries": { + "name": "result_search_queries", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "result_id": { + "name": "result_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "kind": { + "name": "kind", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "position": { + "name": "position", + "type": "integer", + "primaryKey": false, + "notNull": false + }, + "query": { + "name": "query", + "type": "text", + "primaryKey": false, + "notNull": true + } + }, + "indexes": { + "result_search_queries_result_idx": { + "name": "result_search_queries_result_idx", + "columns": [ + { + "expression": "result_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "result_search_queries_kind_idx": { + "name": "result_search_queries_kind_idx", + "columns": [ + { + "expression": "kind", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "result_search_queries_result_id_results_id_fk": { + "name": "result_search_queries_result_id_results_id_fk", + "tableFrom": "result_search_queries", + "tableTo": "results", + "columnsFrom": [ + "result_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.result_sources": { + "name": "result_sources", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "result_id": { + "name": "result_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "kind": { + "name": "kind", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "position": { + "name": "position", + "type": "integer", + "primaryKey": false, + "notNull": false + }, + "url": { + "name": "url", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "domain": { + "name": "domain", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "label": { + "name": "label", + "type": "text", + "primaryKey": false, + "notNull": false + } + }, + "indexes": { + "result_sources_result_idx": { + "name": "result_sources_result_idx", + "columns": [ + { + "expression": "result_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "result_sources_domain_idx": { + "name": "result_sources_domain_idx", + "columns": [ + { + "expression": "domain", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "kind", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "result_sources_result_id_results_id_fk": { + "name": "result_sources_result_id_results_id_fk", + "tableFrom": "result_sources", + "tableTo": "results", + "columnsFrom": [ + "result_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + }, + "public.results": { + "name": "results", + "schema": "", + "columns": { + "id": { + "name": "id", + "type": "uuid", + "primaryKey": true, + "notNull": true, + "default": "gen_random_uuid()" + }, + "prompt_id": { + "name": "prompt_id", + "type": "uuid", + "primaryKey": false, + "notNull": true + }, + "engine": { + "name": "engine", + "type": "text", + "primaryKey": false, + "notNull": true + }, + "task_id": { + "name": "task_id", + "type": "text", + "primaryKey": false, + "notNull": false + }, + "status": { + "name": "status", + "type": "text", + "primaryKey": false, + "notNull": true, + "default": "'pending'" + }, + "response": { + "name": "response", + "type": "jsonb", + "primaryKey": false, + "notNull": false + }, + "error": { + "name": "error", + "type": "text", + "primaryKey": false, + "notNull": false + }, + "credits_charged": { + "name": "credits_charged", + "type": "integer", + "primaryKey": false, + "notNull": true, + "default": 0 + }, + "created_at": { + "name": "created_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": true, + "default": "now()" + }, + "completed_at": { + "name": "completed_at", + "type": "timestamp with time zone", + "primaryKey": false, + "notNull": false + }, + "extraction_revision": { + "name": "extraction_revision", + "type": "integer", + "primaryKey": false, + "notNull": false + } + }, + "indexes": { + "results_task_id_idx": { + "name": "results_task_id_idx", + "columns": [ + { + "expression": "task_id", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": true, + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_prompt_created_idx": { + "name": "results_prompt_created_idx", + "columns": [ + { + "expression": "prompt_id", + "isExpression": false, + "asc": true, + "nulls": "last" + }, + { + "expression": "created_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_created_idx": { + "name": "results_created_idx", + "columns": [ + { + "expression": "created_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_pending_idx": { + "name": "results_pending_idx", + "columns": [ + { + "expression": "created_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "where": "\"results\".\"status\" = 'pending'", + "concurrently": false, + "method": "btree", + "with": {} + }, + "results_unextracted_idx": { + "name": "results_unextracted_idx", + "columns": [ + { + "expression": "completed_at", + "isExpression": false, + "asc": true, + "nulls": "last" + } + ], + "isUnique": false, + "where": "\"results\".\"status\" = 'completed' AND \"results\".\"extraction_revision\" IS NULL", + "concurrently": false, + "method": "btree", + "with": {} + } + }, + "foreignKeys": { + "results_prompt_id_prompts_id_fk": { + "name": "results_prompt_id_prompts_id_fk", + "tableFrom": "results", + "tableTo": "prompts", + "columnsFrom": [ + "prompt_id" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "policies": {}, + "checkConstraints": {}, + "isRLSEnabled": false + } + }, + "enums": {}, + "schemas": {}, + "sequences": {}, + "roles": {}, + "policies": {}, + "views": {}, + "_meta": { + "columns": {}, + "schemas": {}, + "tables": {} + } +} \ No newline at end of file diff --git a/drizzle/meta/_journal.json b/drizzle/meta/_journal.json index fb4d677..7bad239 100644 --- a/drizzle/meta/_journal.json +++ b/drizzle/meta/_journal.json @@ -8,6 +8,27 @@ "when": 1786829773996, "tag": "0000_init", "breakpoints": true + }, + { + "idx": 1, + "version": "7", + "when": 1788077219316, + "tag": "0001_hot_moonstone", + "breakpoints": true + }, + { + "idx": 2, + "version": "7", + "when": 1788089565249, + "tag": "0002_freezing_sumo", + "breakpoints": true + }, + { + "idx": 3, + "version": "7", + "when": 1788213732745, + "tag": "0003_vengeful_havok", + "breakpoints": true } ] } \ No newline at end of file diff --git a/grafana/README.md b/grafana/README.md index 73cceac..441429b 100644 --- a/grafana/README.md +++ b/grafana/README.md @@ -39,9 +39,76 @@ Click **Save & test**. **Dashboards → New → Import**, upload [`dashboard.json`](./dashboard.json), and pick the datasource you just created when prompted. +## The brand-visibility dashboard + +[`geo-visibility.json`](./geo-visibility.json) is the second dashboard: +which brands the engines name, which pages they cite, and where you are +losing. Import it the same way as the first one. + +Its panel order and geometry match the internal GEO dashboard cloro runs +on its own data, so the two read the same way. Three panels of that one +have no counterpart here: two are keyed on a prompt-set concept this repo +does not have, and the third needs a list of pages you are already listed +on, which there is nowhere to keep yet. + +It needs at least one brand configured with `is_own = true` — several +panels are written against "your" brand, and are empty without one: + +```bash +curl -X POST https://.vercel.app/api/brands \ + -H "Authorization: Bearer $CRON_SECRET" \ + -H "content-type: application/json" \ + -d '{"name":"Acme","domains":["acme.io"],"isOwn":true}' +``` + +It also reads `result_sources` and `result_brand_mentions`, which the +scheduler tick fills. A freshly imported dashboard is empty until a tick +has run — that is a queue waiting, not a broken panel. + +### Watching brands you do not track yet + +`lib/brand-candidates.json` holds names to look for **without** tracking +them. Anything listed there that an answer names, and that is not in your +`brands` table, appears in the "Named but not tracked" panel — the +shortlist worth promoting. + +Each entry carries the spellings an engine might write, and they fold into +one row: + +```json +{ "name": "Acme", "aliases": ["Acme, Inc", "Acme Corp"] } +``` + +A bare string works when a name needs no aliases. What ships is a list of +**fictional placeholder companies** — replace every one of them with the +real vendors in your category. Editing the file re-derives the whole +history on the next tick, so a name added today is scored against answers +already stored. + +It cannot discover a brand nobody wrote down. That is deliberate: finding +unknown names in prose means entity extraction, and geo-tracker does not +interpret answers — it records whether a name you chose is present. + +Two things the panels are built to keep apart: + +- **Named and cited are different outcomes.** An answer can recommend you + without linking you, or link you without naming you. No panel pools them. +- **The misses are counted.** A brand that is never named holds rows at 0% + rather than dropping out, so every percentage has an honest denominator. + +Google is expected to sit low on "Named %": it returns a page of links and +only writes prose when an AI Overview was served. Its "Cited %" is the +meaningful number there. + ## Building your own panels -Every panel's SQL lives standalone in [`queries.sql`](./queries.sql) — -copy, tweak, add. The schema is just two tables (`prompts`, `results`); -`results.response` holds the full raw engine response as `jsonb`, so -Postgres JSON operators (`response -> 'field'`) work in panels too. +Every panel's SQL from both dashboards lives standalone in +[`queries.sql`](./queries.sql) — copy, tweak, add. The Grafana macros are +replaced there with plain predicates, so each block runs as-is in `psql`. + +Four tables: `prompts` and `results` hold what was asked and what came +back, and `result_sources` and `result_brand_mentions` are derived from +`results.response` by the scheduler tick. `results.response` still holds +the full raw payload as `jsonb`, so Postgres JSON operators +(`response -> 'field'`) work in panels too when the derived tables do not +have what you need. diff --git a/grafana/geo-visibility.json b/grafana/geo-visibility.json new file mode 100644 index 0000000..3521621 --- /dev/null +++ b/grafana/geo-visibility.json @@ -0,0 +1,1080 @@ +{ + "__inputs": [ + { + "name": "DS_POSTGRES", + "label": "Postgres", + "description": "The Postgres database geo-tracker writes to", + "type": "datasource", + "pluginId": "grafana-postgresql-datasource", + "pluginName": "PostgreSQL" + } + ], + "__requires": [ + { + "type": "datasource", + "id": "grafana-postgresql-datasource", + "name": "PostgreSQL", + "version": "1.0.0" + } + ], + "annotations": { + "list": [] + }, + "editable": true, + "id": null, + "links": [], + "refresh": "30m", + "schemaVersion": 39, + "tags": ["geo-tracker"], + "templating": { + "list": [ + { + "name": "urlsearch", + "label": "Search URL", + "type": "textbox", + "query": "", + "current": { + "selected": false, + "text": "", + "value": "" + }, + "options": [] + }, + { + "name": "engine", + "label": "Engine", + "type": "query", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "query": "SELECT DISTINCT engine FROM results ORDER BY 1;", + "refresh": 1, + "multi": true, + "includeAll": true, + "allValue": null, + "current": { + "selected": true, + "text": ["All"], + "value": ["$__all"] + }, + "options": [] + } + ] + }, + "time": { + "from": "now-30d", + "to": "now" + }, + "timezone": "utc", + "title": "geo-tracker — brand visibility", + "uid": null, + "version": 1, + "panels": [ + { + "type": "row", + "title": "How to read this — ratios and denominators (click to expand)", + "gridPos": { + "h": 1, + "w": 24, + "x": 0, + "y": 0 + }, + "id": 14, + "collapsed": true, + "panels": [ + { + "type": "text", + "title": "How to read this", + "gridPos": { + "h": 11, + "w": 24, + "x": 0, + "y": 1 + }, + "id": 7, + "options": { + "mode": "markdown", + "content": "**Denominators.** Every percentage divides by *completed* answers in the selected range. Failed scrapes are excluded everywhere — an errored run is not an answer that declined to name you. A brand that is never named holds rows at 0% rather than dropping out, so a blank line means no data, not zero.\n\n**Named is not cited.** *Named* means the prose wrote your name. *Cited* means a link pointed at one of your domains. An answer can recommend you without linking you, or link you without naming you, so no panel pools them.\n\n**Google is different.** It returns a page of links and only writes prose when an AI Overview was served, so its Named % is structurally low. Read its Cited % instead.\n\n**Search queries come from four engines.** ChatGPT, Copilot, Grok and Perplexity report what they searched; Gemini and the Google engines do not. That panel is a sample of what a model types, not a ranking — the strings are mostly unique, so ordering by count is close to arbitrary.\n\n**Candidates are names you listed.** The untracked-brands panel can only find names written in `lib/brand-candidates.json`. It does not discover a brand nobody wrote down." + } + } + ] + }, + { + "type": "row", + "title": "Brands, URLs and prompts", + "gridPos": { + "h": 1, + "w": 24, + "x": 0, + "y": 1 + }, + "id": 15, + "collapsed": false, + "panels": [] + }, + { + "type": "table", + "title": "Named but not tracked — candidates worth adding", + "gridPos": { + "h": 13, + "w": 10, + "x": 0, + "y": 2 + }, + "id": 9, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "table", + "rawQuery": true, + "rawSql": "SELECT c.name AS \"Name\",\n sum(c.mention_count) AS \"Mentions\",\n count(DISTINCT r.engine) AS \"Engines\",\n count(DISTINCT r.prompt_id) AS \"Prompts\",\n max(r.completed_at)::date::text AS \"Last seen\"\n FROM result_candidate_mentions c\n JOIN results r ON r.id = c.result_id\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring})\n AND lower(c.name) NOT IN (SELECT lower(name) FROM brands)\n GROUP BY c.name\n ORDER BY sum(c.mention_count) DESC\n LIMIT 30;" + } + ], + "fieldConfig": { + "defaults": {}, + "overrides": [ + { + "matcher": { + "id": "byRegexp", + "options": "Mentions|Engines|Prompts" + }, + "properties": [ + { + "id": "custom.width", + "value": 100 + } + ] + } + ] + }, + "options": { + "showHeader": true + }, + "description": "Names from lib/brand-candidates.json that answers mentioned and that you are NOT tracking yet. This is a proposal, not a gap: a name here may be genuinely out of scope, or a competitor worth promoting into the brands table. Only you can tell those apart, which is why nothing is added automatically. Empty until you fill in the candidates file." + }, + { + "type": "table", + "title": "Brand ranking — named and cited", + "gridPos": { + "h": 13, + "w": 14, + "x": 10, + "y": 2 + }, + "id": 1, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "table", + "rawQuery": true, + "rawSql": "SELECT b.name || CASE WHEN b.is_own THEN ' (us)' ELSE '' END AS \"Brand\",\n count(*) AS \"Answers\",\n round(100.0 * count(*) FILTER (WHERE m.mentioned) / greatest(count(*), 1), 1) AS \"Named %\",\n round(100.0 * count(*) FILTER (WHERE m.cited) / greatest(count(*), 1), 1) AS \"Cited %\"\n FROM result_brand_mentions m\n JOIN brands b ON b.id = m.brand_id\n JOIN results r ON r.id = m.result_id\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring})\n GROUP BY b.id, b.name, b.is_own\n ORDER BY \"Named %\" DESC;" + } + ], + "fieldConfig": { + "defaults": {}, + "overrides": [ + { + "matcher": { + "id": "byRegexp", + "options": "Answers.*|Prompts" + }, + "properties": [ + { + "id": "custom.width", + "value": 160 + } + ] + }, + { + "matcher": { + "id": "byRegexp", + "options": ".*%" + }, + "properties": [ + { + "id": "custom.width", + "value": 150 + } + ] + }, + { + "matcher": { + "id": "byRegexp", + "options": ".* %" + }, + "properties": [ + { + "id": "unit", + "value": "percent" + }, + { + "id": "max", + "value": 100 + }, + { + "id": "min", + "value": 0 + }, + { + "id": "thresholds", + "value": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } + }, + { + "id": "custom.cellOptions", + "value": { + "type": "gauge", + "mode": "gradient" + } + } + ] + } + ] + }, + "options": { + "showHeader": true + }, + "description": "Every tracked brand, including the ones never named — they hold a row at 0% rather than disappearing. A brand with a high Cited % and a low Named % is being used as a source without being recommended." + }, + { + "type": "timeseries", + "title": "Visibility over time — every tracked brand", + "gridPos": { + "h": 10, + "w": 24, + "x": 0, + "y": 15 + }, + "id": 2, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "time_series", + "rawQuery": true, + "rawSql": "SELECT date_trunc('day', r.completed_at) AS \"time\",\n b.name AS metric,\n round(100.0 * count(*) FILTER (WHERE m.mentioned) / greatest(count(*), 1), 1) AS value\n FROM result_brand_mentions m\n JOIN brands b ON b.id = m.brand_id\n JOIN results r ON r.id = m.result_id\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring})\n GROUP BY 1, b.name\n ORDER BY 1;" + } + ], + "fieldConfig": { + "defaults": { + "unit": "percent", + "min": 0, + "max": 100, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 5, + "showPoints": "auto", + "spanNulls": true + } + }, + "overrides": [] + }, + "options": { + "legend": { + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "description": "Daily share of answers naming each brand. A day with few runs swings hard — check 'Answers analysed' before reading a spike as a trend." + }, + { + "type": "timeseries", + "title": "Brand visibility by engine", + "gridPos": { + "h": 10, + "w": 24, + "x": 0, + "y": 25 + }, + "id": 3, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "time_series", + "rawQuery": true, + "rawSql": "SELECT date_trunc('day', r.completed_at) AS \"time\",\n r.engine AS metric,\n round(100.0 * count(*) FILTER (WHERE m.mentioned) / greatest(count(*), 1), 1) AS value\n FROM result_brand_mentions m\n JOIN brands b ON b.id = m.brand_id\n JOIN results r ON r.id = m.result_id\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring}) AND m.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1)\n GROUP BY 1, r.engine\n ORDER BY 1;" + } + ], + "fieldConfig": { + "defaults": { + "unit": "percent", + "min": 0, + "max": 100, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 5, + "showPoints": "auto", + "spanNulls": true + } + }, + "overrides": [] + }, + "options": { + "legend": { + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "description": "Your mention rate per engine per day, drawn whatever the Engine filter says — this panel deliberately ignores it, so one engine going quiet is visible while you are looking at another. Google sits low by construction: it returns a page of links and only writes prose when an AI Overview was served." + }, + { + "type": "table", + "title": "Pages to get listed on", + "gridPos": { + "h": 13, + "w": 24, + "x": 0, + "y": 35 + }, + "id": 5, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "table", + "rawQuery": true, + "rawSql": "SELECT s.domain AS \"Domain\",\n count(DISTINCT s.result_id) AS \"Answers citing it\",\n count(DISTINCT r.prompt_id) AS \"Prompts\",\n round(100.0 * count(DISTINCT s.result_id) FILTER (WHERE own.mentioned)\n / greatest(count(DISTINCT s.result_id), 1), 1) AS \"Named us %\"\n FROM result_sources s\n JOIN results r ON r.id = s.result_id\n LEFT JOIN result_brand_mentions own\n ON own.result_id = s.result_id AND own.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1)\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring})\n AND s.domain NOT IN (SELECT unnest(domains) FROM brands)\n AND s.url ILIKE '%' || ${urlsearch:sqlstring} || '%'\n GROUP BY s.domain\n ORDER BY \"Answers citing it\" DESC\n LIMIT 30;" + } + ], + "fieldConfig": { + "defaults": {}, + "overrides": [ + { + "matcher": { + "id": "byRegexp", + "options": "Answers.*|Prompts" + }, + "properties": [ + { + "id": "custom.width", + "value": 160 + } + ] + }, + { + "matcher": { + "id": "byRegexp", + "options": ".*%" + }, + "properties": [ + { + "id": "custom.width", + "value": 150 + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "Named us %" + }, + "properties": [ + { + "id": "unit", + "value": "percent" + }, + { + "id": "max", + "value": 100 + }, + { + "id": "min", + "value": 0 + }, + { + "id": "thresholds", + "value": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } + }, + { + "id": "custom.cellOptions", + "value": { + "type": "gauge", + "mode": "gradient" + } + } + ] + } + ] + }, + "options": { + "showHeader": true + }, + "description": "Third-party pages the engines retrieved for your prompts, with how often you were named in those same answers. Your own domains are excluded — they are the panel below. A domain cited constantly where 'Named us %' is low is the clearest place to go earn a listing." + }, + { + "type": "table", + "title": "Our own pages — retrieved, and did the answer name us?", + "gridPos": { + "h": 11, + "w": 24, + "x": 0, + "y": 48 + }, + "id": 6, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "table", + "rawQuery": true, + "rawSql": "SELECT s.domain AS \"Domain\",\n s.kind AS \"Cited as\",\n count(DISTINCT s.result_id) AS \"Answers citing it\",\n round(100.0 * count(DISTINCT s.result_id) FILTER (WHERE own.mentioned)\n / greatest(count(DISTINCT s.result_id), 1), 1) AS \"Named us %\"\n FROM result_sources s\n JOIN results r ON r.id = s.result_id\n LEFT JOIN result_brand_mentions own\n ON own.result_id = s.result_id AND own.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1)\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring})\n AND s.domain IN (\n SELECT unnest(domains) FROM brands WHERE is_own AND enabled\n )\n GROUP BY s.domain, s.kind\n ORDER BY \"Answers citing it\" DESC;" + } + ], + "fieldConfig": { + "defaults": {}, + "overrides": [ + { + "matcher": { + "id": "byRegexp", + "options": "Answers.*|Prompts" + }, + "properties": [ + { + "id": "custom.width", + "value": 160 + } + ] + }, + { + "matcher": { + "id": "byRegexp", + "options": ".*%" + }, + "properties": [ + { + "id": "custom.width", + "value": 150 + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "Named us %" + }, + "properties": [ + { + "id": "unit", + "value": "percent" + }, + { + "id": "max", + "value": 100 + }, + { + "id": "min", + "value": 0 + }, + { + "id": "thresholds", + "value": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } + }, + { + "id": "custom.cellOptions", + "value": { + "type": "gauge", + "mode": "gradient" + } + } + ] + } + ] + }, + "options": { + "showHeader": true + }, + "description": "Your pages that engines actually retrieved. 'Cited as' matters: a link that only ever appears as an ad is not the same win as one in the source rail. Empty means no engine has retrieved your site for these prompts." + }, + { + "type": "table", + "title": "Top YouTube videos — retrieved for your prompts", + "gridPos": { + "h": 11, + "w": 12, + "x": 0, + "y": 59 + }, + "id": 10, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "table", + "rawQuery": true, + "rawSql": "SELECT coalesce(s.label, regexp_replace(s.url, '^https?://', '')) AS \"Video\",\n count(DISTINCT s.result_id) AS \"Retrievals\",\n count(DISTINCT r.engine) AS \"Engines\",\n count(DISTINCT s.result_id) FILTER (WHERE own.mentioned) AS \"Named with us\",\n regexp_replace(s.url, '^https?://', '') AS \"URL\"\n FROM result_sources s\n JOIN results r ON r.id = s.result_id\n LEFT JOIN result_brand_mentions own\n ON own.result_id = s.result_id AND own.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1)\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring})\n AND (s.domain = 'youtube.com' OR s.domain = 'youtu.be'\n OR s.domain LIKE '%.youtube.com')\n GROUP BY s.url, s.label\n ORDER BY 2 DESC, 3 DESC\n LIMIT 25;" + } + ], + "fieldConfig": { + "defaults": {}, + "overrides": [ + { + "matcher": { + "id": "byName", + "options": "Video" + }, + "properties": [ + { + "id": "custom.width", + "value": 355 + }, + { + "id": "links", + "value": [ + { + "title": "Open", + "url": "https://${__data.fields.URL}", + "targetBlank": true + } + ] + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "Retrievals" + }, + "properties": [ + { + "id": "custom.width", + "value": 105 + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "Engines" + }, + "properties": [ + { + "id": "custom.width", + "value": 90 + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "Named with us" + }, + "properties": [ + { + "id": "custom.width", + "value": 130 + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "URL" + }, + "properties": [ + { + "id": "custom.hidden", + "value": true + } + ] + } + ] + }, + "options": { + "showHeader": true, + "footer": { + "show": false + }, + "sortBy": [ + { + "displayName": "Retrievals", + "desc": true + } + ] + }, + "description": "Video the engines pulled in while answering, titled where the payload carried a title and linked so you can open it. A video cited often for your prompts is a placement worth pursuing; 'Named with us' says how often the same answer also named you." + }, + { + "type": "table", + "title": "Top Reddit posts — retrieved for your prompts", + "gridPos": { + "h": 11, + "w": 12, + "x": 12, + "y": 59 + }, + "id": 11, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "table", + "rawQuery": true, + "rawSql": "SELECT coalesce(s.label, regexp_replace(s.url, '^https?://', '')) AS \"Post\",\n count(DISTINCT s.result_id) AS \"Retrievals\",\n count(DISTINCT r.engine) AS \"Engines\",\n count(DISTINCT s.result_id) FILTER (WHERE own.mentioned) AS \"Named with us\",\n regexp_replace(s.url, '^https?://', '') AS \"URL\"\n FROM result_sources s\n JOIN results r ON r.id = s.result_id\n LEFT JOIN result_brand_mentions own\n ON own.result_id = s.result_id AND own.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1)\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring})\n AND (s.domain = 'reddit.com' OR s.domain = 'redd.it'\n OR s.domain LIKE '%.reddit.com')\n GROUP BY s.url, s.label\n ORDER BY 2 DESC, 3 DESC\n LIMIT 25;" + } + ], + "fieldConfig": { + "defaults": {}, + "overrides": [ + { + "matcher": { + "id": "byName", + "options": "Post" + }, + "properties": [ + { + "id": "custom.width", + "value": 355 + }, + { + "id": "links", + "value": [ + { + "title": "Open", + "url": "https://${__data.fields.URL}", + "targetBlank": true + } + ] + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "Retrievals" + }, + "properties": [ + { + "id": "custom.width", + "value": 105 + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "Engines" + }, + "properties": [ + { + "id": "custom.width", + "value": 90 + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "Named with us" + }, + "properties": [ + { + "id": "custom.width", + "value": 130 + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "URL" + }, + "properties": [ + { + "id": "custom.hidden", + "value": true + } + ] + } + ] + }, + "options": { + "showHeader": true, + "footer": { + "show": false + }, + "sortBy": [ + { + "displayName": "Retrievals", + "desc": true + } + ] + }, + "description": "Reddit threads the engines retrieved. Forum posts carry disproportionate weight in AI answers, and unlike a publisher's listicle a thread is something you can take part in. 'Named with us' says how often the same answer also named you." + }, + { + "type": "table", + "title": "Every prompt — how we do on each", + "gridPos": { + "h": 14, + "w": 24, + "x": 0, + "y": 70 + }, + "id": 4, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "table", + "rawQuery": true, + "rawSql": "SELECT p.name AS \"Prompt\",\n count(*) AS \"Answers\",\n round(100.0 * count(*) FILTER (WHERE m.mentioned) / greatest(count(*), 1), 1) AS \"Named %\",\n round(100.0 * count(*) FILTER (WHERE m.cited) / greatest(count(*), 1), 1) AS \"Cited %\",\n round(avg(m.first_position) FILTER (WHERE m.mentioned)) AS \"Avg position\"\n FROM result_brand_mentions m\n JOIN brands b ON b.id = m.brand_id\n JOIN results r ON r.id = m.result_id\n JOIN prompts p ON p.id = r.prompt_id\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring}) AND m.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1)\n GROUP BY p.id, p.name\n ORDER BY \"Named %\" ASC;" + } + ], + "fieldConfig": { + "defaults": {}, + "overrides": [ + { + "matcher": { + "id": "byRegexp", + "options": "Answers.*|Prompts" + }, + "properties": [ + { + "id": "custom.width", + "value": 160 + } + ] + }, + { + "matcher": { + "id": "byRegexp", + "options": ".*%" + }, + "properties": [ + { + "id": "custom.width", + "value": 150 + } + ] + }, + { + "matcher": { + "id": "byRegexp", + "options": ".* %" + }, + "properties": [ + { + "id": "unit", + "value": "percent" + }, + { + "id": "max", + "value": 100 + }, + { + "id": "min", + "value": 0 + }, + { + "id": "thresholds", + "value": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } + }, + { + "id": "custom.cellOptions", + "value": { + "type": "gauge", + "mode": "gradient" + } + } + ] + } + ] + }, + "options": { + "showHeader": true + }, + "description": "Sorted worst first: the prompts at the top are the ones you are losing. Joined through the prompt, so a prompt you disabled still shows its history. Avg position is the character offset of your first mention, averaged over the answers that named you — engines write prose, not a numbered list, so how early you appear is the strongest claim the text supports. Blank means never named." + }, + { + "type": "table", + "title": "Prompt redundancy — prompts that retrieve the same pages", + "gridPos": { + "h": 11, + "w": 24, + "x": 0, + "y": 84 + }, + "id": 12, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "table", + "rawQuery": true, + "rawSql": "WITH per_prompt AS (\n SELECT r.prompt_id, s.domain\n FROM result_sources s\n JOIN results r ON r.id = s.result_id\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring})\n GROUP BY r.prompt_id, s.domain\n )\n SELECT p1.name AS \"Prompt\",\n p2.name AS \"Overlaps with\",\n round(100.0 * count(*) / greatest(\n (SELECT count(*) FROM per_prompt x WHERE x.prompt_id = a.prompt_id), 1\n ), 1) AS \"Shared domains %\"\n FROM per_prompt a\n JOIN per_prompt b ON b.domain = a.domain AND b.prompt_id <> a.prompt_id\n JOIN prompts p1 ON p1.id = a.prompt_id\n JOIN prompts p2 ON p2.id = b.prompt_id\n GROUP BY p1.name, p2.name, a.prompt_id\n ORDER BY 3 DESC\n LIMIT 30;" + } + ], + "fieldConfig": { + "defaults": {}, + "overrides": [ + { + "matcher": { + "id": "byName", + "options": "Shared domains %" + }, + "properties": [ + { + "id": "unit", + "value": "percent" + }, + { + "id": "max", + "value": 100 + }, + { + "id": "min", + "value": 0 + }, + { + "id": "custom.width", + "value": 170 + }, + { + "id": "thresholds", + "value": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + } + }, + { + "id": "custom.cellOptions", + "value": { + "type": "gauge", + "mode": "gradient" + } + } + ] + } + ] + }, + "options": { + "showHeader": true + }, + "description": "Reported, not enforced. A near-duplicate prompt is a second SAMPLE of the same question, which is often worth more than the denominator purity it costs — so a high number here is information, not a defect to fix. A prompt that retrieved nothing is absent rather than shown at 0%: it is unmeasured, not maximally distinctive." + }, + { + "type": "timeseries", + "title": "Data quality — answers that retrieved nothing", + "gridPos": { + "h": 10, + "w": 24, + "x": 0, + "y": 95 + }, + "id": 13, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "time_series", + "rawQuery": true, + "rawSql": "SELECT date_trunc('day', r.completed_at) AS \"time\",\n r.engine AS metric,\n round(100.0 * count(*) FILTER (WHERE s.result_id IS NULL)\n / greatest(count(*), 1), 1) AS value\n FROM results r\n LEFT JOIN (SELECT DISTINCT result_id FROM result_sources) s\n ON s.result_id = r.id\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring})\n GROUP BY 1, r.engine\n ORDER BY 1;" + } + ], + "fieldConfig": { + "defaults": { + "unit": "percent", + "min": 0, + "max": 100, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 5, + "showPoints": "auto", + "spanNulls": true + } + }, + "overrides": [] + }, + "options": { + "legend": { + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "description": "Share of completed answers that came back with no links at all, per engine per day. Some of this is real — an engine can answer from memory — but a line that climbs is usually a payload shape that changed under the extractor, and every source panel above goes quiet with it." + }, + { + "type": "table", + "title": "Top search queries the engines issued", + "gridPos": { + "h": 11, + "w": 24, + "x": 0, + "y": 105 + }, + "id": 8, + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "targets": [ + { + "refId": "A", + "datasource": { + "type": "grafana-postgresql-datasource", + "uid": "${DS_POSTGRES}" + }, + "format": "table", + "rawQuery": true, + "rawSql": "SELECT q.query AS \"Query\",\n count(*) AS \"Times\",\n count(DISTINCT r.engine) AS \"Engines\",\n count(DISTINCT r.prompt_id) AS \"Prompts\"\n FROM result_search_queries q\n JOIN results r ON r.id = q.result_id\n WHERE r.status = 'completed'\n AND $__timeFilter(r.completed_at)\n AND r.engine IN (${engine:sqlstring}) AND q.kind = 'issued'\n GROUP BY q.query\n ORDER BY count(*) DESC, q.query\n LIMIT 40;" + } + ], + "fieldConfig": { + "defaults": {}, + "overrides": [ + { + "matcher": { + "id": "byRegexp", + "options": "Times|Engines|Prompts" + }, + "properties": [ + { + "id": "custom.width", + "value": 100 + } + ] + } + ] + }, + "options": { + "showHeader": true + }, + "description": "What the model typed before it retrieved anything — a stage upstream of every other panel. It does not search your prompt; it rewrites the prompt, often with a vendor list already attached. Only ChatGPT, Copilot, Grok and Perplexity report this, so picking any other engine empties the panel. Ordering by count is nearly arbitrary: these strings are mostly unique, so read it as a sample with repeats surfaced first." + } + ] +} diff --git a/grafana/queries.sql b/grafana/queries.sql index 198d9dd..54e9d9e 100644 --- a/grafana/queries.sql +++ b/grafana/queries.sql @@ -52,3 +52,213 @@ LIMIT 20; -- Retention: delete results older than 90 days (run manually or schedule -- it; the free tier holds months-to-years of daily runs regardless) DELETE FROM results WHERE created_at < now() - interval '90 days'; + + +-- ============================================================ +-- Brand visibility (geo-visibility.json) +-- +-- Panel order and geometry match the internal GEO dashboard in +-- cloro-dev/infra. These read the tables the scheduler tick derives from +-- results.response: result_sources, result_brand_mentions, +-- result_search_queries and result_candidate_mentions. The Grafana macros +-- are replaced here with plain predicates so each block runs in psql. +-- ============================================================ + +-- Named but not tracked — candidates worth adding +SELECT c.name AS "Name", + sum(c.mention_count) AS "Mentions", + count(DISTINCT r.engine) AS "Engines", + count(DISTINCT r.prompt_id) AS "Prompts", + max(r.completed_at)::date::text AS "Last seen" + FROM result_candidate_mentions c + JOIN results r ON r.id = c.result_id + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') + AND lower(c.name) NOT IN (SELECT lower(name) FROM brands) + GROUP BY c.name + ORDER BY sum(c.mention_count) DESC + LIMIT 30; + +-- Brand ranking — named and cited +SELECT b.name || CASE WHEN b.is_own THEN ' (us)' ELSE '' END AS "Brand", + count(*) AS "Answers", + round(100.0 * count(*) FILTER (WHERE m.mentioned) / greatest(count(*), 1), 1) AS "Named %", + round(100.0 * count(*) FILTER (WHERE m.cited) / greatest(count(*), 1), 1) AS "Cited %" + FROM result_brand_mentions m + JOIN brands b ON b.id = m.brand_id + JOIN results r ON r.id = m.result_id + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') + GROUP BY b.id, b.name, b.is_own + ORDER BY "Named %" DESC; + +-- Visibility over time — every tracked brand +SELECT date_trunc('day', r.completed_at) AS "time", + b.name AS metric, + round(100.0 * count(*) FILTER (WHERE m.mentioned) / greatest(count(*), 1), 1) AS value + FROM result_brand_mentions m + JOIN brands b ON b.id = m.brand_id + JOIN results r ON r.id = m.result_id + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') + GROUP BY 1, b.name + ORDER BY 1; + +-- Brand visibility by engine +SELECT date_trunc('day', r.completed_at) AS "time", + r.engine AS metric, + round(100.0 * count(*) FILTER (WHERE m.mentioned) / greatest(count(*), 1), 1) AS value + FROM result_brand_mentions m + JOIN brands b ON b.id = m.brand_id + JOIN results r ON r.id = m.result_id + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') AND m.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1) + GROUP BY 1, r.engine + ORDER BY 1; + +-- Pages to get listed on +SELECT s.domain AS "Domain", + count(DISTINCT s.result_id) AS "Answers citing it", + count(DISTINCT r.prompt_id) AS "Prompts", + round(100.0 * count(DISTINCT s.result_id) FILTER (WHERE own.mentioned) + / greatest(count(DISTINCT s.result_id), 1), 1) AS "Named us %" + FROM result_sources s + JOIN results r ON r.id = s.result_id + LEFT JOIN result_brand_mentions own + ON own.result_id = s.result_id AND own.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1) + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') + AND s.domain NOT IN (SELECT unnest(domains) FROM brands) + AND s.url ILIKE '%' || '' || '%' + GROUP BY s.domain + ORDER BY "Answers citing it" DESC + LIMIT 30; + +-- Our own pages — retrieved, and did the answer name us? +SELECT s.domain AS "Domain", + s.kind AS "Cited as", + count(DISTINCT s.result_id) AS "Answers citing it", + round(100.0 * count(DISTINCT s.result_id) FILTER (WHERE own.mentioned) + / greatest(count(DISTINCT s.result_id), 1), 1) AS "Named us %" + FROM result_sources s + JOIN results r ON r.id = s.result_id + LEFT JOIN result_brand_mentions own + ON own.result_id = s.result_id AND own.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1) + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') + AND s.domain IN ( + SELECT unnest(domains) FROM brands WHERE is_own AND enabled + ) + GROUP BY s.domain, s.kind + ORDER BY "Answers citing it" DESC; + +-- Top YouTube videos — retrieved for your prompts +SELECT coalesce(s.label, regexp_replace(s.url, '^https?://', '')) AS "Video", + count(DISTINCT s.result_id) AS "Retrievals", + count(DISTINCT r.engine) AS "Engines", + count(DISTINCT s.result_id) FILTER (WHERE own.mentioned) AS "Named with us", + regexp_replace(s.url, '^https?://', '') AS "URL" + FROM result_sources s + JOIN results r ON r.id = s.result_id + LEFT JOIN result_brand_mentions own + ON own.result_id = s.result_id AND own.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1) + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') + AND (s.domain = 'youtube.com' OR s.domain = 'youtu.be' + OR s.domain LIKE '%.youtube.com') + GROUP BY s.url, s.label + ORDER BY 2 DESC, 3 DESC + LIMIT 25; + +-- Top Reddit posts — retrieved for your prompts +SELECT coalesce(s.label, regexp_replace(s.url, '^https?://', '')) AS "Post", + count(DISTINCT s.result_id) AS "Retrievals", + count(DISTINCT r.engine) AS "Engines", + count(DISTINCT s.result_id) FILTER (WHERE own.mentioned) AS "Named with us", + regexp_replace(s.url, '^https?://', '') AS "URL" + FROM result_sources s + JOIN results r ON r.id = s.result_id + LEFT JOIN result_brand_mentions own + ON own.result_id = s.result_id AND own.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1) + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') + AND (s.domain = 'reddit.com' OR s.domain = 'redd.it' + OR s.domain LIKE '%.reddit.com') + GROUP BY s.url, s.label + ORDER BY 2 DESC, 3 DESC + LIMIT 25; + +-- Every prompt — how we do on each +SELECT p.name AS "Prompt", + count(*) AS "Answers", + round(100.0 * count(*) FILTER (WHERE m.mentioned) / greatest(count(*), 1), 1) AS "Named %", + round(100.0 * count(*) FILTER (WHERE m.cited) / greatest(count(*), 1), 1) AS "Cited %", + round(avg(m.first_position) FILTER (WHERE m.mentioned)) AS "Avg position" + FROM result_brand_mentions m + JOIN brands b ON b.id = m.brand_id + JOIN results r ON r.id = m.result_id + JOIN prompts p ON p.id = r.prompt_id + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') AND m.brand_id = (SELECT id FROM brands WHERE is_own AND enabled LIMIT 1) + GROUP BY p.id, p.name + ORDER BY "Named %" ASC; + +-- Prompt redundancy — prompts that retrieve the same pages +WITH per_prompt AS ( + SELECT r.prompt_id, s.domain + FROM result_sources s + JOIN results r ON r.id = s.result_id + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') + GROUP BY r.prompt_id, s.domain + ) + SELECT p1.name AS "Prompt", + p2.name AS "Overlaps with", + round(100.0 * count(*) / greatest( + (SELECT count(*) FROM per_prompt x WHERE x.prompt_id = a.prompt_id), 1 + ), 1) AS "Shared domains %" + FROM per_prompt a + JOIN per_prompt b ON b.domain = a.domain AND b.prompt_id <> a.prompt_id + JOIN prompts p1 ON p1.id = a.prompt_id + JOIN prompts p2 ON p2.id = b.prompt_id + GROUP BY p1.name, p2.name, a.prompt_id + ORDER BY 3 DESC + LIMIT 30; + +-- Data quality — answers that retrieved nothing +SELECT date_trunc('day', r.completed_at) AS "time", + r.engine AS metric, + round(100.0 * count(*) FILTER (WHERE s.result_id IS NULL) + / greatest(count(*), 1), 1) AS value + FROM results r + LEFT JOIN (SELECT DISTINCT result_id FROM result_sources) s + ON s.result_id = r.id + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') + GROUP BY 1, r.engine + ORDER BY 1; + +-- Top search queries the engines issued +SELECT q.query AS "Query", + count(*) AS "Times", + count(DISTINCT r.engine) AS "Engines", + count(DISTINCT r.prompt_id) AS "Prompts" + FROM result_search_queries q + JOIN results r ON r.id = q.result_id + WHERE r.status = 'completed' + AND r.completed_at > now() - interval '30 days' + AND r.engine IN ('chatgpt','perplexity','gemini','aimode','google') AND q.kind = 'issued' + GROUP BY q.query + ORDER BY count(*) DESC, q.query + LIMIT 40; diff --git a/lib/brand-candidates.json b/lib/brand-candidates.json new file mode 100644 index 0000000..f2d2624 --- /dev/null +++ b/lib/brand-candidates.json @@ -0,0 +1,42 @@ +{ + "note": "Names to WATCH FOR but not track. Anything listed here that an answer names, and that is not in your `brands` table, shows up in the dashboard's 'Named but not tracked' panel — the shortlist of competitors worth adding. Matching is literal and case-insensitive, the same rule the brands table uses. Give each candidate the spellings an engine would actually write: the bare name, the legal form, and any domain people say out loud. Every alias counts toward the same candidate, so 'Acme' and 'Acme, Inc' report as one row. A plain string works as shorthand when a name needs no aliases. Editing this file re-derives the whole history on the next scheduler tick, so a name added today is scored against answers already stored. THE NAMES BELOW ARE PLACEHOLDERS — fictional companies, there to show the shape and to give the seeded demo something to find. Replace every one of them with the real vendors in your category.", + "candidates": [ + { + "name": "Initrode", + "aliases": ["Initrode LLC"] + }, + { + "name": "Hooli", + "aliases": ["Hooli Inc", "Hooli XYZ", "hooli.com"] + }, + { + "name": "Vandelay", + "aliases": ["Vandelay Industries", "Vandelay Inc"] + }, + { + "name": "Soylent", + "aliases": ["Soylent Corp", "Soylent Corporation"] + }, + { + "name": "Cyberdyne", + "aliases": ["Cyberdyne Systems"] + }, + { + "name": "Tyrell", + "aliases": ["Tyrell Corp", "Tyrell Corporation"] + }, + { + "name": "Aperture", + "aliases": ["Aperture Science"] + }, + { + "name": "Massive Dynamic", + "aliases": ["MassiveDynamic"] + }, + { + "name": "Gringotts", + "aliases": [] + }, + "Nakatomi" + ] +} diff --git a/lib/db/schema.ts b/lib/db/schema.ts index cf8222e..71929ea 100644 --- a/lib/db/schema.ts +++ b/lib/db/schema.ts @@ -46,6 +46,11 @@ export const results = pgTable( .notNull() .defaultNow(), completedAt: timestamp("completed_at", { withTimezone: true }), + // Which version of the extractor last derived sources and brand + // mentions from this row. Null means never. Bumping EXTRACTION_REVISION + // (or clearing this column, as a brand-list edit does) is what makes the + // cron pick a finished result up again. + extractionRevision: integer("extraction_revision"), }, (table) => [ uniqueIndex("results_task_id_idx").on(table.taskId), @@ -54,9 +59,214 @@ export const results = pgTable( index("results_pending_idx") .on(table.createdAt) .where(sql`${table.status} = 'pending'`), + // The refresh queue, and it really is a queue: the predicate matches + // only rows still waiting, so a caught-up deployment has an EMPTY + // index and the tick's queue scan costs nothing. + // + // This works only because "needs extraction" is `IS NULL` rather than + // "stamp differs from the current one". A `<>` against a value the + // index cannot know is not sargable, and Postgres answered it with a + // sequential scan of every completed row, every tick, forever. + index("results_unextracted_idx") + .on(table.completedAt) + .where( + sql`${table.status} = 'completed' AND ${table.extractionRevision} IS NULL`, + ), ], ); export type Prompt = typeof prompts.$inferSelect; export type Result = typeof results.$inferSelect; export type NewResult = typeof results.$inferInsert; + +/** + * Brands to look for in answers. User-configured: geo-tracker ships no + * brand list and makes no judgement about who competes with whom. The + * extractor only does literal, case-insensitive matching of the names and + * aliases declared here. + */ +export const brands = pgTable( + "brands", + { + id: uuid("id").primaryKey().defaultRandom(), + name: text("name").notNull(), + // Extra spellings that count as the same brand ("Acme Corp", "acme.io"). + aliases: text("aliases") + .array() + .notNull() + .default(sql`'{}'::text[]`), + // Domains that count as this brand being cited, not just named. + domains: text("domains") + .array() + .notNull() + .default(sql`'{}'::text[]`), + // Your own brand vs the rest. Drives the "us vs them" panels; purely a + // label, the extractor treats every brand identically. + isOwn: boolean("is_own").notNull().default(false), + enabled: boolean("enabled").notNull().default(true), + createdAt: timestamp("created_at", { withTimezone: true }) + .notNull() + .defaultNow(), + }, + (table) => [uniqueIndex("brands_name_idx").on(sql`lower(${table.name})`)], +); + +/** + * One row per link an engine returned, flattened out of `results.response`. + * + * Derived data: every row is reproducible from the raw response, and the + * refresh rebuilds a result's rows wholesale rather than patching them. + */ +export const resultSources = pgTable( + "result_sources", + { + id: uuid("id").primaryKey().defaultRandom(), + resultId: uuid("result_id") + .notNull() + .references(() => results.id, { onDelete: "cascade" }), + // Where in the payload the link came from: sources, citation_pill, + // organic, news, ad, video, people_also_ask. Kept because a citation + // and an ad are not the same evidence, and no panel should pool them. + kind: text("kind").notNull(), + position: integer("position"), + url: text("url").notNull(), + // Registrable-ish host, lowercased and stripped of "www.". Denormalised + // so panels can GROUP BY it without parsing URLs in SQL. + domain: text("domain").notNull(), + label: text("label"), + }, + (table) => [ + index("result_sources_result_idx").on(table.resultId), + index("result_sources_domain_idx").on(table.domain, table.kind), + ], +); + +/** + * One row per (completed result × enabled brand) — including the brands + * that were NOT mentioned. + * + * The absent rows are the point: "share of voice" needs a denominator, and + * a table that only holds hits cannot tell "never named" apart from "never + * asked". Panels count rows, not nulls. + */ +export const resultBrandMentions = pgTable( + "result_brand_mentions", + { + id: uuid("id").primaryKey().defaultRandom(), + resultId: uuid("result_id") + .notNull() + .references(() => results.id, { onDelete: "cascade" }), + brandId: uuid("brand_id") + .notNull() + .references(() => brands.id, { onDelete: "cascade" }), + mentioned: boolean("mentioned").notNull(), + mentionCount: integer("mention_count").notNull().default(0), + // Character offset of the first mention in the answer text. Null when + // absent. A rank would be more useful and is not honest here: engines + // return prose, not a numbered list, so "first named" is the strongest + // claim the text supports. + firstPosition: integer("first_position"), + // A brand domain appeared in the links, whether or not the prose named + // it. Being cited and being named are different outcomes. + cited: boolean("cited").notNull().default(false), + citedSourceCount: integer("cited_source_count").notNull().default(0), + }, + (table) => [ + uniqueIndex("result_brand_mentions_result_brand_idx").on( + table.resultId, + table.brandId, + ), + index("result_brand_mentions_brand_idx").on(table.brandId), + ], +); + +/** + * One row, holding the extraction stamp the stored rows were built at. + * + * The tick compares this to the current stamp and, when they differ, + * clears `results.extraction_revision` in one statement. That is what lets + * the queue be an `IS NULL` test — a sargable one — instead of a + * comparison against a constant no index can hold. + */ +export const extractionState = pgTable("extraction_state", { + // A single row, pinned: `true` is the only value the check allows. + id: boolean("id").primaryKey().default(true), + stamp: integer("stamp").notNull(), + updatedAt: timestamp("updated_at", { withTimezone: true }) + .notNull() + .defaultNow(), +}); + +export type ExtractionState = typeof extractionState.$inferSelect; + +export type Brand = typeof brands.$inferSelect; +export type NewBrand = typeof brands.$inferInsert; +export type ResultSource = typeof resultSources.$inferSelect; +export type NewResultSource = typeof resultSources.$inferInsert; +export type ResultBrandMention = typeof resultBrandMentions.$inferSelect; +export type NewResultBrandMention = typeof resultBrandMentions.$inferInsert; + +/** + * The literal queries an engine typed, flattened out of the payload. + * + * A stage upstream of every other derived table: `result_brand_mentions` + * records that a brand was named and `result_sources` that a page was + * retrieved; this is what the model searched for before either happened. + * It does not search your prompt — it rewrites the prompt, often with a + * vendor list already attached. + */ +export const resultSearchQueries = pgTable( + "result_search_queries", + { + id: uuid("id").primaryKey().defaultRandom(), + resultId: uuid("result_id") + .notNull() + .references(() => results.id, { onDelete: "cascade" }), + // `issued` is what the model actually searched. `suggested` is the + // follow-up chips it offered the reader. Kept apart because they are + // different acts: one is the engine's own reasoning, the other is + // navigation furniture, and pooling them would misread the first. + kind: text("kind").notNull(), + position: integer("position"), + query: text("query").notNull(), + }, + (table) => [ + index("result_search_queries_result_idx").on(table.resultId), + index("result_search_queries_kind_idx").on(table.kind), + ], +); + +/** + * Names from `lib/brand-candidates.json` that an answer mentioned. + * + * Separate from `result_brand_mentions` because these are not tracked + * brands: there is no denominator to hold, no citation to check, and a + * candidate that never appears needs no row. This table only ever says + * "this name was named this often", which is the evidence for deciding + * whether to promote it into `brands`. + */ +export const resultCandidateMentions = pgTable( + "result_candidate_mentions", + { + id: uuid("id").primaryKey().defaultRandom(), + resultId: uuid("result_id") + .notNull() + .references(() => results.id, { onDelete: "cascade" }), + name: text("name").notNull(), + mentionCount: integer("mention_count").notNull(), + }, + (table) => [ + uniqueIndex("result_candidate_mentions_result_name_idx").on( + table.resultId, + table.name, + ), + index("result_candidate_mentions_name_idx").on(table.name), + ], +); + +export type ResultSearchQuery = typeof resultSearchQueries.$inferSelect; +export type NewResultSearchQuery = typeof resultSearchQueries.$inferInsert; +export type ResultCandidateMention = + typeof resultCandidateMentions.$inferSelect; +export type NewResultCandidateMention = + typeof resultCandidateMentions.$inferInsert; diff --git a/lib/extract.test.ts b/lib/extract.test.ts new file mode 100644 index 0000000..d95d629 --- /dev/null +++ b/lib/extract.test.ts @@ -0,0 +1,379 @@ +import { describe, expect, it } from "vitest"; + +import { + BRAND_CANDIDATES, + EXTRACTION_STAMP, + answerText, + extractCandidates, + extractMentions, + extractSearchQueries, + extractSources, + toDomain, + type ExtractedSource, +} from "./extract"; +import type { Brand } from "./db/schema"; + +function brand(overrides: Partial & { name: string }): Brand { + return { + id: overrides.name, + aliases: [], + domains: [], + isOwn: false, + enabled: true, + createdAt: new Date(), + ...overrides, + }; +} + +function source(domain: string): ExtractedSource { + return { + kind: "source", + position: 1, + url: `https://${domain}/x`, + domain, + label: null, + }; +} + +describe("toDomain", () => { + it("lowercases and drops a www prefix", () => { + expect(toDomain("https://WWW.Example.com/a?b=c")).toBe("example.com"); + }); + + it("keeps subdomains distinct", () => { + expect(toDomain("https://docs.example.com/a")).toBe("docs.example.com"); + }); + + it("returns null for something that is not a URL", () => { + expect(toDomain("not a url")).toBeNull(); + }); +}); + +describe("answerText", () => { + it("unwraps the async envelope", () => { + expect(answerText({ success: true, result: { text: "hi" } })).toBe("hi"); + }); + + it("accepts a bare result object", () => { + expect(answerText({ text: "hi" })).toBe("hi"); + }); + + it("falls back to the AI Overview on a Google payload", () => { + expect( + answerText({ + result: { organicResults: [], aioverview: { text: "ov" } }, + }), + ).toBe("ov"); + }); + + it("is empty for a plain SERP, which wrote no prose", () => { + expect( + answerText({ result: { organicResults: [], aioverview: null } }), + ).toBe(""); + }); + + it("is empty rather than throwing on junk", () => { + expect(answerText(null)).toBe(""); + expect(answerText("nonsense")).toBe(""); + }); +}); + +describe("extractSources", () => { + it("tags each link with where it came from", () => { + const sources = extractSources({ + result: { + sources: [{ position: 1, url: "https://a.com/1", label: "A" }], + citationPills: [{ position: 2, url: "https://b.com/2", label: "B" }], + }, + }); + + expect(sources).toEqual([ + { + kind: "source", + position: 1, + url: "https://a.com/1", + domain: "a.com", + label: "A", + }, + { + kind: "citation_pill", + position: 2, + url: "https://b.com/2", + domain: "b.com", + label: "B", + }, + ]); + }); + + it("keeps a link that is both a source and a pill, once per kind", () => { + const sources = extractSources({ + result: { + sources: [{ url: "https://a.com/1" }], + citationPills: [{ url: "https://a.com/1" }], + }, + }); + expect(sources.map((s) => s.kind)).toEqual(["source", "citation_pill"]); + }); + + it("does not count one link twice within a kind", () => { + const sources = extractSources({ + result: { + sources: [{ url: "https://a.com/1" }, { url: "https://a.com/1" }], + }, + }); + expect(sources).toHaveLength(1); + }); + + it("reads Google's organic list, ads and people-also-ask", () => { + const sources = extractSources({ + result: { + organicResults: [{ position: 1, link: "https://o.com", title: "O" }], + ads: [{ position: 1, url: "https://ad.com" }], + peopleAlsoAsk: [ + { link: "https://paa.com", sources: [{ url: "https://deep.com" }] }, + ], + }, + }); + expect(sources.map((s) => [s.kind, s.domain])).toEqual([ + ["organic", "o.com"], + ["ad", "ad.com"], + ["people_also_ask", "paa.com"], + ["people_also_ask", "deep.com"], + ]); + }); + + it("skips entries with no usable URL instead of failing the batch", () => { + const sources = extractSources({ + result: { sources: [{ url: "" }, { url: "javascript:void" }, null, 7] }, + }); + expect(sources).toEqual([]); + }); +}); + +describe("extractMentions", () => { + it("returns a row per brand, including the ones never named", () => { + const mentions = extractMentions( + "Acme is great.", + [], + [brand({ name: "Acme" }), brand({ name: "Globex" })], + ); + + expect(mentions.map((m) => [m.brandId, m.mentioned])).toEqual([ + ["Acme", true], + ["Globex", false], + ]); + }); + + it("matches case-insensitively and counts every occurrence", () => { + const [mention] = extractMentions( + "acme, then ACME.", + [], + [brand({ name: "Acme" })], + ); + expect(mention.mentionCount).toBe(2); + }); + + it("records where the brand was first named", () => { + const [mention] = extractMentions( + "First Globex, then Acme.", + [], + [brand({ name: "Acme" })], + ); + expect(mention.firstPosition).toBe(19); + }); + + it("takes the earliest position across name and aliases", () => { + const [mention] = extractMentions( + "acme.io beats Acme Corp", + [], + [brand({ name: "Acme Corp", aliases: ["acme.io"] })], + ); + expect(mention.firstPosition).toBe(0); + expect(mention.mentionCount).toBe(2); + }); + + it("does not match a brand inside a longer word", () => { + const [mention] = extractMentions( + "acmecorp and acme-killer", + [], + [brand({ name: "acme" })], + ); + expect(mention.mentioned).toBe(false); + }); + + it("matches a name that ends at a dot, which \\b would miss", () => { + const [mention] = extractMentions( + "We use acme.io daily.", + [], + [brand({ name: "acme.io" })], + ); + expect(mention.mentioned).toBe(true); + }); + + it("treats a subdomain of a brand domain as a citation", () => { + const [mention] = extractMentions( + "", + [source("blog.acme.io")], + [brand({ name: "Acme", domains: ["acme.io"] })], + ); + expect(mention.cited).toBe(true); + expect(mention.citedSourceCount).toBe(1); + }); + + it("does not let a lookalike domain count as a citation", () => { + const [mention] = extractMentions( + "", + [source("notacme.io")], + [brand({ name: "Acme", domains: ["acme.io"] })], + ); + expect(mention.cited).toBe(false); + }); + + it("separates being cited from being named", () => { + const [mention] = extractMentions( + "The answer names nobody.", + [source("acme.io")], + [brand({ name: "Acme", domains: ["acme.io"] })], + ); + expect(mention.mentioned).toBe(false); + expect(mention.cited).toBe(true); + }); + + it("does not count a name and its own alias twice in one phrase", () => { + const [mention] = extractMentions( + "Acme Corp is the vendor.", + [], + [brand({ name: "Acme", aliases: ["Acme Corp"] })], + ); + expect(mention.mentionCount).toBe(1); + }); + + it("still counts the bare name where it stands alone", () => { + const [mention] = extractMentions( + "Acme Corp shipped it. Acme is well known.", + [], + [brand({ name: "Acme", aliases: ["Acme Corp"] })], + ); + expect(mention.mentionCount).toBe(2); + }); + + it("ignores an empty alias rather than matching everything", () => { + const [mention] = extractMentions( + "anything at all", + [], + [brand({ name: "Acme", aliases: ["", " "] })], + ); + expect(mention.mentioned).toBe(false); + }); +}); + +describe("extractSearchQueries", () => { + it("separates what the engine searched from what it suggested", () => { + const queries = extractSearchQueries({ + result: { + searchQueries: ["best crm 2026"], + related_queries: ["crm for startups"], + }, + }); + expect(queries.map((q) => [q.kind, q.query])).toEqual([ + ["issued", "best crm 2026"], + ["suggested", "crm for startups"], + ]); + }); + + it("reads Perplexity's differently-named field as the same thing", () => { + const queries = extractSearchQueries({ + result: { search_model_queries: ["find the best crm"] }, + }); + expect(queries).toEqual([ + { kind: "issued", position: 1, query: "find the best crm" }, + ]); + }); + + it("unwraps Google's related searches, which are objects not strings", () => { + const queries = extractSearchQueries({ + result: { relatedSearches: [{ query: "cheap crm", link: "https://g" }] }, + }); + expect(queries).toEqual([ + { kind: "suggested", position: 1, query: "cheap crm" }, + ]); + }); + + it("drops blanks and repeats, case-insensitively", () => { + const queries = extractSearchQueries({ + result: { searchQueries: ["Best CRM", "best crm", " ", ""] }, + }); + expect(queries).toHaveLength(1); + }); + + it("is empty for an engine that reports no queries", () => { + expect(extractSearchQueries({ result: { text: "an answer" } })).toEqual([]); + }); +}); + +describe("extractCandidates", () => { + const cand = (name: string, aliases: string[] = []) => ({ name, aliases }); + + it("ships a worked example list, not an empty one", () => { + expect(BRAND_CANDIDATES.length).toBeGreaterThan(0); + // A blank or untrimmed entry would be silently dropped and the panel + // would quietly under-report. + for (const candidate of BRAND_CANDIDATES) { + expect(candidate.name.trim()).toBe(candidate.name); + expect(candidate.name.length).toBeGreaterThan(0); + for (const alias of candidate.aliases) { + expect(alias.trim()).toBe(alias); + expect(alias.length).toBeGreaterThan(0); + } + } + }); + + it("accepts a bare string as shorthand for a candidate with no aliases", () => { + const shorthand = BRAND_CANDIDATES.find((c) => c.aliases.length === 0); + expect(shorthand).toBeDefined(); + }); + + it("returns only the hits, never a row of zeroes", () => { + const found = extractCandidates("Globex leads, Globex again", [ + cand("Globex"), + cand("Initech"), + ]); + expect(found).toEqual([{ name: "Globex", mentionCount: 2 }]); + }); + + it("folds an alias into the canonical name rather than splitting it", () => { + // Two mentions, not three: the bare "Acme", then "Acme, Inc" claiming + // the phrase the shorter term would otherwise have double-counted. + const found = extractCandidates( + "Acme is good. Acme, Inc is the same firm.", + [cand("Acme", ["Acme, Inc"])], + ); + expect(found).toEqual([{ name: "Acme", mentionCount: 2 }]); + }); + + it("finds a candidate named only by an alias", () => { + const found = extractCandidates("We evaluated Vandelay Industries.", [ + cand("Vandelay", ["Vandelay Industries"]), + ]); + expect(found).toEqual([{ name: "Vandelay", mentionCount: 1 }]); + }); + + it("uses the same word-boundary rule as a tracked brand", () => { + expect(extractCandidates("globexcorp ships", [cand("Globex")])).toEqual([]); + expect(extractCandidates("we use globex.com", [cand("Globex")])).toEqual([ + { name: "Globex", mentionCount: 1 }, + ]); + }); + + it("finds nothing in an answer with no prose", () => { + expect(extractCandidates("", [cand("Globex")])).toEqual([]); + }); +}); + +describe("EXTRACTION_STAMP", () => { + it("is a stable non-negative integer that fits the column", () => { + expect(Number.isInteger(EXTRACTION_STAMP)).toBe(true); + expect(EXTRACTION_STAMP).toBeGreaterThanOrEqual(0); + expect(EXTRACTION_STAMP).toBeLessThanOrEqual(2147483647); + }); +}); diff --git a/lib/extract.ts b/lib/extract.ts new file mode 100644 index 0000000..984652e --- /dev/null +++ b/lib/extract.ts @@ -0,0 +1,471 @@ +import candidatesFile from "./brand-candidates.json"; +import type { Brand } from "./db/schema"; + +/** + * Turns one raw engine response into rows: the links it returned, and + * which of the tracked brands its answer named. + * + * Everything here is mechanical — unnest an array, lowercase a host, match + * a literal string. Nothing scores or ranks an answer; the judgement of + * which brands matter is the user's, declared in the `brands` table. + */ + +/** + * Bump when the rules below change, so finished results are re-derived. + * + * 3: matching became non-overlapping. A brand or candidate whose alias + * contains its own name ("Vandelay" inside "Vandelay Industries") was + * counting one phrase twice, so every stored mention_count derived + * before this is potentially inflated and has to be rebuilt. + * 2: search queries and brand candidates began to be extracted. + */ +export const EXTRACTION_REVISION = 3; + +/** + * A name to watch for without tracking it, plus the other spellings that + * count as the same one. See `brand-candidates.json`. + */ +export interface BrandCandidate { + name: string; + aliases: string[]; +} + +/** + * The candidate list, normalised. + * + * Read through `unknown` rather than trusting the import: this file is + * meant to be hand-edited, so a stray number, a missing `aliases` or a + * bare string in place of an object are all expected inputs, not + * corruption. A plain string is shorthand for a candidate with no aliases. + */ +export const BRAND_CANDIDATES: BrandCandidate[] = ( + (candidatesFile as { candidates?: unknown[] }).candidates ?? [] +) + .map((entry): BrandCandidate | null => { + if (typeof entry === "string") { + const name = entry.trim(); + return name.length > 0 ? { name, aliases: [] } : null; + } + if (typeof entry !== "object" || entry === null) return null; + const record = entry as { name?: unknown; aliases?: unknown }; + const name = typeof record.name === "string" ? record.name.trim() : ""; + if (name.length === 0) return null; + const aliases = (Array.isArray(record.aliases) ? record.aliases : []) + .filter((alias): alias is string => typeof alias === "string") + .map((alias) => alias.trim()) + .filter((alias) => alias.length > 0); + return { name, aliases }; + }) + .filter((candidate): candidate is BrandCandidate => candidate !== null); + +/** FNV-1a, folded to 31 bits so it fits the integer column. */ +function fingerprint(text: string): number { + let hash = 0x811c9dc5; + for (let i = 0; i < text.length; i += 1) { + hash ^= text.charCodeAt(i); + hash = Math.imul(hash, 0x01000193); + } + return hash & 0x7fffffff; +} + +/** + * What `results.extraction_revision` stores: a fingerprint of everything + * that decides the output, not a version number. + * + * The candidate list is a FILE, so nothing calls + * `markAllForReextraction()` when it changes — unlike the brands table, + * which has a write path that can. Folding the list into the stamp means + * an edited file simply no longer matches what is stored, and the next + * tick re-derives the history on its own. Sorted first so reordering the + * file is not a change. + */ +export const EXTRACTION_STAMP = fingerprint( + `${EXTRACTION_REVISION}:${BRAND_CANDIDATES.map( + (candidate) => + `${candidate.name}|${[...candidate.aliases].sort().join(",")}`, + ) + .sort() + .join("\u0000")}`, +); + +export interface ExtractedSource { + kind: string; + position: number | null; + url: string; + domain: string; + label: string | null; +} + +export interface ExtractedMention { + brandId: string; + mentioned: boolean; + mentionCount: number; + firstPosition: number | null; + cited: boolean; + citedSourceCount: number; +} + +export interface ExtractedQuery { + kind: string; + position: number | null; + query: string; +} + +export interface ExtractedCandidate { + name: string; + mentionCount: number; +} + +export interface Extraction { + text: string; + sources: ExtractedSource[]; + mentions: ExtractedMention[]; + queries: ExtractedQuery[]; + candidates: ExtractedCandidate[]; +} + +/** + * Host of a URL, lowercased and stripped of a leading "www.". + * + * Deliberately not a public-suffix lookup: that needs a list dependency + * that goes stale, and every panel here groups by what the engine actually + * linked to. "docs.example.com" and "example.com" stay distinct, which is + * the honest answer — they are different pages to get listed on. + */ +export function toDomain(url: string): string | null { + try { + const parsed = new URL(url); + // `new URL` happily parses "javascript:void" and "mailto:x", which have + // no hostname at all — without this they land in the table as rows with + // an empty domain that every GROUP BY then reports as a real site. + if (parsed.protocol !== "http:" && parsed.protocol !== "https:") { + return null; + } + const host = parsed.hostname.toLowerCase(); + if (host.length === 0) return null; + return host.startsWith("www.") ? host.slice(4) : host; + } catch { + return null; + } +} + +function pushSource( + into: ExtractedSource[], + seen: Set, + kind: string, + raw: unknown, + urlKey = "url", + labelKey = "label", +): void { + if (typeof raw !== "object" || raw === null) return; + const item = raw as Record; + const url = item[urlKey]; + if (typeof url !== "string" || url.length === 0) return; + + const domain = toDomain(url); + if (domain === null) return; + + // One link can arrive twice — as a source rail entry and again as an + // inline citation pill. Both are real, but a panel counting "pages the + // engine retrieved" must not count the page twice for one answer. + const key = `${kind} ${url}`; + if (seen.has(key)) return; + seen.add(key); + + const position = item.position; + const label = item[labelKey] ?? item.title; + into.push({ + kind, + position: typeof position === "number" ? position : null, + url, + domain, + label: typeof label === "string" ? label : null, + }); +} + +function asArray(value: unknown): unknown[] { + return Array.isArray(value) ? value : []; +} + +/** + * The response body geo-tracker stored, unwrapped. + * + * The async API nests the payload under `result`, but a webhook redelivery + * or a hand-inserted row may hold the result object directly. Accepting + * both costs one line and avoids an extraction that silently finds nothing. + */ +function unwrap(response: unknown): Record | null { + if (typeof response !== "object" || response === null) return null; + const body = response as Record; + const inner = body.result; + if (typeof inner === "object" && inner !== null) { + return inner as Record; + } + return body; +} + +/** + * The prose an engine wrote, or "" for the engines that write none. + * + * Google returns a page of links, not an answer, so its only text is the + * AI Overview when one was served. A brand cannot be "named in the answer" + * on a plain SERP, and reporting 0 mentions there is correct rather than + * missing data. + */ +export function answerText(response: unknown): string { + const result = unwrap(response); + if (result === null) return ""; + + if (typeof result.text === "string") return result.text; + + const overview = result.aioverview; + if (typeof overview === "object" && overview !== null) { + const text = (overview as Record).text; + if (typeof text === "string") return text; + } + return ""; +} + +/** Every link in the payload, tagged by where it came from. */ +export function extractSources(response: unknown): ExtractedSource[] { + const result = unwrap(response); + if (result === null) return []; + + const sources: ExtractedSource[] = []; + const seen = new Set(); + + // Chat engines: the source rail and the inline pills. + for (const item of asArray(result.sources)) { + pushSource(sources, seen, "source", item); + } + for (const item of asArray(result.citationPills)) { + pushSource(sources, seen, "citation_pill", item); + } + + // Google: organic links, the news list, ads, and the AI Overview's own + // rail. Kept apart by `kind` — an ad is not a citation. + for (const item of asArray(result.organicResults)) { + pushSource(sources, seen, "organic", item, "link", "title"); + } + for (const item of asArray(result.newsResults)) { + pushSource(sources, seen, "news", item, "link", "title"); + } + for (const item of asArray(result.ads)) { + pushSource(sources, seen, "ad", item); + } + for (const item of asArray(result.videos)) { + pushSource(sources, seen, "video", item, "link", "title"); + } + for (const item of asArray(result.peopleAlsoAsk)) { + if (typeof item !== "object" || item === null) continue; + const entry = item as Record; + pushSource(sources, seen, "people_also_ask", entry, "link", "title"); + for (const nested of asArray(entry.sources)) { + pushSource(sources, seen, "people_also_ask", nested); + } + } + + const overview = result.aioverview; + if (typeof overview === "object" && overview !== null) { + const entry = overview as Record; + for (const item of asArray(entry.sources)) { + pushSource(sources, seen, "ai_overview", item); + } + for (const item of asArray(entry.citationPills)) { + pushSource(sources, seen, "ai_overview", item); + } + } + + return sources; +} + +function escapeRegExp(value: string): string { + return value.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"); +} + +/** + * Match a brand name as a whole word, case-insensitively. + * + * `\b` is wrong at both ends for the names this deals with: it does not + * fire next to a dot, so "acme.io" would never match its own alias, and it + * fires inside a hyphenated word, so "acme" would match "acme-killer". + * + * The lookarounds below count letters, digits, underscores AND hyphens as + * word characters, so "acme-killer" and "acmecorp" are both left alone, + * while a dot stays a boundary — which is what lets a brand tracked as + * "acme" match the "acme.io" an engine wrote. + */ +function matchSpans(haystack: string, needle: string): [number, number][] { + const trimmed = needle.trim(); + if (trimmed.length === 0) return []; + + const pattern = new RegExp( + `(?= 0) spans.push([start, start + match[0].length]); + } + return spans; +} + +/** + * Count how often ANY of `terms` appears, without counting the same words + * twice, and report where the first one starts. + * + * A name and its own alias overlap constantly — "Vandelay Industries" + * contains "Vandelay", and both are terms for the same company. Counting + * each term separately reported that sentence as two mentions. Longer + * terms win, so the most specific spelling claims the text and the shorter + * one only counts where it stands alone. + */ +function countTerms(haystack: string, terms: string[]): [number, number] { + const candidates = terms + .map((term) => ({ term, spans: matchSpans(haystack, term) })) + .sort((a, b) => b.term.trim().length - a.term.trim().length); + + const taken: [number, number][] = []; + for (const { spans } of candidates) { + for (const span of spans) { + const overlaps = taken.some( + ([start, end]) => span[0] < end && start < span[1], + ); + if (!overlaps) taken.push(span); + } + } + + if (taken.length === 0) return [0, -1]; + return [taken.length, Math.min(...taken.map(([start]) => start))]; +} + +/** + * How each tracked brand fared in one answer. + * + * Returns a row for every enabled brand, mentioned or not — see the table + * comment in `lib/db/schema.ts` for why the misses have to be stored. + */ +export function extractMentions( + text: string, + sources: ExtractedSource[], + brandList: Brand[], +): ExtractedMention[] { + const domains = sources.map((source) => source.domain); + + return brandList.map((brand) => { + const [mentionCount, firstIndex] = countTerms(text, [ + brand.name, + ...brand.aliases, + ]); + const firstPosition = firstIndex === -1 ? null : firstIndex; + + // A brand's own domain counts, and so does a subdomain of it: an + // answer citing "blog.acme.io" cited Acme. Suffix-matched on a dot so + // "notacme.io" cannot match "acme.io". + const brandDomains = brand.domains.map((d) => d.toLowerCase()); + const citedSourceCount = domains.filter((domain) => + brandDomains.some( + (owned) => domain === owned || domain.endsWith(`.${owned}`), + ), + ).length; + + return { + brandId: brand.id, + mentioned: mentionCount > 0, + mentionCount, + firstPosition, + cited: citedSourceCount > 0, + citedSourceCount, + }; + }); +} + +/** + * The queries an engine typed, and the follow-ups it offered. + * + * Only ChatGPT, Copilot, Grok and Perplexity report what they searched; + * the rest return nothing here, and an empty result for those is the + * honest answer rather than missing data. + */ +export function extractSearchQueries(response: unknown): ExtractedQuery[] { + const result = unwrap(response); + if (result === null) return []; + + const queries: ExtractedQuery[] = []; + const seen = new Set(); + + const push = (kind: string, value: unknown) => { + const query = typeof value === "string" ? value.trim() : ""; + if (query.length === 0) return; + const key = `${kind} ${query.toLowerCase()}`; + if (seen.has(key)) return; + seen.add(key); + queries.push({ kind, position: queries.length + 1, query }); + }; + + // What the model actually searched. Perplexity names the same thing + // differently; both are the engine's own reformulation of the prompt. + for (const item of asArray(result.searchQueries)) push("issued", item); + for (const item of asArray(result.search_model_queries)) { + push("issued", item); + } + + // Follow-up chips shown to the reader. Navigation furniture, not the + // engine's reasoning, so it never pools with the above. + for (const item of asArray(result.related_queries)) push("suggested", item); + for (const item of asArray(result.relatedSearches)) { + if (typeof item === "object" && item !== null) { + push("suggested", (item as Record).query); + } else { + push("suggested", item); + } + } + + return queries; +} + +/** + * Candidate names the answer mentioned. + * + * Same literal matching as a tracked brand, and deliberately no more: this + * cannot discover a brand nobody listed. It tells you which of the names + * YOU wrote down are turning up, so promoting one into `brands` is a + * decision you make on evidence rather than a guess the code made for you. + * + * Only the hits are returned. A candidate is not tracked, so it has no + * denominator to preserve and a row of zeroes would say nothing. + */ +export function extractCandidates( + text: string, + candidates: BrandCandidate[] = BRAND_CANDIDATES, +): ExtractedCandidate[] { + const found: ExtractedCandidate[] = []; + for (const candidate of candidates) { + // Aliases fold into the canonical name: "Acme" and "Acme, Inc" are one + // company, and reporting them as two rows would split the evidence you + // are meant to weigh. + const [mentionCount] = countTerms(text, [ + candidate.name, + ...candidate.aliases, + ]); + if (mentionCount > 0) found.push({ name: candidate.name, mentionCount }); + } + return found; +} + +/** Everything derived from one stored response, in one pass. */ +export function extractResult( + response: unknown, + brandList: Brand[], +): Extraction { + const text = answerText(response); + const sources = extractSources(response); + return { + text, + sources, + mentions: extractMentions(text, sources, brandList), + queries: extractSearchQueries(response), + candidates: extractCandidates(text), + }; +} diff --git a/lib/http.ts b/lib/http.ts index 6209229..6dcd13b 100644 --- a/lib/http.ts +++ b/lib/http.ts @@ -25,6 +25,23 @@ export function parseOr400(schema: ZodType, data: unknown): T { return parsed.data; } +/** + * Postgres 23505 (unique violation), used to turn a duplicate insert into + * a 409 instead of a 500 whose message is raw constraint text. + * + * Walks the `cause` chain: Drizzle wraps the driver error in a + * DrizzleQueryError, so the `code` is never on the error it actually + * throws. Checking only the top level silently never matches. + */ +export function isUniqueViolation(error: unknown): boolean { + for (let current = error; current != null;) { + if (typeof current !== "object") return false; + if ((current as { code?: string }).code === "23505") return true; + current = (current as { cause?: unknown }).cause; + } + return false; +} + type RouteContext = { params: Promise> }; /** diff --git a/lib/refresh.test.ts b/lib/refresh.test.ts new file mode 100644 index 0000000..f3c0785 --- /dev/null +++ b/lib/refresh.test.ts @@ -0,0 +1,271 @@ +import { eq } from "drizzle-orm"; +import { afterAll, beforeAll, beforeEach, describe, expect, it } from "vitest"; + +import { + applyMigrations, + closeDatabase, + hasDatabase, + resetTables, +} from "../test/db"; +import { getDb } from "./db"; +import { + brands, + extractionState, + prompts, + resultBrandMentions, + results, + resultSources, +} from "./db/schema"; +import { EXTRACTION_STAMP } from "./extract"; +import { markAllForReextraction, refreshDerived } from "./refresh"; + +const describeDb = hasDatabase ? describe : describe.skip; + +/** A ChatGPT-shaped payload: prose plus a source rail. */ +function answer(text: string, urls: string[] = []) { + return { + success: true, + result: { + text, + sources: urls.map((url, index) => ({ + position: index + 1, + url, + label: `S${index}`, + })), + }, + }; +} + +async function seedPrompt(): Promise { + const [row] = await getDb() + .insert(prompts) + .values({ name: "p", prompt: "best crm", engines: ["chatgpt"] }) + .returning({ id: prompts.id }); + return row.id; +} + +async function seedResult( + promptId: string, + response: unknown, + status: "completed" | "failed" | "pending" = "completed", +): Promise { + const [row] = await getDb() + .insert(results) + .values({ + promptId, + engine: "chatgpt", + status, + response, + completedAt: new Date(), + }) + .returning({ id: results.id }); + return row.id; +} + +async function seedBrand( + name: string, + extra: { domains?: string[]; aliases?: string[] } = {}, +): Promise { + const [row] = await getDb() + .insert(brands) + .values({ + name, + domains: extra.domains ?? [], + aliases: extra.aliases ?? [], + }) + .returning({ id: brands.id }); + return row.id; +} + +describeDb("refreshDerived", () => { + beforeAll(applyMigrations); + beforeEach(resetTables); + afterAll(closeDatabase); + + it("writes sources and one mention row per brand", async () => { + const promptId = await seedPrompt(); + const resultId = await seedResult( + promptId, + answer("Acme leads the field.", ["https://acme.io/a", "https://b.com/x"]), + ); + await seedBrand("Acme", { domains: ["acme.io"] }); + await seedBrand("Globex"); + + const summary = await refreshDerived(); + expect(summary).toMatchObject({ extracted: 1, sources: 2, skipped: false }); + + const db = getDb(); + const sources = await db + .select() + .from(resultSources) + .where(eq(resultSources.resultId, resultId)); + expect(sources.map((s) => s.domain).sort()).toEqual(["acme.io", "b.com"]); + + const mentions = await db + .select() + .from(resultBrandMentions) + .where(eq(resultBrandMentions.resultId, resultId)); + expect(mentions).toHaveLength(2); + expect(mentions.find((m) => m.mentioned)?.cited).toBe(true); + }); + + it("does nothing at all when no brand is configured", async () => { + const promptId = await seedPrompt(); + await seedResult(promptId, answer("Acme leads.", ["https://acme.io/a"])); + + expect(await refreshDerived()).toMatchObject({ + extracted: 0, + skipped: true, + }); + + // Crucially the result is NOT marked extracted, so configuring a brand + // later still picks it up. + const [row] = await getDb().select().from(results); + expect(row.extractionRevision).toBeNull(); + }); + + it("leaves pending and failed results alone", async () => { + const promptId = await seedPrompt(); + await seedResult(promptId, null, "pending"); + await seedResult(promptId, answer("Acme"), "failed"); + await seedBrand("Acme"); + + expect(await refreshDerived()).toMatchObject({ extracted: 0 }); + }); + + it("is idempotent: a second pass has nothing to do", async () => { + const promptId = await seedPrompt(); + await seedResult(promptId, answer("Acme", ["https://acme.io/a"])); + await seedBrand("Acme"); + + await refreshDerived(); + expect(await refreshDerived()).toMatchObject({ extracted: 0 }); + + const sources = await getDb().select().from(resultSources); + expect(sources).toHaveLength(1); + }); + + it("stamps the revision it derived a row at", async () => { + const promptId = await seedPrompt(); + await seedResult(promptId, answer("Acme")); + await seedBrand("Acme"); + await refreshDerived(); + + const [row] = await getDb().select().from(results); + expect(row.extractionRevision).toBe(EXTRACTION_STAMP); + }); + + it("rebuilds rather than accumulating when a result is re-extracted", async () => { + const promptId = await seedPrompt(); + await seedResult(promptId, answer("Acme", ["https://acme.io/a"])); + await seedBrand("Acme"); + await refreshDerived(); + + await markAllForReextraction(); + await refreshDerived(); + + expect(await getDb().select().from(resultSources)).toHaveLength(1); + expect(await getDb().select().from(resultBrandMentions)).toHaveLength(1); + }); + + it("reopens the history once when the extraction rules change", async () => { + const promptId = await seedPrompt(); + await seedResult(promptId, answer("Acme")); + await seedBrand("Acme"); + await refreshDerived(); + + // The stamp now matches, so nothing is reopened on later ticks. + expect(await refreshDerived()).toMatchObject({ reopened: 0 }); + + // Simulate a deploy that changed the rules: the stored stamp no longer + // agrees with the code's. + await getDb().update(extractionState).set({ stamp: -1 }); + + const summary = await refreshDerived(); + expect(summary.reopened).toBe(1); + expect(summary.extracted).toBe(1); + + // And it settles: the next tick reopens nothing. + expect(await refreshDerived()).toMatchObject({ reopened: 0, extracted: 0 }); + }); + + it("re-opens history so a newly added brand is scored on old answers", async () => { + const promptId = await seedPrompt(); + await seedResult(promptId, answer("Acme and Globex both rank.")); + await seedBrand("Acme"); + await refreshDerived(); + + await seedBrand("Globex"); + // Adding a brand alone changes nothing — the results are already + // stamped. This is exactly why the brand-writing path has to call it. + expect(await refreshDerived()).toMatchObject({ extracted: 0 }); + + expect(await markAllForReextraction()).toBe(1); + await refreshDerived(); + + const mentions = await getDb().select().from(resultBrandMentions); + expect(mentions).toHaveLength(2); + expect(mentions.every((m) => m.mentioned)).toBe(true); + }); + + it("ignores a disabled brand", async () => { + const promptId = await seedPrompt(); + await seedResult(promptId, answer("Acme")); + const brandId = await seedBrand("Acme"); + await getDb() + .update(brands) + .set({ enabled: false }) + .where(eq(brands.id, brandId)); + + expect(await refreshDerived()).toMatchObject({ skipped: true }); + }); + + it("stops on the time budget and reports there is more to do", async () => { + const promptId = await seedPrompt(); + await seedResult(promptId, answer("Acme")); + await seedResult(promptId, answer("Acme")); + await seedBrand("Acme"); + + // Clock jumps past the budget right after the batch starts: the first + // reading is the start time, every later one is well past it. + let calls = 0; + const summary = await refreshDerived({ + now: () => (calls++ === 0 ? 0 : 60_000), + }); + + // One result still got done: the budget defers work, it never stalls + // the queue completely. + expect(summary).toMatchObject({ extracted: 1, more: true }); + + // And the rest is still queued for the next tick. + expect(await refreshDerived()).toMatchObject({ extracted: 1 }); + expect(await refreshDerived()).toMatchObject({ extracted: 0 }); + }); + + it("survives a result whose response is not the shape we expect", async () => { + const promptId = await seedPrompt(); + await seedResult(promptId, { unexpected: true }); + await seedResult(promptId, answer("Acme")); + await seedBrand("Acme"); + + expect(await refreshDerived()).toMatchObject({ extracted: 2 }); + const mentions = await getDb().select().from(resultBrandMentions); + expect(mentions).toHaveLength(2); + expect(mentions.filter((m) => m.mentioned)).toHaveLength(1); + }); + + it("drops derived rows with the result they came from", async () => { + const promptId = await seedPrompt(); + const resultId = await seedResult( + promptId, + answer("Acme", ["https://a.io/x"]), + ); + await seedBrand("Acme"); + await refreshDerived(); + + await getDb().delete(results).where(eq(results.id, resultId)); + + expect(await getDb().select().from(resultSources)).toHaveLength(0); + expect(await getDb().select().from(resultBrandMentions)).toHaveLength(0); + }); +}); diff --git a/lib/refresh.ts b/lib/refresh.ts new file mode 100644 index 0000000..e1362f2 --- /dev/null +++ b/lib/refresh.ts @@ -0,0 +1,308 @@ +import { and, asc, eq, isNotNull, isNull, sql } from "drizzle-orm"; + +import { getDb } from "./db"; +import { + brands, + extractionState, + resultBrandMentions, + resultCandidateMentions, + results, + resultSearchQueries, + resultSources, + type Brand, + type NewResultBrandMention, + type NewResultCandidateMention, + type NewResultSearchQuery, + type NewResultSource, +} from "./db/schema"; +import { EXTRACTION_STAMP, extractResult } from "./extract"; + +/** + * Derives `result_sources` and `result_brand_mentions` from finished + * results, a batch at a time, from inside the scheduler tick. + * + * This is the "materialised view" for a deployment that has nowhere to run + * one. Neon's free tier has no `pg_cron`, so there is no scheduler in the + * database to refresh anything; the one Vercel Cron job is the only clock + * this app owns. Doing the work here keeps the free tier genuinely free. + */ + +// Bounded so a backlog cannot run the tick past Vercel's function timeout. +// Extraction is CPU and database only — no scrape is ever awaited — so a +// batch is fast; the cap is here for the first tick after an import, when +// the backlog is every result at once. +const BATCH_SIZE = 250; +// Leaves room in a 60s function for the submissions and the sweep that ran +// before this. A batch that runs out of budget simply resumes next tick: +// the queue is a column, not an in-memory cursor. +const TIME_BUDGET_MS = 20_000; + +export interface RefreshSummary { + /** Results whose derived rows were rebuilt this tick. */ + extracted: number; + /** + * Results reopened this tick because the extraction rules changed. + * Non-zero on the first tick after a deploy that changes them, and zero + * on every tick after. + */ + reopened: number; + /** Links written across those results. */ + sources: number; + /** True when the budget or the batch cap stopped us short of the queue. */ + more: boolean; + /** Skipped entirely, with nothing written, when no brand is configured. */ + skipped: boolean; +} + +/** + * The brand fields the extractor actually reads. + * + * `isOwn` is absent on purpose: it labels a brand as yours for the panels, + * and the extractor treats every brand identically, so flipping it changes + * no derived row. Re-deriving the whole history for a cosmetic edit would + * be pure waste. + */ +const EXTRACTION_INPUTS = ["name", "aliases", "domains", "enabled"] as const; + +/** Whether a brand edit changes what the extractor would produce. */ +export function affectsExtraction(changed: Record): boolean { + return EXTRACTION_INPUTS.some((field) => field in changed); +} + +/** + * Queue every completed result for re-extraction. + * + * Call this after any change to the brand list. A brand's mentions are + * computed against the whole history, not just answers that arrive next, + * so adding a competitor has to reopen the past — otherwise its chart + * starts at the day someone remembered to add it, which reads as a brand + * that suddenly appeared. + */ +export async function markAllForReextraction(): Promise { + const db = getDb(); + // Only the rows that still carry a stamp need writing: setting NULL over + // NULL is a no-op that still writes a dead tuple, and this statement runs + // against the whole table. + await db + .update(results) + .set({ extractionRevision: null }) + .where( + and( + eq(results.status, "completed"), + isNotNull(results.extractionRevision), + ), + ); + + // Report the depth of the queue, not the number of rows written. Callers + // ask "how many answers will be re-derived", and a result that was + // already waiting counts toward that just as much as one just reopened. + const [row] = await db + .select({ count: sql`count(*)::int` }) + .from(results) + .where( + and(eq(results.status, "completed"), isNull(results.extractionRevision)), + ); + return row.count; +} + +/** + * Clear the queue flag on every completed result when the extraction rules + * changed since the stored rows were built. + * + * This is what keeps the queue predicate sargable. The alternative — asking + * for rows whose stamp differs from the current one — cannot use an index, + * because no index can hold a constant the query supplies at runtime, and + * Postgres answered it by scanning every completed row on every tick even + * when nothing was waiting. + * + * Returns how many rows were reopened, which is zero on the overwhelming + * majority of ticks. + */ +async function syncExtractionStamp( + db: ReturnType, +): Promise { + const [state] = await db.select().from(extractionState); + if (state?.stamp === EXTRACTION_STAMP) return 0; + + const reopened = await markAllForReextraction(); + await db + .insert(extractionState) + .values({ id: true, stamp: EXTRACTION_STAMP, updatedAt: new Date() }) + .onConflictDoUpdate({ + target: extractionState.id, + set: { stamp: EXTRACTION_STAMP, updatedAt: new Date() }, + }); + return reopened; +} + +/** + * Rebuild the derived rows for one result inside a transaction. + * + * Delete-then-insert rather than upsert: the extractor's output for a + * result is a whole set, and a source that disappears between revisions + * has to disappear from the table too. Upserting would leave it behind + * with no way to tell it from a current row. + */ +async function extractOne( + db: ReturnType, + row: { id: string; response: unknown }, + brandList: Brand[], +): Promise { + const { sources, mentions, queries, candidates } = extractResult( + row.response, + brandList, + ); + + await db.transaction(async (tx) => { + await tx.delete(resultSources).where(eq(resultSources.resultId, row.id)); + await tx + .delete(resultBrandMentions) + .where(eq(resultBrandMentions.resultId, row.id)); + await tx + .delete(resultSearchQueries) + .where(eq(resultSearchQueries.resultId, row.id)); + await tx + .delete(resultCandidateMentions) + .where(eq(resultCandidateMentions.resultId, row.id)); + + if (sources.length > 0) { + const sourceRows: NewResultSource[] = sources.map((source) => ({ + resultId: row.id, + kind: source.kind, + position: source.position, + url: source.url, + domain: source.domain, + label: source.label, + })); + await tx.insert(resultSources).values(sourceRows); + } + + if (mentions.length > 0) { + const mentionRows: NewResultBrandMention[] = mentions.map((mention) => ({ + resultId: row.id, + brandId: mention.brandId, + mentioned: mention.mentioned, + mentionCount: mention.mentionCount, + firstPosition: mention.firstPosition, + cited: mention.cited, + citedSourceCount: mention.citedSourceCount, + })); + await tx.insert(resultBrandMentions).values(mentionRows); + } + + if (queries.length > 0) { + const queryRows: NewResultSearchQuery[] = queries.map((entry) => ({ + resultId: row.id, + kind: entry.kind, + position: entry.position, + query: entry.query, + })); + await tx.insert(resultSearchQueries).values(queryRows); + } + + if (candidates.length > 0) { + const candidateRows: NewResultCandidateMention[] = candidates.map( + (entry) => ({ + resultId: row.id, + name: entry.name, + mentionCount: entry.mentionCount, + }), + ); + await tx.insert(resultCandidateMentions).values(candidateRows); + } + + // Inside the transaction: a crash between writing the rows and marking + // the result would otherwise leave it looking extracted when it is + // half-written, and nothing would ever revisit it. + await tx + .update(results) + .set({ extractionRevision: EXTRACTION_STAMP }) + .where(eq(results.id, row.id)); + }); + + return sources.length; +} + +/** + * One refresh pass: extract up to a batch of results that are finished but + * not yet derived at the current revision. + */ +export async function refreshDerived( + options: { now?: () => number } = {}, +): Promise { + const db = getDb(); + const now = options.now ?? Date.now; + const startedAt = now(); + + const brandList = await db + .select() + .from(brands) + .where(eq(brands.enabled, true)); + + // No brands, no extraction — not even the sources. Deriving links for a + // deployment that has configured nothing would fill a table nobody is + // reading and mark the results done, so adding the first brand later + // would need a full re-extraction to notice them. + if (brandList.length === 0) { + return { + extracted: 0, + reopened: 0, + sources: 0, + more: false, + skipped: true, + }; + } + + const reopened = await syncExtractionStamp(db); + + const pending = await db + .select({ id: results.id, response: results.response }) + .from(results) + .where( + and(eq(results.status, "completed"), isNull(results.extractionRevision)), + ) + .orderBy(asc(results.completedAt)) + .limit(BATCH_SIZE); + + let extracted = 0; + let sources = 0; + for (const row of pending) { + // `extracted > 0` guarantees forward progress. Without it a tick whose + // submissions already ate the budget extracts nothing, and a deployment + // with a slow tick never derives a single row — the queue would look + // busy forever while staying exactly the same size. + if (extracted > 0 && now() - startedAt > TIME_BUDGET_MS) { + return { extracted, reopened, sources, more: true, skipped: false }; + } + sources += await extractOne(db, row, brandList); + extracted += 1; + } + + return { + extracted, + reopened, + sources, + more: pending.length === BATCH_SIZE, + skipped: false, + }; +} + +/** + * Rows that no longer have a parent result, for a deployment that pruned + * old results by hand. The foreign keys cascade, so this is only ever a + * safety net; it exists because a silently growing derived table is the + * classic way a free-tier database fills up. + */ +export async function countDerived(): Promise<{ + sources: number; + mentions: number; +}> { + const db = getDb(); + const [sourceCount] = await db + .select({ count: sql`count(*)::int` }) + .from(resultSources); + const [mentionCount] = await db + .select({ count: sql`count(*)::int` }) + .from(resultBrandMentions); + return { sources: sourceCount.count, mentions: mentionCount.count }; +} diff --git a/lib/runner.ts b/lib/runner.ts index 0798d24..ad423d7 100644 --- a/lib/runner.ts +++ b/lib/runner.ts @@ -3,6 +3,7 @@ import { and, asc, eq, isNotNull, isNull, lt, lte, or } from "drizzle-orm"; import { fetchTask, submitTask } from "./cloro"; import { getDb } from "./db"; import { prompts, results, type NewResult, type Prompt } from "./db/schema"; +import { refreshDerived, type RefreshSummary } from "./refresh"; import { webhookCallbackUrl } from "./webhooks"; import type { Engine } from "./engines"; @@ -76,6 +77,7 @@ export async function submitPromptOnce( export interface TickSummary { submitted: { promptId: string; name: string; results: SubmittedResult[] }[]; sweep: { checked: number; updated: number; timedOut: number }; + refresh: RefreshSummary; } /** @@ -83,6 +85,12 @@ export interface TickSummary { * pending rows by polling. Due-logic: a prompt with runsPerDay=N is due when * its last run is at least 24h/N (minus slack) old, so any external tick * frequency works — granularity is simply capped by how often ticks arrive. + * + * The derived-table refresh runs last, on purpose: it is the only step that + * can be cut short without losing anything. Submissions are time-sensitive + * and the sweep closes out rows the webhook missed, so if the function runs + * out of budget it must run out here, where the leftover work is still + * queued in a column and the next tick resumes it. */ export async function runTick(): Promise { const db = getDb(); @@ -120,7 +128,8 @@ export async function runTick(): Promise { }); } - return { submitted, sweep: await sweepPending() }; + const sweep = await sweepPending(); + return { submitted, sweep, refresh: await refreshDerived() }; } /** diff --git a/lib/validation.ts b/lib/validation.ts index 9068663..0cdb93a 100644 --- a/lib/validation.ts +++ b/lib/validation.ts @@ -32,6 +32,54 @@ export const updatePromptSchema = z message: "Provide at least one field to update", }); +/** + * A domain, not a URL: "acme.io", never "https://acme.io/path". + * + * Accepting a URL here and quietly parsing it out would be friendlier for + * about a week, until someone entered "acme.io/blog" and wondered why the + * citations never matched. The extractor compares against a hostname, so + * that is what this asks for. + */ +const domainSchema = z + .string() + .min(1) + .max(253) + .transform((value) => + value + .trim() + .toLowerCase() + .replace(/^www\./, ""), + ) + .refine((value) => /^[a-z0-9.-]+\.[a-z]{2,}$/.test(value), { + message: "Must be a bare domain such as acme.io, not a URL", + }); + +// Aliases are matched literally against answer text. A blank one would +// match nothing (the extractor skips it), so reject it here rather than +// storing a value that silently does nothing. +const aliasSchema = z.string().min(1).max(200); + +export const createBrandSchema = z.object({ + name: z.string().min(1).max(200), + aliases: z.array(aliasSchema).max(50).default([]), + domains: z.array(domainSchema).max(50).default([]), + isOwn: z.boolean().default(false), + enabled: z.boolean().default(true), +}); + +export const updateBrandSchema = z + .object({ + name: z.string().min(1).max(200), + aliases: z.array(aliasSchema).max(50), + domains: z.array(domainSchema).max(50), + isOwn: z.boolean(), + enabled: z.boolean(), + }) + .partial() + .refine((value) => Object.keys(value).length > 0, { + message: "Provide at least one field to update", + }); + export const resultsQuerySchema = z.object({ promptId: z.uuid().optional(), engine: z.enum(ENGINES).optional(), @@ -46,4 +94,6 @@ export const idSchema = z.uuid(); export type CreatePromptInput = z.infer; export type UpdatePromptInput = z.infer; +export type CreateBrandInput = z.infer; +export type UpdateBrandInput = z.infer; export type ResultsQuery = z.infer; diff --git a/scripts/seed.mjs b/scripts/seed.mjs new file mode 100644 index 0000000..693c1e6 --- /dev/null +++ b/scripts/seed.mjs @@ -0,0 +1,344 @@ +/** + * Fills a local database with synthetic answers, so the Grafana panels can + * be built and looked at without waiting a month for real data. + * + * Development only. It writes prompts, brands and completed results — the + * raw material — and derives nothing. Run the scheduler tick afterwards + * (`curl -H "Authorization: Bearer $CRON_SECRET" localhost:3000/api/cron`) + * and the app itself fills the derived tables, through exactly the code + * that runs in production. A seed script that wrote those tables directly + * would prove only that the seed script works. + * + * The prompts are seeded DISABLED on purpose: an enabled prompt is due the + * moment it exists, so the tick would try to submit it to the real cloro + * API before getting to the refresh. + * + * node scripts/seed.mjs # add to whatever is there + * node scripts/seed.mjs --reset # delete existing rows first + */ +import { Pool } from "pg"; + +const DAYS = 30; +const RESET = process.argv.includes("--reset"); + +const connectionString = process.env.DATABASE_URL; +if (!connectionString) { + console.error("DATABASE_URL is not set"); + process.exit(1); +} + +/** + * Deterministic PRNG, so two runs produce the same database and a panel + * that looks wrong stays wrong while you fix it. + */ +let seed = 1337; +function random() { + seed = (seed * 1103515245 + 12345) % 2147483648; + return seed / 2147483648; +} + +function pick(items) { + return items[Math.floor(random() * items.length)]; +} + +const BRANDS = [ + { name: "Acme", aliases: ["Acme Corp"], domains: ["acme.io"], isOwn: true }, + { name: "Globex", aliases: [], domains: ["globex.com"], isOwn: false }, + { name: "Initech", aliases: ["Init Tech"], domains: ["initech.io"] }, + // Never named in prose, but its pages get cited. This is the case that a + // single "mentioned" boolean would hide, and the reason `cited` is stored + // separately — check a panel can still see it. + { name: "Umbrella", aliases: [], domains: ["umbrella.co"] }, +]; + +const PROMPTS = [ + "best CRM for startups", + "top project management tools 2026", + "how to choose a help desk platform", + "best analytics tools for small teams", + "cheapest CRM with an API", + "which CRM integrates with Slack", +]; + +const ENGINES = ["chatgpt", "perplexity", "gemini", "aimode", "google"]; + +// Third-party pages the engines cite. Two are strong (cited constantly), +// the rest thin out — a realistic long tail for "pages to get listed on". +const PUBLISHER_PAGES = [ + ["g2.com", "/categories/crm"], + ["g2.com", "/categories/help-desk"], + ["capterra.com", "/crm-software"], + ["capterra.com", "/project-management"], + ["reddit.com", "/r/sales/comments/best-crm"], + ["reddit.com", "/r/startups/comments/crm-recommendations"], + ["reddit.com", "/r/smallbusiness/comments/help-desk-picks"], + ["techcrunch.com", "/2026/01/crm-roundup"], + ["youtube.com", "/watch?v=crm-review"], + ["youtube.com", "/watch?v=crm-vs-crm-2026"], + ["youtube.com", "/watch?v=help-desk-teardown"], + ["forbes.com", "/advisor/business/software/best-crm"], + ["nytimes.com", "/wirecutter/reviews/best-crm"], + ["stackoverflow.com", "/questions/crm-api-integration"], + ["producthunt.com", "/topics/crm"], +]; + +/** + * How often each brand is named, by engine. Deliberately uneven: a chart + * where every line sits on top of every other proves nothing about whether + * the panel separates them. + */ +const MENTION_RATE = { + chatgpt: { Acme: 0.55, Globex: 0.8, Initech: 0.2, Umbrella: 0 }, + perplexity: { Acme: 0.45, Globex: 0.7, Initech: 0.35, Umbrella: 0 }, + gemini: { Acme: 0.3, Globex: 0.75, Initech: 0.15, Umbrella: 0 }, + aimode: { Acme: 0.4, Globex: 0.6, Initech: 0.25, Umbrella: 0 }, + // Google writes prose only when an AI Overview was served. + google: { Acme: 0.25, Globex: 0.4, Initech: 0.1, Umbrella: 0 }, +}; + +/** + * Vendors the answers name but that nobody tracks. + * + * Drawn from the list shipped in `lib/brand-candidates.json`, so the + * "Named but not tracked" panel has something to find straight after a + * seed with no extra setup. None of these is in the brands table, which is + * the whole point of that panel. + */ +const UNTRACKED_VENDORS = ["Initrode", "Hooli", "Vandelay", "Cyberdyne"]; + +const BRAND_SITE = { + Acme: ["acme.io", "/product"], + Globex: ["globex.com", "/platform"], + Initech: ["initech.io", "/pricing"], + Umbrella: ["umbrella.co", "/solutions"], +}; + +/** Prose that names the given brands, in a plausible listicle voice. */ +function answerProse(query, named) { + if (named.length === 0) { + return `There are many options for "${query}". The right choice depends on your team size, budget and the tools you already use.`; + } + const sentences = named.map( + (brand, index) => + `${index + 1}. ${brand} — a strong option for teams that care about ${pick(["price", "integrations", "onboarding", "reporting", "support"])}.`, + ); + // Some answers also name a vendor nobody is tracking, which is the whole + // point of the candidates panel. + const aside = + random() < 0.35 + ? ` ${pick(UNTRACKED_VENDORS)} comes up often too, though it is less established.` + : ""; + return `Here are the leading tools for "${query}":\n\n${sentences.join("\n")}\n\nEach has a free tier worth trying before you commit.${aside}`; +} + +/** + * Links an answer cites: some of the named brands' own sites, plus + * publishers. + * + * A named brand is linked only about three times in four. Real answers + * recommend tools without linking them, and if naming always implied + * citing then the two columns would agree on every row — which would hide + * a panel that read the same column twice. + */ +function sourceList(named, alsoCited, promptIndex = 0) { + const links = [...named.filter(() => random() < 0.75), ...alsoCited].map( + (brand) => { + const [host, path] = BRAND_SITE[brand]; + return { + url: `https://${host}${path}`, + label: `${brand} — official site`, + }; + }, + ); + + // One answer in twelve retrieves nothing. Real engines do this, and a + // data-quality panel that never sees it cannot be trusted to show it. + if (random() < 0.08) return []; + + // A window over the publisher list, offset per prompt. Neighbouring + // prompts overlap heavily, distant ones barely — which is what the + // redundancy panel is meant to show. One shared pool made every pair + // read 100%, and a panel that always says the same thing tests nothing. + const window = 8; + const offset = (promptIndex * 3) % PUBLISHER_PAGES.length; + const pool = Array.from( + { length: window }, + (_, i) => PUBLISHER_PAGES[(offset + i) % PUBLISHER_PAGES.length], + ); + + const publisherCount = 2 + Math.floor(random() * 4); + for (let i = 0; i < publisherCount; i += 1) { + const [host, path] = pick(pool); + links.push({ + url: `https://${host}${path}`, + label: `${host} — ${path.split("/").pop().replace(/[-?=]/g, " ").trim()}`, + }); + } + + return links.map((link, index) => ({ position: index + 1, ...link })); +} + +/** The queries an engine says it ran, and the follow-ups it offers. */ +function queryFanOut(query) { + const issued = [ + query, + `${query} ${pick(["2026", "reviews", "pricing", "vs"])}`, + ]; + const suggested = [`best ${pick(["free", "cheap", "enterprise"])} ${query}`]; + return { issued, suggested }; +} + +/** The payload shape cloro returns, per engine. */ +function buildResponse(engine, query, named, alsoCited, promptIndex) { + const sources = sourceList(named, alsoCited, promptIndex); + + if (engine !== "google") { + const fanOut = queryFanOut(query); + return { + success: true, + result: { + text: answerProse(query, named), + sources, + citationPills: sources.slice(0, 3).map((source, index) => ({ + position: index + 1, + url: source.url, + label: source.label, + domain: new URL(source.url).hostname, + })), + // Only some engines report what they searched. Perplexity names the + // field differently, and Gemini reports nothing at all — so a panel + // that pooled them would quietly under-count. + ...(engine === "chatgpt" || engine === "copilot" + ? { searchQueries: fanOut.issued } + : {}), + ...(engine === "perplexity" + ? { + search_model_queries: fanOut.issued, + related_queries: fanOut.suggested, + } + : {}), + }, + }; + } + + // Google is a page of links, not an answer. An AI Overview appears on + // roughly half of them; the rest have no prose at all, and a brand + // genuinely cannot be "named in the answer" there. + const hasOverview = random() < 0.5; + return { + success: true, + result: { + organicResults: sources.map((source, index) => ({ + position: index + 1, + title: source.label, + link: source.url, + snippet: "A comparison of the leading tools in this category.", + })), + aioverview: hasOverview + ? { text: answerProse(query, named), sources: sources.slice(0, 4) } + : null, + peopleAlsoAsk: [ + { + question: `What is the best ${query}?`, + link: "https://g2.com/categories/crm", + title: "Category overview", + }, + ], + relatedSearches: [ + { query: `${query} pricing`, link: "https://google.com/search" }, + ], + }, + }; +} + +async function main() { + const pool = new Pool({ connectionString, max: 3 }); + + if (RESET) { + await pool.query("TRUNCATE results, prompts, brands CASCADE"); + console.log("Cleared prompts, results and brands."); + } + + const brandIds = {}; + for (const brand of BRANDS) { + const { rows } = await pool.query( + `INSERT INTO brands (name, aliases, domains, is_own) + VALUES ($1, $2, $3, $4) + ON CONFLICT (lower(name)) DO UPDATE SET aliases = EXCLUDED.aliases + RETURNING id`, + [brand.name, brand.aliases, brand.domains, brand.isOwn ?? false], + ); + brandIds[brand.name] = rows[0].id; + } + console.log(`Tracking ${BRANDS.length} brands.`); + + let results = 0; + let failures = 0; + + for (const [promptIndex, prompt] of PROMPTS.entries()) { + const { rows } = await pool.query( + `INSERT INTO prompts (name, prompt, engines, runs_per_day, enabled) + VALUES ($1, $2, $3, 1, false) + RETURNING id`, + [prompt, `What are the ${prompt}?`, ENGINES], + ); + const promptId = rows[0].id; + + for (let day = DAYS; day > 0; day -= 1) { + const completedAt = new Date(Date.now() - day * 24 * 60 * 60 * 1000); + + for (const engine of ENGINES) { + // A few runs fail, as they do in production. They must stay out of + // every visibility denominator: a failed scrape is not an answer + // that declined to name you. + if (random() < 0.04) { + await pool.query( + `INSERT INTO results (prompt_id, engine, task_id, status, error, created_at, completed_at) + VALUES ($1, $2, $3, 'failed', 'Task failed', $4, $4)`, + [promptId, engine, `seed_${++results}`, completedAt], + ); + failures += 1; + continue; + } + + const rates = MENTION_RATE[engine]; + const named = Object.keys(rates).filter( + (brand) => random() < rates[brand], + ); + // Umbrella is cited without ever being named, and the others get + // linked sometimes without appearing in the prose. + const alsoCited = ["Umbrella"].filter(() => random() < 0.3); + + await pool.query( + `INSERT INTO results (prompt_id, engine, task_id, status, response, credits_charged, created_at, completed_at) + VALUES ($1, $2, $3, 'completed', $4, $5, $6, $6)`, + [ + promptId, + engine, + `seed_${++results}`, + JSON.stringify( + buildResponse(engine, prompt, named, alsoCited, promptIndex), + ), + 1 + Math.floor(random() * 4), + completedAt, + ], + ); + } + } + } + + console.log( + `Seeded ${PROMPTS.length} prompts and ${results} results over ${DAYS} days (${failures} failed).`, + ); + console.log( + "\nNow run the scheduler tick to derive the tables the panels read:\n" + + ' curl -H "Authorization: Bearer $CRON_SECRET" localhost:3000/api/cron', + ); + + await pool.end(); +} + +main().catch((error) => { + console.error(error); + process.exit(1); +}); diff --git a/test/api.test.ts b/test/api.test.ts index c63945e..00551c5 100644 --- a/test/api.test.ts +++ b/test/api.test.ts @@ -14,6 +14,11 @@ import { vi, } from "vitest"; +import { + DELETE as brandDelete, + PATCH as brandPatch, +} from "../app/api/brands/[id]/route"; +import { GET as brandsGet, POST as brandsPost } from "../app/api/brands/route"; import { GET as cronGet } from "../app/api/cron/route"; import { DELETE, GET as promptGet, PATCH } from "../app/api/prompts/[id]/route"; import { POST as runPost } from "../app/api/prompts/[id]/run/route"; @@ -36,6 +41,23 @@ const authed = (path: string, init: RequestInit = {}) => const params = (id: string) => ({ params: Promise.resolve({ id }) }); +const webhookToken = () => + createHmac("sha256", SECRET) + .update("geo-tracker:webhook") + .digest("hex") + .slice(0, 32); + +/** A completed-task delivery carrying a ChatGPT-shaped answer. */ +const completedWebhook = (taskId: string, text: string) => + new Request(`${BASE}/api/webhook?token=${webhookToken()}`, { + method: "POST", + body: JSON.stringify({ + task: { id: taskId, status: "COMPLETED" }, + credits: { creditsCharged: 1 }, + response: { success: true, result: { text, sources: [] } }, + }), + }); + async function createPrompt(body: Record = {}) { const res = await promptsPost( authed("/api/prompts", { @@ -52,6 +74,23 @@ async function createPrompt(body: Record = {}) { return { res, body: (await res.json()) as { prompt: { id: string } } }; } +async function createBrand(body: Record = {}) { + const res = await brandsPost( + authed("/api/brands", { + method: "POST", + body: JSON.stringify({ name: "Acme", ...body }), + }), + params(""), + ); + return { + res, + body: (await res.json()) as { + brand: { id: string }; + queuedForExtraction: number; + }, + }; +} + describe.skipIf(!hasDatabase)("API routes (need a database)", () => { beforeAll(async () => { process.env.CRON_SECRET = SECRET; @@ -219,11 +258,7 @@ describe.skipIf(!hasDatabase)("API routes (need a database)", () => { }); describe("webhook", () => { - const token = () => - createHmac("sha256", SECRET) - .update("geo-tracker:webhook") - .digest("hex") - .slice(0, 32); + const token = webhookToken; async function pendingTaskId() { const { body } = await createPrompt(); @@ -313,6 +348,109 @@ describe.skipIf(!hasDatabase)("API routes (need a database)", () => { updated: 0, timedOut: 0, }); + // Nothing is configured, so the refresh declines to derive anything. + expect(summary.refresh).toMatchObject({ extracted: 0, skipped: true }); + }); + }); + + describe("brands", () => { + it("creates a brand and returns 201", async () => { + const { res, body } = await createBrand({ domains: ["acme.io"] }); + expect(res.status).toBe(201); + expect(body.brand).toMatchObject({ name: "Acme", isOwn: false }); + }); + + it("normalises a domain to a bare lowercase host", async () => { + const { body } = await createBrand({ domains: ["WWW.Acme.IO"] }); + expect(body.brand).toMatchObject({ domains: ["acme.io"] }); + }); + + it("rejects a URL where a domain belongs", async () => { + const { res } = await createBrand({ domains: ["https://acme.io/blog"] }); + expect(res.status).toBe(400); + }); + + it("treats a name that differs only in case as the same brand", async () => { + await createBrand({ name: "Acme" }); + const { res } = await createBrand({ name: "acme" }); + expect(res.status).toBe(409); + }); + + it("lists brands by name", async () => { + await createBrand({ name: "Globex" }); + await createBrand({ name: "Acme" }); + const res = await brandsGet(authed("/api/brands"), params("")); + const body = (await res.json()) as { brands: { name: string }[] }; + expect(body.brands.map((b) => b.name)).toEqual(["Acme", "Globex"]); + }); + + it("queues the stored history when a brand is added", async () => { + await createPrompt(); + await cronGet(authed("/api/cron"), params("")); + await webhookPost(completedWebhook("task_1", "Acme wins."), params("")); + + const { body } = await createBrand(); + expect(body.queuedForExtraction).toBe(1); + + // The next tick derives it, so the brand's chart starts full. + const res = await cronGet(authed("/api/cron"), params("")); + const summary = await res.json(); + expect(summary.refresh).toMatchObject({ extracted: 1, skipped: false }); + }); + + it("re-queues history when an edit changes what is matched", async () => { + await createPrompt(); + await cronGet(authed("/api/cron"), params("")); + await webhookPost(completedWebhook("task_1", "Acme wins."), params("")); + const { body } = await createBrand(); + await cronGet(authed("/api/cron"), params("")); + + const res = await brandPatch( + authed(`/api/brands/${body.brand.id}`, { + method: "PATCH", + body: JSON.stringify({ aliases: ["Acme Corp"] }), + }), + params(body.brand.id), + ); + expect(res.status).toBe(200); + await expect(res.json()).resolves.toMatchObject({ + queuedForExtraction: 1, + }); + }); + + it("does not re-queue history for a label-only edit", async () => { + await createPrompt(); + await cronGet(authed("/api/cron"), params("")); + await webhookPost(completedWebhook("task_1", "Acme wins."), params("")); + const { body } = await createBrand(); + + const res = await brandPatch( + authed(`/api/brands/${body.brand.id}`, { + method: "PATCH", + body: JSON.stringify({ isOwn: true }), + }), + params(body.brand.id), + ); + await expect(res.json()).resolves.toMatchObject({ + brand: { isOwn: true }, + queuedForExtraction: 0, + }); + }); + + it("404s on an unknown brand and 400s on a malformed id", async () => { + const missing = await brandDelete( + authed("/api/brands/00000000-0000-4000-8000-000000000000", { + method: "DELETE", + }), + params("00000000-0000-4000-8000-000000000000"), + ); + expect(missing.status).toBe(404); + + const malformed = await brandDelete( + authed("/api/brands/nope", { method: "DELETE" }), + params("nope"), + ); + expect(malformed.status).toBe(400); }); }); }); diff --git a/test/db.ts b/test/db.ts index d86f9d7..d857f7c 100644 --- a/test/db.ts +++ b/test/db.ts @@ -19,8 +19,11 @@ export async function applyMigrations(): Promise { } export async function resetTables(): Promise { + // `brands` is listed even though CASCADE would reach the derived tables + // anyway: truncating results alone leaves brand rows behind, and a test + // that adds a brand would then see the previous test's brands too. await getDb().execute( - sql`TRUNCATE results, prompts RESTART IDENTITY CASCADE`, + sql`TRUNCATE results, prompts, brands, extraction_state RESTART IDENTITY CASCADE`, ); }