No claim without evidence: an AI research engine that can't invent facts
This page describes the engine's design for engineers. Market Intelligence has
no public API: /api/market-intel is the embedded app's own endpoint and is
authenticated with the Shopify admin session, not an API key. For what
merchants see, read the Market Intelligence overview.
The problem: fluent is not the same as true
Ask a model for "the size of the UK pet-food market" and you get a number, a growth rate and a trend. Some of it came from a real publication. Some of it is a blend of several publications from different years. Some of it is invented. The prose reads the same in all three cases, and the reader can't tell which is which.
We didn't have to imagine this failure. Auditing our own codebase before
building the engine, we found that SerpCompetitor, the table seven features
read to answer "who ranks in the top 10 for this keyword", had exactly one
writer. That writer asked a model to recall the results "grounded in your
training knowledge, not a live crawl". Competitor discovery, share of voice, a
Decision Intelligence input and an image-analysis job all read those rows as if
they were measured. The image job went further and fetched URLs the model had
made up.
So the engine is built around one invariant: nothing reaches a merchant without the evidence it rests on, and code, not the model, decides whether the evidence is good enough.
Architecture
The code lives in app/lib/market-intel/, in four layers:
| Layer | Modules | Rule |
|---|---|---|
| Admission | admission.server.js | The only code that creates a run. No provider calls |
| Orchestration | orchestrator.server.js, workers/research.worker.js | Runs stages in order, checkpoints each, meters every call |
| Evidence | provider.server.js, evidence.js, ledger.server.js, signals/ | Turns pages, crawls and store data into sources and claims |
| Judgement | verify.js, report-guard.js, kinds/swot-rules.js, kinds/xray-rules.js | Pure functions. Decide what's admissible. No I/O, no model |
The judgement layer being pure is the most important structural choice in the engine. The rules that decide what a merchant sees are reproducible, and they're unit-tested with fixtures rather than with a live model.
Three products run on it. Each is a "kind" that declares which stages it needs:
| Kind | Stages | Web research |
|---|---|---|
industry (market report) | all seven | yes |
xray (Competitor X-Ray) | all seven, with three fixed questions | yes |
swot | collect → verify → synthesize → publish | no |
Admission: one door, checked in order
A request is admitted in well under a second and returns 202 with a run ID.
Checks run in a fixed order, cheapest and most operator-facing first:
- Kill switch and readiness →
503. "Research is not configured" is an operator problem, so it's reported as one, not as a failed run. - Plan feature and allowed depth →
403. - Kind-specific input validation →
400/404. Inputs are normalised once (control characters stripped, length-bounded, enumerations snapped), because they feed prompts, subject keys and idempotency keys. - Idempotency. The key is a SHA-256 of shop, kind, subject, depth, the key-sorted inputs and the UTC day. A duplicate returns the existing run.
- In-flight cap (3), monthly run quota, monthly USD cap with the estimate
reserved →
429. - Insert and enqueue.
Steps 4 to 6 run inside one transaction holding a per-shop advisory lock:
await prisma.$transaction(async (tx) => {
await tx.$executeRaw`SELECT pg_advisory_xact_lock(hashtext(${`market-intel:${shop}`}))`;
// idempotency lookup → month usage → quota + budget checks → create
});
Two details matter:
- The lock is transaction-scoped (
pg_advisory_xact_lock), so it's safe behind PgBouncer in transaction pooling mode, where a session-level lock could be released onto another client's connection. - The database is the final arbiter. A unique index on
(shop, idempotencyKey)catches the race the lock can't (two requests reading before either locks). AP2002is caught and resolved by returning the winner.
Quota and spend are derived from the ResearchRun table itself, not from a
counter that can drift. A run that failed before spending anything doesn't
consume quota. A failed or cancelled run releases its idempotency key, so the
same request can be retried the same day.
Errors carry a stable code alongside a readable message, so the UI never
string-matches:
| Code | Status |
|---|---|
market_intel_disabled, research_not_configured, queue_unavailable | 503 |
plan_feature_required, plan_depth_not_allowed | 403 |
invalid_input | 400 |
not_found | 404 |
monthly_quota_exhausted, monthly_budget_exhausted, too_many_runs_in_flight | 429 |
not_cancellable | 409 |
Stages, checkpoints and resumption
A deep run makes dozens of searches and model calls over several minutes. If a
retry repeated the whole run, a transient 529 near the end would double the
bill. So every stage writes its checkpoint (stageState) before the next
starts, and the gather stage checkpoints after every question:
for (const stage of stages) {
if (ctx.state.completed.includes(stage)) continue; // resume past finished stages
await assertNotCancelled(ctx);
await STAGE_HANDLERS[stage](ctx);
ctx.state.completed.push(stage);
await checkpoint(ctx, stage);
}
Errors are classified by what a retry would buy:
| Disposition | Examples | Behaviour |
|---|---|---|
| retry | throttling, overload, timeouts | Re-throw. BullMQ backs off (3 attempts, exponential) and the orchestrator resumes from the checkpoint |
| terminal | provider not configured, model doesn't support the tools | UnrecoverableError. Retrying buys the same failure |
| soft | one question refused or malformed | Recorded against that question. The rest of the plan runs |
The run row is the source of truth, not the queue:
- Each run gets its own job ID (
intel_run_<runId>). Re-adding a BullMQ job whose ID is still retained is a silent no-op, so IDs are never reused. - A heartbeat is written on every metered call and checkpoint. A 15-minute maintenance job fails runs with no heartbeat for 30 minutes, and queued runs that never started within 6 hours, each with a readable reason.
- Worker locks are long (15 minutes, renewed every 5), because one research question can legitimately run for minutes.
- Publish re-checks cancellation inside its transaction, so a cancel that lands during synthesis can't be overwritten by a publish.
The provider port: uncited prose is not evidence
researchQuestion() takes one question and returns one normalised shape,
whatever the provider:
{ sources, findings, usage, costUsd, toolErrors, searchQueries, stopReason, model, provider }
A finding is a short passage bound to the URLs that support it. The port drops any passage without a citation. This is the first gate: text the provider wrote without citing anything is the model talking, not evidence.
The primary adapter uses Anthropic's server-side web_search_20260209 and
web_fetch_20260209 tools, with citations enabled, through a governed
callAIResearch() in ai-client.server.js. It sits beside callAI and uses
the same rate-limit and budget gates, so the no-direct-SDK lint rule needed no
exemption. The adapter handles the protocol's details:
pause_turnis resumed by re-sending the conversation, up to a per-depth continuation bound.- Server-tool failures arrive as result blocks in an HTTP 200 response, not
as exceptions. They're collected into
toolErrorsrather than lost. - Each response is priced at the model that served it, and web searches are metered separately ($10 per 1,000).
A Perplexity Sonar adapter can serve the gather stage instead. The plan, extract and synthesize stages always use schema-constrained output, so they stay on Anthropic.
Sources are de-duplicated by canonical URL (lower-cased host, no www., no
fragment, tracking parameters removed, sorted query, no trailing slash). The
richest record wins: a fetched document beats a search hit, which beats a bare
citation URL.
Verification is deterministic
A model extracts candidate claims (statement, kind, metric key, value, unit,
period) and must reference the findings each claim comes from. A claim citing a
finding that isn't in its batch is dropped. Then verify.js, with no model and
no I/O, decides how far each claim can be believed:
if (validRefs.length === 0) support = UNSUPPORTED;
else if (claim.provenance === MEASURED) support = SINGLE_SOURCE;
else if (independentPublishers.size >= 2) support = CORROBORATED;
else support = SINGLE_SOURCE;
- Independence is by publisher, not URL.
publisherKey()reducesnews.example.co.ukandshop.example.co.ukto one publisher, handling common two-level suffixes. It deliberately errs towards under-claiming corroboration. - Contested: figures sharing a metric key, unit and year, more than 25% apart, are all marked contested and shown side by side. The engine never picks a winner.
- Stale: the period described, or the newest supporting publication, is more than 24 months old.
- Value not in quote: a figure whose number doesn't appear in any cited
passage (checked against several textual forms, such as
1.2 billion,1.2bnand1,200,000,000) is kept but capped at low confidence.
Confidence is derived, and a model's self-reported confidence can only lower it (simplified):
export function confidenceFor(claim, avgReliability) {
const ceiling = claim.confidence ?? HIGH; // the model's own view: a ceiling, never a floor
let level = /* from support, provenance and source reliability */;
if (claim.stale || claim.valueUnverified || claim.provenance === ESTIMATED) level = downgrade(level);
return min(level, ceiling);
}
Source reliability is a stated heuristic by publisher type: official and statistical sources high, established press next, syndicated market-size vendors and press-release wires low, user-generated content lowest. It ranks source types. It doesn't grade individual articles.
The citation guard runs after the model
The synthesis model writes paragraphs and names the claim references each one
rests on. report-guard.js doesn't trust that naming:
const refs = paragraph.claimRefs.filter((r) => {
const c = claimsByRef.get(r);
return c && c.support !== SUPPORT.UNSUPPORTED;
});
if (!text || refs.length === 0) { removed++; continue; } // deleted, not softened
A paragraph whose references are all unknown or unsupported is removed. A paragraph resting only on contested claims is kept and flagged. The count of removed paragraphs is stored with the report's coverage figures, so a report that lost half its draft says so.
This matters more than any prompt instruction. Prompts ask the model to cite. The guard makes an uncited paragraph impossible to publish whether the model complied or not.
SWOT: admissibility rules in code
The SWOT's model call proposes items and TOWS strategies. swot-rules.js
decides which survive:
- Every item cites a supported claim.
- A strength or weakness must cite one of the store's own measured signals,
and its comparator must be one that signal carries:
history(the store's trailing window),competitor(matched competitors) orexposure(money or stock at risk). A figure with nothing to compare it to is rejected. - An opportunity or threat must cite external evidence (a market-report claim) or a competitor-comparator signal.
- Confidence never exceeds the weakest cited claim.
- A TOWS strategy survives only if it pairs the right quadrants from items that themselves survived.
Each item gets an evidenceHash: an FNV-1a hash of its quadrant plus the
sorted fingerprints of its claims, with numeric values bucketed to two
significant digits. When a SWOT is rebuilt, an item with the same quadrant and
hash inherits the merchant's earlier decision (accepted, dismissed with reason,
actioned). A dismissal therefore persists until the evidence moves materially,
not on every rounding wobble.
Store signals follow the same discipline. A collector returns a signal only
when it has enough data to mean something. Otherwise it returns a note ("too
few reviews in 90 days to trend") that the UI shows. Absence is reported, never
rendered as a zero. ProductIntelligence counters are all-time totals, so no
signal derives a rate from them.
Research never acts. A SWOT item can become a proposal that is evaluated by the platform action policy as a model-originated, material-tier change. Low confidence stays a recommendation. Anything else becomes an approval request on the shared approval spine, where the decider must not be the requester. The spine compares RBAC user IDs, so the requester is passed as a user ID, never an email.
Cost: reserve, meter, stop, settle
- The estimate is deliberately high-side. It's reserved against the shop's monthly cap before any call, and an under-estimate would let the cap cut a run short.
- Metering uses atomic increments (
costUsd: { increment: cost }), never read-modify-write. With concurrent calls,x += await f()loses updates. - The hard stop is a stage boundary. Gather checks the cap between questions and publishes what it has, marked partial with the reason.
- Month-to-date "committed" spend is settled cost plus the unspent remainder of in-flight estimates. A burst of admissions can't collectively exceed the cap.
Every call also goes through the shop's general AI budget. A BUDGET_EXCEEDED
during gather ends research and publishes partial, rather than failing the run.
Provenance as data, with a fail-safe default
The SERP fix is small and worth copying. SerpCompetitor gained a
provenance column:
ALTER TABLE "SerpCompetitor"
ADD COLUMN "provenance" TEXT NOT NULL DEFAULT 'ai_estimate';
- The default is the backfill. Every existing row really was an AI estimate.
- The default fails safe. A future writer that forgets to stamp provenance makes a real measurement look like an estimate (under-trusted), never an estimate look real.
The writer now fetches the live Google top 10 from DataForSEO first and stamps
dataforseo. The model's job changes from inventing the results to
interpreting the measured ones (content gaps, title patterns, angle). A
measured row never carries a model-invented position, domain or score. The
estimate remains only as a labelled fallback when live data is unavailable.
Live SERPs are cached for 24 hours per keyword, market and device, and a
billing failure trips a 15-minute breaker so an unfunded account doesn't log
hundreds of identical failures an hour.
One module states the rule every reader follows:
// seo-engine/serp-provenance.js
// A reader that computes a NUMBER from top10 filters to MEASURED rows.
// A reader that DISPLAYS an estimate labels it as one.
export const MEASURED_SERP_WHERE = Object.freeze({ provenance: "dataforseo" });
Competitor discovery, share-of-voice snapshots, the SERP rollup and the
Decision Intelligence signal fusion now filter with it. Signal fusion also
stopped substituting 50 for a missing difficulty score: each component
contributes only when it has a value, weights are renormalised, and with no
components the score is null, not a middling number. The image analysis
fetches competitor pages through a helper that re-validates every redirect hop
against private address ranges.
Web content is untrusted input
- The research model has no tools that write. Its only tools are search and fetch.
- Third-party text is wrapped and labelled as data in every prompt that carries it. The structural defences (no write tools, schema-constrained output, code-side citation enforcement) don't depend on the model obeying that label.
- Search engines and link shorteners are blocked domains for the web tools. They launder provenance: the page that made the claim is one hop away. Operators can add more.
- Every URL the engine fetches itself (the X-Ray's homepage read, competitor
image pages) goes through
fetchPublicPage. It runs SSRF validation on the first URL and each redirect, follows redirects manually to a bound, caps the body, and returns a reason code instead of throwing. - Stored errors are redacted (API keys, Shopify tokens,
password=style pairs) and length-bounded. - Source excerpts are purged after 180 days. The URL and hash are kept, so a citation still resolves.
One gotcha worth knowing: Claude 5 thinks by default
Claude Opus 5 and Sonnet 5 run adaptive thinking when a request doesn't specify
thinking. Earlier models didn't. Any code that reads content[0].text gets a
thinking block first and silently returns an empty string. A small
max_tokens sized for the answer can also be spent entirely on thinking.
Building this engine surfaced it across our plain completion paths. The fix
lives in one pure module, ai/anthropic-thinking.js. Plain completions send
thinking: {"type": "disabled"} on models that would otherwise think, and
response text is read from every text block, never content[0]. Research
calls deliberately keep thinking on and budget max_tokens for it.
Testing
Because the judgement layer is pure, most of the engine is tested without a database or a provider:
| Suite | Covers |
|---|---|
market-intel-evidence | URLs and publishers, deterministic verification, the citation guard, input normalisation and the cost estimate |
market-intel-rules | SWOT admissibility, X-Ray dimensions, exports, provider findings |
market-intel-admission | Admission checks, usage and idempotency keys, cancellation |
market-intel-orchestrator | A full market-report run including resume-from-checkpoint, the cost cap and cancellation; a SWOT run; stage normalisers |
market-intel-xray, market-intel-schedules, market-intel-actions-signals | An X-Ray run; schedule timing and back-off; action proposals, decisions and store signal collectors |
ai-research-client, anthropic-thinking | Server tools and pause_turn resumption, the Perplexity adapter; text extraction on models that think by default |
serp-provenance-readers, serp-analysis-measured | Discovery, share of voice and Decision Intelligence read measured rows only; the live-first writer and its labelled fallback |
public-page-fetch, dataforseo-balance-monitor | Per-hop SSRF validation; balance classification and alerting |
Provider calls are stubbed with fixture responses. Nothing calls a live provider in CI.
Design choices we rejected
| Rejected | Why |
|---|---|
| One large prompt producing a finished report | Can't be verified, diffed or audited claim by claim |
| Running research in the request | Gateway timeouts, and a retry repeats the whole spend |
| Using the merchant's chosen copywriting model for research | Citation quality would depend on a copy preference. Research is a platform choice, like the SERP provider |
| Scraping Google directly | Terms of service and blocking. Measured SERPs are bought from a provider |
| A new "approved" AI file outside the governed client | Past audits found exemptions are where cost-tracking drift hides |
| A cross-shop cache of research evidence | A market query reveals strategy. Deferred until a data-classification review |
See also
- 43 → 0: Governing AI in production: the governed AI gateway research calls go through
- Rate limits: platform rate limiting
- Data export & GDPR: how shop data is handled