Skip to main content

No claim without evidence: an AI research engine that can't invent facts

Scope

This page describes the engine's design for engineers. Market Intelligence has no public API: /api/market-intel is the embedded app's own endpoint and is authenticated with the Shopify admin session, not an API key. For what merchants see, read the Market Intelligence overview.

The problem: fluent is not the same as true​

Ask a model for "the size of the UK pet-food market" and you get a number, a growth rate and a trend. Some of it came from a real publication. Some of it is a blend of several publications from different years. Some of it is invented. The prose reads the same in all three cases, and the reader can't tell which is which.

We didn't have to imagine this failure. Auditing our own codebase before building the engine, we found that SerpCompetitor, the table seven features read to answer "who ranks in the top 10 for this keyword", had exactly one writer. That writer asked a model to recall the results "grounded in your training knowledge, not a live crawl". Competitor discovery, share of voice, a Decision Intelligence input and an image-analysis job all read those rows as if they were measured. The image job went further and fetched URLs the model had made up.

So the engine is built around one invariant: nothing reaches a merchant without the evidence it rests on, and code, not the model, decides whether the evidence is good enough.

Architecture​

The code lives in app/lib/market-intel/, in four layers:

LayerModulesRule
Admissionadmission.server.jsThe only code that creates a run. No provider calls
Orchestrationorchestrator.server.js, workers/research.worker.jsRuns stages in order, checkpoints each, meters every call
Evidenceprovider.server.js, evidence.js, ledger.server.js, signals/Turns pages, crawls and store data into sources and claims
Judgementverify.js, report-guard.js, kinds/swot-rules.js, kinds/xray-rules.jsPure functions. Decide what's admissible. No I/O, no model

The judgement layer being pure is the most important structural choice in the engine. The rules that decide what a merchant sees are reproducible, and they're unit-tested with fixtures rather than with a live model.

Three products run on it. Each is a "kind" that declares which stages it needs:

KindStagesWeb research
industry (market report)all sevenyes
xray (Competitor X-Ray)all seven, with three fixed questionsyes
swotcollect → verify → synthesize → publishno

Admission: one door, checked in order​

A request is admitted in well under a second and returns 202 with a run ID. Checks run in a fixed order, cheapest and most operator-facing first:

  1. Kill switch and readiness → 503. "Research is not configured" is an operator problem, so it's reported as one, not as a failed run.
  2. Plan feature and allowed depth → 403.
  3. Kind-specific input validation → 400 / 404. Inputs are normalised once (control characters stripped, length-bounded, enumerations snapped), because they feed prompts, subject keys and idempotency keys.
  4. Idempotency. The key is a SHA-256 of shop, kind, subject, depth, the key-sorted inputs and the UTC day. A duplicate returns the existing run.
  5. In-flight cap (3), monthly run quota, monthly USD cap with the estimate reserved → 429.
  6. Insert and enqueue.

Steps 4 to 6 run inside one transaction holding a per-shop advisory lock:

await prisma.$transaction(async (tx) => {
await tx.$executeRaw`SELECT pg_advisory_xact_lock(hashtext(${`market-intel:${shop}`}))`;
// idempotency lookup → month usage → quota + budget checks → create
});

Two details matter:

  • The lock is transaction-scoped (pg_advisory_xact_lock), so it's safe behind PgBouncer in transaction pooling mode, where a session-level lock could be released onto another client's connection.
  • The database is the final arbiter. A unique index on (shop, idempotencyKey) catches the race the lock can't (two requests reading before either locks). A P2002 is caught and resolved by returning the winner.

Quota and spend are derived from the ResearchRun table itself, not from a counter that can drift. A run that failed before spending anything doesn't consume quota. A failed or cancelled run releases its idempotency key, so the same request can be retried the same day.

Errors carry a stable code alongside a readable message, so the UI never string-matches:

CodeStatus
market_intel_disabled, research_not_configured, queue_unavailable503
plan_feature_required, plan_depth_not_allowed403
invalid_input400
not_found404
monthly_quota_exhausted, monthly_budget_exhausted, too_many_runs_in_flight429
not_cancellable409

Stages, checkpoints and resumption​

A deep run makes dozens of searches and model calls over several minutes. If a retry repeated the whole run, a transient 529 near the end would double the bill. So every stage writes its checkpoint (stageState) before the next starts, and the gather stage checkpoints after every question:

for (const stage of stages) {
if (ctx.state.completed.includes(stage)) continue; // resume past finished stages
await assertNotCancelled(ctx);
await STAGE_HANDLERS[stage](ctx);
ctx.state.completed.push(stage);
await checkpoint(ctx, stage);
}

Errors are classified by what a retry would buy:

DispositionExamplesBehaviour
retrythrottling, overload, timeoutsRe-throw. BullMQ backs off (3 attempts, exponential) and the orchestrator resumes from the checkpoint
terminalprovider not configured, model doesn't support the toolsUnrecoverableError. Retrying buys the same failure
softone question refused or malformedRecorded against that question. The rest of the plan runs

The run row is the source of truth, not the queue:

  • Each run gets its own job ID (intel_run_<runId>). Re-adding a BullMQ job whose ID is still retained is a silent no-op, so IDs are never reused.
  • A heartbeat is written on every metered call and checkpoint. A 15-minute maintenance job fails runs with no heartbeat for 30 minutes, and queued runs that never started within 6 hours, each with a readable reason.
  • Worker locks are long (15 minutes, renewed every 5), because one research question can legitimately run for minutes.
  • Publish re-checks cancellation inside its transaction, so a cancel that lands during synthesis can't be overwritten by a publish.

The provider port: uncited prose is not evidence​

researchQuestion() takes one question and returns one normalised shape, whatever the provider:

{ sources, findings, usage, costUsd, toolErrors, searchQueries, stopReason, model, provider }

A finding is a short passage bound to the URLs that support it. The port drops any passage without a citation. This is the first gate: text the provider wrote without citing anything is the model talking, not evidence.

The primary adapter uses Anthropic's server-side web_search_20260209 and web_fetch_20260209 tools, with citations enabled, through a governed callAIResearch() in ai-client.server.js. It sits beside callAI and uses the same rate-limit and budget gates, so the no-direct-SDK lint rule needed no exemption. The adapter handles the protocol's details:

  • pause_turn is resumed by re-sending the conversation, up to a per-depth continuation bound.
  • Server-tool failures arrive as result blocks in an HTTP 200 response, not as exceptions. They're collected into toolErrors rather than lost.
  • Each response is priced at the model that served it, and web searches are metered separately ($10 per 1,000).

A Perplexity Sonar adapter can serve the gather stage instead. The plan, extract and synthesize stages always use schema-constrained output, so they stay on Anthropic.

Sources are de-duplicated by canonical URL (lower-cased host, no www., no fragment, tracking parameters removed, sorted query, no trailing slash). The richest record wins: a fetched document beats a search hit, which beats a bare citation URL.

Verification is deterministic​

A model extracts candidate claims (statement, kind, metric key, value, unit, period) and must reference the findings each claim comes from. A claim citing a finding that isn't in its batch is dropped. Then verify.js, with no model and no I/O, decides how far each claim can be believed:

if (validRefs.length === 0) support = UNSUPPORTED;
else if (claim.provenance === MEASURED) support = SINGLE_SOURCE;
else if (independentPublishers.size >= 2) support = CORROBORATED;
else support = SINGLE_SOURCE;
  • Independence is by publisher, not URL. publisherKey() reduces news.example.co.uk and shop.example.co.uk to one publisher, handling common two-level suffixes. It deliberately errs towards under-claiming corroboration.
  • Contested: figures sharing a metric key, unit and year, more than 25% apart, are all marked contested and shown side by side. The engine never picks a winner.
  • Stale: the period described, or the newest supporting publication, is more than 24 months old.
  • Value not in quote: a figure whose number doesn't appear in any cited passage (checked against several textual forms, such as 1.2 billion, 1.2bn and 1,200,000,000) is kept but capped at low confidence.

Confidence is derived, and a model's self-reported confidence can only lower it (simplified):

export function confidenceFor(claim, avgReliability) {
const ceiling = claim.confidence ?? HIGH; // the model's own view: a ceiling, never a floor
let level = /* from support, provenance and source reliability */;
if (claim.stale || claim.valueUnverified || claim.provenance === ESTIMATED) level = downgrade(level);
return min(level, ceiling);
}

Source reliability is a stated heuristic by publisher type: official and statistical sources high, established press next, syndicated market-size vendors and press-release wires low, user-generated content lowest. It ranks source types. It doesn't grade individual articles.

The citation guard runs after the model​

The synthesis model writes paragraphs and names the claim references each one rests on. report-guard.js doesn't trust that naming:

const refs = paragraph.claimRefs.filter((r) => {
const c = claimsByRef.get(r);
return c && c.support !== SUPPORT.UNSUPPORTED;
});
if (!text || refs.length === 0) { removed++; continue; } // deleted, not softened

A paragraph whose references are all unknown or unsupported is removed. A paragraph resting only on contested claims is kept and flagged. The count of removed paragraphs is stored with the report's coverage figures, so a report that lost half its draft says so.

This matters more than any prompt instruction. Prompts ask the model to cite. The guard makes an uncited paragraph impossible to publish whether the model complied or not.

SWOT: admissibility rules in code​

The SWOT's model call proposes items and TOWS strategies. swot-rules.js decides which survive:

  • Every item cites a supported claim.
  • A strength or weakness must cite one of the store's own measured signals, and its comparator must be one that signal carries: history (the store's trailing window), competitor (matched competitors) or exposure (money or stock at risk). A figure with nothing to compare it to is rejected.
  • An opportunity or threat must cite external evidence (a market-report claim) or a competitor-comparator signal.
  • Confidence never exceeds the weakest cited claim.
  • A TOWS strategy survives only if it pairs the right quadrants from items that themselves survived.

Each item gets an evidenceHash: an FNV-1a hash of its quadrant plus the sorted fingerprints of its claims, with numeric values bucketed to two significant digits. When a SWOT is rebuilt, an item with the same quadrant and hash inherits the merchant's earlier decision (accepted, dismissed with reason, actioned). A dismissal therefore persists until the evidence moves materially, not on every rounding wobble.

Store signals follow the same discipline. A collector returns a signal only when it has enough data to mean something. Otherwise it returns a note ("too few reviews in 90 days to trend") that the UI shows. Absence is reported, never rendered as a zero. ProductIntelligence counters are all-time totals, so no signal derives a rate from them.

Research never acts. A SWOT item can become a proposal that is evaluated by the platform action policy as a model-originated, material-tier change. Low confidence stays a recommendation. Anything else becomes an approval request on the shared approval spine, where the decider must not be the requester. The spine compares RBAC user IDs, so the requester is passed as a user ID, never an email.

Cost: reserve, meter, stop, settle​

  • The estimate is deliberately high-side. It's reserved against the shop's monthly cap before any call, and an under-estimate would let the cap cut a run short.
  • Metering uses atomic increments (costUsd: { increment: cost }), never read-modify-write. With concurrent calls, x += await f() loses updates.
  • The hard stop is a stage boundary. Gather checks the cap between questions and publishes what it has, marked partial with the reason.
  • Month-to-date "committed" spend is settled cost plus the unspent remainder of in-flight estimates. A burst of admissions can't collectively exceed the cap.

Every call also goes through the shop's general AI budget. A BUDGET_EXCEEDED during gather ends research and publishes partial, rather than failing the run.

Provenance as data, with a fail-safe default​

The SERP fix is small and worth copying. SerpCompetitor gained a provenance column:

ALTER TABLE "SerpCompetitor"
ADD COLUMN "provenance" TEXT NOT NULL DEFAULT 'ai_estimate';
  • The default is the backfill. Every existing row really was an AI estimate.
  • The default fails safe. A future writer that forgets to stamp provenance makes a real measurement look like an estimate (under-trusted), never an estimate look real.

The writer now fetches the live Google top 10 from DataForSEO first and stamps dataforseo. The model's job changes from inventing the results to interpreting the measured ones (content gaps, title patterns, angle). A measured row never carries a model-invented position, domain or score. The estimate remains only as a labelled fallback when live data is unavailable. Live SERPs are cached for 24 hours per keyword, market and device, and a billing failure trips a 15-minute breaker so an unfunded account doesn't log hundreds of identical failures an hour.

One module states the rule every reader follows:

// seo-engine/serp-provenance.js
// A reader that computes a NUMBER from top10 filters to MEASURED rows.
// A reader that DISPLAYS an estimate labels it as one.
export const MEASURED_SERP_WHERE = Object.freeze({ provenance: "dataforseo" });

Competitor discovery, share-of-voice snapshots, the SERP rollup and the Decision Intelligence signal fusion now filter with it. Signal fusion also stopped substituting 50 for a missing difficulty score: each component contributes only when it has a value, weights are renormalised, and with no components the score is null, not a middling number. The image analysis fetches competitor pages through a helper that re-validates every redirect hop against private address ranges.

Web content is untrusted input​

  • The research model has no tools that write. Its only tools are search and fetch.
  • Third-party text is wrapped and labelled as data in every prompt that carries it. The structural defences (no write tools, schema-constrained output, code-side citation enforcement) don't depend on the model obeying that label.
  • Search engines and link shorteners are blocked domains for the web tools. They launder provenance: the page that made the claim is one hop away. Operators can add more.
  • Every URL the engine fetches itself (the X-Ray's homepage read, competitor image pages) goes through fetchPublicPage. It runs SSRF validation on the first URL and each redirect, follows redirects manually to a bound, caps the body, and returns a reason code instead of throwing.
  • Stored errors are redacted (API keys, Shopify tokens, password= style pairs) and length-bounded.
  • Source excerpts are purged after 180 days. The URL and hash are kept, so a citation still resolves.

One gotcha worth knowing: Claude 5 thinks by default​

Claude Opus 5 and Sonnet 5 run adaptive thinking when a request doesn't specify thinking. Earlier models didn't. Any code that reads content[0].text gets a thinking block first and silently returns an empty string. A small max_tokens sized for the answer can also be spent entirely on thinking.

Building this engine surfaced it across our plain completion paths. The fix lives in one pure module, ai/anthropic-thinking.js. Plain completions send thinking: {"type": "disabled"} on models that would otherwise think, and response text is read from every text block, never content[0]. Research calls deliberately keep thinking on and budget max_tokens for it.

Testing​

Because the judgement layer is pure, most of the engine is tested without a database or a provider:

SuiteCovers
market-intel-evidenceURLs and publishers, deterministic verification, the citation guard, input normalisation and the cost estimate
market-intel-rulesSWOT admissibility, X-Ray dimensions, exports, provider findings
market-intel-admissionAdmission checks, usage and idempotency keys, cancellation
market-intel-orchestratorA full market-report run including resume-from-checkpoint, the cost cap and cancellation; a SWOT run; stage normalisers
market-intel-xray, market-intel-schedules, market-intel-actions-signalsAn X-Ray run; schedule timing and back-off; action proposals, decisions and store signal collectors
ai-research-client, anthropic-thinkingServer tools and pause_turn resumption, the Perplexity adapter; text extraction on models that think by default
serp-provenance-readers, serp-analysis-measuredDiscovery, share of voice and Decision Intelligence read measured rows only; the live-first writer and its labelled fallback
public-page-fetch, dataforseo-balance-monitorPer-hop SSRF validation; balance classification and alerting

Provider calls are stubbed with fixture responses. Nothing calls a live provider in CI.

Design choices we rejected​

RejectedWhy
One large prompt producing a finished reportCan't be verified, diffed or audited claim by claim
Running research in the requestGateway timeouts, and a retry repeats the whole spend
Using the merchant's chosen copywriting model for researchCitation quality would depend on a copy preference. Research is a platform choice, like the SERP provider
Scraping Google directlyTerms of service and blocking. Measured SERPs are bought from a provider
A new "approved" AI file outside the governed clientPast audits found exemptions are where cost-tracking drift hides
A cross-shop cache of research evidenceA market query reveals strategy. Deferred until a data-classification review

See also​