Lodestar Stamp Chicago HVAC edition Methodology New edition monthly

How the record is made.

Every published receipt has to be re-derivable from the method below and the raw artifacts we publish with it. Change the method and the version changes. Old receipts keep their version.

pilot · Chicago HVAC · September 2026 edition Pilot edition. Dated receipts on covered markets. Monthly editions. Read the method

1.0-draft + 1.0.1 probe protocol
Methodology version
58
Businesses this edition
29of 58
Licensed on the register this edition
1.7%
Probe flip rate · publish gate under 5%

The short version

We read the public register for a named local business and verify the facts that change a booking against those sources, then date them: is it real, at that address, licensed, and whatever else this market’s instruments cover. Each field is a dated envelope — state, source, verified_on, expires_on. That receipt is what an assistant reads before it books, calls, or pays. Dated public receipt. We never approve.

The Chicago HVAC deep-record also ran a query battery (ChatGPT, Gemini and Perplexity) and a dual-pass site probe. Those are inputs to this edition, not a score and not the product. Reliability is a published number, not a promise. This edition’s probe flip rate is 1.7% against a hard publish gate of 5%. An edition that fails the gate is held and re-probed, never shipped.

The table further down titled “Internal measurement rubric” is how we check our own instruments in the lab. It is not a public rating, not a rank, and not what an assistant reads. The published product is the dated receipt.

Watching does not change the method. A subscriber gets the baseline briefing after verified payment and exact business matching, the monthly readout, the change alert for published changes, and the weekly digest for the selected business. The public receipt and seal remain free. See /watch-sample.

Edition notes

These caveats travel with this edition. They are here rather than on the front page, but nothing is hidden: the same text ships in /index.json and /llms.txt.

Pilot. Dated receipts are public at lodestarstamp.com. Chicago restaurants is the flagship record; Chicago HVAC carries a deeper visibility pilot. On the record. We do not approve the booking.

Everything on this page in machine-readable form: /index.json and /llms.txt, plus the Trust API.


Lodestar Stamp — Methodology v1.0 (draft for comment)

Status: Draft v1.0 · August 10, 2026 · This document is the public /methodology page on lodestarstamp.com. Every published Lodestar Stamp must be re-derivable from what is written here. When the method changes, the version number changes, and old stamps keep their version.

Weights 2026-09-03 — addendum v1.1: llms.txt drops from 5 to 2 readiness points and structured data rises from 5 to 8, effective from the next edition (draft-v0.8). Records already published keep 1.0 / 1.0.1 and their numbers. The reasons are in the addendum below.

Errata 2026-08-27 — naming only: the product noun is now Lodestar Stamp, and the per-business page is the record; no method, threshold or pillar changed, and methodology_version is unchanged at 1.0-draft + 1.0.1 probe protocol.


1. What we measure

A Lodestar Stamp answers one question: what did we verify about this named entity on this date? The full methodology is built from three pillars:

Visibility (40 points). When AI assistants are asked the questions real customers ask, does this business get recommended — and how often, and how prominently?

Readiness (35 points). Can automated agents actually read and interact with the business's digital surface? Crawl permissions, machine-readable summaries, structured data, and (increasingly) transactable agent endpoints.

Ground truth (25 points). Is what an agent would learn accurate? License status verified against official registries, review-pattern integrity, and consistency of the business's core facts across the web.

Draft editions publish only the pillars fielded in that edition. In draft-v0.7-multimodel-pilot, Visibility (40) and Readiness (25 fielded points from the 35-point pillar) are shown as the fielded receipt this edition; Ground Truth is pending. Deeper ground truth (pricing honesty, booking follow-through) arrives later and will be version-stamped when it does.

2. The query battery

Each vertical/metro pair gets a battery of ~100 queries across five intent categories, drawn from how real customers actually ask:

  1. Discovery — "Who's the best HVAC company in Evanston?"
  2. Problem-driven — "My furnace is blowing cold air — who should I call near Oak Park?"
  3. Comparison & reputation — "Is [company] reputable? Who's better, X or Y?"
  4. Transactional — "Book me a furnace tune-up near 60614." (This category is weighted highest: it's where money changes hands.)
  5. Constraint-based — "Emergency 24/7 furnace repair with financing in Naperville."

Queries are parameterized by neighborhood/suburb and rotated. ~70% of the battery is published; ~30% is held out and rotated each cycle so records cannot be gamed by optimizing against a fixed list.

Models tested: ChatGPT, Claude, Gemini, and Perplexity — the consumer products, because that is what customers actually use. Each query runs 3 times per model per cycle (AI answers vary; we measure recommendation frequency, not a single lucky answer). We capture: every business mentioned, its position, qualifying language ("highly rated," "mixed reviews"), and the sources the model cited.

3. Internal measurement rubric v1.0

This table is how we check our own instruments in the lab. It is not a public rating, not a rank, and not a score we issue on a business. The published product is the dated receipt. We publish dated facts; we never rate or rank.

PillarComponentPoints
Visibility (40)Recommendation share across the battery, weighted by intent category (transactional highest)30
Prominence & sentiment of mentions10
Readiness (35)Agent/transactable surface (MCP endpoint, machine-readable booking)15
AI crawl policy (explicit robots.txt treatment of AI crawlers)5
llms.txt / machine-readable site summary5
Structured data (schema.org business markup)5
Sitemap & basic discoverability5
Ground truth (25)License verified against official state/municipal registry10
Review-pattern integrity (velocity anomalies, distribution shape)10
NAP consistency (name/address/phone identical across major surfaces)5

Public shop pages show the dated receipt and the fielded component totals for the current edition. While ground truth is unfielded, no summary mark beyond the measured total is published; a business should never wonder which measured signals produced the receipt it sees.

4. Reproducibility standard

For every published cycle we publish: the query list (public portion), capture dates, model names and versions, run counts, and per-business raw mention counts. Anyone with the same tools can re-run our public battery and get materially the same results. Claims we cannot make reproducible, we do not publish.

The publish gate: a receipt that flips on re-probe is not a receipt. Every signal is probed twice with agreement required before it publishes, and no market edition ships while its measured flip rate is 5% or higher. The flip rate itself is published with every cycle — our reliability is a number you can check, not a promise you have to take.

5. Cadence & versioning

Full query battery: monthly per market. Readiness probes: weekly. Ground-truth checks: monthly. Records carry their date and methodology version forever. The methodology itself is versioned semantically: component weight changes bump the minor version; pillar changes bump the major version; every change is logged publicly.

6. Anti-gaming

Held-out rotating queries (30%), outcome-weighted scoring (being genuinely recommended matters more than checkbox compliance), anomaly detection on sudden score jumps, and the founding rule that makes gaming pointless at the source: payment never touches the record. No paid placement, no lead-gen, no exceptions.

7. Fairness & corrections

Any business can claim its listing free. Disputes with evidence get investigated within one cycle. A business that fixes a deficiency gets re-tested free in the next cycle — improvement should show up fast. Corrections and their reasons are logged publicly.

8. Known limitations (v1.0)

Consumer AI answers vary by session, location, and account history — our 3-run sampling approximates, not exhausts, this variance. Readiness signals are proxies for agent usability, not guarantees. Ground truth v1 covers license status, review-pattern heuristics, and NAP consistency only — it does not yet verify pricing or service quality. We publish these limitations because a receipt that hides its own uncertainty doesn't deserve the name.


Lodestar Stamp · lodestarstamp.com · Methodology v1.0 draft · Comments to watch@lodestarstamp.com


Methodology Addendum v1.0.1 — Probe Reliability Protocol

Aug 10, 2026 · Adopted after an independent external stress test re-ran our day-one probes and observed ~28% signal instability. A receipt that flips on re-probe is not a receipt. This addendum defines the stabilization protocol; it merges into methodology v1.1.

Why probes flip (observed causes)

Soft-404s (sites returning a 200 homepage for any path, making llms.txt look "present"); WordPress plugins serving llms.txt intermittently; sitemaps that exist at standard paths but aren't declared in robots.txt; WAFs and timeouts blocking probes; redirect chains; classifier edge cases on non-standard directives. Some instability is genuinely site-side variance — which is itself a finding — but we only publish what is stable.

The stabilization protocol (all published signals)

  1. Two-probe agreement. Every signal is probed at least twice, hours apart. Published value = the agreeing value. Disagreement → a third probe; still unstable → published as "variable" (its own state, scored as not-present but labeled honestly).
  2. Soft-404 detection. A file "exists" only if its content is plausibly that file type: llms.txt must be text/markdown-like and must materially differ from the homepage; a response matching the site's known-404 or homepage fingerprint = absent.
  3. Sitemap multi-path. Checked in robots.txt declaration AND at standard paths (/sitemap.xml, /sitemap_index.xml). Existing-but-undeclared = present, sub-noted.
  4. MCP verification, not MCP advertisement. An agent endpoint counts only if it returns valid, schema-conformant JSON at a standard path (or a working declared endpoint). Text that merely mentions MCP tools = "advertised — unverified," never "offers."
  5. Unreachable ≠ zero. Fetch failures publish as "unreachable," excluded from the fielded denominators, retried next cycle.
  6. Retry discipline. Transient failures get one retry after a delay before classification.
  7. Published classifier. Exact classification rules and thresholds for every signal state publish with methodology v1.1 — a business must be able to derive its own classification.

The publish gate (non-negotiable)

No market edition publishes while its measured signal flip rate is 5% or higher across consecutive same-week probe passes. This is a gate, not a target: an edition that fails it is held, re-probed, and fixed — never shipped. The flip rate itself is measured every cycle and published alongside the records — our reliability is a number, not a promise.

A receipt that flips on re-probe is not a receipt.

Also adopted from the stress test


Methodology Addendum v1.1 — Readiness weights: llms.txt down-weighted

Sep 3, 2026 · Adopted for the next edition (draft-v0.8 onward). Records published under 1.0 / 1.0.1 keep their version and their numbers; nothing already on the record is restated. Component weight changes bump the minor version, as section 5 of v1.0 says they must.

The change

Readiness componentv1.0v1.1
Agent/transactable surface (MCP endpoint, machine-readable booking)1515
AI crawl policy (explicit robots.txt treatment of AI crawlers)55
llms.txt / machine-readable site summary52
Structured data (schema.org business markup)58 (partial markup: 3)
Sitemap & basic discoverability55

The pillar still totals 35, and the fielded partial (crawl + llms.txt + sitemap + structured data + MCP) still totals 25. Three points move from llms.txt to structured data.

Why

  1. Nothing we fielded cited one. Across every capture in the current battery, 41 of 82 answers carried cited sources; none of those sources was an llms.txt. We score what assistants demonstrably use, and on our own evidence they are not using this file.
  2. It was the signal that flipped. Addendum 1.0.1 exists because a re-probe found ~28% instability, and llms.txt was the leading cause: soft-404 homepages served under the path, plugins serving the file intermittently. A signal that needed a soft-404 fingerprint check to be trusted is not a five-point signal.
  3. It is self-declared and unverifiable. A business writes its own llms.txt and nothing external checks it. Structured data is validated by any schema.org parser; an MCP endpoint must answer with conformant JSON. Weight should follow what can be verified.
  4. It is the cheapest signal to game. Dropping one text file at a path is a smaller effort than any other readiness signal. Under "outcome-weighted scoring" (section 6), a checkbox that costs nothing should be worth close to nothing.

Structured data receives the three points because it is the machine-readable surface on the same page an assistant already reads, it is verifiable by anyone with a parser, and 32 of the 58 businesses on the current record already publish it, so the move rewards existing good practice rather than inventing a new hoop.

What does not change

llms.txt stays on the probe set and on every record, at 2 points: it is still a real signal of intent, and a business that publishes one is still told so. The probe protocol, the two-pass agreement rule, the soft-404 detection and the 5% flip gate are unchanged. The Lodestar Stamp's own /llms.txt stays published for agents that do read it; this addendum is about what we score, not what we serve.

Where it lands

probe-fleet/reprobe_v101.py carries both weight tables and defaults to 1.0.1 so the pinned draft-v0.7 numbers reproduce byte for byte; the next probe run passes --methodology 1.1. merge_scores.py stamps the new methodology_version when the next edition merges. Every record carries the version it was scored under.

Payment never touches the record. A receipt that flips on re-probe is not a receipt. The record can’t be bought, only earned.

Found something wrong? Send evidence to watch@lodestarstamp.com. Corrections are re-checked in the next edition and the method is versioned in public.

Where to next

See the Chicago HVAC record

The receipts this method produced, for every business in the market.

Edition history

Every frozen edition, append-only.

The charter

Why payment never touches the record.