Methodology 1.2 · 2026-07-20

Evidence first. Uncertainty stays visible.

A high WebsiteIQ score means the current rubric found more of the evidence it tests. It does not mean a search engine, customer, court, regulator, or security assessor agrees.

Four claim types

Type Meaning Example
Observed Directly present in a fetched public response. A Content-Security-Policy header was present.
Derived Calculated from disclosed observations. Average HTML bytes across permitted pages.
Heuristic A versioned interpretation with known limits. E-E-A-T evidence coverage or locality-volume review.
Unknown The crawl does not provide enough evidence. Whether a change will improve rankings or revenue.

Scoring

Scores are deterministic for the same fetched inputs and rubric version, subject to network timing and target-site changes. Weighted modules summarize tested evidence; informational modules such as query intent and locality review do not silently change the core score.

Version 1.2 rejects HTML or homepage fallbacks at /llms.txt, reduces the score weight of that ecosystem proposal, and recognizes an API catalog only when the RFC 9727/RFC 9264 Linkset is valid. RFC 9116 security.txt evidence is validated and reported without inventing unregistered well-known names.

The RFC discovery panel also validates whether a homepage response advertises its API catalog through an RFC 8288 Link header using RFC 9727’s registered api-catalog relation. Draft and provider-specific discovery mechanisms are not scored as RFC conformance.

The report exposes module scores, weights, findings, page evidence, crawl time, and limitations. Rubric changes require a methodology-version change.

E-E-A-T and YMYL

Google describes experience, expertise, authoritativeness, and trust as concepts used in evaluating helpful content, but does not publish a public numeric E-E-A-T score for a page. WebsiteIQ therefore maps observable signals—authorship, organization identity, contact and policy pages, credentials, citations, and structured data—into its own documented heuristic.

YMYL classification is conservative and topical. It indicates that mistakes could matter more and that evidence should be reviewed more carefully; it is not a legal classification or statement of search treatment.

Locality architecture review

The module samples a bounded sitemap tree and groups geographic route structures such as /texas/cities/, /locations/, full state-name paths, and postal-code shapes. Arbitrary slugs are not declared to be cities without geographic path evidence.

When more than half of a sufficiently large sitemap sample appears locality-oriented, WebsiteIQ distinguishes a homepage that establishes corporate or multi-location scope from one that does not. The result is client-facing refactor guidance, not a spam verdict or ranking prediction.

The module prioritizes addressLocality evidence over broad areaServed data, then considers visible City, ST patterns. It reports locality links, distinct service areas, repetition, and density. Threshold crossings produce “human review,” never a spam verdict. Statewide directories and legitimate multi-location businesses are expected false-positive cases.

HTTP errors and problem details

WebsiteIQ preserves RFC 9110 status semantics. Browser requests may receive a themed recovery page while API consumers receive RFC 9457 application/problem+json; the body status and actual HTTP response agree. MCP GET remains 405 with Allow: POST because this deployment does not expose a standalone SSE stream.

See the HTTP response library.

Crawl contract

WebsiteIQBot fetches robots.txt before content, applies exact product-token groups before the wildcard group, merges matching groups, uses longest-match semantics, lets Allow win an equal-specificity tie, and honors matching crawl-delay as a conservative extension. Network, authentication, rate-limit, and server failures on robots stop the crawl fail-closed.

See the complete crawler policy and the Robots Exclusion Protocol (RFC 9309).

Primary references

Changelog

1.2 — 2026-07-20: single-page beta sample, mandatory requester email and audit-processing consent, per-audit DNS validation cache, strict llms.txt validation, RFC well-known discovery, sitemap scope estimates, and expanded agent repair briefs. Prior draft pricing was withdrawn pending the production x402 cost structure.

1.1 — 2026-07-20: named robots-enforcing crawler, SSRF/redirect/402 controls, unlisted retained reports, channel-aware beta limits, locality review, and explicit E-E-A-T/YMYL caveats.