Sample report — a real check of our own website (43sunsets.com), generated 16 August 2026 with the same process you would receive. Only internal file paths were removed; findings are unedited.  ·  Get one for your site  ·  43 Sunsets Services

AI Visibility Audit — 43sunsets.com

2026-08-16 19:56Z · 3 of 3 known URLs deep-inspected · robots.txt respected · recommended next check: 2026-09-15
Crawler registry updated 2026-08-06 (bundled snapshot); token presence checked against official pages this run (this is a presence check, not semantic verification).

100/100
ACCESS — search & retrieval crawlers allowed in (6 bots scored)
67%
TECHNICAL PAGE READINESS — unweighted mean of 4 pass rates (uncalibrated heuristic; see table)
100%
DEEP-INSPECTION COVERAGE — confidence: High
What this report can and can't tell you. It can tell you, checkably: which crawlers your robots.txt lets in, which of your pages are discoverable, and what technical signals each inspected page carries — with the evidence and an acceptance test for every fix. It cannot measure how AI systems understand or cite your content, and it does not promise traffic, rankings or mentions — nobody can verify those, so we don't sell them. Blocking training bots can be the right business choice; this report records your current policy so it is a decision, not an accident.

Coverage — what this audit did and did not see

Inventory sourcessitemap (1 declared in robots.txt) + link crawl + 0 owner-provided URL(s)
Declared in sitemaps3
Known URLs (union)3
Deep-inspected3 (100%) — 3 request(s) attempted against the 60-request budget; 0 failed or non-HTML
Liveness-checked only0 (alive 0 / confirmed dead 404-410 0 / access blocked 4xx 0 / server error 5xx 0 / no response 0)
Redirect aliases consolidated0 known URL(s) resolved to already-inspected pages (not re-probed); 0 duplicate fetch(es) consumed request budget
Crawl redirects refused (robots/host guard)0 known URL(s) — final classification, each refusal named in the robots-skipped list; not probed again
Not probed (robots.txt disallows this audit)0
Not checked (caps)0
Sitemap URLs confirmed dead (GET-confirmed 404/410 only)0
Reachable pages missing from sitemap0
Owner-provided pages not in sitemapn/a
ConfidenceHigh

Pages in no inventory above (e.g. unlinked landing pages) are invisible to this audit — and to crawlers. Provide them as additional URLs to include them; if intentionally hidden, make sure they carry noindex. Internal consistency: all internal consistency checks passed.

Who can get in today (robots.txt policy, RFC 9309 evaluation)

TokenOperatorTypeStatusWhyOfficial-page token check
GPTBotOpenAITraining / data-use control🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://developers.openai.com/api/docs/bots
OAI-SearchBotOpenAISearch & retrieval🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://developers.openai.com/api/docs/bots
ChatGPT-UserOpenAIUser-triggered fetch🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://developers.openai.com/api/docs/bots
OAI-AdsBotOpenAIOther🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://developers.openai.com/api/docs/bots
ClaudeBotAnthropicTraining / data-use control🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
Claude-UserAnthropicUser-triggered fetch🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
Claude-SearchBotAnthropicSearch & retrieval🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
PerplexityBotPerplexitySearch & retrieval🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://docs.perplexity.ai/guides/bots
Perplexity-UserPerplexityUser-triggered fetch🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://docs.perplexity.ai/guides/bots
GooglebotGoogleSearch & retrieval🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
Google-ExtendedGoogleTraining / data-use control (control token)🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
CCBotCommon CrawlTraining / data-use control🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://commoncrawl.org/ccbot
BytespiderByteDanceTraining / data-use control🟢 Allowed inall 3 known URL(s) allowed— no official documentation page (unverified)
HTTP — · 2026-08-16 19:56Z · no official source URL on file
AmazonbotAmazonTraining / data-use control🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://developer.amazon.com/amazonbot
Amzn-SearchBotAmazonSearch & retrieval🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://developer.amazon.com/amazonbot
Amzn-UserAmazonUser-triggered fetch🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://developer.amazon.com/amazonbot
ApplebotAppleSearch & retrieval🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://support.apple.com/en-us/119829
Applebot-ExtendedAppleTraining / data-use control (control token)🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://support.apple.com/en-us/119829
meta-externalagentMetaTraining / data-use control🟢 Allowed inall 3 known URL(s) allowed✓ token present on official page
HTTP 200 · 2026-08-16 19:56Z · source: https://developers.facebook.com/docs/sharing/webmasters/web-crawlers
anthropic-aiAnthropicLegacy token🟢 Allowed inall 3 known URL(s) allowed✓ absent from current official page (consistent with legacy status)
HTTP 200 · 2026-08-16 19:56Z · source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
Claude-WebAnthropicLegacy token🟢 Allowed inall 3 known URL(s) allowed✓ absent from current official page (consistent with legacy status)
HTTP 200 · 2026-08-16 19:56Z · source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
Applied robots.txt rule text per bot (evidence)

GPTBot (wildcard(*))

User-agent: *
Allow: /

OAI-SearchBot (wildcard(*))

User-agent: *
Allow: /

ChatGPT-User (wildcard(*))

User-agent: *
Allow: /

OAI-AdsBot (wildcard(*))

User-agent: *
Allow: /

ClaudeBot (wildcard(*))

User-agent: *
Allow: /

Claude-User (wildcard(*))

User-agent: *
Allow: /

Claude-SearchBot (wildcard(*))

User-agent: *
Allow: /

PerplexityBot (wildcard(*))

User-agent: *
Allow: /

Perplexity-User (wildcard(*))

User-agent: *
Allow: /

Googlebot (wildcard(*))

User-agent: *
Allow: /

Google-Extended (wildcard(*))

User-agent: *
Allow: /

CCBot (wildcard(*))

User-agent: *
Allow: /

Bytespider (wildcard(*))

User-agent: *
Allow: /

Amazonbot (wildcard(*))

User-agent: *
Allow: /

Amzn-SearchBot (wildcard(*))

User-agent: *
Allow: /

Amzn-User (wildcard(*))

User-agent: *
Allow: /

Applebot (wildcard(*))

User-agent: *
Allow: /

Applebot-Extended (wildcard(*))

User-agent: *
Allow: /

meta-externalagent (wildcard(*))

User-agent: *
Allow: /

Access score counts search & retrieval crawlers only; each contributes its allowed share of known URLs (ALLOWED=1, PARTIAL=allowed/(allowed+disallowed), BLOCKED/UNAVAILABLE=0). Allowing/blocking training bots is a business trade-off (protection from training vs presence in AI products) — recorded, not scored. Perplexity-User and Amzn-User officially may not honor robots.txt.

Technical page readiness (3 pages inspected)

CheckPass rateNote
🔴Parseable JSON-LD with a declared @type33%syntax-level check, NOT schema.org validation; unparseable blocks count as fail
🔴Title & meta description in range33%heuristic ranges, not official requirements
🟢Canonical present100%URL-consolidation signal
🟢Raw-HTML text not thin100%~50 words, CJK-aware — heuristic

Findings by URL family

FindingPagesURL families (shared template suspected)
meta description length outside 50–160 chars (heuristic, not an official requirement)2 (67%)sitemap (2)
no parseable JSON-LD structured data (syntax-level check)2 (67%)sitemap (2)

"URL family" groups by sitemap section or first path segment — a shared template is suspected, not confirmed; verify in your CMS. Per-page detail: (included in the delivered report) (included in the delivered report)

Fix plan — ticket-grade, do this then re-run

MEDIUMA1-FIX-01 — 2 of 3 page(s) have no parseable JSON-LD

Certainty: suspected (shared templates inferred from URL families — verify in your CMS) · Suggested owner: whoever owns the page templates · Affected: 2 pages (67%)

Evidence:

URL families affected: sitemap (2)

Samples: https://43sunsets.com/skills/n8n-production-readiness-auditor/ · https://43sunsets.com/skills/api-breaking-change-detector/

Requested change:

Start with the homepage: add an Organization JSON-LD to <head> (template below is prefilled from your own pages — EDIT it and make sure every statement matches your visible content before publishing):
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "43 Sunsets — Small, production-grade automations",
  "url": "https://43sunsets.com",
  "description": "43 Sunsets builds small, production-grade automations for real businesses. Every template ships with error handling, idempotency guards, and clear setup notes."
}
Then add the matching type per family template (Product, Article, FAQPage …) — again, only stating what the visible page already says.

Acceptance test: Re-run this audit: the JSON-LD pass rate rises family by family.

Risk / rollback: Structured data that contradicts visible content can be treated as spam by search engines — keep them consistent.

MEDIUMA1-FIX-02 — 2 page(s) with missing/out-of-range title or meta description

Certainty: confirmed (ranges are heuristics, not official requirements) · Suggested owner: content owner per family · Affected: 2 pages

Evidence:

Families: sitemap (2) — per-page detail in the dataset

Samples: https://43sunsets.com/skills/n8n-production-readiness-auditor/ · https://43sunsets.com/skills/api-breaking-change-detector/

Requested change:

Write each title as the page's one-line answer and each description as the two-sentence version; for templated families, fix the generating template.

Acceptance test: Re-run this audit: title/description flags disappear.

Risk / rollback: None.

Appendix A: llms.txt starter draft (optional)

Show draft
# 43 Sunsets — Small, production-grade automations

> 43 Sunsets builds small, production-grade automations for real businesses. Every template ships with error handling, idempotency guards, and clear setup notes.

## Pages
- [43 Sunsets — Small, production-grade automations](https://43sunsets.com/): 43 Sunsets builds small, production-grade automations for real businesses. Every template ships with
- [API Breaking-Change Detector — documentation · 43 Sunsets](https://43sunsets.com/skills/api-breaking-change-detector/): How the API Breaking-Change Detector skill works: give it two OpenAPI/Swagger spec versions, get a v
- [n8n Production-Readiness Auditor — documentation · 43 Sunsets](https://43sunsets.com/skills/n8n-production-readiness-auditor/): How the n8n Production-Readiness Auditor skill works: what to paste, what the graded report contains

Appendix B: discovered URL inventory

First 100 of 3 known URLs (the dataset holds one row per deep-inspected page; URLs beyond this list and the probe outcomes appear in OUTPUT's counts, not row by row)
- https://43sunsets.com/
- https://43sunsets.com/skills/n8n-production-readiness-auditor/
- https://43sunsets.com/skills/api-breaking-change-detector/

Method & limits

Raw HTML only (no JavaScript rendering) — what many crawlers see. Sitemaps are the site's claim, never truth: every URL reported with an HTTP result was actually fetched or probed (robots-skipped, capped, and inventory-only URLs are labelled as such), and unanswered probes are "unverifiable", not "dead". robots.txt evaluated per RFC 9309 (merged groups, longest match, Allow wins ties, * and $). The official-page check is token presence, not semantic verification. Heuristic thresholds are labeled. /cdn-cgi/ excluded. Hostnames are resolved once per run, every address must be publicly routable, and connections are made to the vetted address only. The one pre-robots request is the initial GET of the URL you provided (canonical host resolution); everything after is robots-checked pre-fetch. The audit reads only the site you provided (canonical apex/www resolution disclosed above) plus allowlisted official crawler-documentation pages, following their redirects. Crawl redirects refused by the robots/host guard are classified once (named in the robots-skipped list), not probed again; crawl-failed URLs keep their liveness check. Credential screening: inputs with credential-looking query parameters are refused; stored outputs mask the values of known secret-named parameters (token, key, signature, password, …), nested URLs included; values past the safety caps (length/decode/nesting) are masked wholesale as unstorable — known names only, so treat any URL embedding credentials as exposed wherever it is linked.