2026-08-16 19:56Z · 3 of 3 known URLs deep-inspected · robots.txt respected · recommended next check: 2026-09-15
Crawler registry updated 2026-08-06 (bundled snapshot); token presence checked against official pages this run (this is a presence check, not semantic verification).
| Inventory sources | sitemap (1 declared in robots.txt) + link crawl + 0 owner-provided URL(s) |
| Declared in sitemaps | 3 |
| Known URLs (union) | 3 |
| Deep-inspected | 3 (100%) — 3 request(s) attempted against the 60-request budget; 0 failed or non-HTML |
| Liveness-checked only | 0 (alive 0 / confirmed dead 404-410 0 / access blocked 4xx 0 / server error 5xx 0 / no response 0) |
| Redirect aliases consolidated | 0 known URL(s) resolved to already-inspected pages (not re-probed); 0 duplicate fetch(es) consumed request budget |
| Crawl redirects refused (robots/host guard) | 0 known URL(s) — final classification, each refusal named in the robots-skipped list; not probed again |
| Not probed (robots.txt disallows this audit) | 0 |
| Not checked (caps) | 0 |
| Sitemap URLs confirmed dead (GET-confirmed 404/410 only) | 0 |
| Reachable pages missing from sitemap | 0 |
| Owner-provided pages not in sitemap | n/a |
| Confidence | High |
Pages in no inventory above (e.g. unlinked landing pages) are invisible to this audit — and to crawlers. Provide them as additional URLs to include them; if intentionally hidden, make sure they carry noindex. Internal consistency: all internal consistency checks passed.
| Token | Operator | Type | Status | Why | Official-page token check |
|---|---|---|---|---|---|
| GPTBot | OpenAI | Training / data-use control | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://developers.openai.com/api/docs/bots |
| OAI-SearchBot | OpenAI | Search & retrieval | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://developers.openai.com/api/docs/bots |
| ChatGPT-User | OpenAI | User-triggered fetch | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://developers.openai.com/api/docs/bots |
| OAI-AdsBot | OpenAI | Other | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://developers.openai.com/api/docs/bots |
| ClaudeBot | Anthropic | Training / data-use control | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler |
| Claude-User | Anthropic | User-triggered fetch | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler |
| Claude-SearchBot | Anthropic | Search & retrieval | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler |
| PerplexityBot | Perplexity | Search & retrieval | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://docs.perplexity.ai/guides/bots |
| Perplexity-User | Perplexity | User-triggered fetch | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://docs.perplexity.ai/guides/bots |
| Googlebot | Search & retrieval | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers | |
| Google-Extended | Training / data-use control (control token) | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers | |
| CCBot | Common Crawl | Training / data-use control | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://commoncrawl.org/ccbot |
| Bytespider | ByteDance | Training / data-use control | 🟢 Allowed in | all 3 known URL(s) allowed | — no official documentation page (unverified) HTTP — · 2026-08-16 19:56Z · no official source URL on file |
| Amazonbot | Amazon | Training / data-use control | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://developer.amazon.com/amazonbot |
| Amzn-SearchBot | Amazon | Search & retrieval | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://developer.amazon.com/amazonbot |
| Amzn-User | Amazon | User-triggered fetch | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://developer.amazon.com/amazonbot |
| Applebot | Apple | Search & retrieval | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://support.apple.com/en-us/119829 |
| Applebot-Extended | Apple | Training / data-use control (control token) | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://support.apple.com/en-us/119829 |
| meta-externalagent | Meta | Training / data-use control | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ token present on official page HTTP 200 · 2026-08-16 19:56Z · source: https://developers.facebook.com/docs/sharing/webmasters/web-crawlers |
| anthropic-ai | Anthropic | Legacy token | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ absent from current official page (consistent with legacy status) HTTP 200 · 2026-08-16 19:56Z · source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler |
| Claude-Web | Anthropic | Legacy token | 🟢 Allowed in | all 3 known URL(s) allowed | ✓ absent from current official page (consistent with legacy status) HTTP 200 · 2026-08-16 19:56Z · source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler |
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
User-agent: * Allow: /
Access score counts search & retrieval crawlers only; each contributes its allowed share of known URLs (ALLOWED=1, PARTIAL=allowed/(allowed+disallowed), BLOCKED/UNAVAILABLE=0). Allowing/blocking training bots is a business trade-off (protection from training vs presence in AI products) — recorded, not scored. Perplexity-User and Amzn-User officially may not honor robots.txt.
| Check | Pass rate | Note | |
|---|---|---|---|
| 🔴 | Parseable JSON-LD with a declared @type | 33% | syntax-level check, NOT schema.org validation; unparseable blocks count as fail |
| 🔴 | Title & meta description in range | 33% | heuristic ranges, not official requirements |
| 🟢 | Canonical present | 100% | URL-consolidation signal |
| 🟢 | Raw-HTML text not thin | 100% | ~50 words, CJK-aware — heuristic |
| Finding | Pages | URL families (shared template suspected) |
|---|---|---|
| meta description length outside 50–160 chars (heuristic, not an official requirement) | 2 (67%) | sitemap (2) |
| no parseable JSON-LD structured data (syntax-level check) | 2 (67%) | sitemap (2) |
"URL family" groups by sitemap section or first path segment — a shared template is suspected, not confirmed; verify in your CMS. Per-page detail: (included in the delivered report) (included in the delivered report)
URL families affected: sitemap (2)
Start with the homepage: add an Organization JSON-LD to <head> (template below is prefilled from your own pages — EDIT it and make sure every statement matches your visible content before publishing):
{
"@context": "https://schema.org",
"@type": "Organization",
"name": "43 Sunsets — Small, production-grade automations",
"url": "https://43sunsets.com",
"description": "43 Sunsets builds small, production-grade automations for real businesses. Every template ships with error handling, idempotency guards, and clear setup notes."
}
Then add the matching type per family template (Product, Article, FAQPage …) — again, only stating what the visible page already says.
Families: sitemap (2) — per-page detail in the dataset
Write each title as the page's one-line answer and each description as the two-sentence version; for templated families, fix the generating template.
# 43 Sunsets — Small, production-grade automations > 43 Sunsets builds small, production-grade automations for real businesses. Every template ships with error handling, idempotency guards, and clear setup notes. ## Pages - [43 Sunsets — Small, production-grade automations](https://43sunsets.com/): 43 Sunsets builds small, production-grade automations for real businesses. Every template ships with - [API Breaking-Change Detector — documentation · 43 Sunsets](https://43sunsets.com/skills/api-breaking-change-detector/): How the API Breaking-Change Detector skill works: give it two OpenAPI/Swagger spec versions, get a v - [n8n Production-Readiness Auditor — documentation · 43 Sunsets](https://43sunsets.com/skills/n8n-production-readiness-auditor/): How the n8n Production-Readiness Auditor skill works: what to paste, what the graded report contains
- https://43sunsets.com/ - https://43sunsets.com/skills/n8n-production-readiness-auditor/ - https://43sunsets.com/skills/api-breaking-change-detector/
Raw HTML only (no JavaScript rendering) — what many crawlers see. Sitemaps are the site's claim, never truth: every URL reported with an HTTP result was actually fetched or probed (robots-skipped, capped, and inventory-only URLs are labelled as such), and unanswered probes are "unverifiable", not "dead". robots.txt evaluated per RFC 9309 (merged groups, longest match, Allow wins ties, * and $). The official-page check is token presence, not semantic verification. Heuristic thresholds are labeled. /cdn-cgi/ excluded. Hostnames are resolved once per run, every address must be publicly routable, and connections are made to the vetted address only. The one pre-robots request is the initial GET of the URL you provided (canonical host resolution); everything after is robots-checked pre-fetch. The audit reads only the site you provided (canonical apex/www resolution disclosed above) plus allowlisted official crawler-documentation pages, following their redirects. Crawl redirects refused by the robots/host guard are classified once (named in the robots-skipped list), not probed again; crawl-failed URLs keep their liveness check. Credential screening: inputs with credential-looking query parameters are refused; stored outputs mask the values of known secret-named parameters (token, key, signature, password, …), nested URLs included; values past the safety caps (length/decode/nesting) are masked wholesale as unstorable — known names only, so treat any URL embedding credentials as exposed wherever it is linked.