Local SEO
Site-wide audit findings
What the whole-site audit reports about a domain's structure (broken internal links, redirect chains and loops, orphan pages, deep pages, thin, duplicate and templated content, canonical conflicts), and when orphan and depth findings are held back.
Updated · 10 min read
The site-wide audit is the second pass over a site after the per-page crawl. Where the page-by-page scan says what is wrong on each page, the site-wide findings say what is wrong with the site's structure as a whole. Each finding names the page to edit and the fix, in the same shape the per-page scan uses, with what is there and what to change.
Where the audit lives
- Web: a domain's SEO & search tab, Audit every page, below the page-by-page crawl results.
- iOS: the app's deep-SEO crawl (
/api/deep-seo/jobs) reports the same findings on each page in the group. - API/MCP:
POST /api/site-audit/start,GET /api/site-audit/status,GET /api/site-audit/results; toolsstart_site_audit,get_site_audit_status,get_site_audit_results. The app's crawl uses/api/deep-seo/jobsandfetch_sitemapandstart_deep_seo_job.
A run reads from https://<domain> after a landing check (a redirect to another host or a registrar parking page is reported as one issue, with no pages audited). It covers up to 25 pages on Starter, up to 300 on Pro and Enterprise, plus the plan's own scan allowance (the deep-SEO crawl the app uses covers up to 5,000). The site audit's own ceiling is 300, so a Pro or Enterprise run that asks for more stops at 300 and the response says so.
The findings
Each finding has a type, a severity (error, warning, info), the URL it is filed under, a short message, and an evidence block with the fix.
Broken internal links (broken_internal_link)
A link on a listed page points at an address that does not load. The audit reports it two ways:
- One per page with few sources: a small number of pages link to one bad address, so each page gets one row, listing the broken links on it.
- One per target with many sources: when more than a few pages link to one bad address (almost always a header, footer or menu link), the finding is filed under the first source and lists every page that links to it. One edit to the template fixes all of them.
The target's state is read from the crawl itself when it was opened, or from a short probe at the end of the audit. A 404 or 410 says "remove or 301-redirect"; a redirecting target says "point the link at the final address"; an unreachable target says "open it in a browser; if it loads there, the host is blocking our crawler".
Redirect chains and loops
Two findings:
| Type | Severity | When |
|---|---|---|
redirected |
Info | One redirect: a sitemap or a link points at an address that 301s or 302s to another. The fix is to point straight at the final address. |
redirect_chain |
Warning | Two or more redirects: each hop costs a crawler a request (Google follows at most ten), and the sitemap should list final URLs only. The fix is the same: replace the source address with the final one everywhere. |
redirect_loop |
Error | The chain never lands. Browsers show an error and search engines give up. The fix is the redirect rules that undo each other (an http/https or www rule in the host fighting one in a plugin or CDN, or two trailing-slash rules). |
Orphan pages (orphan_page)
A page that is in the sitemap, but no other page on the site links to it. The visitor who lands on it from search has nowhere to go; Google treats a page the site itself never links to as one that does not matter much. The fix is to link to it from a related page (a service page, the menu) or 301-redirect it to its replacement if it is retired.
A page is not orphaned when:
- another page in the audit landed on it (a 301 → /decommissioned/ that /about/ also linked to);
- it is the homepage;
- it is noindex, or answered a non-2xx status.
Deep pages (deep_page)
A page whose shortest path from the homepage is 5 or more clicks (DEEP_PAGE_DEPTH). Search engines crawl pages near the homepage more often and treat them as more important, and visitors rarely dig that deep. The fix is to link to it from a page nearer the home: the menu, a hub page, a related page.
Thin content (thin_content)
A page with under 150 words of its own text once the menu, header, footer, sidebars and forms are left out. A page that says this little rarely ranks for anything; the fix is to write what a customer would ask about (what you do, where, what it costs, what happens next), or to merge it into a page that already has the answer. Contact and thank-you pages are often short on purpose; this finding is skipped for those.
The check is on the page's own HTML, not on a sample. A client-rendered shell, a page that defers to another, and a page whose noindex is expected are left alone.
Duplicate content (duplicate_content)
A pair of pages whose main text matches by an estimated 90% or more, compared five words at a time. Each page is masked by its own URL, so the comparison is on the body, not on the address. The fix is to keep one and 301-redirect the other, or to point the copy's canonical at the one to keep.
Pairs at 60-90% are reported as a note (the message reads "Much of the main text matches…"). A paraphrase check would accuse honest pages (nhmohio.com's independently written city pages share 38-46% of their five-word runs), so the threshold is set where a human reading would agree.
A pair inside the same templated group (see below) is not raised twice.
Templated pages (templated_pages)
A group of three or more pages that share an estimated 80% or more of their main text. This is the city page copied for every town with the name swapped. Google's spam policies call it scaled content or doorway pages; the deep-SEO crawl stores one per group, with all the pages in the locations list.
Product pages, noindex pages, and pages that point their canonical elsewhere are left out. Pages with too little text are not compared. Pages rewritten for each town, with their own sentences, are not flagged; the check would rather miss a pattern than accuse a site of one. What to do about a group is in AI-written content and location pages.
Canonical conflicts
Two findings, on different halves of the same problem.
| Type | Where it is raised | When raised |
|---|---|---|
canonical_conflict (per-page, from the shared checks) |
On the page itself | Several canonical tags on one page naming different URLs. Google ignores the lot when they disagree. |
canonical_conflict (cross-page, from the audit) |
On the page whose canonical points elsewhere | The named canonical is itself broken (does not load, redirects, is noindex, or names a third page). Google cannot use it. |
The fix is <link rel="canonical" href="<this page's address>"> on each affected page, or naming the canonical's final, indexable address.
Other things the audit reports
These are not new with the site-wide pass; the per-page scan reports them too, and the audit confirms them at scale:
- HTTP errors (
http_error) — a listed URL that answered 4xx or 5xx. A 404 or 410 should be removed from the sitemap and the links that point at it, or 301-redirected to its replacement. - Noindex pages (
noindex) — a page that asks not to be indexed. An error when the sitemap lists it (the sitemap contradicts the page); a note when only a link reaches it (the choice is usually deliberate). - Missing sitemap (
missing_sitemap) — no readable XML sitemap from robots.txt or the four common paths; pages were found by following links, which misses pages nothing links to. The fix names the file to create and the line to add. - Site redirects (
site_redirects) / site parked (site_parked) — the domain itself answers from somewhere else, or shows a registrar parking page. One issue, no pages audited.
When orphan and depth findings are withheld
Orphan and depth need the whole site: the audit must have reached every page, and the home page's links must have been read. The audit reports them only when the graph is "usable":
- the crawl covered the whole site (no
truncatedpage list, no partial link walk); - the home page was read (its HTML was opened, not just probed);
- at least 90% of the audited pages were read (
GRAPH_MIN_READ_SHARE); - fewer than 20% of them are client-rendered shells (their links are drawn later and not in the HTML).
When any of these is not true, orphan_page and deep_page are held back and the audit's response says the matching reason ("the crawl did not cover the whole site", "the home page was not read", "too many pages could not be read", or "too many pages draw their links with JavaScript"). The other findings (broken links, redirects, thin, duplicate, templated, canonical) are still reported. The page-by-page scan covers the rest.
A partial walk is rare. It can happen when the link walk stopped on its clock (the audit moves the work to a slower cron tick and finishes) or when the site's own page count went past the plan's cap. The audit's truncated and partial flags say so in its status.
Where the audit keeps this
A finished audit holds:
site_audits(the run): id, root URL, discovery method (sitemaporcrawl), pages found and checked, pages discovered, truncated, issue count, started and completed timestamps, and the error if it did not finish.site_audit_issues: each finding, ordered by severity and URL, withevidence_jsoncarrying the fix.site_audit_page_meta: per-page title, description and indexability, kept for the duplicate check at the end.
A run that the reaper ended (no progress for 30 minutes) is moved to status: 'error' with error: "This audit did not finish. The scan stopped before it reached the end of the site, so anything it did check has been kept. Run it again to cover the rest.". A finished audit and its issues are kept for 90 days; older ones are pruned.
Common questions
The audit says 200 issues on a 50-page site. Many findings come from the same few broken links — a footer link on 80 pages is one finding under "links from 80 pages", not 80 findings. Filter the list by type and read the message; the audit groups sources for you.
My orphan pages are ones I just launched. A page that is in the sitemap but has no inbound links is orphaned by definition, however new it is. Link to it from a related page on launch.
A canonical conflict says the named canonical is noindex. Both pages want to be in search. Pick the page to rank and 301-redirect the other to it, or remove its noindex and let both compete (only if the content is genuinely different).
Keep reading
Related articles
- Whole-site crawlHow the crawl finds pages (sitemap plus links), what it checks on each, the 25-page Starter cap, the 300-page site audit and the 5,000-page deep crawl, progress, cancelling and per-page.Website and SEO ·Updated
- DMARC reportsThe report address DomainGuard hosts for you, what aggregate reports reveal about who sends as your domain, the 30-day summary, rotation, and alerts.Email authentication ·Updated
- AI-written content and location pagesWhat Google says about AI-written pages, city pages and structured data, how DomainGuard's checks and AI features follow it, and when a page for a town is worth making.Website and SEO ·Updated
