If your crawler shows 400 warnings, most of them won’t cost you a single visit. The ones that do stop search engines from crawling, rendering or indexing the pages that make money. Below are 35 checks, each tagged by traffic impact and effort, so your team fixes those first and backlogs the rest.
How to prioritize technical fixes
Crawlers report everything they can measure, not everything that matters. A slightly long title and a noindex tag on your pricing page sit side by side in the export. Only one of them removes a page from Google.
I tag every finding on two scales:
- Impact. High: can remove pages or whole sections from search. Medium: affects how well indexed pages rank or how efficiently they’re crawled. Low: hygiene.
- Effort. Low: a config or CMS change, under a day. Medium: a template change within one sprint. High: rendering or architecture work across several sprints.
Scope by template, not by URL. The same canonical bug on 3,000 product pages is one ticket, not 3,000 warnings.
Warnings you can usually batch
- Titles or meta descriptions slightly over a length limit
- Multiple H1 tags on a page
- “Low word count” on login, contact and other pages not meant to rank
- Single-hop internal redirects
- Missing alt text on decorative images (fix it for accessibility, not rankings)
The fix-first shortlist
Start with the High-impact, Low-effort checks: 1, 7, 8, 9, 10 and 29. Each is quick to verify, most failures are configuration fixes, and any one can keep key pages out of search.
Crawlability and robots.txt
| # | Check | Impact | Effort |
|---|---|---|---|
| 1 | robots.txt returns 200 and doesn’t block revenue sections or keep staging rules | High | Low |
| 2 | Crawl Stats in Search Console show no spikes in 5xx errors or timeouts | High | Medium |
| 3 | Filter, sort and tracking parameters don’t create endless crawlable URLs | High | Medium |
| 4 | XML sitemaps list only indexable, canonical, 200-status URLs | Medium | Low |
| 5 | No redirect chains or loops; internal links point to final URLs | Medium | Low |
| 6 | CSS and JavaScript files needed for rendering aren’t blocked | Medium | Low |
Check 1 often fails right after a launch, when a staging Disallow: / ships to production. The website redesign SEO checklist covers that window.
On check 2, Google slows crawling when a server struggles, and a robots.txt file that returns a server error can pause crawling of the whole site. Check 3 mostly hits ecommerce and large catalogs, where every color, size and sort combination becomes a URL. Fix it with a written policy on which facets deserve indexable pages, not one blanket robots rule.
Indexing, canonicals and duplicates
| # | Check | Impact | Effort |
|---|---|---|---|
| 7 | Priority templates (product, category, pricing) show as indexed in the Page indexing report | High | Low |
| 8 | No stray noindex in meta robots tags or X-Robots-Tag headers | High | Low |
| 9 | http, www and trailing-slash variants redirect in one hop to one version | High | Low |
| 10 | Indexable pages have an absolute, self-referencing canonical to a 200 URL | High | Low |
| 11 | Sitemaps, internal links and hreflang all use the canonical URL | Medium | Medium |
| 12 | Soft 404s and thin templates (empty categories, tag archives, site search) are consolidated or noindexed | Medium | Medium |
| 13 | hreflang is reciprocal, uses valid codes and points to canonicals | Medium | Medium |
Start check 7 in Search Console, not your crawler: the Page indexing report says why Google excluded each URL. Filter by sitemap or URL prefix, one template at a time. A large share of one template stuck in “Crawled - currently not indexed” usually signals thin or duplicate content, not a technical bug.
A canonical is a hint, not a directive. If internal links, sitemaps and redirects point somewhere else, Google may choose its own, hence check 11. And don’t block URLs in robots.txt that you want noindexed: Google can’t see a noindex tag on a page it isn’t allowed to crawl.
Site architecture and internal linking
| # | Check | Impact | Effort |
|---|---|---|---|
| 14 | Links are real <a href> elements, not JavaScript click handlers | High | Medium |
| 15 | Commercial pages get contextual links from related high-traffic pages | Medium | Low |
| 16 | No important page is orphaned | Medium | Low |
| 17 | Money pages sit within about three clicks of the homepage | Medium | Medium |
| 18 | Paginated lists use crawlable page URLs, not only a “load more” button | Medium | Medium |
| 19 | Internal links don’t point to 404 pages | Low | Low |
Check 14 breaks often on modern front ends. A menu built from buttons with click handlers works for users but gives Google no links to follow.
Check 15 is often the audit’s cheapest ranking lever: make sure each of your top 20 organic pages links to its closest commercial page with descriptive anchor text. Check 19 is Low: a few broken links rarely cost traffic, however loudly crawlers flag them.
Page speed and Core Web Vitals
| # | Check | Impact | Effort |
|---|---|---|---|
| 20 | Core Web Vitals judged on field data (CrUX) by template, not one lab score | Medium | Low |
| 21 | The LCP image isn’t lazy-loaded and loads early at the right size | Medium | Low |
| 22 | Server response time is stable, with caching and a CDN | Medium | Medium |
| 23 | INP: long JavaScript tasks and third-party tags are cut or deferred | Medium | High |
| 24 | CLS: images, embeds, banners and fonts load without shifting layout | Low | Low |
Core Web Vitals are a ranking signal, but a modest one next to relevance, so I rate them Medium unless mobile pages fail badly. The fastest win is usually check 21: a theme that lazy-loads every image, hero included, delays the largest element on every page. For INP, open your tag manager before touching code. Tags for abandoned tools often still load on every page.
Structured data
| # | Check | Impact | Effort |
|---|---|---|---|
| 25 | Eligible templates use matching schema (Product, BreadcrumbList, Article, VideoObject) | Medium | Low |
| 26 | Markup matches visible prices, availability and ratings | Medium | Low |
| 27 | No errors in Search Console’s rich result reports for key templates | Medium | Low |
| 28 | Organization schema has name, logo and sameAs profile links | Low | Low |
Structured data doesn’t raise rankings by itself; it makes pages eligible for rich results and states entity details plainly. Generate it from template data so a price change updates the markup automatically. Skip FAQ and HowTo markup as a rich-results play: in 2023 Google limited FAQ rich results to a small set of authoritative sites and dropped HowTo rich results.
JavaScript rendering and AI crawler access
| # | Check | Impact | Effort |
|---|---|---|---|
| 29 | CDN and firewall rules don’t block Googlebot, Bingbot or AI crawlers you intend to allow | High | Low |
| 30 | Main content and internal links are in the server-rendered HTML | High | High |
| 31 | URL Inspection’s rendered HTML shows the full page; JavaScript doesn’t rewrite titles, canonicals or meta robots | High | Medium |
| 32 | Mobile pages match desktop content, links and structured data | High | Medium |
| 33 | Tab, accordion and infinite-scroll content is in the HTML or has its own URL | Medium | Medium |
| 34 | robots.txt doesn’t block AI crawlers (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot) by accident | Medium | Low |
| 35 | The site is verified in Bing Webmaster Tools and key pages are indexed in Bing | Medium | Low |
Check 29 catches a silent failure: CDN bot protection that blocks crawlers before robots.txt is ever read. Look for 403s or challenge pages served to crawler user agents in your logs.
For check 30, compare “view source” with the rendered DOM on one page per template. Google renders JavaScript, but rendering can lag and errors fail quietly. AI search raises the stakes: many AI crawlers fetch raw HTML without running JavaScript, so a client-rendered page can look almost empty to them.
On check 34, Google-Extended is a robots.txt token, not a separate crawler, and blocking it doesn’t remove your pages from Google Search. Check 35 matters because Bing’s index feeds Microsoft Copilot and some other AI search tools.
This audit only confirms crawlers can reach you; whether to allow or block specific AI crawlers is a policy decision covered in the llms.txt guide.
Turning findings into dev tickets
An audit spreadsheet doesn’t ship fixes; tickets do. Each one should let a developer fix the issue and QA verify it without a meeting. When I run technical SEO work, every finding leaves the audit as a ticket like this made-up example:
Title: Product pages: canonical points to parameter URL, splitting signals
Priority: High impact / Low effort
Affected: /products/* (about 2,400 URLs), e.g. /products/wool-throw?color=gray
Current behavior: canonical uses the first variant URL, including ?color=
Expected behavior: canonical is the clean product URL on every variant
Acceptance criteria:
- Canonical on every /products/* URL is absolute, https, no query string
- Variant URLs canonicalize to the parameter-free product URL
- Canonical target returns 200 and appears in the product sitemap
How to verify: crawl /products/ after deploy; zero canonicals containing "?"
Three rules make tickets like this land:
- Write acceptance criteria a crawler can check. “Improve canonicals” can’t be verified. “Zero canonicals containing a query string” can.
- One ticket per root cause, grouped by template. Assign it to the team that owns the template, not to “SEO.”
- Re-crawl after release and log the deploy date, so you can read the effect in Search Console.
Get it built
If your crawler report is long and your dev queue is short, I can run the audit and hand your team ranked, ready-to-build tickets. The entry point is the fixed-price Growth Audit: $1,500, credited if we continue. Get in touch with your domain, or see pricing.