All guides
SEO

Technical SEO Audit Checklist: 48 Checks for Small Businesses

A plain-English technical SEO audit checklist with 48 checks for small business sites: crawling, indexing, redirects, speed, mobile, schema and AI bots.

By Sahil Aggarwal, Founder, Growvia · September 13, 2026 · 15 min read

You can write the best service page in your city and still get zero traffic from it. If Google can't crawl the page, won't index it, or picks a duplicate copy instead, your content never gets a chance to rank. Technical SEO is the plumbing underneath everything else.

The good news is that most small business websites have the same dozen or so problems. A dental clinic site that loads on both http:// and https://. A bakery whose new website still has a leftover noindex tag from the developer's staging copy. A CA firm whose old blog URLs all lead to a 404 page after a redesign. None of these need an agency to find. They need a checklist and an hour.

Here are 48 checks grouped by area, with how to test each one using free tools and how to decide what to fix first. New to audits? Start with the 60-minute walkthrough near the end.

How Google actually finds and ranks a page

Every page goes through three steps, and each part of the audit maps to one of them.

How a page gets from your server to Google's results
  1. 1Discover
  2. 2Crawl
  3. 3Render
  4. 4Index
  5. 5Rank
  1. Crawling. Googlebot, Google's automated visitor, discovers your URL through links or your sitemap and downloads it.
  2. Rendering. Google runs the page's JavaScript to see what a real browser would show.
  3. Indexing. Google decides whether the page is useful and unique enough to store, and which version of it is the main one.

Ranking only happens after all three. A technical audit checks for anything that blocks or confuses those steps. Keywords and content come later; our guide to keyword research for small businesses picks up there.

Crawling and indexing checks

One wrong line here can hide your whole site from Google.

robots.txt: the "do not enter" sign

Your robots.txt file lives at yourdomain.com/robots.txt and tells crawlers which areas they may visit. Disallow: /cart/ sensibly keeps bots out of checkout. Disallow: / blocks everything, and it's surprisingly common after a launch.

Many people miss this: Google says robots.txt is not a mechanism for keeping a page out of Google. A blocked page can still appear in results without a description. To hide a page, use noindex or a password.

noindex: the "don't list me" label

A noindex tag, <meta name="robots" content="noindex">, tells Google not to show a page in results. Fine on thank-you pages, a disaster on your homepage. WordPress's "Discourage search engines" setting does the same thing site-wide.

XML sitemap: your list of pages

An XML sitemap is a file listing the URLs you want indexed, usually at /sitemap.xml. Google's sitemap guidelines set a limit of 50,000 URLs or 50MB per file. Google ignores the priority and changefreq fields, and only trusts lastmod dates that are consistently accurate.

Search Console: Page indexing report and URL Inspection

Google Search Console is free, and it's where Google tells you directly what it thinks of your site. Two features matter most:

  • Page indexing report (Indexing, then Pages). It groups non-indexed URLs by reason, such as "Excluded by noindex tag" or "Duplicate, Google chose different canonical than user".
  • URL Inspection. Paste any URL to see whether it's indexed, which canonical Google chose and when it was last crawled.
#CheckHow to testWhat good looks like
1robots.txt exists and loadsVisit /robots.txtReturns a 200 status, no Disallow: / for all bots
2CSS and JavaScript aren't blockedRead robots.txt rulesGoogle can fetch the files it needs to render pages
3XML sitemap existsVisit /sitemap.xmlLoads and lists your real pages
4Sitemap lists only good URLsSpot-check URLs in a crawlerOnly canonical, indexable, 200-status pages
5Sitemap submitted and referencedSearch Console Sitemaps; Sitemap: line in robots.txtStatus "Success", URL count matches reality
6No accidental noindexCrawler or URL Inspection on key pagesMoney pages are indexable
7Page indexing report reviewedSearch Console, Indexing, PagesEvery excluded reason is understood and intentional
8Key pages inspectedURL Inspection on homepage and top services"URL is on Google", correct canonical

Duplicate versions: HTTPS, www and trailing slashes

To Google, http://clinic.com, https://clinic.com and https://www.clinic.com are different addresses. If all of them load the same page, you split your signals and let Google guess which to show.

Pick one HTTPS version and permanently redirect every other version to it. Do the same for trailing slashes: /services/ or /services, never both.

Also watch for mixed content: an HTTPS page loading images or scripts over plain HTTP.

#CheckHow to testWhat good looks like
9Valid HTTPS certificateOpen the site in a browserPadlock shows, no warnings, certificate not near expiry
10HTTP redirects to HTTPSType the http:// versionOne 301 hop to the HTTPS URL
11One host versionTry www and non-wwwThe other version redirects to your chosen one
12Consistent trailing slashesTry a URL with and without the slashOne version redirects to the other
13No mixed contentBrowser console or a crawlerAll images, scripts and fonts load over HTTPS

Canonicals and redirects

Canonical tags: "this is the main version"

A canonical tag, <link rel="canonical" href="...">, tells Google which URL is the original when similar versions exist, such as /shirts/blue and /shirts/blue?colour=navy.

It's a hint, not a command. If it conflicts with your redirects, sitemap or internal links, Google may choose its own canonical.

301 vs 302 redirects

A redirect sends visitors and bots from an old URL to a new one. The type matters. Google's redirects documentation says a permanent redirect (301 or 308) signals that the target should be canonical. A temporary one (302 or 307) doesn't, so the old URL may stay in results.

Use 301 for anything permanent, like a redesign or domain move. Keep 302 for genuinely short-term cases.

Redirect chains and loops

A chain is A redirecting to B, which redirects to C. Each hop slows things down. A loop, where B sends you back to A, breaks the page. Point every old URL straight to its final destination.

#CheckHow to testWhat good looks like
14Every indexable page has a canonicalCrawler reportUsually self-referencing on unique pages
15Canonicals point to live, indexable URLsCrawler reportTargets return 200 and aren't noindexed
16Only one canonical per pageView source or crawlerNo duplicate tags from theme plus plugin
17Permanent moves use 301 or 308Crawler status codesNo long-lived 302s on moved pages
18No redirect chains or loopsCrawler redirect reportEvery redirect resolves in one hop
19Internal links point to final URLsCrawler "redirected internal links"Menus and body links skip redirects entirely

Status codes: 404, soft 404 and 5xx

Every page load returns a three-digit status code that Google reads.

  • 200 means OK.
  • 404 means not found. Google says it drops previously indexed URLs that return 404 or 410. That's fine for pages you removed on purpose.
  • Soft 404 is a page that's empty or says "not found" while sending 200. Empty category pages are common culprits.
  • 5xx means a server error. Google slows its crawling when it sees these, and persistent server errors can get URLs removed from the index.

A few 404s from old, unlinked pages are harmless. Fix 404s your own pages still link to, and redirect removed pages that had backlinks.

#CheckHow to testWhat good looks like
20No internal links to 404 pagesCrawler broken-link reportZero broken internal links
21No soft 404s on real pagesSearch Console Page indexing reportReal pages have real content
22No recurring 5xx errorsSearch Console Crawl stats; crawlerServer errors are rare and short-lived
23Custom 404 page returns a 404 codeVisit a made-up URL, check statusHelpful page, but status is 404, not 200

Click depth and orphan pages

Google finds pages by following links and treats often-linked pages as more important. Keep every important page within about three clicks of the homepage.

An orphan page has no internal links pointing to it. Ad landing pages often end up orphaned, then never rank.

Anchor text

Anchor text is a link's clickable text. "Root canal treatment in Mohali" tells Google far more than "click here".

URL design

Good URLs are short, lowercase and hyphenated: /services/teeth-whitening/ beats /index.php?page_id=482. Never change a live URL without a 301 redirect.

If ?sort=price and ?sort=newest each create an indexable URL, a small store can produce hundreds of near-duplicates.

#CheckHow to testWhat good looks like
24Key pages within three clicksCrawler "crawl depth" columnServices, products and contact are shallow
25No orphan pagesCompare sitemap URLs against crawled URLsEvery important page has internal links
26Descriptive anchor textReview menus and body linksAnchors describe the destination
27Clean, readable URLsScan the crawl listLowercase, hyphens, no random IDs
28Parameters don't create duplicatesCrawl and look for ? URLsCanonicals or noindex handle filtered views

Titles, descriptions, headings and images

Title tags and meta descriptions

The title tag is the clickable headline in search results; the meta description is the text beneath it. Google doesn't always use them. Google's title link documentation says it may generate a different title when yours is missing, vague, outdated or doesn't match the page's main heading. For descriptions, Google's snippet guidance says it mostly builds snippets from page content, and uses your meta description when it describes the page better. So write clear, specific ones, and make your title and H1 agree.

Headings

Use one H1 that states the page's topic, then H2s and H3s like chapters and sub-chapters. Never pick heading levels for font size.

Images

Images are usually the heaviest part of a page. Give each one alt text, such as "dentist fitting clear aligners". Compress and size files to how they display. Prefer WebP or AVIF, and lazy-load only images below the fold.

#CheckHow to testWhat good looks like
29Every page has a titleCrawler "missing titles"No blanks, no "Home" or "Untitled"
30Titles are unique and specificCrawler "duplicate titles"Service plus location or benefit, brand at the end
31Meta descriptions are uniqueCrawler reportAccurate summary with a reason to click
32One clear H1 per pageCrawler H1 reportMatches the page's topic and title
33Logical heading orderBrowser heading extensionH2s and H3s in a sensible outline
34Images have descriptive alt textCrawler "missing alt"Decorative images excepted
35Images compressed and sizedPageSpeed Insights opportunitiesNo oversized image warnings
36Modern formats and correct lazy-loadingPageSpeed InsightsWebP or AVIF; hero image loads immediately

Mobile, speed and JavaScript

Mobile-first indexing

Google indexes and ranks the mobile version of your site, crawled with a smartphone bot. If your mobile layout hides your services or reviews, Google may not count them. Keep content and structured data the same on both.

Page speed and Core Web Vitals

Core Web Vitals are Google's three user-experience metrics. According to web.dev, "good" means Largest Contentful Paint (loading) within 2.5 seconds, Interaction to Next Paint (responsiveness) of 200 milliseconds or less, and Cumulative Layout Shift (visual stability) of 0.1 or less, measured at the 75th percentile of real visits. For what each metric means and how to fix it, read Core Web Vitals explained.

JavaScript rendering

Google can run JavaScript, but its JavaScript SEO basics guide flags real traps. It only follows links that are <a> elements with an href, so script-driven buttons may never be followed. And if the raw HTML says noindex, Google may skip rendering, so removing it with JavaScript may not work.

Quick test: in URL Inspection, run "Test live URL" and view the rendered page. If prices or reviews are missing there, Google can't see them.

#CheckHow to testWhat good looks like
37Responsive design works on phonesOpen key pages on your own phoneReadable text, tappable buttons, no sideways scroll
38Mobile content matches desktopCompare both versionsNothing important hidden or removed on mobile
39Core Web Vitals passPageSpeed Insights field data; Search ConsoleLCP, INP and CLS all "Good" on mobile
40Links are real <a href> linksCrawler finds all pages; inspect menu codeNavigation is crawlable without clicks
41Key content is in the rendered HTMLURL Inspection, Test live URLText, prices and reviews visible to Google
42No noindex in raw HTML that JS later removesView source vs rendered HTMLRaw HTML has the correct robots tag

Structured data, hreflang and AI crawlers

Structured data

Structured data, or schema markup, labels your content for machines: your business, hours and reviews. It can make pages eligible for rich results. Test it with Google's Rich Results Test, and make sure the markup matches what visitors actually see. Our schema markup guide for local businesses shows exactly which types to use.

hreflang (only if you have multiple languages)

If you run English and Hindi versions, hreflang tags tell Google which suits which audience. Each version must reference all others, reciprocally, plus an x-default fallback. Single-language sites can skip this.

AI crawlers in robots.txt

AI companies run their own crawlers, which you can allow or block in robots.txt. Some collect training data. Others fetch pages to cite in AI answers, and blocking those can remove you from them.

CrawlerCompanyWhat it does
GPTBotOpenAICollects content for training models
OAI-SearchBotOpenAISurfaces sites in ChatGPT search answers
ChatGPT-UserOpenAIVisits pages when a user asks; robots.txt may not apply
Google-ExtendedGoogleControls use of content for Gemini training
ClaudeBotAnthropicCollects content for model training
Claude-SearchBotAnthropicIndexes content for Claude's search results
Claude-UserAnthropicFetches pages when a user asks Claude
PerplexityBotPerplexitySurfaces and links sites in Perplexity answers
Perplexity-UserPerplexityUser-initiated fetches; generally ignores robots.txt

These names come from each company's own documentation: OpenAI, Google, Anthropic and Perplexity. Google states that Google-Extended doesn't affect inclusion or ranking in Google Search.

Most businesses should allow the search-type bots and decide on training bots by preference. The bigger risk is a security plugin or CDN blocking every unfamiliar bot. Our guide to getting recommended by ChatGPT, Gemini and Perplexity covers the rest of AI visibility.

#CheckHow to testWhat good looks like
43Structured data is validRich Results TestNo errors on key page types
44Schema matches visible contentCompare markup to the pageSame name, address, hours and prices
45hreflang is reciprocal (if used)Crawler hreflang reportEvery version links to every other and back
46hreflang has x-default (if used)View sourceFallback version declared
47AI search crawlers not blocked by accidentRead robots.txt; check CDN bot settingsOAI-SearchBot, PerplexityBot and Claude-SearchBot allowed if you want visibility
48AI training crawlers set deliberatelyRead robots.txtGPTBot, Google-Extended and ClaudeBot rules reflect your choice

How to prioritise your fixes

Even a small site can surface dozens of issues. Score each one by impact and effort rather than fixing them in tool order.

IssueImpactEffortPriority
Site-wide noindex or Disallow: /SevereMinutesFix today
HTTP and HTTPS both liveHighLowFix this week
Key pages not indexedHighVariesFix this week
Broken internal links to service pagesHighLowFix this week
Redirect chains after a redesignMediumLowThis month
Duplicate or missing titlesMediumLowThis month
Poor Core Web Vitals on mobileMediumMedium to highPlan it in
Missing alt textLow to mediumLowBatch it
Missing schemaMediumMediumThis month
hreflang errorsOnly if multilingualMediumAs needed

Anything stopping Google indexing your money pages comes first, then duplicates and redirects, then on-page polish, then speed and schema.

A 60-minute audit with free tools

You need Search Console, PageSpeed Insights, the Rich Results Test and a crawler. Screaming Frog's free version crawls up to 500 URLs, plenty for most small sites.

Minutes 0–10: the basics

Open /robots.txt and /sitemap.xml. Type the http:// and non-preferred www versions of your domain and confirm they redirect.

Minutes 10–25: Search Console

Note each "not indexed" reason in the Page indexing report. Inspect your homepage and top three service pages, confirming the expected canonical. Check your sitemap shows "Success".

Minutes 25–40: the crawl

Crawl from your homepage. Sort by status code for 404s, 5xx errors and redirects, then check titles, H1s, alt text, noindex and canonicals.

Minutes 40–50: speed and mobile

Run PageSpeed Insights on mobile for your homepage and one key page. Note Core Web Vitals and the top suggestions, then try both pages on your phone.

Minutes 50–60: schema, AI and your fix list

Run the Rich Results Test on two key pages and re-read robots.txt for AI crawler rules. Then score every issue by impact and effort and pick the top five.

Growvia's SEO audit automates much of this list: it crawls up to 25 pages, checks titles, redirects, canonicals, schema, images, sitemap and robots.txt, and ranks issues by impact. It also shows PageSpeed and Search Console data alongside. If you're also chasing local customers, pair this audit with our local SEO checklist.

Frequently asked questions

How often should a small business run a technical SEO audit?

A full audit every three to six months suits most small sites, with a monthly glance at the Page indexing report. Always audit again after a redesign or domain move.

What's the difference between a technical SEO audit and a full SEO audit?

A technical audit checks whether search engines can crawl, render and index your site properly. A full SEO audit adds content quality, keyword targeting, backlinks and local signals like your Google Business Profile. Technical issues come first, because they can block every other effort.

Can I do a technical SEO audit without coding skills?

Yes. Free tools show most problems in plain language. Some fixes, like server speed, may need your developer, but finding the issues doesn't.

Why is Google showing a different title from the one I wrote?

Google may rewrite title links when your title tag is vague, outdated, stuffed with keywords or doesn't match the page's main heading. Make your title specific, accurate and consistent with your H1. That usually raises the chance Google keeps it.

Should I block AI crawlers like GPTBot in robots.txt?

It depends on your goal. Blocking training bots such as GPTBot or Google-Extended limits use of your content for model training. Blocking search-type bots such as OAI-SearchBot or PerplexityBot can remove you from AI answers. Most businesses that want customers should keep the search bots allowed.

Does every 404 error hurt my rankings?

No. A 404 for a page you deliberately removed is normal, and Google simply drops it from the index. The problems are broken links on your own site and removed pages that had backlinks or traffic. Redirect those to the closest relevant live page.

Put this into practice with Growvia

Audit your website, track Google rankings and AI mentions, send follow-ups and manage every lead and message in one place. Free forever plan, 14-day Pro trial, no card.

Start growing free

Related guides