LaunchRanked

Technical SEO checklist for startups (2026)

Technical SEO checklist in the order Google works: crawling, indexing, rendering, structure, speed, markup, monitoring. A free checker for each step.

Published 10 min readBy Piyush Aaryan, founder

  • technical seo
  • checklist
  • indexing
  • seo

Short answer: Work down the list in order, because each step depends on the one before it. Google says a page is eligible for indexing when Googlebot isn't blocked, the page returns HTTP 200 and the page has indexable content. So first make sure crawlers can reach your pages, then that each page you want in search can be indexed, then that your content and links are in the HTML Google fetches. Speed, structured data and hreflang come after, and Search Console tells you when something breaks.

Each item says what to check, how to check it and what good looks like, with a free checker where we have one. For AI search (llms.txt, answer-first writing, citations), use our AI SEO launch checklist alongside this one.

One honest limit first. Meeting those requirements, Google says, "doesn't mean that a page will be indexed; indexing isn't guaranteed." Technical SEO removes reasons a page can't rank. It doesn't make a thin page good (see technical SEO).

The checklist at a glance

Check Good looks like Free tool
robots.txt Returns 200, blocks no page you want found, lists your sitemap robots.txt tester
AI crawlers and firewall Search crawlers allowed in robots.txt and by your CDN AI crawler checker
Status codes 200 on real pages, 301 for moved, 404 or 410 for gone Redirect checker
noindex Absent on every page you want indexed noindex checker
Canonical One absolute, self-referencing URL per page Canonical checker
Redirects One 301 hop from every variant to one HTTPS address Redirect checker
Sitemap Only canonical, indexable URLs, honest lastmod Sitemap checker
Rendering Title, H1, text and links are in the server HTML Heading checker
Links and structure Every page linked from another page with a real <a href> Internal link checker
Titles One distinct title and one H1 per page Title tag checker
Core Web Vitals LCP 2.5 s or less, INP under 200 ms, CLS under 0.1 PageSpeed Insights
Structured data Valid JSON-LD that matches the visible page Structured data validator
hreflang (only with translated versions) Every version lists itself and all the others hreflang checker
Search Console Page indexing and Sitemaps reviewed weekly Google index checker

1. Crawling: can Google reach your pages?

robots.txt

robots.txt says what crawlers may fetch. It isn't an indexing control. Google calls it "not a mechanism for keeping a web page out of Google": a blocked URL can still be indexed from links elsewhere, without a description.

  • Check: /robots.txt exists on every host you serve (www and non-www are separate), returns 200, and has no Disallow: / or rule that matches pages you want found.
  • How: open it in a browser, then test key URLs in the robots.txt tester, which shows the deciding rule.
  • Good looks like: a short file that blocks only crawl traps (admin, API, internal search), leaves the CSS and scripts pages need to render, points to your sitemap, and stays under the 500 KiB Google reads.

Errors matter more than rules. Google treats a 4xx (except 429) as if no robots.txt existed. On a 5xx it stops crawling for 12 hours, then uses its last good copy for up to 30 days.

AI crawlers and your firewall

  • Check: that the crawlers behind AI search (OAI-SearchBot for ChatGPT search, PerplexityBot for Perplexity) can fetch your pages, and that your CDN or firewall isn't turning bots away first.
  • How: the AI crawler checker reads your robots.txt and reports crawler by crawler. For the firewall, check your CDN's bot settings. OpenAI and Perplexity both publish IP ranges and recommend allowing them.
  • Good looks like: search crawlers allowed, and your stance on training crawlers chosen on purpose. OpenAI says sites opted out of OAI-SearchBot won't be shown in ChatGPT search answers, though they can still appear as navigational links.

The three bot groups and sample files are in our robots.txt, llms.txt and AI crawlers post.

2. Indexing: can each page be indexed?

Status codes

noindex

  • Check: no noindex in a robots meta tag or an X-Robots-Tag header on any page you want indexed. The classic cause is a staging setting that shipped to production.
  • How: the noindex checker reads the robots meta tag, the X-Robots-Tag header and robots.txt for up to 20 URLs. Or view source and search for "noindex." URL Inspection shows the HTML Googlebot received.
  • Good looks like: none on indexable pages, deliberate on thin or duplicate ones. Don't also block those pages in robots.txt. Google's noindex guide says a blocked crawler "will never see the noindex rule," and that noindex inside robots.txt isn't supported.

Canonical

  • Check: each indexable page declares one canonical URL: absolute, in the <head>, pointing at itself, and matching what your sitemap and internal links use.
  • How: the canonical checker reads the HTML tag and HTTP header and flags http/https, www and trailing-slash conflicts. In URL Inspection, compare "User-declared canonical" with "Google-selected canonical."
  • Good looks like: the two match. Google's canonicalization guide calls redirects and rel="canonical" strong signals and sitemap inclusion a weak one, and says not to give one page different canonicals through different methods.

Redirects and one address per page

  • Check: http, www, non-www and trailing-slash variants all end at one HTTPS address, in one hop, and moved pages use a 301.
  • How: run each variant through the redirect checker.
  • Good looks like: a single permanent redirect. Google follows up to 10 hops, but its crawl budget guide says long chains have "a negative effect on crawling." Use a 302 only for temporary moves: per Google, it keeps the source page in results, while permanent redirects show the target.

Sitemap

  • Check: the sitemap lists only canonical, indexable URLs (200, no noindex, no redirects) as absolute URLs, at most 50,000 URLs and 50 MB uncompressed per file.
  • How: the sitemap checker finds the file, follows an index and checks lastmod and limits. Submit it in Search Console and add a Sitemap: line to robots.txt.
  • Good looks like: Search Console says Success and the discovered count roughly matches what you expect. Google's sitemap guide says it ignores priority and changefreq, uses lastmod only if it's "consistently and verifiably" accurate, and treats a submission as "merely a hint."

3. Rendering: is your content in the HTML?

  • Check: the title, H1, main text, navigation links, canonical and robots meta are in the HTML your server returns, not added later by JavaScript.
  • How: use View source, not Inspect, which shows the page after scripts run. The heading checker reads the HTML as fetched, so an empty outline on a page that looks fine in your browser points to client-side rendering. For Google's view, run URL Inspection's live test and open the tested page for rendered HTML and a screenshot.
  • Good looks like: server-rendered or pre-rendered pages, so nothing important depends on a script.

Google is relaxed about JavaScript. Its guide says Google Search runs JavaScript with an evergreen Chromium, yet "server-side or pre-rendering is still a great idea because it makes your website faster for users and crawlers, and not all bots can run JavaScript." Two details bite single-page apps:

  • Use real <a href> links, not #/ fragment routes. Google can only discover links that are <a> elements with an href.
  • Return a real 404 for missing pages, or add noindex to the error view, to avoid soft 404s.
  • Check: every page you care about has a link from at least one other page, as an <a> element with an href, and URLs are readable and consistent.
  • How: the internal link checker lists one page's links and anchors, or scans up to 50 pages for inbound link counts and orphan candidates.
  • Good looks like: Google's link guide says "every page you care about should have a link from at least one other page on your site." Anchors describe the target. URLs use hyphens, one case (Google treats /APPLE and /apple as different pages) and few parameters, with no endless filter or calendar combinations, which its URL guide says create "unnecessarily high numbers of URLs."

Our internal linking system turns this into four rules.

5. Titles and headings

  • Check: one <title> and one H1 per page, distinct across pages, describing that page.
  • How: the title tag checker measures titles in characters and pixels and flags duplicates and H1 mismatches. The heading checker shows the H1 outline.
  • Good looks like: Google's title link guide says to avoid "repeated or boilerplate text," and that there's no length limit, only truncation to fit the screen. Make the first visible <h1> the page's main title, since Google reads headings as a source too. How we write them is in title tags that get clicked.

6. Page experience: Core Web Vitals, mobile and HTTPS

Google's targets are LCP within 2.5 seconds, INP under 200 milliseconds and CLS under 0.1. web.dev says to judge them at the 75th percentile of page loads, split by mobile and desktop. PageSpeed Insights rates LCP over 4 seconds, INP over 500 ms and CLS over 0.25 as poor.

  • Check: your key pages on mobile, with real-user (field) data first and lab data second.
  • How: run each page through PageSpeed Insights. It shows 28 days of Chrome field data when there is enough. A new or low-traffic page may have none, and PSI then falls back to site-wide data or shows none. Use lab results for debugging only.
  • Good looks like: all three metrics passing at the 75th percentile.

Keep it in proportion. Google says there is no single page experience signal and that "trying to get a perfect score just for SEO reasons may not be the best use of your time."

Two basics sit next to speed. Google indexes the mobile version of your pages, so mobile needs the same content, headings, structured data and robots meta tags as desktop. Google also prefers the HTTPS version of a page; the HTTPS report confirms it.

7. Structured data

  • Check: JSON-LD that describes what's visible on the page, using the types that fit (Organization on the homepage, Article on posts, BreadcrumbList on inner pages).
  • How: the structured data validator pulls every JSON-LD block, flags broken JSON and checks the properties Google lists as required. Google's Rich Results Test is the second opinion. After deploys, watch the rich result reports in Search Console, because template changes can break markup.
  • Good looks like: valid markup that matches the page. Google's guidelines say not to mark up content visitors can't see, and that even valid markup isn't guaranteed to produce a rich result.

One 2026 change: Google's documentation updates say FAQ rich results stopped appearing on May 7, 2026. Don't add FAQ markup expecting a different-looking result.

8. International: hreflang, only if you need it

Skip this on a single-language site.

  • Check: each version lists itself and every other version, with full URLs, and the pages point back at each other.
  • How: the hreflang checker validates the tags and fetches the alternates to confirm they link back.
  • Good looks like: Google's localized versions guide says "if two pages don't both point to each other, the tags will be ignored," and that Google doesn't use hreflang or the lang attribute to detect a page's language.

9. Monitoring: what to watch in Search Console

Let Search Console tell you what changed.

  • Page indexing: watch for new reason types and for "Crawled - currently not indexed" or "Discovered - currently not indexed" growing. Google's help says "don't expect every URL on your site to be indexed": key pages should be, and the rest should be out for reasons you chose. After a fix, Validate fix "typically takes up to about two weeks."
  • Sitemaps: status should stay Success.
  • URL Inspection: after a deploy, inspect your homepage, pricing page and top landing page.
  • Skip Crawl Stats on a small site: Google says under a thousand pages you shouldn't need it.

For a quick look at one URL, use the Google index checker. It runs a live site: search, which is an estimate, so confirm important pages in Search Console. The two statuses founders hit most have guides: fix Discovered - currently not indexed and fix Crawled - currently not indexed. The other reports are in Search Console for founders.

What we did on Versely

On Versely, our own product, we built this base before scaling: a canonical helper, JSON-LD, a sitemap index with honest lastmod, a generated llms.txt and AI-crawler tracking. It didn't stop Google throttling indexing when we then published about 2,500 posts in a month. Technical SEO makes pages indexable, not instantly indexed.

If you work with Claude, Cursor or Codex, LaunchRanked's MCP server can run sitemap, structured data, link and AI-search checks on your live site, as we did in We pointed Claude Code at our site. After the base, content and links decide what ranks: see the 90-day plan for a new domain, or launch free on LaunchRanked for a permanent page and a followed link after review.

Frequently asked questions

What should I fix first in technical SEO?

Anything that stops a page being indexed: Googlebot blocked, a non-200 status, or no indexable content. Check robots.txt, status codes and stray noindex tags before you look at speed or structured data.

Do I need to worry about crawl budget on a small site?

Usually not. Google's guide is for sites with 1 million or more pages changing weekly, 10,000 or more changing daily, or many URLs in Discovered - currently not indexed. For everyone else, it says an up-to-date sitemap and a regular look at the Page Indexing report are adequate.

How much do Core Web Vitals matter for rankings?

Google says its ranking systems use them and recommends good scores, but there is no single page experience signal, passing guarantees no top result, and chasing a perfect score just for SEO may not be the best use of your time.

Tools and pages for this topic

Free tools, templates, lists and glossary entries on the same subject.