Technical SEO starts in the server response because a crawler decides what to do with a URL before it reads a single word of content. The HTTP status code tells it whether the page exists, moved or failed. Redirects tell it where the authority should go. The canonical, the robots directives and the HTML in the first response tell it what to index. Response time affects how much of the site gets crawled at all. If those signals are wrong, no amount of on-page optimisation fixes it, and AI crawlers that read only the raw HTML and never run JavaScript are even less forgiving.
What does a crawler see first?
When Googlebot or any other bot requests a URL, it gets three things before anything else: a status code, a set of headers and a body. Most SEO work focuses on the body. Most of the problems I find in audits are in the first two.
If the status code says the page is gone, the crawler never sees your design, your sliders or your carefully written introduction. If the headers and the HTML disagree about the canonical, it may index the wrong version of the page. So read the response the way a bot reads it: with curl -I, the network panel, or a crawler like Screaming Frog. A browser hides all of this behind the rendered page.
Status codes are instructions
Each status code tells search engines what to do with a URL. Send the wrong one and the crawler follows the wrong instruction.
| Code | What it tells a crawler | Common WordPress mistake |
|---|---|---|
| 200 | This page exists; consider indexing it | Returning 200 for “not found” pages (soft 404) |
| 301 / 308 | Moved permanently; transfer signals to the new URL | Using 302 for permanent changes |
| 302 / 307 | Temporary move; keep the old URL in mind | Leaving temporary redirects in place for years |
| 404 / 410 | Gone; drop it from the index over time | Redirecting every 404 to the homepage |
| 503 | Temporarily unavailable; come back later | Maintenance mode that returns 200 with a “be right back” page |
Two of these need a closer look. The first is soft 404s: an empty search result, a deleted product page or a category with no posts that returns 200 with almost no content. Google will often classify them as soft 404s anyway, but by then you have wasted crawl and created noise. The second is maintenance. A 503 with a Retry-After header tells crawlers the outage is temporary. A 200 maintenance page tells them your whole site has been replaced by one sentence.
Redirects: chains, loops and where authority goes
A permanent redirect tells search engines that a URL has a new home. Done cleanly, in one hop from old to new, the signals consolidate on the destination. Done badly, you get chains.
Chains pile up on WordPress sites without anyone noticing. HTTP goes to HTTPS, then non-www to www, then the old slug to the new slug, then to the trailing-slash version. That is four hops for one URL. Googlebot follows up to ten, but every hop costs a request, adds latency for users and raises the chance that something breaks along the way.
- Redirect directly to the final URL, in one hop.
- Use 301 or 308 for permanent changes, not 302.
- Update internal links to point at the destination, so you are not relying on the redirect at all.
- Do redirects at the server or edge where possible, before WordPress boots, rather than through a plugin that loads the whole stack first.
Canonicals and robots signals must agree
The canonical tag is a hint, not a command. Google weighs it against other signals: redirects, internal links, sitemaps, hreflang. When those disagree, Google picks its own canonical, and it may not be the one you wanted.
In practice this comes down to consistency. The URL in the canonical should return 200, not redirect. It should be the URL you link to internally and list in the sitemap, and it should not be blocked by robots.txt or marked noindex. You can also send a canonical as a Link HTTP header, which helps with PDFs and other non-HTML files. For the same reason, noindex can go out as an X-Robots-Tag header.
Search engines do not index what you meant. They index what the server said.
TTFB is an SEO problem too
Time to first byte is the time between the request and the first byte of the response. The web.dev guidance treats 0.8 seconds or less at the 75th percentile as a good TTFB. It is not a Core Web Vital itself, but it sits under all of them: a slow first byte delays everything that follows, LCP included.
It affects crawling as well. Google adjusts how much it crawls based on how the server responds. A host that answers quickly and consistently can be crawled more, and one that slows down or throws 5xx errors gets crawled less. On a small site this rarely matters. On a large WooCommerce catalogue or a news site that publishes daily, it decides how fast new and updated pages get discovered.
On WordPress the fixes are the same ones you find in any performance audit: a working page cache, a persistent object cache for uncacheable requests, a cleaned-up options table, and removing plugins that do heavy work on every request.
Server-rendered HTML for search engines and AI crawlers
Google can render JavaScript, but it does it in a second stage that may run later than the initial crawl. Content that only shows up after scripts run gets indexed later, and sometimes not the way you expected.
AI crawlers are stricter. The bots that collect content for AI assistants and answer engines, such as GPTBot, ClaudeBot and PerplexityBot, generally fetch the HTML and do not execute JavaScript. If your main content, headings, prices or FAQs are injected client-side, those systems do not see them, and they cannot cite something that was never in the response.
WordPress starts with an advantage here, because by default it renders HTML on the server. The risk comes from what gets added on top: page builders that load content through AJAX, tabs and accordions filled by scripts, or headless front ends that ship an empty shell. A quick test I run: fetch the URL with curl and search the output for your main heading and a key paragraph. If they are not there, crawlers that do not render JavaScript will never find them.
So get the response right first: correct status, one clean redirect, a consistent canonical, a fast first byte and real content in the HTML. The rest of your SEO work builds on that.