Crawl budget is one of those SEO terms that gets more attention than it deserves on small sites and far too little on large ones. A 50-page local business site will not see a difference if it “optimises crawl budget”. A store with two million product and filter URLs, on the other hand, can have new products waiting weeks to be discovered because Googlebot is busy crawling parameter combinations nobody searches for.
This guide explains what crawl budget actually is according to Google, how to tell whether it matters for your site, how to diagnose crawl waste with Search Console and log files, and the practical fixes that free up crawling for the pages you care about.
Key Takeaways
- Crawl budget is the time and resources Google allocates to crawling a site, shaped by crawl capacity limit and crawl demand.
- It mainly matters for large or fast-changing sites; most small sites can ignore it.
- Faster, more reliable servers raise the capacity limit; popular and fresh content raises demand.
- The biggest crawl waste comes from faceted navigation, parameters, duplicates, soft errors and redirect chains.
- Use robots.txt to stop crawling of worthless URLs; noindex does not save crawl budget.
What Is Crawl Budget?
The web is effectively infinite, so Google cannot crawl every URL it knows about as often as it would like. It has to decide how much time and how many requests to spend on each site.
“A site's crawl budget is determined by two main elements: crawl capacity limit and crawl demand.”
— Google, Large site owner's guide to managing your crawl budget, Crawl budget management
Crawl capacity limit
This is how many simultaneous connections Googlebot can use, and how long it waits between fetches, without overloading your server. Google's guide explains that if your site responds consistently and response times remain stable or improve, the limit goes up. If the site slows down or returns server errors or rate-limiting signals such as 429, the limit goes down and Google crawls less.
Crawl demand
This is how much Google wants to crawl your site. The guide lists factors including the URLs Google perceives as worth crawling, popularity (more popular URLs tend to be crawled more often) and staleness (Google wants to recrawl documents often enough to pick up changes). Large events such as site moves can also increase demand temporarily.
Does Crawl Budget Matter for Your Site?
Google's guide is explicit about its audience. It is written for:
- Large sites with 1 million+ unique pages whose content changes moderately often (weekly).
- Medium or larger sites with 10,000+ unique pages whose content changes very rapidly (daily).
- Sites with a large share of URLs classified in Search Console as “Discovered – currently not indexed”.
The guide also says that if your site does not have a large number of rapidly changing pages, or your pages seem to be crawled the same day they are published, you don't need to read it. For most small business sites, indexing problems are about quality and duplication, not crawl budget. Our guide to “Crawled – currently not indexed” covers those cases.
| Site profile | Crawl budget priority | Focus instead on |
|---|---|---|
| Local business site, under 500 pages | Very low | Content quality, internal links, local signals |
| Blog with a few thousand posts | Low | Pruning thin content, sitemaps, internal linking |
| Ecommerce store with faceted navigation | High | Parameter control, duplicates, server speed |
| Marketplace, classifieds or news site with daily changes | High | Freshness signals, sitemaps with lastmod, removals |
| Enterprise site with 1M+ URLs | Very high | Log analysis, URL inventory, infrastructure |
What Wastes Crawl Budget
- Faceted navigation and filters generating near-endless combinations such as
?color=red&size=9&sort=price. - Session IDs and tracking parameters that create unique URLs for identical content.
- Duplicate content across protocol, host, trailing-slash and case variants. See our guide to duplicate content.
- Soft 404s that return 200 for empty or error pages; see how to fix 404 errors.
- Infinite spaces such as calendars with “next month” links forever.
- Redirect chains, where each hop is an extra request.
- Slow server responses, which lower the capacity limit.
- Hacked or spam pages injected into the site.
How to Diagnose Crawl Budget Issues: Step by Step
- Open the Crawl Stats report in Search Console (Settings → Crawl stats). Review total crawl requests, average response time, and breakdowns by response code, file type and Googlebot type.
- Check host status in the same report for robots.txt fetch, DNS or server connectivity problems.
- Review the Page indexing report for large counts of “Discovered – currently not indexed”, duplicates and soft 404s.
- Analyse server logs. Filter for verified Googlebot requests and group by URL pattern. Work out what share of crawling goes to parameters, redirects or errors versus your key templates.
- Crawl the site yourself with Screaming Frog, Sitebulb or Semrush Site Audit to map parameter URLs, chains and orphan pages.
- Compare crawled URLs against your list of pages you want indexed. The gap is your crawl waste.
How to Optimise Crawl Budget
Google's guide groups its advice around managing your URL inventory, keeping the site efficient and monitoring crawling. Here is how that translates into actions:
| Problem | Fix | Why it works |
|---|---|---|
| Worthless parameter and filter URLs | Disallow patterns in robots.txt; avoid linking to them | Stops Googlebot requesting URLs you never want crawled |
| Duplicate URL variants | 301 redirects and rel=canonical; consistent internal links | Helps Google consolidate and crawl duplicates less often |
| Removed pages returning 200 | Return 404 or 410 | Google crawls 4xx URLs less over time |
| Redirect chains | Redirect straight to the final URL | Removes extra hops |
| Slow responses | Caching, CDN, faster TTFB, efficient databases | Raises the crawl capacity limit |
| New or updated pages discovered slowly | Accurate XML sitemaps with lastmod; strong internal links | Raises discovery and recrawl demand |
| Heavy pages and resources | Smaller pages, fewer render-blocking resources | Fewer bytes and requests per page |
Two examples of robots.txt rules for faceted navigation:
User-agent: *
# Block sort and session parameters anywhere in the URL
Disallow: /*?*sort=
Disallow: /*?*sessionid=
Sitemap: https://www.example.com/sitemap_index.xml
Test any pattern carefully before deploying; an overly broad rule can block important pages. Our guide to robots.txt vs meta robots explains which tool to use when, and our XML sitemap guide shows how to keep sitemaps lean.
Server speed matters too. Improvements to time to first byte help users as well as crawlers; see our guide on how to improve page speed, or let our page speed optimization team handle it.
What Doesn't Help (and Common Mistakes)
“Don't use noindex, as Google will still request, but then drop the page.”
— Google, Large site owner's guide to managing your crawl budget, Crawl budget management
- Using noindex to save crawl budget. Google still requests the page.
- Toggling robots.txt to shift budget. Google advises against using robots.txt to temporarily reallocate crawl budget to other pages; use it only for URLs you never want crawled.
- Blocking pages you want de-indexed. If a URL is blocked, Google cannot see a noindex on it, and it may still be indexed without content.
- Relying on nofollow internal links. They are not a dependable way to control crawling; remove the links or block the URLs instead.
- Returning 5xx or 429 to slow Googlebot for long periods. This is an emergency measure only; prolonged errors can lead to URLs being dropped.
- Expecting crawl rate to boost rankings. More crawling does not mean better rankings; it only ensures your content can be indexed and refreshed.
Related Guides
- Mobile SEO: How to Optimize for Mobile-First Indexing
- Schema Markup: A Practical Guide to Structured Data
- Image SEO: How to Optimize Images for Search
Frequently Asked Questions
Does my small website need to worry about crawl budget?
Almost certainly not. Google's guide is aimed at very large sites (1 million+ pages changing weekly) or medium-to-large sites (10,000+ pages changing daily), plus sites with many URLs stuck in 'Discovered – currently not indexed'.
Is crawl budget a ranking factor?
Being crawled more often does not by itself improve rankings. Crawl budget matters because pages that are not crawled cannot be indexed or updated in search results.
Does noindex save crawl budget?
No. Google still has to request the page to see the noindex tag, then drops it. Use robots.txt for URLs you never want crawled, and 404/410 for removed pages.
Can I ask Google to crawl my site more?
Not directly. Google increases crawling when your server responds quickly and reliably and when your content is popular and frequently updated. You can, however, temporarily reduce crawling in emergencies.
Where can I see how Google crawls my site?
In Search Console under Settings, the Crawl Stats report shows total requests, download size, response times, status codes and file types. Server logs give an even more detailed view.
Do redirects waste crawl budget?
Each redirect hop is an extra request. Google's guide advises avoiding long redirect chains because they have a negative effect on crawling.
Conclusion
Crawl budget is a real constraint, but only for sites big or dynamic enough to hit it. If that is you, the playbook is straightforward: shrink your URL inventory to pages worth crawling, make your server fast and reliable, keep sitemaps and internal links pointing at canonical URLs, and monitor Crawl Stats and logs. For everyone else, effort is better spent on content quality and a regular SEO audit, as our technical SEO guide explains.
Running a large ecommerce or enterprise site? Our enterprise SEO team specialises in log file analysis and crawl optimisation at scale. Get a free quote to find out where your crawl budget is going.
References
- Google Crawling Infrastructure: Large site owner's guide to managing your crawl budget
- Google Search Central: How HTTP status codes and network errors affect Google Search
- Google Search Central: Site moves with URL changes
- Google Search Central: Introduction to robots.txt
- Google Search Central: Build and submit a sitemap
- Semrush: Crawl Budget — What Is It and Why Is It Important for SEO?



