"Crawl budget" is talked about a lot in SEO, usually in an alarmed tone. The reality is calmer: for most business websites it is not a problem. But when it is, it can explain why important pages take weeks to show up on Google.
Let's look at what it is exactly, who it affects and what to do to avoid wasting it.
What crawl budget is
It is the number of URLs that Googlebot (Google's crawler) can and wants to visit on your site within a given period. According to Google's documentation, it depends on two factors. The first is the crawl capacity limit: how many requests it can make without overloading your server; if the server answers slowly or returns 5xx errors, Google slows down. The second is crawl demand: how interested Google is in your URLs, based on their popularity and how often they change.
Do not confuse crawling with indexing. Crawling means Googlebot visits the page; indexing means it stores the page so it can show it in results. A page can be crawled and not indexed, and that almost never has anything to do with budget.
Which sites it really affects (and which it does not)
Google itself aims its crawl budget guide at large sites: those with more than a million unique pages whose content changes with some regularity, or more than 10,000 pages that change very quickly (daily). It also mentions sites with many URLs sitting in the "Discovered - currently not indexed" status in Search Console.
If you have a thirty-page company website, a blog or even a shop with a few hundred products, crawl budget is normally not your bottleneck. If one of your pages is not being indexed, look first at content quality, internal linking or a technical error. Where the problem does appear, even with mid-sized catalogues, is in online shops whose filters generate thousands of URL combinations.
What wastes it
Googlebot spends requests on any URL it finds, useful or not. The usual culprits are:
- Duplicate URLs: tracking and session parameters, http and https, with and without www.
- Faceted navigation and filters: size, colour and price combined can produce thousands of near-identical URLs.
- Infinite spaces: calendars with an endless "next month" link, or internal search results.
- Redirect chains (A leads to B, B to C, C to D). Googlebot follows up to 10 hops, but every hop is one more request.
- Errors: lots of internal links returning 404, "not found" pages that return a 200 code (soft 404s), or 5xx failures.
- A sitemap full of URLs that redirect, error out or are marked noindex.
How to improve it
In order of usual impact:
- Always link to the final URL. Internal links and the sitemap should point to an address that answers 200, with no redirects in between.
- Shorten redirect chains so each old URL jumps straight to its final destination.
- Control filters and duplicates. With robots.txt you can stop whole sections from being crawled; with noindex or a canonical you stop them being indexed, but Google has to crawl the page to see those instructions. And do not combine the two: if you block a URL in robots.txt, Google will never see its noindex.
- Improve server speed and stability. A fast server can handle more requests.
- Keep the sitemap clean, with indexable URLs only and reliable modification dates.
- Remove, or return 404/410 for, useless content you are not going to recover.
How to check where your crawling goes
In Search Console, under Settings, the Crawl stats report shows requests per day, average response time, and a breakdown by response code and file type. If you see a high share of 404s, redirects or odd parameters, that is your clue.
To dig deeper there are your server logs, which show which URLs Googlebot actually requests. If Google spends most of its visits on filter URLs you do not care about, you know where to act. A desktop SEO crawler also helps you see your site the way a robot sees it.
Frequently asked questions
Can I ask Google to crawl my site more?
There is no setting to raise the budget. What you can do is improve the two factors: a fast, stable server, content worth visiting and good internal linking. Requesting indexing in Search Console works for specific URLs, not for increasing crawling.
Does my new website have a small crawl budget?
It is normal for Google to take a while to discover and visit a new site, because it does not know it yet and no links point to it. That is not a budget problem but a discovery problem: submitting a sitemap and earning a few external links helps.
Does blocking pages in robots.txt free up budget for the rest?
It reduces requests to those URLs, but there is no guarantee that Google will reallocate that crawling to your other pages, unless you were hitting your server's capacity limit. Use it to avoid crawling what adds nothing, not as a trick to speed up everything else.