Indexing
Indexing is the process by which Google crawls, analyses and stores your site's pages in its database so they can appear in search results.
Why it matters
Without indexing, your site is invisible on Google. It is the precondition for any kind of search visibility. Silent indexing problems can make entire pages disappear from Google without you noticing, costing thousands of lost visits.
What is indexing?
Indexing is the process by which search engines (Google, Bing and the rest) discover, analyse and store your site's pages in their index, a vast database of billions of web pages. Until a page is indexed it simply does not exist as far as Google is concerned, and it cannot appear in search results.
The process runs in three steps: crawling (exploration), indexing (analysis and storage), and ranking (ordering). Googlebot, Google's crawler, works its way across the web by following links from page to page. When it finds a new page it analyses it (content, tags, structure) and decides whether or not to add it to the index. An indexed page is then ranked according to how relevant it is to each query.
How Google's crawl works
Googlebot has a limited crawl budget for each site. That budget depends on the size, the popularity and the technical health of your site. For a small brochure site (10 to 50 pages), crawl budget is never a problem. For an online shop with 100 000 pages or more, optimising crawl budget is critical.
The robots.txt file (at the root of your site) tells Googlebot which pages it may and may not crawl. The XML sitemap (sitemap.xml) lists all the pages you want indexed, with their last modified date and their priority. These two files are fundamental to guiding the crawl.
Internal linking (links between your pages) helps Googlebot find every page on your site. An orphan page (with no internal link pointing to it) risks never being crawled at all. Click depth (the number of clicks from the homepage) also affects crawling: pages more than 3 clicks deep get crawled less often.
Why is a page not indexed?
Several things can stop a page being indexed. A noindex tag in the HTML header or the meta tags explicitly tells Google not to index the page. The robots.txt file may be blocking the crawl. The page may be returning a 404 error or a 500 error. The content may be too thin or duplicated from another page.
Content quality is a factor too. Google actively decides not to index pages it considers low value: copied content, pages of no interest to the user, boilerplate pages (identical terms and conditions everywhere). That is the "crawled but not indexed" status you see in Google Search Console.
Making sense of the indexing statuses in Search Console
The indexing report in Google Search Console shows statuses that often leave site owners puzzled. Decoding them is what lets you act precisely:
- Crawled, currently not indexed: Google has seen the page but judged it not useful enough to index. This usually signals content that is too thin or too close to another page. - Discovered, currently not indexed: Google knows the URL exists but has not crawled it yet, generally for lack of crawl budget or authority. - Duplicate without user-selected canonical: several versions of the same page exist and Google has picked one other than yours. - Blocked by robots.txt: you are unintentionally preventing the crawl. - Excluded by noindex tag: a directive is explicitly asking not to index.
Working through these statuses one by one turns an intimidating report into a concrete to-do list.
Google Search Console: your indexing tool
Google Search Console (GSC) is Google's free tool for monitoring and managing your site's indexing. The Indexing > Pages report shows how many pages are indexed, which ones are not, and why. It is the single most important diagnostic tool in technical SEO.
The URL Inspection tool lets you check whether a specific page is indexed and request manual indexing (a recrawl). Useful after publishing a new page or fixing a problem. Indexing after a manual request usually takes 24 to 72 hours.
Improving your site's indexing
Submit your XML sitemap in Google Search Console. Build solid internal linking: every page should be reachable within 3 clicks of the homepage. Fix the crawl errors (404, 500) that GSC reports. Avoid duplicate content by using canonical tags. Make sure your site is fast (Google crawls more pages on fast sites).
On large sites, optimise the crawl budget: block pages with no SEO value (filter pages, sort pages, pagination pages) using robots.txt or the noindex tag, and keep the crawl budget for the pages that matter.
Mobile-first indexing
Since 2019, Google has used mobile-first indexing: it is the mobile version of your page that gets crawled and indexed first. If your mobile version has less content than the desktop version, the stripped-back version is the one that ends up in the index. Make sure the content is identical on mobile and desktop.
Indexing and new sites: patience and signals
A common trap for young businesses: believing a freshly launched site appears on Google instantly. In reality, a new domain with no authority can take several weeks to get its pages indexed, because Googlebot visits it rarely at first. Several signals help speed things up: submit the sitemap as soon as you launch, pick up a few quality external links (a Google Business Profile listing, trade directories, partners), and publish fresh content regularly so Google has a reason to come back more often.
A local small business is also better off watching the indexing of its key pages (services, contact, town pages) than worrying about the total number of indexed URLs. Ten useful, well-indexed pages beat a hundred weak ones that dilute the site's authority.
At ConvertiLab, every site we build is set up for indexing: automatic XML sitemap, configured robots.txt, structured internal linking, and monitoring through Google Search Console so indexing problems get spotted and fixed before they cost you traffic.
Practical examples
An online shop discovers in Search Console that 60% of its product pages are not indexed because of duplicate content. After rewriting the descriptions, 85% of the pages are indexed within 6 weeks.
A blog submits an XML sitemap to Google Search Console and adds internal links: the average time to index a new article drops from 2 weeks to 48 hours.
A brochure site was unintentionally blocking its service pages in robots.txt. Fixing it allows indexing and the site goes from 0 to 500 organic visits a month in 3 months.
Frequently asked questions
How do I know whether my pages are indexed by Google?
Type 'site:yourdomain.com' into Google to see all the indexed pages. For a precise check, use Google Search Console > Indexing > Pages, which lists every URL with its indexing status and any errors.
How long does Google take to index a new page?
From a few hours to several weeks, depending on your site's authority and how often it is crawled. To speed things up, submit the URL in Google Search Console with the URL Inspection tool. A site with strong authority gets crawled daily, a new site far less often.
What is an XML sitemap and do I need one?
An XML sitemap is a file listing all the pages on your site that you want indexed. It helps Google find every page, especially the deeper ones. And yes, every site needs one. Next.js and most CMS platforms generate the sitemap automatically.
Need help with Indexing?
Our experts work with you to put an effective strategy in place. Get a free, tailored quote within 24 hours.
Go further
Other definitions
Last updated: 6 April 2026


