📘 Guide General information

Technical website audit using five search operators

Five queries in the search bar that find test copies of a site in the index, junk URLs with parameters, an error page in search results, forgotten exports, and an open file listing on the server.

You can find half of the technical problems with your website’s indexing in five minutes using an ordinary search bar, without crawlers or paid services. Below are five queries that show what the search engine knows about the site beyond what you showed it: forgotten test copies, junk URLs with parameters, an error page in search results, publicly accessible exports and logs, and, most unpleasantly, an open file listing on the server.

A test copy of the site that got indexed

After a migration or redesign, a development or test copy often remains accessible and gets indexed alongside the live site. These are duplicates on the scale of the entire site.

Query to check:

site:*.вашдомен.ru -www

The first part asks the search engine to show everything it knows about the subdomains; the second removes the main version from the results. What remains is something you may have forgotten. In the example examined, this query immediately pulled up a test copy of the website of a company that sells plugins: its pages contained the meta tag index, follow, while the canonical URL pointed to the test copy itself rather than the live domain. To the search engine, it was simply a second independent site with the same content.

If you have subdomains that belong in the index, exclude them with a chain of minus signs:

site:*.вашдомен.ru -www -blog -support -help

An empty result means there is nothing extra.

URLs with parameters

Filters, sorting, and pagination generate URLs that differ by the part after the question mark. When thousands of such URLs get indexed, you end up with thin duplicates.

site:вашдомен.ru inurl:?

This is usually how internal search pages and endless combinations of catalog filters appear. The correct behavior can be seen on large stores: a page with a sorting parameter opens normally, but its code contains a canonical URL pointing to the clean version without the parameter. The user gets a convenient filter, while the search engine gets one URL instead of a hundred.

One qualification should be made here, because videos about technical audits scare everyone with crawl budget. Crawl budget becomes a real problem on large sites: Google’s own guideline is at least ten thousand pages updated daily or at least one million pages updated roughly once a week. For a blog with a couple hundred articles or a corporate site with a thousand pages, crawl budget is almost never the bottleneck, while duplicates in the index are harmful regardless of size. So parameters are worth fixing, but because of duplicates, not crawl budget.

An error page in search results

The page a visitor sees at a nonexistent URL should not appear in search results. Check it like this:

site:вашдомен.ru intitle:"страница не найдена"

Replace the text in quotation marks with the title of your own error page—“404,” “Nothing found,” or whatever you use. In practice, such a page has been found in the index even on a bank’s website.

The advice usually given next is to add noindex. This treats the symptom. If the error page has entered the index at all, it almost certainly returns response code 200 instead of 404, meaning that to the search engine it is an ordinary existing page. The response code needs to be fixed: 404 for “not found” or 410 for “permanently removed.” With the correct response code, noindex is unnecessary—Google representatives have repeated this for years. The reverse is also true: if the response code cannot be fixed, then noindex remains the only option.

Files no one intended to publish

Files that were never meant to be searchable end up in the media library and on the server for years, and the search engine indexes them just like pages.

site:вашдомен.ru filetype:csv
site:вашдомен.ru filetype:xls
site:вашдомен.ru filetype:log
site:вашдомен.ru filetype:sql

What turns up in practice: a CSV containing product inventory was sitting in the index of a store selling construction kits and was available for download. If it is an export for customers, there is no issue; if it is an internal document, that is already a problem. On another site, indexed log files were found that redirected to the “About Us” page when clicked—in other words, they definitely were not meant to be published. Database dumps are the worst: they should not enter the index, either for security reasons or because of crawl considerations.

An open file listing on the server

This point deserves a separate section because it is no longer about SEO.

site:вашдомен.ru intitle:"index of"

When a server directory has no index file (index.html or index.php), the web server’s default configuration displays a raw list of its contents instead of a page: file and folder names, sizes, and dates. Any visitor can browse this list and download anything—backups, configuration files, database exports, and internal documents. The search engine indexes this list too.

An empty result for this query is good news. If there are results, do not put it off: disable directory listing in the server configuration and determine exactly what was publicly accessible and for how long.

The same thing in Yandex

Yandex has its own syntax, and there is no exact equivalent:

TaskGoogleYandex
Everything on the site, including subdomainssite:*.домен.rusite:домен.ru
One host onlysite:www.домен.rurhost:ru.домен.www
A word in the page URLinurl:url:
A word in the titleintitle:title:
Files of a specific typefiletype:mime:

Operator site: in Yandex includes subdomains by default, so an asterisk is unnecessary. With mime: the list of formats is limited to documents: pdf, doc, xls, ppt, rtf, and open equivalents. You cannot search for a database dump or log this way; Google is the remaining option.

What to do with the findings

What was foundWhat to do
A test copy of the site in the indexProtect it with authentication, return 404, or at least set the canonical to the live domain and remove it from the index
URLs with parametersUse the canonical URL without the parameter; block service search pages from indexing
An error page in search resultsChange the response code to 404 or 410; noindex— only if the response code cannot be changed
Exports, logs, and dumpsRemove them from the web directory rather than merely blocking them from indexing: a file available via a direct link remains accessible
An open file listingDisable directory listing on the server and check what has already leaked

One final note about the order of operations. Hiding a finding from the search engine and fixing the underlying cause are different things. A file that disappears from the index has not disappeared from the server and is still served to anyone who knows the URL. Therefore, always start with the server, and clean up the search results afterward.

search operators technical website audit site: operator website indexing Google dorks Yandex operators index of

SEO Mind42 editorial team

We explore SEO and neural networks in practice: test services on our own projects, verify prices and limits against primary sources, and share things you can put to use the same day.

📚 Reference guide to SEO and AI 🔄 Materials are updated 🕐 Updated: 3 October 2026

Related reading

All in this section →