For twenty years, SEO was built on one assumption: search evaluates a page and compares it with other similar pages. Hence rewrites, increasing content volume, and covering keywords. Research on selecting data subsets describes a different logic: the system assembles a set of sources, with each new item adding something the previous ones don’t. The difference is fundamental. Here’s what happens to generic content under this kind of selection, and how to restructure pages when your competitor is an aggregator with a big budget.
GIST: what it really is—and what it isn’t
GIST stands for Greedy Independent Set Thresholding. The algorithm was developed by Google Research researchers Morteza Zadimoghaddam and Matthew Farbach. The preprint has been on arXiv under number 2405.18754 since May 2024, and the paper was presented at NeurIPS 2025. Google Research published a detailed analysis on 23 January 2026.
The algorithm’s task is to select a compact subset from a large dataset that is both useful and free of near-duplicate items. The publication explicitly states its application: thinning training datasets, where thousands of examples must be retained from millions without reducing model quality.
This is where retellings usually distort the source. GIST guarantees that the result is no worse than half the optimal solution. It’s also proven that guaranteeing more than 0,56 of the optimum is NP-hard. Both figures refer to the proportion of the ideal set, not to “content usefulness of 56 percent.” The paper’s acknowledgments mention Google DeepMind and YouTube, and the YouTube Home team used a related principle of maximizing the minimum distance between recommendations: after someone watches a video, they’re shown something other than ten nearly identical videos.
What the publication doesn’t contain: evidence that this specific algorithm is used in web search results. There’s enough evidence that search and generative answers select sources with an eye to how different they are from one another, but it would be inaccurate to present GIST as an active ranking factor. Its practical value lies elsewhere. The logic of selecting a set is described mathematically, so you can use it as a working model and test it on your own pages.
Why it makes economic sense for search to cut duplicates
Crawling, rendering, and storing pages in the index cost money at every stage, and duplicates are a pure loss in that equation. This drives a shift toward more stringent selection of what gets processed in the first place.
On 3 February 2026, Google rewrote its documentation on crawlers and specified a file size limit for HTML and text formats of 2 MB, up from 15 MB. The limit for PDFs is 64 MB; for crawlers without a separately specified value, the default remains 15 MB. The company framed this as a documentation clarification, not a change in behavior, but the practical consequence is the same: inline images, embedded scripts, and bloated markup eat into the limit, after which reading cuts off without warning—and search simply won’t see the text below the cutoff.
The second cost layer is rendering. Parsing clean HTML is cheap, while running a browser engine to render a page is expensive, so not everything makes it to the second pass. The third layer concerns generative answers: the model has a limited context window, and if five selected sources all say the same thing, that budget is wasted instead of covering different aspects of the question.
Practitioners also point to demotion at the pattern level: a crawler parses some of a batch of similar cards, finds no value in them, and extrapolates that conclusion to the rest of the sample. It’s difficult to verify this from the outside, but the behavior is consistent with the logic of saving resources.
Rewrites are over—and here’s why
Text uniqueness checks compare letters, words, and phrases. Meaning-based selection compares vectors: the text becomes a point in multidimensional space, and similarity is measured as the distance between points. “Mom washed the frame” and “the mother cleaned the frame” will be neighbors in this space, even though none of the words match.
This leads to several things that are bad news for traditional content production:
- Replacing words with synonyms and rearranging paragraphs don’t create new meaning. A uniqueness score of 95% and a semantic duplicate can easily coexist in the same text.
- Translating someone else’s article from another language produces roughly the same set of vectors as the original.
- Length is no longer an argument. A page of 5000 words with 200 worth of unique meaning loses to five hundred words where every paragraph adds something.
- A well-written text that repeats the top-ranking pages’ content is still replaceable. Good wording benefits readers and engagement metrics, but it doesn’t give search a reason to show this page in addition to those already selected.
Here’s a working criterion to use instead of “how many keywords did we cover?”: if you remove this page from the results and replace it with any competitor from the top results, exactly what will the user lose? If the answer is “nothing specific,” there’s no point optimizing further.
How to find empty coordinates
The method requires nothing but search results and a careful eye.
Collect 10–20 results for the target query. You don’t need the full text—snippets and the heading structure are enough: H1, H2, H3. The outline shows how competitors break the topic into parts.
List every point on a separate line and count how many sources cover each one. The sample will then sort itself into layers:
- Redundancy core — topics with a frequency of 7–15 out of 20. Everyone has written about depreciation and the weight of running shoes, so there’s no point going after those topics: the set is already covered, and the spots will go to the most authoritative domains based on signals unrelated to the text.
- Empty coordinates — topics with a frequency of 0–2 that are still relevant to the original query. How to tell a fake from the original, what happens to a product after a year of use, when standard advice doesn’t work, and what a bad choice can lead to.
Different ways of phrasing the same meaning count as one. “Technical specifications,” “Parameter overview,” and “What’s inside” are one point, not three. When three quarters of the sample collapse into one point, they look diverse only at a glance.
Real questions missing from SEO articles can be found in forum discussions and topical threads, product-card reviews, and messages to your own support team. That’s where people ask what they genuinely want to know before buying—not what’s convenient to fit into a subheading.
The trap of diversity for diversity’s sake
Being different from competitors isn’t enough, because selection takes both factors into account: distinctiveness and usefulness. Here’s the difference in practice:
| Unique but useless | Unique and useful |
|---|---|
| Brown cord | 1,2 m cord, with a compartment in the base for excess cord |
| Convenient red button | The button mechanism is rated for 15 000 presses, with a 10-year warranty on the assembly |
| Stainless steel body | The double walls act like a thermos, keeping the outside of the body below 40 °C |
| Stylish, modern design | A 200 ml mug boils in 40 seconds; the full 1,7 l capacity takes 4 minutes |
The left column is also different from competitors, but it doesn’t resolve any uncertainty involved in choosing. Signs of a useful section: it helps people make a decision, distinguishes between options, spells out limitations and exceptions, and explains when standard advice doesn’t work. Signs of an empty section: it states the obvious, repeats the query in different words, or starts with a generic introduction and a definition of the term.
How to beat aggregators
Marketplaces use templated presentation and have money to spend on everything else. There’s no point trying to outspend them, but their template is a vulnerability: millions of product cards follow the same pattern and cluster together in vector space.
Visuals
The aggregator standard is a product on a white background, studio lighting, three-quarter angle. You can stand out with photography they don’t have: macro shots showing material texture and weld quality; the product in a real interior instead of on a white cyclorama; or a teardown showing what’s inside—what kind of heating element it uses, how the soldering is done, whether it’s sealed.
How specifications are organized
The numbers are the same everywhere because they come from the manufacturer’s spec sheet. You can’t change them, but you can group them differently: not in a “power, capacity, material” column, but by use case. One model for people who brew green tea at 80 °C, another for families with small children, and a third for people who want a single mug ready in 40 seconds before heading out. Accuracy isn’t compromised, but the document’s semantic structure no longer matches the aggregator’s template. The product card becomes an expert analysis instead of a showroom.
Being cited in generative answers
There are only two or three spots, not ten, and sources compete for the right to be cited. Aggregators relay manufacturer data that the model has already seen on hundreds of sites, so the value of another identical source is close to zero. Two approaches work.
The first is comparative claims instead of descriptions. Not “the kettle has a gooseneck spout,” but “unlike Model A, Model B has a gooseneck spout that produces a laminar stream for pour-over brewing.” The model gets a ready-made connection that doesn’t appear in the manufacturer’s source data.
The second is original measurements. Use a stopwatch to time how long a mug takes to boil; use a smartphone sound meter to measure decibel levels at different stages of heating; calculate the cost of a hundred boiling cycles using your local rate. These figures don’t physically exist anywhere else, can’t be created by rewriting, and are more useful to someone choosing a product than a line saying “2200 W power.”
Competition within your own page
Duplicates don’t occur only between websites. A page can collapse in on itself when several sections serve the same informational purpose: the introductory paragraph repeats what the table will say, the comparison duplicates the list of differences, and a paragraph for “uncovered keywords” at the bottom repeats the same points as the middle of the text.
The rule is simple: each idea appears once on the page, in its strongest form. Delete the rest. Once you’ve cleaned things up, there’s room for the missing nodes—the empty coordinates. Move the most valuable information to the top of the page.
How to run this through a language model
Manually analyzing twenty competitors takes hours, so it makes sense to offload the process to an LLM. There are a few details to get right, or the result falls apart.
Give the model files, not links
It’s best to save a competitor’s page as Markdown and attach it as a file. According to analyses shared among practitioners, when a chatbot follows a link, it reads around 5–6 thousand characters. If the page header and menu contain dozens of links, the amount read drops to a third of that. These third-party measurements can’t be verified from the outside, but the direction matches experience: clean Markdown is parsed much more thoroughly than a live URL. Browser extensions can save pages as Markdown with one click and can be configured to capture the sections you need.
Attach what competitors don’t have
The manufacturer’s PDF manual, specifications, certificates, support questions, and reviews. Running the analysis with competitors alone produces a different result from running it with the manual attached: the documentation reveals limits on minimum and maximum fill levels, the initial setup procedure, and rules for reheating and cleaning. Competitors don’t have any of this for one reason: nobody opened the manual.
The first answer isn’t the final one
The workflow consists of four modes: analyzing competitors with an overcoverage map, designing the structure, writing the copy, and auditing the finished result section by section. The audit almost always finds both unnecessary content and omissions, after which the structure is rebuilt. A single pass without an audit produces a noticeably weaker page.
Proofread it yourself
An automated pipeline won’t work here. The model confidently invents plausible product features that the product doesn’t have, and the cost of that mistake on a commercial product page is greater than any time saved. One more thing about language: the same prompt with the same data performs differently in Russian-, English-, and Spanish-language niches, so specify the niche, language, and region explicitly.
What to do on Monday
- Pick one page that’s stuck in the third to fifth dozen of results, and ten top-ranking competitors for its query.
- Save all eleven as Markdown, list their headings, and count how often each point appears.
- Mark the core of the overcoverage—it makes up almost the entire content of your page.
- Find three to five empty coordinates and assess each one for usefulness: does it reduce uncertainty when making a choice?
- Think of one measurement you can take in five minutes, and take it. A stopwatch, scales, a sound meter on your phone.
- Condense repetitions to a single mention, use the space you free up for new elements, and move the most valuable content higher up.
- Check the size of the finished page: if the markup with inline resources is approaching two megabytes, move out anything unnecessary.
- Check your rankings again in two to three weeks. Practitioners note that changes show up earlier in Yandex than in Google.
Now you have to count not keyword coverage, but the meaning-bearing elements your competitors lack. And for each one, ask a second question: does anyone besides you need it?