- Two keywords belong to the same group when their top 10s share enough URLs: 3 to 4 in common is the most common threshold.
- The method relies on what Google actually shows, not on word similarity: two close-sounding keywords may or may not end up together.
- With the Semscraper API, 1,000 keywords on page 1 cost €0.30, and the Python script in this article does the rest.
The principle
SERP similarity clustering starts from a simple idea: if Google shows the same pages for two queries, it understands them as the same intent. A single page of your site can then rank for both. Conversely, two keywords that look alike but whose SERPs have nothing in common call for two separate pages.
The method is more reliable than grouping on the words themselves, because it relies on Google's judgment, not on string similarity. “Women's running shoes” and “ladies jogging sneakers” share almost no words, yet their SERPs often overlap widely.
The steps
- Collect the top 10 of every keyword, in the right country, language and device.
- Extract the organic result URLs, ignoring ads, videos and other blocks.
- Compare keywords pairwise: how many URLs do they share?
- Group the pairs above a threshold.
Choosing the threshold
| Shared URLs in the top 10 | Effect |
|---|---|
| 2 | Wide groups, at the risk of mixing neighboring intents |
| 3 or 4 | The most common setting: consistent groups without too much fragmentation |
| 5 and more | Very tight groups, many isolated keywords |
Start at 3, review a dozen groups by hand, then adjust. The right threshold depends on your topic: e-commerce SERPs, dominated by a few big sites, overlap more easily than editorial ones.
Strict or loose groups
There are two approaches. In strict clustering, every keyword in a group must share the threshold with a pivot keyword, usually the most searched one in the group. Groups are clean, and it is the approach of the script below. In loose clustering, a keyword only needs to be linked to one member of the group: groups grow by chaining, at the risk of drifting from one intent to another.
To decide which pages to create or merge, strict clustering is almost always the right choice.
The complete script
This script sends your keywords to the API, fetches the SERPs, then builds the groups. Replace the API key, the keyword list and, if needed, the threshold. Keywords should be sorted from most to least searched, so that each group's pivot is its main query.
import time
import requests
API_KEY = "API_KEY"
API = "https://api.semscraper.com/v1/serp"
HEADERS = {"Authorization": "Bearer " + API_KEY}
THRESHOLD = 3
# from most to least searched
keywords = ["womens running shoes", "ladies jogging sneakers", "womens trail shoes"]
# 1. create the collections: top 10, google.com, desktop
items = [{"search_engine": "google_search", "keyword": k, "device": "desktop",
"location": "en", "language": "en", "depth": 1} for k in keywords]
created = requests.post(API, json=items, headers=HEADERS).json()
ids = [row["id"] for row in created["data"]]
# 2. wait, then fetch the SERPs
serps = {}
while len(serps) != len(ids):
time.sleep(10)
res = requests.get(API, params={"ids": ",".join(ids), "output": "json"}, headers=HEADERS).json()
for row in res["data"]:
if row["status"] == "done":
urls = set()
for block in row["results"]:
if block["type"] == "organic":
urls.update(item["url"].rstrip("/") for item in block["items"])
serps[row["keyword"]] = urls
# 3. strict clustering around a pivot
clusters, seen = [], set()
for pivot in keywords:
if pivot in seen or pivot not in serps:
continue
group = [pivot]
seen.add(pivot)
for other in keywords:
if other not in seen and other in serps:
if len(serps[pivot] & serps[other]) >= THRESHOLD:
group.append(other)
seen.add(other)
clusters.append(group)
for group in clusters:
print(" | ".join(group))The IDs returned on creation are then passed to the fetch call, separated by commas. For large volumes, use a callback URL instead: the API sends you the results as soon as they are ready, with no polling loop.
Pitfalls to avoid
- Mixing locations or devices: compare SERPs collected under the same conditions, otherwise the gaps come from context, not intent. For a local topic, geolocate every query in the same city.
- Comparing SERPs from different dates: collect the whole batch on the same day.
- Keeping raw URLs: normalize at least the trailing slash, and possibly parameters, so identical URLs aren't missed.
- Counting non-organic blocks: videos, news or ads change from hour to hour and skew the comparison.
What it costs
A single results page is enough to compare top 10s: with Semscraper, 1,000 keywords cost €0.30, 10,000 keywords €3. The 1,000 free requests at sign-up let you test the method on a real keyword batch. Collection parameters are detailed on the Google Search API page.
