hushwork
1259 wordsDrafted by the Hushwork engine · human-reviewedupdated

Keyword Clustering for Ecommerce SEO

How to decide which queries share a page and which deserve their own

The question clustering actually answers

Keyword clustering is the practice of deciding which search queries deserve one shared page and which need pages of their own. That framing matters: the goal is not sorting keywords into tidy folders, it is settling a page-count question before you write anything.

Take a store selling home fitness equipment. Searches for "adjustable dumbbells", "adjustable dumbbell set", and "dumbbells with adjustable weight" are three phrasings of one buying decision — a single page can rank for all of them. But "dumbbells vs kettlebells" is a different job entirely: that searcher has not picked a category yet, and a product listing will not answer them.

Getting the grouping wrong costs you in both directions. Split one topic across three pages and each version ends up thin, competing against its own siblings. Cram genuinely different intents onto one page and it answers nobody well enough to rank. Clustering replaces gut feel with two measurable signals — semantic similarity and SERP overlap — and this guide walks through both, then covers what to do with the clusters once you have them.

Measuring meaning with embeddings

The first signal is whether two queries mean the same thing. Clustering tools measure this with embeddings — numerical representations of meaning, where each query becomes a long list of numbers and queries with similar meanings land close together in that number space. "Handmade ceramic mugs" and "artisan pottery coffee cups" share almost no words, yet their embeddings sit close, because they describe the same product.

That is the strength: embeddings catch groupings that word matching misses. A handmade homeware store would want "stoneware dinner set" and "ceramic dinnerware set" in one cluster, and embeddings put them there without anyone maintaining a synonym list.

The known failure mode is queries that are about the same thing while wanting different things. "Walking pad" and "walking pad belt lubricant" sit close in embedding space — both concern the same machine — but one comes from a buyer and one from an owner with a maintenance problem. Cluster on embeddings alone and you will occasionally merge a product query with a support query. That is why a serious clustering pass pairs embeddings with a second signal that measures intent directly.

The SERP overlap test

The second signal comes from Google itself. If two queries return mostly the same top-ten results, Google has already decided that one kind of page can satisfy both — which means one page of yours can too. If the result sets barely overlap, Google sees two different intents, and you need two pages no matter how similar the wording looks.

You can run the check by hand for your most valuable terms:

  1. Search each query in a private browser window so personalization does not skew the results.
  2. Write down the ten ranking URLs for each query.
  3. Count how many URLs appear on both lists. A common rule of thumb: four or more shared results is a strong same-cluster signal; one or zero means separate pages.

Try it on "stoneware dinner set" and "handmade dinnerware" and you will likely find heavy overlap — one collection page can carry both. Try "rowing machine" against "rowing machine workout plan" and the results diverge sharply: product listings for the first, training content for the second. Same object, different pages.

Label every cluster with an intent tier

Once queries are grouped, tag each cluster with the intent tier it belongs to, because the tier determines what kind of page can win it.

  • Informational — the searcher wants an answer, not a product yet: "how to repair a chipped ceramic mug", "treadmill belt keeps slipping". Guides win these.
  • Commercial — the searcher is comparing options before buying: "best compact rowing machines", "walking pad vs under-desk treadmill". Comparison pages and buying guides win these.
  • Transactional — the searcher is ready to purchase something specific: "buy handmade ceramic dinner set", "adjustable dumbbells 25 kg". Product and collection pages win these.

The rule that follows: never merge clusters across tiers, even when embeddings and SERPs both look close. A guide about fixing chipped mugs and a collection selling replacement mugs serve neighbouring moments in the same customer's life, but a page that tries to do both — half tutorial, half buy buttons — ranks for neither. Keep the tiers separate and connect them with links instead, so the reader who finishes the repair guide can step into the collection when they are ready.

Mapping clusters to Shopify page types

Shopify's URL structure is fixed — products live under /products/, collections under /collections/, posts under /blogs/ — so each intent tier maps naturally onto a page type you already have.

Transactional clusters belong to collection pages. A cluster around "compact home gym equipment" becomes a collection with that exact framing in its title and description, not a blog post. Commercial clusters belong to comparison guides published as blog posts, where you can lay out honest trade-offs between models. Informational clusters belong to how-to and care guides, also on the blog.

Two follow-on decisions come out of this mapping. First, when one commercial cluster repeats across many audiences or use cases — a buying guide per space, per goal, per budget — you are looking at the entry point to programmatic SEO done with real data per page. Second, a cluster's pages must point at each other: the care guide links to the collection it feeds, the comparison links to the products it compares. A deliberate internal linking structure is what turns a clustering spreadsheet into actual site architecture.

Cannibalization and how to fix it

Cannibalization is what unclustered publishing eventually produces: two of your own URLs competing for the same query. The telltale symptom is rank instability. In Google Search Console, open Performance, filter to a single query, and switch to the Pages tab. If two of your URLs swap in and out of the rankings — one holds the position for a stretch, then the other replaces it — Google cannot decide which page is your answer, and typically neither ranks as well as one consolidated page would.

There are two fixes, and the intent tiers tell you which applies:

  1. Consolidate. If both pages serve the same intent, merge the content into the stronger URL, delete the weaker page, and redirect it under Online Store → Navigation → URL Redirects. Shopify redirects only fire when the source URL would otherwise 404, so the losing page must actually be removed first.
  2. Differentiate. If the intents genuinely differ, rewrite each page so it commits to its own tier — retitle the guide as a guide, strip comparison content out of the collection description, and adjust the internal anchors pointing at each page.

Make clustering a routine, not a project

Clustering decays. New products create new query spaces, seasons shift language (gift clusters balloon before December), and Google re-interprets intent as the market changes. Re-run the SERP overlap check on your most important clusters every quarter, and re-cluster whenever a new product line launches.

Keep one artifact at the center: a map of cluster to owning URL. Before any new page is briefed, check it against the map; if the cluster already has an owner, the work goes into improving that page instead. That single habit prevents most cannibalization before it starts.

This is also the discipline automated content systems have to get right — Hushwork, for instance, uses clustering to size its output, grouping the queries it discovers for a store and generating one page per topic rather than one page per keyword. Whether you cluster with embeddings and a spreadsheet or with software, the deliverable is the same: a page count that matches your topic count, with every query knowing exactly which page it belongs to.

Frequently asked questions

How many keywords should one page target?

As many as belong to one cluster — often five to thirty phrasings of the same intent. The count is an output of clustering, not a target you set in advance. If queries share meaning and their top-ten results overlap heavily, one page carries them all; if not, they belong on separate pages.

What is the fastest way to check whether two keywords need separate pages?

Search both in a private browser window and compare the top ten results. If four or more of the ranking URLs are the same, one page can plausibly win both queries. If the results barely overlap, Google is treating them as different intents, and you should build separate pages.

How do I know if my store has a cannibalization problem?

In Google Search Console, open the Performance report, filter to a single query you care about, then switch to the Pages tab. If two of your URLs alternate in the rankings over time — each taking turns holding the position — those pages are competing with each other and need consolidating or differentiating.

Are embeddings alone enough to build reliable clusters?

No. Embeddings group queries by meaning, so they occasionally merge a buying query with a support query about the same product — semantically close, but needing different pages. Pair semantic similarity with a SERP overlap check, which reflects how Google actually interprets the intent behind each query.

Keep reading