Stackmerit

Method

How we test and score

Every review on Stackmerit is scored against the same five-part rubric, so a score means the same thing across the site. This page sets out exactly what we assess, how the 0 to 10 score is derived, and where AI is and is not used. It is the standard our about page summarises, written out in full.

Last updated

What we assess, and what we do not

Our reviews are research-based. We build each one from three sources: the vendor's own documentation and live pricing page, checked and dated; publicly available user reviews across the major platforms; and, where they exist, independent benchmarks and named third-party reviews. Every figure we publish is attributed to a source and dated, and if a claim cannot be verified from a live source we leave it out.

We do not run first-hand product trials. We have not installed and operated each tool through a full billing cycle, and we never write as though we have. Where a review would benefit from hands-on measurement we say so, and the score for that tool stays an editorial judgement built from the research above, not a lab result. A review we have not finished carries an "In testing" badge and no score.

The test environment

There is no first-hand test lab behind these reviews today, and we do not pretend otherwise. The "environment" is the documented evidence: current pricing pages read on a stated date, feature and limit documentation, published independent benchmarks for speed and deliverability, and the body of verified user reviews. For prices that vary by region we read them in a consistent context and note when a vendor page geo-redirects, because a regional price is not a list price. Every review and comparison page carries a visible "Prices checked" date so you can see how current the pricing is.

The five sub-scores

Each tool is scored from 0 to 10 on five criteria. The criteria are the same in every category; what shifts is how much each one counts.

  • Data quality measures the size, freshness and accuracy of the numbers a tool runs on: keyword and backlink databases, index coverage, deliverability data, or the engines an AI-visibility tracker watches. We check claims against the vendor's own documentation and against what independent sources report.
  • Feature depth is how much of the real job the tool does, and how well, rather than the length of the feature list. We weigh the parts that decide day-to-day work more heavily than the ones that look good in a demo.
  • Usability is how quickly someone gets useful output, how the interface holds up under regular use, and where it slows people down. We draw this from published walkthroughs and from what verified users report.
  • Value weighs what the plan costs against what it includes, read from the vendor's live pricing page and dated. We count the extras that push the real bill up: seats, project limits, tracked keywords or prompts, and paid add-ons.
  • Support covers the channels offered, the documentation quality, and what users report about response times and how support behaves when something breaks.

How the emphasis shifts by category

A raw average would treat every criterion as equally decisive, which is not how a real buying decision works. So the overall score weights the criteria that actually decide a purchase in that category:

  • In SEO tools, data quality and feature depth carry the most, because the size and accuracy of the keyword and backlink data is what everything else runs on.
  • In hosting, real-world speed, uptime and support carry the most, because downtime and slowness on a revenue site cost real money.
  • In email marketing, deliverability, automation depth and the pricing math carry the most, because an email in spam is wasted and contact-based pricing decides the true cost.
  • In AI visibility tools, engine coverage and data quality carry the most, because a tracker that watches one engine or counts a single mention tells you little. This is a meta-review category, so we weigh it with extra caution where independent sources are thin.

How the 0 to 10 score is derived

The overall score is an editorial judgement built from the five sub-scores, weighted by the category emphasis above. It is not a mechanical average, and we do not publish fixed percentage weights, because the point of the emphasis is to let the criteria that decide a purchase carry more than the ones that do not. The number reflects the whole picture the research supports, on a 0 to 10 scale where best-in-class sits near the top and a serious, disqualifying weakness pulls it down regardless of a strong feature list. The same editor sets every score, so scores stay comparable within a category. A score ships to search engines as a machine-readable rating only once it is a settled verdict the site stands behind; until then the number is shown to readers but withheld from structured data.

Where AI is, and is not, used

AI tools help us gather and summarise published information faster: pulling together pricing tiers, documentation and user reports so a reviewer can work from one place. AI does not decide scores, rankings or verdicts, and it does not get the final word on any published figure. A human editor is responsible for every score and every number, checks each figure against a named primary source, and dates it. We do not publish AI-generated numbers, invented benchmarks, screenshots that do not exist, or review videos that were never made. Where AI assistance helped draft a page, the facts in it are verified against primary sources before it goes live.

Independence and corrections

Rankings are decided on merit first, before we check which tools have affiliate programs, and a commission never moves a tool up a list. How we make money is set out on our disclosure page. If you spot a figure that is wrong or out of date, tell us on the contact page and we will verify and fix it, and note any material change. You can also see the full field we rank in our best SEO tools guide and browse every SEO tool review we have published.