AIToolsy

How We Test and Compare AI Tools

Our testing and comparison methodology — what we verify, what we won't claim, and how to read our scores.

Trust is the whole product here, so we’re explicit about how we evaluate tools and what our labels mean.

Curation first

We don’t list every AI tool. We curate a smaller set that matters — tools that are widely used, genuinely distinct, or clearly best for a specific job. A smaller catalog that’s accurate beats a huge catalog that’s stale. (If a tool isn’t listed, it’s usually because we haven’t verified it yet — not because we’re being paid to exclude it.)

Evidence levels

Every listing is labeled one of two ways:

  • Hands-on tested — we used the product, recorded the date, version, test, and result.
  • Research / specifications only — based on vendor documentation and public information; not independently tested by us.

We never claim to have tested what we haven’t. If you see “research-only,” that’s us being honest rather than inflating credibility.

What we verify (and how)

For every tool we aim to confirm, with a date stamp:

  • Status — does it still exist and work? (active / maintenance / acquired / discontinued)
  • Pricing — free plan, entry/pro/team price, usage limits, credit system.
  • Platforms — web, desktop, mobile, API, CLI, IDE, self-host.
  • Privacy — data retention, training on user data, self-hosting, SSO, compliance claims.
  • Commercial rights — can you use output commercially, and under what terms?
  • Integrations — what it connects to, and via what (native, API, Zapier).

Where we can’t verify something, we write “not disclosed” or “not verified.” We do not invent data to fill gaps.

Scoring is dimensional, not a single number

A single “9.2/10” hides more than it reveals. Where we score, we score dimensions separately — quality, ease, speed, reliability, privacy, integrations, value — and let your profile weight them. A developer weights quality + API + integrations; a beginner weights ease + value. One number can’t serve both.

Community sentiment ≠ verified fact

We read Reddit, GitHub issues, and review sites to understand real-world complaints — billing problems, feature regressions, reliability. We summarize patterns and clearly separate community sentiment from facts we’ve verified. We do not scrape star ratings and republish them as if they were evidence.

What we refuse to do

  • Never rank by commission.
  • Never claim testing that didn’t happen.
  • Never call a trial “free.”
  • Never keep a dead tool listed as active.
  • Never fabricate benchmarks or prices.

If something here looks wrong or outdated, tell us — accuracy is the point.