r/SalesOps

▲ 8 r/SalesOps+1 crossposts

Independent, open benchmarks of company lookalike APIs & providers - raw responses and judge prompts published

We ran a lookalike / similar-company benchmark across 7 vendors: Parallel, Extruct, Ocean.io, Exa, PredictLeads, Discolike and CUFinder.

Here - https://openbenchmarks.com/lookalikes

48 seed companies across 13 categories — b2b-saas, devtools, ecommerce, healthtech, home services, trades, real estate, fintech, cybersecurity, industrial, logistics, hospitality, energy. We asked each API for up to 100 lookalikes per seed, then had an LLM judge (gpt-5.6) score every returned company on whether it's genuinely a lookalike of that seed.

Different vendors win at different K

Vendor          P@10    P@25   P@100
-------------------------------------
PredictLeads   95.83   77.17       -
Exa            95.00   80.83   52.40
Parallel       74.79   75.00   67.54
Extruct        74.79   70.00   61.21
Ocean.io       73.19   67.06   56.53
Discolike      42.98   41.53   35.19
CUFinder       37.45       -       -

Average precision, 48 seeds, judged by gpt-5.6. Higher is better. A - means unscored, not zero — see caveats.

At K=10, PredictLeads leads at 95.8 and Exa is second at 95.0. At K=100, Exa is fourth at 52.4 and Parallel — third at K=10 — is first at 67.5.

Precision drop from P@10 to P@100

Exa        95.0 → 52.4    −42.6
Ocean.io   73.2 → 56.5    −16.7
Extruct    74.8 → 61.2    −13.6
Discolike  43.0 → 35.2     −7.8
Parallel   74.8 → 67.5     −7.3

Exa drops 42.6 points between K=10 and K=100. Parallel drops 7.3.

Which one to use

  • A list of 10–25 accounts: PredictLeads (95.8 at K=10) or Exa (95.0 at K=10, 80.8 at K=25). The drop-off never reaches you.
  • A list of 100+ accounts: Parallel, at 67.5. A P@10 comparison would point you at Exa, which scores 52.4 at that depth.

Two caveats

  • PredictLeads and CUFinder return fewer than 100 results per seed, so they have no P@100. Read the - as no data, not as a low score.
  • "Lookalike" is judged, not ground truth. An LLM decided what counts. The judge prompt is published, so you can read it and disagree with it.

Reproducing it

Every cell is backed by the HTTP request and response we sent to each vendor, plus the judge prompt and its response, stored per seed and per vendor.

Happy to add a vendor or run more seeds. If you think the judging is wrong on a specific pair, the raw file shows what the judge saw.

Disclosure: I run Openbenchmarks. We run independent benchmarks and publish the raw artifacts and results for agents.

No vendor paid for placement or inclusion.

reddit.com
u/-GeneX- — 12 days ago