Convert
Vendor list-precision benchmarks are sales theater
On Sep 29 Landbase launched GTM-3 Omni and posted its own benchmark claiming 76.7% precision vs Clay, Apollo, and ZoomInfo. Treat those as Landbase's claimed results, not audited buyer truth.
September 30, 2026

Vendor list-precision benchmarks are sales theater
A vendor marking its own list-precision exam is not buyer proof.
On Sep 29, Landbase announced GTM-3 Omni, an agentic GTM model for discovery, qualification, and outreach, available in its platform and natively in Claude Code, Codex, and Gemini CLI. The same Business Wire and Yahoo Finance coverage carried Landbase's own AI Lab benchmark: on 26 identical list prompts, Landbase claims 76.7% precision, Clay 47.9%, Apollo 36.0%, and ZoomInfo 23.2%, with Landbase winning 18 of 26. Those are Landbase's claimed results from a study Landbase ran. They are not an audited third-party truth about your pipeline.
Vendor-run list precision is a claim. Meetings on your ICP are the score.
What did the Landbase numbers actually say?
They said Landbase scored itself highest on a private set of 26 list prompts against named competitors.
I am not here to litigate their methodology in public without the full prompt pack, judge rubric, and holdout set. I am here to name the category error. When the seller designs the prompts, runs the lab, and publishes the leaderboard, you are reading marketing. Useful as a hypothesis. Dangerous as a buying decision. Same pattern I already covered for private AI benchmarks that are slogans instead of proof.
If Landbase's precision is real for your ICP, you will see it in meetings booked and junk rows cut. If it is not, a slide with 76.7% will not save Q4.
Why do GTM teams fall for precision theater?
Because precision feels like a grown-up metric and volume feels dirty.
Leadership wants a number that says "our lists are clean." A vendor hands them a percentage with three competitors underneath. The deck writes itself. Nobody asks who defined a true positive, whether the ICP matched yours, or whether the winning rows ever produced a meeting. Enrichment vendors have played this game for years. Agentic list builders just put a model name on the same move.
Precision on someone else's prompts is not Convert. Convert is a reply that becomes a meeting with a buyer who can buy.
How should you score a list tool instead?
Run your own bake-off on your ICP. Keep the rubric boring.
- Freeze 26 of your prompts, not theirs. Same roles, same segments, same exclusions you actually use.
- Define a true positive before you run. Fit plus reachable contact plus not on exclusion. Write it down.
- Blind the review if you can. Have someone who did not pick the vendor label the rows.
- Score meetings, not vibes. Meetings from signal-matched sends beat list precision every time.
- Track DQ rate. How many rows should never have been touched? A tool that looks precise while flooding bad fit is expensive theater.
- Hold cost constant. Credits, seats, and human review hours. A "win" that triples spend for the same meetings is not a win.
Do that for two weeks. Keep the vendor benchmark in the appendix if you want. Do not let it run the PO.
What about GTM-3 Omni in Claude Code and Codex?
Native CLI access matters for operators who already live in coding agents. It does not change the proof standard.
You can discover, qualify, and draft from the same window. Great. The Convert question is unchanged: does the path produce meetings with ICP buyers, and can you stop it when DQ fires? If the agent can build a list but cannot respect your exclusions, you bought speed without Convert.
What should I do Monday?
Take the last list a vendor or an agent built for you. Sample 100 rows. Mark fit, exclusion hits, and whether anyone booked. Put that table next to any vendor precision claim in the deck. If the table is empty, you are buying a story.
Adapt or fail. Claimed precision is not a meeting.
FAQ
What precision did Landbase claim on Sep 29?
In Business Wire and Yahoo Finance coverage of the GTM-3 Omni launch, Landbase's AI Lab reported claimed precision of 76.7% for Landbase versus Clay 47.9%, Apollo 36.0%, and ZoomInfo 23.2% on 26 identical list prompts, with Landbase winning 18 of 26. Treat those as Landbase's claimed results.
Are vendor benchmarks useless?
No. They are a starting claim. They become useful only when you reproduce the test on your ICP with a frozen rubric and meeting outcomes.
What should I measure instead of list precision?
Meetings booked from ICP-matched sends, DQ rate on generated rows, cost per meeting, and exclusion-list violations. Precision without those is a slide.
Does agentic list building change the rules?
It changes speed. It does not change proof. Faster wrong lists fail faster.
How do I compare Clay, Apollo, ZoomInfo, and newer agent tools fairly?
Same prompts, same ICP definition, same exclusion lists, same human labeling rules, and the same meeting scoreboard. Anything else is brand preference dressed as science.
Frequently asked questions
- What precision did Landbase claim on Sep 29?
- In Business Wire and Yahoo Finance coverage of the GTM-3 Omni launch, Landbase's AI Lab reported claimed precision of 76.7% for Landbase versus Clay 47.9%, Apollo 36.0%, and ZoomInfo 23.2% on 26 identical list prompts, with Landbase winning 18 of 26. Treat those as Landbase's claimed results.
- Are vendor benchmarks useless?
- No. They are a starting claim. They become useful only when you reproduce the test on your ICP with a frozen rubric and meeting outcomes.
- What should I measure instead of list precision?
- Meetings booked from ICP-matched sends, DQ rate on generated rows, cost per meeting, and exclusion-list violations. Precision without those is a slide.
- Does agentic list building change the rules?
- It changes speed. It does not change proof. Faster wrong lists fail faster.
- How do I compare Clay, Apollo, ZoomInfo, and newer agent tools fairly?
- Same prompts, same ICP definition, same exclusion lists, same human labeling rules, and the same meeting scoreboard. Anything else is brand preference dressed as science.