Skip to content

How we measure S1 research

Where every S1 figure on this site comes from: what each term means, the samples and dates, how the tests were run, where the other side won and what has not been tested yet.

Updated October 7, 2026. SalesOne research team. Back to S1 agents

What do the terms mean?

The same words are used the same way on every page that quotes these figures.

Usable account
In the blind tests: an account the checkers judged could be worked as delivered or after small corrections. “As is” means no correction was needed.
Clear misfit
A record that plainly falls outside the profile it was pulled for, such as an airline, a law firm or a university on a list for manufacturers and financial services firms.
Dated
A buying signal dated to the day the event happened, or to the month when the source gives no day.
Kept
A company S1 agents accepted against the profile and delivered as an account. Also called accepted or delivered.
Reject reason
The rule a company failed and the evidence, stored with the company. The server refuses a reject without one. Rule codes at the start of each reason appear only on the latest run.
Ready person
A person with a public source dated within the last 12 months and a profile match. Only ready people go to contact lookup.

Which production runs are the figures from?

The depth figures describe what S1 agents recorded in real research runs with the current pipeline: a lead agent, parallel research agents, one shared page reader and, on three runs, a review pass. A fifth run on Oct 3, 2026 used an older pipeline and is left out. These figures were not blind-checked.

Runs
Four production runs, Oct 4 and 5, 2026: three on one B2B consulting client’s profile, one on SalesOne’s own profile
Companies decided
1,051: 249 kept and 802 turned down (24% kept, between 22% and 26% per run)
Time per run
44 to 82 minutes for 27 to 102 accounts kept, with at most 10 research agents at once
Searches per account kept
31.2 (7,780 searches over 249 accounts), counted from the run logs
Pages read per account kept
6.7 (1,667 over 249). 81% of page reads succeeded; when a page could not be read, the agent switched source
Buying signals
375 on 249 accounts: all 375 link their source and 97.9% are dated to the day. The newest signal on a typical account was 59 days old at delivery (median)
People
1,057, or 4.2 per account: 95.3% with a source and a date, 84.7% from a source beyond professional-network profiles
Rejects
802 of 802 keep a reason; 88% of reasons contain a date. Top reason: no qualifying event inside its window (67%)
Review pass
Three of the four runs, 147 accounts: 40 signals corrected, 58 fresher ones added, 3 accounts pulled (from the review agents’ own reports)
Companies remembered
About 7,000 so far across SalesOne (6,954 registered on Oct 7, 2026), matched by website and by name

How was the blind test run?

Two tests against a leading AI research API. The first is the one quoted across the site; the second is too small to settle anything, and it tied.

Test one: a client profile

Date and profile
Oct 3, 2026, one B2B consulting client’s profile: US mid-market and enterprise manufacturers, financial services and insurance
Samples
61 accounts from S1 agents against 39 from a leading AI research API (7 runs). Both sides excluded the same 5,771 known companies
Effort
S1 had about an hour and 15 agents. The API ran at its default effort, 5 to 11 minutes and about $1 a run
Checking
Records were mixed and anonymized, then checked against their cited sources by 10 AI checker agents that did not know which side produced each one. The checkers came from the same model family as S1’s agents
Usable
58 of 61 (95%) against 27 of 39 (69%); as is, 37 against 6
Passed every profile rule
67% against 38%; the API delivered 2 companies that had already been acquired
People per company
3.9 against 1.6; companies with nobody found, 0 against 13
People backed beyond professional-network profiles
83% against 30%
Accounts with 2 or more signals
46 against 8

Where the API won

Its buying signals held up more often when checked: 46 of 48 (96%) against 99 of 132 (75%) for S1. It was also faster and much cheaper. The S1 side of this test ran on Oct 3, before the review pass that now re-checks every signal was added to production runs on Oct 4; the review pass has not yet been through a blind test.

Test two: our own profile

Date and profile
Oct 2 and 3, 2026, SalesOne’s own profile, 10 accounts a side, checked the same way
Usable
9 of 10 each: a tie
Passed every profile rule
7 against 4
Effort
The API took 6.3 minutes and $1.27; S1 took 34.5 minutes, 75 searches and 95 page reads

What do the database comparisons show?

Two pulls built the way a team would build a list from data, then sorted on their own returned fields. They show what a filtered list looks like before anyone qualifies it. They are not a head-to-head.

Contact database, Sept 28, 2026
70 people from a filtered search (industry, size and title) for the client’s profile, each labelled by an AI classifier: 44% on target, 36% borderline, 20% off. 6 of 10 people listed as new in role still showed an older role
Company-data API, Sept 29, 2026
499 companies from industry and size filters, each sorted on its returned data: 28.9% clear misfits. Records were a median 20 months since their last refresh (7 to 25); 87% were over a year old
S1 research
71 accounts checked blind against sources in the two tests above: about 1 in 20 clear misfits

Read the pooled comparison, about 1 in 4 clear misfits against about 1 in 20, as a direction: 569 database and API records against 71 researched accounts, different samples and methods, and a small S1 sample. The researched list overlapped the 499 API companies on only 37. No named contact database was tested.

How were the models chosen?

A blind test in two rounds, Oct 2, 2026. Every model got the same research brief and 5 batches of 10 accounts. Five AI judges, one per batch, saw only “Set X” and “Set Y”, re-opened every source and scored yield, accuracy, fit, signals, people and file hygiene out of 10.

Round 1, accuracy
Frontier model 8.6, small model 2.6. The small model invented people and companies and delivered 32 of 50 accounts, so it is never used for research
Round 2, accuracy
Frontier model 8.0, mid-size model 7.2. The mid-size model invented nothing but was weaker on fit and completeness

Judges compare two sets side by side, so a score is relative to its opponent in that round: read the gap, not the number. One run per batch makes it a practical comparison, not a formal benchmark.

Which outside benchmarks are quoted?

WideSearch (arXiv 2508.07999, August 2025) asks agents to collect every entity in a set into a table; the best agent tested fully completed 5.1% of tasks. Dedicated deep research products were left out because they returned reports instead of tables. DeepWideSearch (arXiv 2510.20168, October 2025) adds multi-step reasoning to find each entry; the best average success was 2.4%. Both measure general AI agents; neither tested S1.

What has not been tested yet?

  • A head-to-head with a contact database on the same sample. The database figures and the S1 figures come from different samples and methods.
  • A human-reviewed audit. Every checked figure so far comes from AI checker agents working from the cited sources.
  • Monitoring between runs. S1 does not watch your accounts between research requests, so there is nothing of that kind to measure.
  • Whether later runs get measurably better. Fit ratings go into the next research brief, but too few have been recorded to measure an effect.
  • More than one client profile in a blind test of this size. The second test, on our own profile, had 10 accounts a side.

Who is the client?

A B2B consulting firm whose profile targets US mid-market and enterprise manufacturers, financial services and insurance. Its name, its target companies and its people are withheld. The second profile in the runs and tests is SalesOne’s own.

See your own market, researched.

Bring your offer and one customer you would like more of. We will research a sample of your market and show you the accounts S1 agents accept, the ones they turn down, and why.