How we measure S1 research
Where every S1 figure on this site comes from: what each term means, the samples and dates, how the tests were run, where the other side won and what has not been tested yet.
Updated October 7, 2026. SalesOne research team. Back to S1 agents
What do the terms mean?
The same words are used the same way on every page that quotes these figures.
- Usable account
- In the blind tests: an account the checkers judged could be worked as delivered or after small corrections. “As is” means no correction was needed.
- Clear misfit
- A record that plainly falls outside the profile it was pulled for, such as an airline, a law firm or a university on a list for manufacturers and financial services firms.
- Dated
- A buying signal dated to the day the event happened, or to the month when the source gives no day.
- Kept
- A company S1 agents accepted against the profile and delivered as an account. Also called accepted or delivered.
- Reject reason
- The rule a company failed and the evidence, stored with the company. The server refuses a reject without one. Rule codes at the start of each reason appear only on the latest run.
- Ready person
- A person with a public source dated within the last 12 months and a profile match. Only ready people go to contact lookup.
Which production runs are the figures from?
The depth figures describe what S1 agents recorded in real research runs with the current pipeline: a lead agent, parallel research agents, one shared page reader and, on three runs, a review pass. A fifth run on Oct 3, 2026 used an older pipeline and is left out. These figures were not blind-checked.
- Runs
- Four production runs, Oct 4 and 5, 2026: three on one B2B consulting client’s profile, one on SalesOne’s own profile
- Companies decided
- 1,051: 249 kept and 802 turned down (24% kept, between 22% and 26% per run)
- Time per run
- 44 to 82 minutes for 27 to 102 accounts kept, with at most 10 research agents at once
- Searches per account kept
- 31.2 (7,780 searches over 249 accounts), counted from the run logs
- Pages read per account kept
- 6.7 (1,667 over 249). 81% of page reads succeeded; when a page could not be read, the agent switched source
- Buying signals
- 375 on 249 accounts: all 375 link their source and 97.9% are dated to the day. The newest signal on a typical account was 59 days old at delivery (median)
- People
- 1,057, or 4.2 per account: 95.3% with a source and a date, 84.7% from a source beyond professional-network profiles
- Rejects
- 802 of 802 keep a reason; 88% of reasons contain a date. Top reason: no qualifying event inside its window (67%)
- Review pass
- Three of the four runs, 147 accounts: 40 signals corrected, 58 fresher ones added, 3 accounts pulled (from the review agents’ own reports)
- Companies remembered
- About 7,000 so far across SalesOne (6,954 registered on Oct 7, 2026), matched by website and by name
How was the blind test run?
Two tests against a leading AI research API. The first is the one quoted across the site; the second is too small to settle anything, and it tied.
Test one: a client profile
- Date and profile
- Oct 3, 2026, one B2B consulting client’s profile: US mid-market and enterprise manufacturers, financial services and insurance
- Samples
- 61 accounts from S1 agents against 39 from a leading AI research API (7 runs). Both sides excluded the same 5,771 known companies
- Effort
- S1 had about an hour and 15 agents. The API ran at its default effort, 5 to 11 minutes and about $1 a run
- Checking
- Records were mixed and anonymized, then checked against their cited sources by 10 AI checker agents that did not know which side produced each one. The checkers came from the same model family as S1’s agents
- Usable
- 58 of 61 (95%) against 27 of 39 (69%); as is, 37 against 6
- Passed every profile rule
- 67% against 38%; the API delivered 2 companies that had already been acquired
- People per company
- 3.9 against 1.6; companies with nobody found, 0 against 13
- People backed beyond professional-network profiles
- 83% against 30%
- Accounts with 2 or more signals
- 46 against 8
Where the API won
Its buying signals held up more often when checked: 46 of 48 (96%) against 99 of 132 (75%) for S1. It was also faster and much cheaper. The S1 side of this test ran on Oct 3, before the review pass that now re-checks every signal was added to production runs on Oct 4; the review pass has not yet been through a blind test.
Test two: our own profile
- Date and profile
- Oct 2 and 3, 2026, SalesOne’s own profile, 10 accounts a side, checked the same way
- Usable
- 9 of 10 each: a tie
- Passed every profile rule
- 7 against 4
- Effort
- The API took 6.3 minutes and $1.27; S1 took 34.5 minutes, 75 searches and 95 page reads
What do the database comparisons show?
Two pulls built the way a team would build a list from data, then sorted on their own returned fields. They show what a filtered list looks like before anyone qualifies it. They are not a head-to-head.
- Contact database, Sept 28, 2026
- 70 people from a filtered search (industry, size and title) for the client’s profile, each labelled by an AI classifier: 44% on target, 36% borderline, 20% off. 6 of 10 people listed as new in role still showed an older role
- Company-data API, Sept 29, 2026
- 499 companies from industry and size filters, each sorted on its returned data: 28.9% clear misfits. Records were a median 20 months since their last refresh (7 to 25); 87% were over a year old
- S1 research
- 71 accounts checked blind against sources in the two tests above: about 1 in 20 clear misfits
Read the pooled comparison, about 1 in 4 clear misfits against about 1 in 20, as a direction: 569 database and API records against 71 researched accounts, different samples and methods, and a small S1 sample. The researched list overlapped the 499 API companies on only 37. No named contact database was tested.
How were the models chosen?
A blind test in two rounds, Oct 2, 2026. Every model got the same research brief and 5 batches of 10 accounts. Five AI judges, one per batch, saw only “Set X” and “Set Y”, re-opened every source and scored yield, accuracy, fit, signals, people and file hygiene out of 10.
- Round 1, accuracy
- Frontier model 8.6, small model 2.6. The small model invented people and companies and delivered 32 of 50 accounts, so it is never used for research
- Round 2, accuracy
- Frontier model 8.0, mid-size model 7.2. The mid-size model invented nothing but was weaker on fit and completeness
Judges compare two sets side by side, so a score is relative to its opponent in that round: read the gap, not the number. One run per batch makes it a practical comparison, not a formal benchmark.
Which outside benchmarks are quoted?
WideSearch (arXiv 2508.07999, August 2025) asks agents to collect every entity in a set into a table; the best agent tested fully completed 5.1% of tasks. Dedicated deep research products were left out because they returned reports instead of tables. DeepWideSearch (arXiv 2510.20168, October 2025) adds multi-step reasoning to find each entry; the best average success was 2.4%. Both measure general AI agents; neither tested S1.
What has not been tested yet?
- A head-to-head with a contact database on the same sample. The database figures and the S1 figures come from different samples and methods.
- A human-reviewed audit. Every checked figure so far comes from AI checker agents working from the cited sources.
- Monitoring between runs. S1 does not watch your accounts between research requests, so there is nothing of that kind to measure.
- Whether later runs get measurably better. Fit ratings go into the next research brief, but too few have been recorded to measure an effect.
- More than one client profile in a blind test of this size. The second test, on our own profile, had 10 accounts a side.
Who is the client?
A B2B consulting firm whose profile targets US mid-market and enterprise manufacturers, financial services and insurance. Its name, its target companies and its people are withheld. The second profile in the runs and tests is SalesOne’s own.
See your own market, researched.
Bring your offer and one customer you would like more of. We will research a sample of your market and show you the accounts S1 agents accept, the ones they turn down, and why.