Proof Over PromisesP>PProof over Promises

Published by BitPlan

Published by BitPlan Loyalty Inc., the company behind BitStage and Barnes Signal Advisor. We publish what we find — including about our own products.

Strand 02 · The Evidence

A research library. Revised figures only.

Results from real engagements and the published research behind them, posted as the data arrives — whether or not it flatters us. The direction is well established: structured, evidence-based methods beat unstructured judgment. The effect sizes circulating on LinkedIn are not. We cite the revised, less flattering figures — including on the method closest to what BitStage does.

Results from real engagements.

00 entries · 00 verified

Client evidence from across the ecosystem — hiring engagements and AI consulting engagements alike — published as the data arrives. A claim stays a promise until it is measured. Then it moves.

Results are published whether or not they flatter us. The moment this section only ever supports us, it stops being evidence.

No engagement results published yet. Results publish as each engagement completes its measurement window, whether or not they flatter us.


The case, in three proofs — 24 of 26 entries primary-sourced; the exception is marked.
StakesAccommodation and food services has the highest quit rate of any sector in the economy: 4.5% of jobs ended in quits in a single month. Those are separation events, not distinct people. Bureau of Labor Statistics, June 2026 →

Pillar A

Demonstration beats description.

Structured interviews predict performance at rho = .42. Work samples sit at .33. The unstructured interview sits at .19. Talking about the job is the weakest of the three. An AI-scored work sample scores consistently (test-retest r = .72) and predicts job performance weakly (r = .24, n = 1,124, and the study is vendor co-authored; Liff et al., Journal of Applied Psychology, 2024).

See the evidence →
Pillar B

Score it by machine, verify it by human.

The algorithm scores consistently (beats holistic judgment, Kuncel 2013); the recruiter verifies authenticity, handles exceptions, watches adverse impact, owns the call. Overriding a valid test hires worse (Hoffman 2018); averaging human+AI often loses (Vaccaro 2024). So we divide labour, not average.

See the evidence →
Pillar C

Neither extreme works.

Gut is weak (unstructured interviews, rho = .19) and biased (Kline, Rose and Walters, 2022). Autopilot fails differently: 85.1% white-name preference across roughly three million comparisons (Wilson and Caliskan, 2024). Nobody has measured the two against each other. Regulators in Ontario, New York City, Illinois, California, Texas and the European Union now impose disclosure duties — a sign of perceived risk, not evidence that any given tool is biased.

See the evidence →
Pillar D

What AI recommends, and who checks it

Businesses are now recommended, or not recommended, by systems nobody audits. This pillar covers what is known about how those systems choose, what is not known, and where the claims being sold outrun the evidence. Barnes Signal Advisor sells work in this area. The entries are marked accordingly.

See the evidence →

Library integrity
entries
26
complicate our case
4
primary-sourced
24 / 26
peer-reviewed
16
Pillar
Confidence
Cuts

ENTRY D5· 2026industryconfidence: strongneutral

Google says optimising for its AI features is the same work as optimising for search.

Google's published guidance states there are no additional requirements to appear in its AI features, and directs site owners to conventional search optimisation.

ENTRY D7· 2026industryconfidence: contestedcuts against us

Nobody has shown that this work produces customers. Including us.

The one study that separated intervention from platform growth found most of the apparent gain was the platform growing, and its own conservative test was not conclusive.

ENTRY D4· 2026industryconfidence: moderateneutral

Adding structured data did not move citations in AI answers.

Across 1,885 pages that added structured markup, matched against roughly 4,000 controls, the change in citations was near zero and slightly negative in one engine.

ENTRY D2· 2026peer-reviewedconfidence: moderatesupports us

Visible prices and current dates decide whether a page gets cited. Formatting does not.

Across 252,000 controlled trials on six language models, four factors predicted citation almost without exception: topic relevance, an explicit price, a recent timestamp, and position in the list.

ENTRY D1· 2025peer-reviewedconfidence: strongsupports us

The strongest independent test found that being findable beats being clever.

Across question-answering and product recommendation, most published techniques for influencing AI answers were ineffective or lowered ranking. Ordinary search optimisation performed significantly better.

ENTRY D6· 2025industryconfidence: strongneutral

When a summary appears, people click a link inside it about one per cent of the time.

Browser-tracked data from 900 adults found users clicked a traditional result on 8% of visits where an AI summary appeared, against 15% where none did. Clicks on links inside the summary occurred on 1% of visits.

ENTRY A5· 2025regulatoryconfidence: moderatecuts against us

Why this matters most on the frontline: the highest churn of any sector — and a smaller per-hire cost than the industry claims.

Accommodation and food services has the highest quit and separation rates in the economy. The cost of each individual hire, however, is far lower than the figures usually quoted at employers.

ENTRY C6· 2025analystconfidence: moderatesupports us

Unsupervised, remote AI assessment is increasingly gamed.

Identity fraud and AI-assisted cheating in remote/automated hiring are rising fast enough that analysts expect a quarter of applicant profiles to be fake by 2028.

ENTRY C4· 2024peer-reviewedconfidence: strongsupports us

Off-the-shelf AI resume screeners preferred white and male names — overwhelmingly.

A peer-reviewed audit of language-model resume rankers found large racial, gender and intersectional bias.

ENTRY B3· 2024peer-reviewedconfidence: strongcuts against us

Bolting a human onto an AI doesn't automatically help — so we split the work, not average it.

A meta-analysis of human-AI teams found combinations often performed worse than the better of human-alone or AI-alone on decision tasks.

ENTRY D3· 2024peer-reviewedconfidence: moderatesupports us

Quotations and statistics change how a page is cited — once the page has already been found.

Adding quotations to a source raised its share of a generated answer by about 41% relative. Keyword stuffing lowered it.

ENTRY C5· 2024newsconfidence: moderatesupports us

A widely used generative model ranked resumes with racial and gender skew.

A reproducible investigation found GPT ranked identical resumes differently by the demographic implied by names, enough to fail standard discrimination benchmarks.

ENTRY A2· 2024peer-reviewedconfidence: moderatesupports us

An AI-scored work sample can measure competencies reliably.

A large study finds automated, AI-scored assessments are highly reliable and predict performance at roughly the level of a human structured interview.

ENTRY C3· 2023regulatoryconfidence: strongsupports us

Automated hiring is now regulated across multiple jurisdictions, not one city.

Bias-audit, disclosure and recordkeeping duties for automated hiring tools are in force in several US states, Canada and the EU, and an automated screen has already been penalised.

ENTRY B5· 2023peer-reviewedconfidence: moderatesupports us

Candidates trust a process more when a human is in it.

Applicants rate fully automated screening as significantly less fair than a human or human-assisted process — regardless of the outcome.

ENTRY C1· 2022peer-reviewedconfidence: strongsupports us

Human screening is biased — and it's concentrated in specific firms.

A field experiment sending 83,000 applications found systematic name-based discrimination in callbacks, driven by a minority of employers.

ENTRY A1· 2022peer-reviewedconfidence: strongsupports us

Doing the task predicts performance. Judging or describing it barely does.

On the corrected modern hierarchy, structured and demonstrated methods top the list; talk-based methods sit near the bottom.

ENTRY A7· 2021peer-reviewedconfidence: moderateneutral

Most of what turnover costs you has already happened by the time someone resigns.

A causal study of retail turnover found 63% of the productivity loss occurs before the departing worker gives notice, not after they leave.

ENTRY B6· 2020industryconfidence: moderatesupports us

Algorithmic screening beat human screeners on quality and diversity.

At one large firm, a machine-learning screen outperformed human resume reviewers on interview success and productivity while selecting more non-traditional candidates.

ENTRY C2· 2018newsconfidence: moderatesupports us

Amazon built an AI recruiter, found it biased against women, and scrapped it.

An internal AI resume-ranking tool taught itself to penalise women and was abandoned.

ENTRY B2· 2018peer-reviewedconfidence: strongsupports us

A valid test improves hires. Overriding it makes them worse.

Introducing a hiring test raised retention; when managers used discretion to override the test, those hires did worse — with no productivity offset.

ENTRY B4· 2018peer-reviewedconfidence: strongsupports us

People distrust algorithms — but will use one they can nudge.

People abandon a superior algorithm after seeing it err, yet will adopt it if allowed even a small adjustment — the case for a human in the loop.

ENTRY B1· 2013peer-reviewedconfidence: strongcuts against us

Let the algorithm do the scoring. Formulas beat human judgment at combining evidence.

Combining candidate data by formula predicts performance markedly better than an expert combining the same data in their head.

ENTRY A4· 2011peer-reviewedconfidence: strongneutral

Showing people the real job lowers turnover — modestly, and we'll say so.

A realistic preview of the actual job reduces quitting, but the effect is small; its driver is perceived honesty.

ENTRY A6· 2008peer-reviewedconfidence: moderatesupports us

The closest study to what we actually do: a job test in retail raised how long people stayed.

Introducing a job test across a national retail chain raised the median tenure of new hires by about 10%, with no adverse effect on minority hiring.

ENTRY A3· 1968peer-reviewedconfidence: strongsupports us

The oldest rule in hiring science: sample the work, don't infer it.

The best predictor of a behaviour is a sample of that behaviour, not a "sign" (a credential or trait) you infer it from.