← XPRIZE Top 100 screen

Methodology · Build with Gemini XPRIZE

How we screened and tested the XPRIZE Top 100

100 listing screens, 14 deeper submission or repository checks, and five hands-on cases. Selection favored minimum effort to a meaningful independent finding.

We screened projects by how cheaply an outsider could reach a meaningful, evidence-backed conclusion about their important claims. This is not a quality ranking, and a Level 5 project may be excellent. The source universe is exactly the 100 projects on the official Build with Gemini XPRIZE Top 100 page; no other list was used.

Screening dimensions

CriterionQuestion we askedFavours
Product accessibilityCan an outsider exercise it today?No login, free, public demo or API
Claim specificityIs the key claim falsifiable?Numbers, mechanisms, named data sources
ReproducibilityCan we run it with our own inputs and repeat?Public modules, repos, APIs
Independent reference dataIs there a source of truth outside the participant?Official datasets, public ledgers, registries
Setup effortWhat stands between us and the first test?Browser or a single command
Founder dependenceHow much needs private data or testimony?Little
Business-claim testabilityAre revenue, customers or usage numbers stated, and can any be checked?Specific numbers, some public trail
AI-role testabilityCan we see what the model decides versus ordinary code?Code, logs, or observable behaviour
Editorial valueWould the result tell a reader something?Tension created by evidence
ConsequenceWould the finding change how someone relies on it?Money, law, safety, purchase decisions

Scores

Validation effort (1–5). 1 = browser, concrete claim, immediate independent test. 2 = signup, small install, or repeated controlled tests. 3 = repo or runtime work, APIs, specialist inputs. 4 = conclusions hinge on private production data or founder cooperation. 5 = physical, clinical or regulated environments, or claims that cannot realistically be reproduced.

Story potential (1–5). 5 = a consequential, falsifiable claim where independent evidence can plausibly confirm, contradict or bound it. 1 = no claim a test could move.

The full screen preserves the dossier’s numeric screening scores. Where a deeper check changed the assessment, the reason column records that later assessment. These are editorial judgments about verification effort and potential, not scores for product quality.

The screening funnel

  1. Listing screen: all 100. Scored from the one-line problem statement on the official page plus the evidence type the domain implies. Live-product and repository fields were marked “Not checked” rather than guessed.
  2. Deeper submission or repository checks: 14. Opened the provisional candidates and additional promising projects. The other 86 remained at listing depth.
  3. Hands-on verification: five. Ran public code, recalculated published results, inspected evidence repositories, fetched live endpoints, and compared claims with official data. The completed cases are Polyfork, Sloane & Pearl, Veritas, Doppelganger and TradePass.

The 14 projects opened were Polyfork, Sloane & Pearl, Veritas by Optume Translations, Doppelganger, TradePass, Ten Minutes Wiser, SpaceSphere, Workopia Hiring Intelligence, Merlin Clips, Kosmo, StockSpade.com, Sarthi Kalyan, Damage Control AI and NoBanks Nearby. Ten Minutes Wiser received a submission check; its lesson audit was not completed and it has no published verification record.

Selection rule: minimum effort to a meaningful, independently supported finding. We did not simply take the five lowest effort scores: StockSpade scored 1 on effort, but a test would mostly show that a web page works. The five selected cases and their rationale show why public modules, code excerpts, response data and official reference documents made these claims comparatively testable.

Different projects received different depths of investigation. The five are not a representative sample of the Top 100. Accessibility of evidence shaped selection; it does not establish which businesses are strongest.

Evidence states

  • Claimed: the participant says it.
  • Observed: sound.fan saw it on a participant-controlled public surface.
  • Reproduced: sound.fan independently reran the behavior or calculation.
  • Corroborated: an independent external source supports it.
  • Not verified: available evidence did not let us independently establish it.

These states apply to individual claims, not entire projects. Participant-controlled code, logs, billing exports and webpages can expose a mechanism or support a reproducible calculation. Reading them does not independently corroborate production deployment, revenue, customer counts or usage. Reproducing a calculation on participant-supplied data does not establish how that data was collected.

“Not verified” does not mean false. Differences are findings to explain, not evidence of wrongdoing. Where an authoritative source directly contradicts a statement, the record identifies the precise statement and source.

Evidence in time

ClassificationMeaningExample
Submission-time claimWhat the participant claimed on Devpost by the 17 August 2026 deadline“204 real customers”
Hackathon-period recordRepository commits, logs and exports dated inside the hackathon periodAn ad-pause log dated 29 July–4 August
Later live observationWhat sound.fan observed on the live product or site on the research date, 16 or 17 SeptemberA pricing endpoint or exam profile
Independent sourceEvidence the participant does not control; its own publication or effective date still mattersOfficial WMT25 data or a PSI candidate bulletin

A live-site difference is not treated as a submission-time contradiction unless the conflicting claim can be shown to have applied then. Git commit dates are set by the author’s machine: they place work in time but are not independent proof. “Independent source” describes control of the evidence, rather than a claim that the source is timeless.

The listing screen and submission checks were conducted on 16 September 2026. TradePass’s live checks were completed on 17 September. The records retain those dates separately from the shared evidence cutoff.

Scope and disclosure

The package uses public evidence. No revenue, customer or usage figure across the five businesses was independently confirmed. Each record names the smallest useful additional artifact that could change a finding; additional evidence goes through the existing public correction mechanism.

sound.fan’s founder entered the same competition; neither entry reached the Top 100. sound.fan sells claim verification to hackathon organizers.