Verification record · Build with Gemini XPRIZE
Doppelganger
The 91.2% result reproduces under the any-model rule; pooling the same responses gives 73.5%. The benchmark cannot measure false alarms.
Claim
Doppelganger offers AI respondents to pilot surveys before researchers pay human participants. Its submission says an ensemble reproduced about 91% of 17 classic behavioral-science effects, “well above what a single model manages.” It reports paying customers, including a repeat buyer, without a count.
What we tested
On 16 September 2026, sound.fan ran the public research repository’s reproduce.py against its published responses, without API keys. We then re-aggregated those same responses with the repository’s own scoring code, inspected the significance and direction rules, and checked whether the benchmark included effects that should come out null. We did not generate fresh model responses or run the commercial product.
Finding
Reproduced: the published matrix and headline match exactly: 15.5 of 17 effects, or 91.2%. The aggregation counts an effect when any one of five models shows it. The repository itself describes this as an optimistic ceiling.
| Aggregation | Effects reproduced | Rate |
|---|---|---|
| Any one of five models (published) | 15.5 / 17 | 91.2% |
| Best single model (Model A) | 13.5 / 17 | 79.4% |
| All five models pooled into one dataset | 12.5 / 17 | 73.5% |
| Majority of five models | 11.5 / 17 | 67.6% |
Four ensemble hits each come from a single model: less-is-more, risk–benefit, omission bias and probability matching. Scoring checks significance (p < 0.05) and direction, not effect size.
What the evidence establishes
The published calculation moved from Claimed to Reproduced on participant-supplied response data. The any-model rule and its caveat were Observed in the hackathon-period repository, whose three commits are dated 10 August. Neither result establishes how fresh runs or the live product would perform.
All 17 experiments have a published human effect. None is a known null. The benchmark tests whether the system finds known effects, but cannot measure how often it reports an effect that is absent. That matters to a product intended to warn researchers that an effect “just wasn’t there” before they pay for humans. Pooling is our approximation of combining results; the product’s actual combination method is not public.
What remains unverified
Not verified: the commercial product’s aggregation method, false-alarm rate, paying customers or repeat purchase, and Qualtrics Sessions API integration. Model names in the research data are redacted as A–E, so the repository cannot establish the submitted claims that Gemini on Vertex AI participates in the ensemble and writes result summaries.
What evidence would change that
A benchmark arm using published failed replications would measure false alarms. Disclosure of the product’s aggregation rule and model identities, or at least Gemini’s column, would connect the research result to the submitted mechanism. Customer and revenue records would be needed to reconcile the business claim.