That verified buyers make better survey respondents than self-reported panelists isn't exactly a hot take — everyone in research nods along. The question nobody puts a number on is how much better: a rounding error you can weight away, or a gap wide enough to change the decision. So we measured it.
We ran the same survey with two groups of a hundred people. One group we knew had bought Pampers — we'd seen it in their order history. The other a panel vendor rounded up because they'd checked a box saying they were moms of kids under three. Same questions, same incentive, same survey platform. Then we let that platform's built-in quality score grade every answer on the same yardstick, and asked the only thing that matters: whose answers can you trust?
It wasn't close.
Survey shown only to respondents with a confirmed, recent Pampers purchase.
Sourced through an audience vendor's demographic filter.
Quality score
Conjointly's authenticity metric
Effort spent
Time + question rushing
Screening
Speeding / low-score flags
Open text
Specificity of written answer
How each cohort was sourced
Ario Verified Audience
The Ario Verified Audience is made up of people who have connected one or more of their retail accounts to Ario. For this study, the survey went only to members whose connected purchase history showed a prior Pampers purchase.
Comparison Audience
The comparison group came from a separate audience vendor, selected through its demographic and self-reported targeting filters: people who reported being female, being parents, and having a child between 0 and 36 months old.
What verified buyers did differently
On the platform's quality score — 0 to 1, built from response time, question rushing, and internal consistency — the verified buyers averaged 0.65. The self-reported panelists averaged 0.49.
But the average buries the point. Look at the floor. Not one verified buyer landed below 0.5. Among the self-reported panelists, 35% did — nearly one in three sitting in the range the platform flags as low-confidence. A self-reported panel isn't uniformly so-so; it's good data cut with a third you'd throw out if you could see it. Verification lets you see it — by never letting it in.
The timing tells the same story. Verified buyers took about twice as long on the survey overall, and fired off snap answers — under two seconds — about three times less often than the self-reported panelists. They weren't just the right people. They were paying attention, because for once someone was asking them about a product they actually use every day.
What people write when they actually care
The clearest evidence isn't in the scores. It's in the writing. Every respondent got the same open question: what's the one thing Pampers could fix to earn more of your business?
“What is the ONE thing Pampers could improve to earn more of your business?”
- “Tighten the sides a bit so no leak.”
- “Making the back higher because bowel movements go up the back.”
- “More diapers in the package for sizes 6+.”
- “Half sizes!”
- “Better fit for blowouts.”
- “Nothing.”
- “No comments.”
- “Everything.”
- “Duration.”
- “Its anuniwue brand for the entire fmsiky to enjoy and such thigns sorund this entire tjme…”The only gibberish answer in either cohort.
Read the verified buyers' answers. Every line is a product brief — leak points, fit, pack sizes, the gap where a half-size should be. Hand them to R&D on Monday and they'd know what to build. The self-reported panelists wrote "Nothing." "Everything." "Duration." And one answer that dissolves into typos halfway through. Nobody with a real complaint writes "Everything." They tell you exactly where it leaked.
What this means for anyone buying sample
Every gap in this test traces back to one difference in how the two cohorts were built. The comparison audience was qualified on a claim — a self-reported demographic the vendor recorded at face value. The Ario Audience was qualified on a confirmed purchase, read from the retailer's own record before the survey opened.
The two approaches also fix quality at different points. A demographic panel screens after the fact: field the survey, score the responses, then weight or discard the ones that come back weak. Verification screens before it — the purchase is confirmed at the door, so the people who reach the questions already belong in the study. It's the same purchase-verified foundation Ario CoreLens reads from: confirmed transactions rather than self-reported behavior.
Methodology
Two cohorts of 100 respondents each answered an identical survey about the Pampers brand. The first was Ario's verified buyers: respondents who consented to connect their Amazon account, from which Ario reads SKU-level order history directly from the retailer's own record. Only those whose real purchase history showed a confirmed, recent Pampers order were placed in the buyer cohort — no self-report, no receipt upload. The second cohort was sourced through a panel vendor's demographic filter: people who self-reported being mothers of children aged 0–36 months.
Both cohorts were fielded on the same survey platform, Conjointly, and graded by its built-in quality score (0–1) — the same standardized measure applied identically to both groups. That score combines response timing, question rushing (answers submitted in under two seconds), and internal-consistency checks. We compared the cohorts on four lenses: overall quality score, effort spent, screening flags, and the specificity of open-text answers.