Post-Purchase Survey Questions That Actually Explain Returns (2026)
TL;DR
Most post-purchase survey questions fail because they ask the customer to classify themselves into a category the merchant wrote in advance — "Didn't fit," "Not as described," "Changed my mind" — which produces reason codes, not reasons. The National Retail Federation and Happy Returns estimate that 15.8% of 2025 US retail sales, roughly $849.9 billion of merchandise, came back as returns, and 19.3% of online sales specifically. A dropdown tells you 34% of returns were tagged "Didn't fit"; it cannot tell you whether the size chart is wrong, the model photo is misleading, or the fabric has no stretch — three problems with three different owners. The fix is not more questions. It is pairing every question with a follow-up probe that converts a category into a cause, then asking at the moment the customer actually knows the answer. This guide is a copy-ready bank of post-purchase survey questions organised by what you are trying to learn — expectation gap, fit and sizing, delivery and packaging, quality and durability, and intent to repurchase — with the failing version of each question, why it fails, and the probing replacement. Nielsen Norman Group's research on survey design is the reason this works: closed questions "eliminate surprises: what you expect is what you get."
What makes a post-purchase survey question useful?
A post-purchase survey question is useful when its answer names something a specific team can change this quarter. That is a higher bar than it sounds, and almost no question on a standard returns form clears it.
Test any question you are about to ship against three criteria:
- Ownership. Can you read the answer and know whose backlog it belongs in — merchandising, photography, packaging, fulfilment, or product? "Didn't fit" fails this test. "The waist matched the size chart but the rise was two inches shorter than the photo suggested" passes it, and lands on the copy and photography team.
- Specificity of the referent. Does the answer point at a thing? A page, a photo, a seam, a courier, a size chart row. Sentiment without a referent ("disappointed") is unactionable no matter how strongly felt.
- Recency of the memory. Nielsen Norman Group's ten best practices for writing survey questions advise asking about specific, recent memories rather than predictions or generalities. A question about what the customer expected when they clicked Buy is answerable for about a week; a question about whether they will buy again is a prediction and is worth roughly what predictions usually are.
The structural trap is that closed questions are genuinely better for measurement — you cannot trend "waist matched, rise didn't" across 40,000 returns — while open questions are the only ones that produce the cause. NN/g's guidance on open-ended versus closed questions resolves this with the funnel technique: begin broad and open, then narrow to closed. Most returns flows do the exact opposite. They lead with the dropdown and, if you are lucky, append an optional "Anything else?" text box that 6% of people fill in with "n/a".
Every question bank below is written for the funnel. The category question stays — you need it for the operational routing that your returns platform performs. The probe is what makes the category mean something. If you want the same logic applied to the cancellation moment instead of the return moment, our guide to customer churn survey questions that surface why customers really leave is the sibling to this one; the structure is identical and the content does not overlap.
The classification trap: why reason dropdowns produce unusable data
Reason dropdowns produce unusable data because the customer is not answering your question — they are completing a task. The task is "get my money back," and the dropdown is an obstacle between them and it.
That single fact explains almost every pathology in returns data:
- Satisficing. The customer picks the first option that is not obviously wrong. Options at the top of the list are structurally over-reported. If "Didn't fit" is item one and "Fabric felt cheap" is item nine, your data will say fit and mean quality.
- The escape hatch effect. "Other / Changed my mind" absorbs everything the list failed to anticipate. NN/g recommends always offering an opt-out to avoid bad data, which is correct — but an opt-out that collects 20% of your volume is not a safety valve, it is your largest and least legible reason category.
- Blame avoidance. Customers who suspect a reason might cost them the refund, the free return label, or their standing with the brand will not select it. Quality complaints get laundered into "changed my mind."
- Category collapse. One code covers several unrelated causes. "Not as described" is simultaneously a photography problem, a copy problem, a colour-calibration problem, and a third-party-listing problem.
- The company wrote the list. This is the root cause of the other four. The taxonomy encodes what the merchant already believes about why things come back, so the data can only ever confirm existing hypotheses. Nothing new can enter the dataset.
This is the gap running through the whole post-purchase software category. Returns platforms — Loop Returns, Narvar, ReturnGO, AfterShip and the rest — are genuinely good at what they are for: portals, labels, exchanges, refunds, and routing. Their "return reason analytics" is built on the dropdown described above, which is why we ranked the category by explanatory depth in the returns management software comparison for 2026 and again in the post-purchase experience platform ranking. Review platforms have the same limitation from the other direction — the review is whatever the customer volunteered, unprobed, which is the theme of our look at product review platforms for DTC brands beyond star ratings.
To be precise about the distinction, because it recurs throughout this guide: a reason code is a label the customer selected from your list. A reason is a causal account, in the customer's own words, of what they expected, what arrived, and where the two diverged. Your returns platform needs the code. Your merchandising team needs the reason.
Post-purchase survey questions: the expectation gap
Expectation-gap questions establish what the customer believed they were buying, which is the only baseline against which "not as described" means anything. Ask these before any question about the product itself — the gap, not the product, is what generated the return.
Ask these:
- "Before it arrived, what were you picturing? Describe it as if you were telling a friend." Probe: "Where did that picture come from — a specific photo, the description, a review, or somewhere else?" This is the question that converts a vague complaint into a URL and an asset ID.
- "When you opened the box, what was the first thing that was different from what you expected?" Probe: "Was that difference a dealbreaker on its own, or did it become one once you'd tried it?" Distinguishes a hard defect from an accumulation of small disappointments.
- "What did you check on the product page before ordering?" Probe: "Was there anything you looked for and couldn't find?" Baymard Institute's product-page research finds that up to 62% of leading ecommerce sites have mediocre or worse product page UX, and missing specification detail is one of the recurring failures — this question finds your specific instance of it.
- "On a 1–5 scale, how close was the item to what you expected?" (keep the scale — it is your trendable metric) Probe, asked immediately after the score: "What would have had to be different for that to be a 5?" A 2 with no probe is noise. A 2 plus "the green was much darker in person" is a colour-calibration ticket.
- "Did anything about the listing make the decision feel risky, and you ordered anyway?" Probe: "What made you go ahead despite that?" Surfaces the ambiguity your page tolerated. Related: Baymard's abandonment data attributes 13% of cart abandonment to an unsatisfactory returns policy, so page-level uncertainty costs you before the order as well as after it.
Replace this question
- The common question: "Was the item as described? Yes / No"
- Why it fails: It is a verdict, not evidence. "No" is true of a wrong colour, a wrong size, a wrong material, and a wrong quantity — four different teams, one indistinguishable data point. It also invites blame avoidance, because "No" reads as an accusation.
- The replacement: "What was different from what you expected?" followed by the probe "Where did that expectation come from?" Same answer volume, and now every response carries a referent.
Post-purchase survey questions: fit and sizing
Fit questions must separate the garment from the size chart from the photography, because "didn't fit" collapses all three and the fix differs completely for each. This is the single highest-value question set for apparel, footwear, and anything with a size dimension.
Ask these:
- "What size did you order, and what size do you normally wear in this category?" Probe: "Which brand were you comparing to when you picked your usual size?" Cross-brand sizing drift is the most common invisible cause of apparel returns, and this pair of answers quantifies it.
- "Where exactly did it not fit?" (offer body-region options and keep a text field open) Probe: "Was it too tight, too loose, or the wrong proportion in that area?" "Too small overall" is a grading problem; "shoulders fine, sleeves four inches long" is a pattern problem. Different factory conversation.
- "Did you use the size guide?" Probe (if yes): "Did the item match the measurements it listed?" This is the fork in the road. Matched-but-wrong means the chart is fine and the fit model or the description is wrong. Didn't-match means you have a manufacturing or spec-accuracy problem.
- "How did the fabric behave compared to what you expected — stretch, weight, drape?" Probe: "Did that change how it fit, or just how it felt?" Fabric masquerades as sizing constantly. Non-stretch denim in a cut designed for stretch reads to the customer as "runs small."
- "If we'd had the size that would have worked, would you have kept it?" Probe: "What would have made you confident enough to order that size instead?" Separates recoverable exchanges from genuine rejections — the difference between a merchandising fix and an inventory fix.
Replace this question
- The common question: "Reason for return: Too small / Too large / Didn't fit"
- Why it fails: It reports the symptom the customer noticed, not the mechanism. "Too small" from a customer who ordered their usual size, used your chart, and got an item matching the chart is a fit-model problem. "Too small" from a customer who guessed is a size-guide-discoverability problem. Identical code, opposite fixes.
- The replacement: Keep "Too small / Too large" for operational routing, then always ask "What size do you normally wear, and did you use our size guide?" as the immediate follow-up. Two extra seconds; a completely different dataset.
Post-purchase survey questions: delivery and packaging
Delivery questions are worth asking separately because delivery-caused returns are misattributed to the product more often than any other category. A crushed box becomes "poor quality" in your dashboard, and merchandising spends a quarter fixing a product that was fine.
Ask these:
- "How did the package arrive — condition, timing, anything notable?" Probe: "Did the condition it arrived in change how you felt about the item itself?" This is the misattribution detector. When the answer is yes, you have found a return your product data is currently blaming on the product.
- "Was the delivery date what you expected when you ordered?" Probe (if no): "What did you expect, and where did you see that?" Checkout promise, confirmation email, and tracking page frequently disagree; the customer remembers whichever they read first.
- "Did you need the item by a particular date?" Probe: "Did the timing affect whether you kept it?" Occasion-driven returns are a distinct segment with a distinct fix — clearer cut-off messaging, not a better product.
- "How was unboxing?" Probe: "Was anything harder than it should have been, or unclear about what to do next?" Missing assembly instructions and unlabelled parts read to the customer as defects.
- "If you've started a return, how has that process been so far?" Probe: "What was the most annoying part?" Returns-experience quality predicts repurchase independently of the return itself, which is why we treat it as a retention lever in the guide to turning one-time buyers into repeat customers.
Replace this question
- The common question: "Rate your delivery experience: 1–5"
- Why it fails: It produces a number that moves for reasons you cannot see and cannot allocate. A 3 could be a late parcel, a rude driver, a damaged box, or a tracking page that never updated. Averaged across a month, it detects nothing and diagnoses less.
- The replacement: Keep the score for trending, and attach one probe: "What's the main thing that kept it from a 5?" A CSAT number with a mandatory cause field is a different instrument from a CSAT number alone — the pattern behind our AI CSAT interview template.
Post-purchase survey questions: quality and durability
Quality questions have to establish when the problem appeared, because a fault out of the box and a fault after three washes point at different stages of your supply chain. Timing is the diagnostic variable that quality surveys almost always omit.
Ask these:
- "When did you notice the problem — immediately, after first use, or after a while?" Probe: "What were you doing when you noticed?" Out-of-box means inspection or transit. First-use means design or spec. Later means materials or durability testing.
- "Which part of it failed, and what did failure look like?" Probe: "Had you used it in a way you'd expect it to handle?" Establishes whether this is a defect, a misuse pattern you should document better, or a durability claim you should stop making.
- "How does it compare to a similar item you already own?" Probe: "What does that one do better?" The comparison item is competitive intelligence you cannot buy. Ask what brand it is.
- "Was there anything on the product page that led you to expect different quality?" Probe: "Which part — the material description, the price, the photos, the reviews?" Price is an implicit quality claim, and premium pricing sets a durability expectation your spec sheet never made.
- "Would you have kept it at a lower price?" Probe: "What price would have felt right for what arrived?" Distinguishes a genuine quality failure from a value-perception mismatch, which is a positioning fix rather than a sourcing one.
Replace this question
- The common question: "Was the item defective or damaged? Yes / No"
- Why it fails: It merges two causes with different owners — "defective" is a supplier and QA problem, "damaged" is a packaging and carrier problem — and it offers no timing dimension at all. It also makes the customer render a technical judgment they are not equipped to make, so genuine material failures get logged as "changed my mind."
- The replacement: "Tell us what went wrong and when you noticed it." One open question with a timing anchor replaces the whole branch, and the coding is straightforward afterwards. Our roundup of thematic analysis software compared by what it can actually code covers the tooling side of turning those answers into a countable taxonomy.
Post-purchase survey questions: intent to repurchase
Repurchase questions are only worth asking if you ask about behaviour and constraints rather than intentions, because stated intent to buy again is weakly related to buying again. Ask what would have to be true, not what they predict they will do.
Ask these:
- "Are you replacing this with something else, or dropping the purchase entirely?" Probe: "What are you looking at instead?" A named competitor is a direct answer to "why did we lose this order," and it is far more reliable than any switching-intent scale.
- "What would we have needed to get right for you to keep this?" Probe: "Is that something you'd have noticed on the page, or only once it arrived?" Splits your fixes into pre-purchase (page, copy, photography) and post-purchase (product, packaging) buckets, which is exactly the split your roadmap needs.
- "Would you order from us again? What would you want to be different next time?" Probe: "What would make you confident enough to try the same category again versus switching to something else?" Category-level confidence is the recoverable asset here.
- "Have you bought from us before?" Probe (if yes): "How did this order compare to the last one?" Repeat customers returning for the first time are your most informative respondents and your highest-value churn risk — the connection we draw in the ecommerce customer lifetime value guide.
- "Is there anything you'd want us to pass on to the team that made this?" Probe: "What would you want them to do about it?" A surprisingly productive closer. Framing the recipient as a person rather than a form changes what people are willing to say.
Replace this question
- The common question: "How likely are you to recommend us? 0–10"
- Why it fails: NPS at the return moment measures the customer's mood about the refund process, not the product decision, and it delivers no cause. A high satisfaction score is also perfectly compatible with never buying again — the disconnect examined in why satisfied customers still leave.
- The replacement: Keep the score if your reporting depends on it, but make the follow-up mandatory and specific: "What's the main thing driving that number?" then probe once more on whatever they name. That is the whole design of our NPS follow-up questions playbook.
Follow-up probes that turn a category into a cause
A follow-up probe works by asking the customer to explain the word they just used, which is what converts a label into a mechanism. NN/g's user-research guidance recommends exactly this move — "Tell me more about that," "What do you mean by that?" — and it is the single technique missing from nearly every post-purchase feedback question set on the market.
Five probe patterns cover most of what you will need:
Two rules make these work in practice. First, probe once, not three times — a single well-chosen follow-up captures most of the available information, and each additional one costs you completions. Second, the probe has to be conditional on the answer, which is precisely what a static form cannot do. A form can show a text box; it cannot read "too expensive" and ask "compared to what?" This is the mechanical reason post-purchase surveys plateau at reason codes, and it is the same limitation that makes on-site polls and exit surveys shallow — covered in the on-site survey tool ranking and in why multi-step forms leak.
An AI interviewer closes that gap by branching on the actual words the customer used. Perspective AI's AI interviewer agent runs the funnel automatically: the category question for your operational routing, then the conditional probe, then a second probe when the first answer is still vague. It does not process the return, issue the refund, or print the label — your returns platform keeps doing that. It replaces the reason-code step inside the flow with a two-minute conversation, and hands back coded themes plus verbatim quotes. For a return in progress where the customer is already frustrated, the advocate agent handles the same conversation with recovery framing rather than research framing.
When to ask: timing by order type
Post-purchase survey timing should be anchored to the moment the customer has formed the judgment you want to measure, which is different for every product category. Sending everything at "delivered + 3 days" is the most common and most costly default in the category.
Three timing rules matter more than the specific windows:
- Ask at return initiation, not after the refund. The customer is in your flow, motivated to be heard, and the memory is intact. Post-refund, the incentive to explain anything drops to nearly zero.
- Anchor to the event, not the order date. Delivery-triggered timing is the minimum. Ask about the return experience only after the return actually completes.
- Never batch a monthly send. A survey four weeks after delivery is asking for a reconstruction, not a memory, and NN/g's survey guidance is explicit that recent, specific recall beats generalisation. Continuous triggered sampling also gives you a moving read on whether last month's page fix worked, which is the operating model in the ecommerce customer experience guide.
Response rate: what to expect and how to lift it
Post-purchase survey response rates vary enormously by trigger moment, and the trigger matters far more than the question wording. Treat the following as planning benchmarks to validate against your own baseline rather than industry constants, because they move with brand affinity, order value, and category.
Typical planning ranges by moment:
- In-flow at return initiation: the highest by a wide margin, because the customer is already in your interface and expects to be asked something. Anything that requires leaving the flow costs you most of it.
- Email, 1–3 days post-delivery: mid-single to low-double digits, heavily dependent on subject line and sender reputation.
- Email, 2+ weeks post-delivery: low single digits, and the responses skew to the delighted and the furious.
- SMS post-delivery: higher open rates than email but much lower tolerance for length — one question and one probe, maximum.
Six levers that actually move the number:
- Move the survey into the return flow. This is the largest single lift available and it costs nothing in incentives. The customer is already there.
- Open with one question, not a grid. A visible matrix of rating scales is the most reliable abandonment trigger in survey design.
- Make every question optional. NN/g's best practices are direct about this: forcing answers produces either bad data or drop-offs.
- Promise a length you keep. "Two questions" that turns into nine trains customers to ignore the next request.
- Ask conversationally rather than in a matrix. A conversation that adapts finishes more often than a form of equivalent length, because the questions the customer has nothing to say about never get asked.
- Close the loop visibly. Tell customers what changed. "We rewrote the size guide because 300 of you told us the rise ran short" is the cheapest response-rate intervention there is, and it feeds the review and repurchase surfaces discussed in the retail customer experience software ranking.
One thing not to do: raise your response rate with a discount code. Incentives at the return moment recruit the incentive-motivated rather than the informative, and the NRF and Happy Returns landscape report notes that 9% of returns show fraud indicators — paying for return-flow survey completions is not a neutral intervention in that context.
For analysis, the volume you generate will exceed what a person can read. Coding open-text at scale is a solved problem now, and the tooling landscape is mapped in our comparisons of text analytics for customer feedback and customer sentiment analysis tools ranked by explanatory power. The pairing that matters is quantified themes for the dashboard plus the verbatim quote for the team that has to act.
Frequently Asked Questions
How many questions should a post-purchase survey have?
Three to five questions plus conditional probes is the right size for a post-purchase survey. The count that matters is questions asked, not questions available — a conversational flow can hold twelve possible questions and ask any given customer four, because it skips the branches their answers made irrelevant. Fixed forms cannot do that, so they have to be short.
What is the best time to send a post-purchase survey?
The best moment is at return initiation, because the customer is already in your flow with the memory intact. Absent a return, trigger 1–2 days after delivery for apparel, 3–5 days for electronics, and 7–10 days for consumables that need a usage window. Never send on a batched monthly schedule — recall degrades quickly and the responses become reconstructions.
Why do return reason codes not explain returns?
Return reason codes do not explain returns because the customer selects from a list the merchant wrote, while completing a task whose goal is a refund rather than accurate reporting. The result is satisficing toward the top options, an overloaded "Other" bucket, and single codes like "Didn't fit" that collapse size-chart errors, photography problems, and fabric behaviour into one indistinguishable label.
What are good open-ended post-purchase feedback questions?
The strongest open-ended post-purchase feedback questions are "Before it arrived, what were you picturing?", "What was the first thing that was different from what you expected?", "When did you notice the problem and what were you doing?", and "What would we have needed to get right for you to keep this?" Each one asks for a specific recent memory with a referent rather than a judgment or a prediction.
Can a post-purchase survey reduce return rates?
A post-purchase survey reduces return rates only when its answers reach the page, the size guide, or the spec sheet that generated the return. Surveys that stop at reason codes tend not to, because "Didn't fit" does not identify a fix. Surveys that capture causes do, most commonly through size-guide corrections, added specification detail, and photography that shows scale and material accurately.
Should I keep rating scales in a post-purchase survey?
Yes — keep rating scales for trending, but never ship one without a cause probe attached. A 2-out-of-5 expectation score tells you something changed; "the green was much darker in person" tells you what to change. The score is the metric and the probe is the diagnosis, and a scale without a probe gives you a graph you cannot act on.
The post-purchase survey questions worth shipping
The post-purchase survey questions in this guide share one property: each is paired with a follow-up probe that turns a category into a cause. That pairing is the whole difference between a returns dashboard that reports 34% "Didn't fit" every month and a merchandising note that says the rise on one style runs two inches short of what the photography implies. With the National Retail Federation and Happy Returns putting 2025 US returns at roughly $849.9 billion — 19.3% of online sales — the gap between a code and a cause is not a research nicety. It is the difference between measuring a cost and reducing it.
Your returns platform should keep doing what it does: portals, labels, exchanges, refunds, routing. What it cannot do is read "too expensive" and ask "compared to what?" Perspective AI replaces the reason-code step inside that flow with a short AI-led conversation that probes on the customer's own words, then hands back coded themes and verbatim quotes to the people who own the fix. It is built for CX teams who are accountable for the return rate but do not own the survey tool.
Start with the post-purchase survey template and adapt the expectation-gap bank above, or use the return and refund advocate template if you want the conversation to run inside a live return. If you would rather start from your own questions, you can set up a study in a few minutes and put the first hundred returns through it before your next merchandising review. The batch also includes rankings for the adjacent moments if you are auditing the whole lifecycle: checkout abandonment tools ranked by why shoppers left and customer journey analytics tools ranked by the why behind the drop-off.
More articles on AI Customer Interviews & Research
Before You Renew Qualtrics: The Negotiation Levers Buyers Actually Have
AI Customer Interviews & Research · 19 min read
Customer Feedback Data: Retention, Privacy, and the Decisions CX Teams Must Make
AI Customer Interviews & Research · 22 min read
Getting Your Data Out of Qualtrics: Exports, Formats, and What Breaks
AI Customer Interviews & Research · 21 min read
Verbatim Analysis: What to Actually Do With 40,000 Open-Ended Responses
AI Customer Interviews & Research · 19 min read
Customer Sentiment Examples: What Customers' Words Signal
AI Customer Interviews & Research · 13 min read
Customer Sentiment Analysis in 2026: Methods, Tools, and the Conversational Edge
AI Customer Interviews & Research · 12 min read