Return reason codes are the shortest lie in your analytics. A buyer who ordered a 180 cm sofa for a 170 cm alcove is handed a drop-down with about ten options, none of which say "the listing did not make the size obvious." They pick "changed my mind," you file it under buyer's remorse, and a fixable listing problem disappears from your dashboard forever.
This is not a theory about buyer psychology. It is a data-model problem you can read in the platform documentation.
The full list, as one major platform defines it
Shopify's Admin API has published the exact enum for years, which makes it the cleanest public specimen of what a return reason taxonomy actually contains. Ten values, with the platform's own descriptions:
| Code | Platform description |
|---|---|
| COLOR | The item is returned because the buyer did not like the color. |
| DEFECTIVE | The item is returned because it is damaged or defective. |
| NOT_AS_DESCRIBED | The item is returned because it was not as described. |
| OTHER | The item is returned for another reason. For this value, a return reason note is also provided. |
| SIZE_TOO_LARGE | The item is returned because the size was too large. |
| SIZE_TOO_SMALL | The item is returned because the size was too small. |
| STYLE | The item is returned because the buyer did not like the style. |
| UNKNOWN | The item is returned because of an unknown reason. |
| UNWANTED | The item is returned because the customer changed their mind. |
| WRONG_ITEM | The item is returned because the customer received the wrong one. |
Count the ones that mention size: two. Count the ones that can quietly absorb a size failure: five — NOT_AS_DESCRIBED, OTHER, UNKNOWN, UNWANTED and, for anything with a footprint, STYLE. A taxonomy where the most common physical failure has two dedicated buckets and five hiding places will always under-report that failure.
A return reason code is a buyer-selected label, not a diagnosis. It records what the customer chose to click at the moment they wanted their money back, which is usually the option with the least friction attached to it, not the most accurate one.
The platforms already know the list is too blunt
This is not a complaint from the seller side. On 1 January 2026, in API version 2026-01, Shopify introduced a new type called ReturnReasonDefinition — with id, handle and name fields — to capture more granular, category-specific return reasons, and said it replaces the previous ReturnReason enum with a richer data model that helps provide merchants with better insights in their return analytics. Three fields that used the old enum are now marked deprecated.
Read that as an admission: a flat ten-item list applied to every category at once was never going to tell a furniture seller anything actionable. "Size was too small" means one thing for a t-shirt and something completely different for a wall cabinet that fouled a light switch.
If you are still reading a dashboard built on the old flat enum, your size numbers are the floor, not the figure.
The code you get decides who pays
The taxonomy is not only an analytics artefact. It routes money, and the routing does not follow the underlying cause.
eBay's seller documentation states the rule plainly: if the buyer returns an item because it is damaged, faulty, or did not match the listing description, the seller pays return shipping — even if the seller's policy says they do not offer free returns. If the buyer ordered the wrong item or changed their mind, the buyer pays, unless the seller offers free returns.
Now put a size failure through that gate:
| What the buyer clicks | Who pays return shipping | What actually happened |
|---|---|---|
| Doesn't match description | Seller | Listing understated or omitted a dimension |
| Changed my mind / ordered wrong item | Buyer | Listing understated or omitted a dimension |
Identical root cause, opposite financial outcome, decided by a drop-down. That asymmetry is also why the codes drift: a buyer who wants free return shipping learns quickly which option produces it, and a buyer who does not care picks the first plausible line. Neither is trying to help you.
The scale you are working against
The National Retail Federation's 2025 Retail Returns Landscape, published 15 October 2025, projects total retail returns of 849.9 billion dollars for 2025 and estimates that 19.3% of online sales will be returned. The same report found 9% of all returns are fraudulent, and that 82% of consumers say free returns are an important consideration when shopping online.
Roughly one in five online orders comes back. If even a third of the reasons you receive are mislabelled, you are steering your listing and photography budget with a compass that is off by a quarter turn. Before you argue about whether it matters, put your own numbers into a return cost calculator — the per-unit cost of a returned oversized item is what turns a 3% taxonomy error into a real annual figure. Category-level context is in the e-commerce returns size statistics reference.
Mapping return reason codes to what actually went wrong
This is the table worth keeping. Left column is what your export gives you; middle column is the failure that most often hides behind it in dimension-driven categories such as furniture, lighting, building materials and hardware; right column is the fix that changes next quarter's number.
| Code received | Most common hidden cause | Fix that moves the number |
|---|---|---|
| SIZE_TOO_LARGE / SIZE_TOO_SMALL | Genuine fit failure, already correctly labelled | Publish the measured dimension on the image, not only in the spec table |
| NOT_AS_DESCRIBED | A dimension was present but ambiguous — no reference edge, no unit, or assembled vs packed confusion | Label both assembled and packed dimensions, with units on every number |
| UNWANTED ("changed their mind") | Buyer measured their space after delivery and it did not fit | Add a clearance or footprint diagram to the image stack |
| OTHER | Something specific enough that no code fit — the free-text note is where your real data lives | Read the notes weekly; they are the only unfiltered channel you get |
| STYLE | Proportion, not taste — the piece was the wrong scale for the room | Show the product in a scene with a labelled reference dimension |
| COLOR | Sometimes finish or material, not hue | Show finish options as separate labelled images |
| WRONG_ITEM | Variant confusion between two sizes of the same SKU | Put the size on the variant image itself |
| DEFECTIVE | Occasionally an install failure caused by a clearance the buyer never saw | State required installation clearances alongside the product size |
| UNKNOWN | Data loss at the platform or 3PL layer | Chase it as an integration bug, not a customer insight |
Every row in the middle column is a listing problem, not a product problem. That is the useful conclusion: most of what gets filed as buyer behaviour is a communication defect with a fixed cost.
How to recover your real size-return number
You do not need a new system. You need three passes over data you already have.
- Pull the free-text notes attached to OTHER. Platforms require a note for that value. Read a hundred of them and count how many contain a measurement word — fit, big, small, tall, wide, deep, space, room, doorway, clearance.
- Cross-tab reason against SKU dimensions. Sort SKUs by longest dimension and compare return rates across the top and bottom quartile. If the biggest items return more under every code, the codes are not measuring what they claim to.
- Compare returns against pre-sale questions. The pre-sale message queue is unlabelled and honest. If 40% of questions are about dimensions and only 12% of returns are coded as size, the gap is your undercount.
Do this once and the number usually moves by a factor, not by a few points. Treat return reason codes as one input among three rather than the answer, and the picture stops flattering you. If you want a defensible baseline before and after, the method in how to calculate return rate matters more than the reason breakdown, because a reason split of a wrong denominator is doubly wrong. For where your category sits to begin with, size return rate by category gives the comparison band.
Next steps, in the order that pays
Pick by where your leak actually is, not by what is easiest to schedule.
- If your OTHER and UNWANTED buckets are large, start with the notes audit above. It costs an afternoon and it usually reframes the whole budget conversation.
- If your top-returning SKUs are your largest ones, the problem is almost never the product. Rebuild the image stack for those SKUs first: one dimension-annotated image showing the measured width, depth and height on the edges they describe, plus one clearance or in-room shot. This is the single highest-yield change in dimension-heavy categories, because it puts the number where the buyer is already looking instead of in a table below the fold.
- If you sell variants that differ only by size, put the size onto the variant thumbnail itself. WRONG_ITEM returns on variant SKUs are usually a picking error by the customer, not by you.
- If you need those diagrams at SKU volume, the constraint is production speed, and this is where tooling decides whether the plan survives contact with a catalogue. Software that snaps a measurement to the product's detected edge — so the printed number is the measured one — and exports at each marketplace's required image size makes this a few-minutes-per-SKU job. That accuracy is the whole point: an AI image generator will happily render a confident-looking "60 cm" onto a photo of something that is 54 cm, and a wrong number on an image is a NOT_AS_DESCRIBED return with your name on it.
- If your reason data comes from an old flat enum, check whether your platform now offers category-specific definitions and migrate. Better raw categories beat better dashboards.
FAQ
What are return reason codes?
Return reason codes are the fixed set of options a buyer selects from when starting a return, stored as a machine-readable value on the return record. One widely documented example runs to ten values — color, defective, not as described, other, size too large, size too small, style, unknown, unwanted and wrong item — and they are chosen by the customer, not assigned after inspection.
Why do my return reasons not match what customers tell support?
Because the drop-down and the conversation have different incentives. The drop-down decides who pays return shipping, so buyers gravitate to the option that gets them a free label; the support conversation has no such consequence and is usually more accurate. When the two disagree, trust the conversation and treat the code as a routing decision.
Which return reason means the seller pays for shipping?
On eBay, returns filed as damaged, faulty, or not matching the listing description put return shipping on the seller even when the seller does not offer free returns. Changed-my-mind and ordered-the-wrong-item returns put it on the buyer unless the seller offers free returns. Exact wording varies by marketplace, but the buyer-fault versus seller-fault split is close to universal.
How do I reduce size-related returns if the codes undercount them?
Stop trying to fix the measurement of the problem and fix the cause: make the dimension impossible to miss before purchase. In practice that means a measured, labelled diagram in the image stack rather than a specification table three scrolls down, plus both assembled and packed dimensions where they differ. The undercount stops mattering once the underlying rate falls.
Is "other" worth reading?
It is the most valuable column you have. Platforms require a free-text note with that value, which makes it the only place a buyer describes the failure in their own words. A hundred of those notes will tell you more than a year of aggregate reason charts.
