Returns and refunds: where automation should stop

Roughly one in five online orders is returned. Some of that queue is mechanical and should be automated today. Some of it is a judgement call wearing a policy question, and automating it costs more than it saves.

The useful split is not "simple versus complex". It is whether the answer is determined or decided. Is this within the returns window is determined — a date, a policy, a lookup. Should we make an exception is decided, and a machine that decides it is either giving away margin or refusing a customer you wanted to keep.

The size of the queue

Ecommerce return rates run around 20% overall, two to three times the in-store rate, and the spread by category is wide: apparel 20–40%, footwear 17–30%, electronics 8–15%, beauty 4–12%. If you sell clothing, returns are not an edge case in your support queue. They are a second queue the same size as the first.

They are also the conversations most likely to end in a refund, a replacement or a churned customer, which is why the instinct to automate them wholesale is worth resisting.

What should be automated

The mechanical majority. These have one right answer that does not depend on who is asking:

  • What is your returns policy — window, condition, who pays postage.
  • Am I still inside the window — a delivery date and a subtraction.
  • How do I return this — the actual steps, not a link to a page of them.
  • Where is my refund — status, and the honest processing time rather than an optimistic one.
  • Has my return arrived — a tracking lookup on the inbound parcel.

Every one of those is a fact your systems already hold, asked at a moment when the customer wants an answer rather than a conversation. Automating them is straightforwardly good: it is faster than a human at 11pm and it never gets the window wrong.

What should not be

The decisions. An agent should recognise these and hand them over, not attempt them:

  • Anything outside policy. Two weeks past the window with a good reason is a judgement about a customer relationship, and it is exactly the kind of judgement a model will make inconsistently — generous to a persuasive message, firm to a terse one. Inconsistent exception-granting is worse than a firm no, because it is unfair in a way customers compare notes about.
  • Damaged, faulty or wrong item. Often a legal question rather than a policy one, since consumer protection in most markets does not care what your returns page says. It also tends to arrive with photographs and an angry opening line.
  • High value. Pick a threshold and hand everything above it to a person. The cost of getting one of these wrong dwarfs the cost of the handover.
  • Repeat returners. Return abuse is real and expensive, and it needs a person who can see the pattern and decide what to do about it.
  • Anyone who is upset. Not because a model handles anger badly in the abstract, but because someone who is already annoyed and gets a bot has been told what their complaint is worth.

The handover is the product

If half of the returns queue is going to a human, the quality of that transfer is what determines whether automating the other half was a good idea.

A bad handover is a message saying a human will be in touch. The customer repeats everything. The agent reads back through the thread. Nothing was saved and the customer has now explained their problem twice.

A good one arrives with the conversation, the order, and what has already been checked — so the person picks up mid-context and answers the actual question. That is also the difference between automation reducing your team's workload and merely rearranging it.

On WhatsApp the handover is not just good practice. Automation inside the 24-hour window is permitted on the condition that a prompt, clear and direct escalation path exists. A returns bot with no way out of it is outside the terms.

Say no properly

When something genuinely falls outside policy, the agent still has a job, and it is not to soften the refusal until it sounds like a yes. Give the reason, name the specific rule, and offer the next step — a person to speak to, or an alternative that is actually available.

The failure mode to watch for is an agent that says "I'll look into that for you" and then nothing happens, because nothing was ever going to happen. That is not politeness, it is a second complaint scheduled for tomorrow.

Where the line moves

Once you have run this for a quarter, the boundary should shift on evidence rather than ambition. The signal is not the resolution rate — an agent can resolve a lot of conversations badly. It is what customers do next: how many come back within a week, and how many of the automated resolutions ended up in a human's queue anyway.

Those two tell you where the line actually is, as opposed to where you hoped it was.

Common questions

What percentage of online orders get returned?

Around 20% overall, roughly two to three times the in-store rate. It varies widely by category: apparel runs 20 to 40%, footwear 17 to 30%, electronics 8 to 15% and beauty 4 to 12%.

Should an AI agent approve refunds?

It can process a refund that is clearly within policy, where the answer is determined by the rules rather than decided by a person. It should not grant exceptions. A model asked to make judgement calls will make them inconsistently, which is unfair to customers and expensive to you.

What should always go to a human?

Anything outside policy, damaged or faulty items, high-value orders above a threshold you set, customers with a pattern of returns, and anyone who is clearly upset. The cost of getting one of these wrong is far higher than the cost of a handover.

How do I stop the handover annoying the customer?

Pass the conversation, the order and what has already been checked, so the person picks up mid-context. If the customer has to repeat their problem, the automation did not save anyone anything and it cost you the goodwill.

How do I tell whether it is working?

Not by resolution rate on its own, since an agent can resolve conversations badly. Look at how many customers come back within a week about the same thing, and how many automated resolutions ended up with a human anyway. Those show where the line really is.

Published 11 August 2026. Return rate figures are 2026 industry benchmarks and vary considerably by category and market — use your own where you have them.

Automate the mechanical part. Hand over the rest.

Thredo answers the policy and status questions itself, and hands the judgement calls to a person with the whole conversation and the order already attached.

Start free — 50 messages, no card How the handover works