Did the Fix Work? Measure Product Changes with Reviews
How to tell if a product fix reduced complaints: pick before and after windows, account for review lag and old stock, and test the difference.
Here is a scene that plays out in product teams every quarter. Reviews complain that a bottle leaks. The team traces it to a gasket, changes the supplier, and ships the new version. Two months later someone asks whether the leak complaints went away. Nobody knows. Someone scrolls the recent reviews, sees one leak complaint, and says "still happening". Someone else sees three happy reviews and says "fixed". The fix is either quietly declared a success or quietly forgotten.
Customer reviews can answer the question properly, if you set the measurement up correctly. This guide covers how: what to measure, which time windows to compare, the two traps that fool almost everyone, and a simple test to separate a real change from noise.
Key takeaways
- Measure issue share (complaining reviews ÷ all reviews), never raw counts.
- Start the "after" window when the new version reaches buyers, not when production changed.
- Expect lag: buyers review days to weeks after delivery, and old stock keeps selling.
- Use a simple two-proportion test before declaring victory. Halving a small number is often noise.
- Every fix ends with one of four verdicts: improved, no change, worse, or not enough data yet.
What to measure
Pick one issue per fix, defined the same way it was when you decided to act, for example "leaks from the bottom". Then measure its share of reviews:
Issue share = reviews mentioning the issue ÷ all reviews in the window.
Do not use the raw number of complaints. Review volume follows sales, and sales move with seasons, promotions and stock-outs. If you ran a sale after the fix, you will get more complaints of every kind simply because you sold more units. Share removes that effect.
Do not use the star rating either. The average includes every review the product ever received and moves far too slowly to reflect a single fix. It is also affected by everything else happening to the product at the same time.
Choosing the before window
The before window is the baseline: how often buyers mentioned the issue under the old version.
- Use the period right before the change, not the all-time history. Products drift, and a two-year-old baseline describes a different product.
- Make it long enough to hold real volume. Aim for at least 100 to 200 reviews. For a high-volume product that might be four weeks; for a slow seller it might be three months.
- Exclude anything unusual. If the before window contains a known bad batch that was already pulled, your baseline will be inflated and any change will look like a win.
Choosing the after window, and the two traps
This is where most measurements go wrong.
Trap 1: Review lag
Buyers do not review on delivery day. They use the product for a while, and many problems (like a seal that fails after daily use) take weeks to appear. So reviews written in the first days after your change still mostly describe the old experience, or not enough usage to show the problem at all.
Trap 2: Old stock in the channel
The day your factory switches to the new gasket is not the day buyers start receiving it. Old units are sitting in your warehouse, in marketplace fulfillment centers and possibly with retailers. Depending on your inventory, the old version can keep shipping for weeks or months.
Put together, the after window should:
- Start when the new version is actually reaching buyers. Estimate when old inventory sold through, or better, track it with a lot code, a packaging change, or a new variant.
- Allow for usage time. If the issue appears after two weeks of use, reviews mentioning it will lag by at least that long.
- Run long enough to collect a comparable number of reviews. Thirty days is a sensible default for most consumer products; slow sellers need longer.
If you cannot tell which version a buyer received, treat the first few weeks after the change as a mixed period and leave it out of both windows.
Testing the difference
Suppose the before window has 250 reviews, 46 of which mention leaking. That is an issue share of 18.4%. The question is whether the after window is lower by more than chance would explain.
A two-proportion z-test is simple enough for a spreadsheet:
- p1 = before share, p2 = after share.
- p = pooled share = (complaints before + complaints after) ÷ (reviews before + reviews after).
- standard error = sqrt( p × (1 − p) × (1/n1 + 1/n2) ), where n1 and n2 are the review counts.
- z = (p1 − p2) ÷ standard error.
As a rule of thumb, z above about 2 means the drop is unlikely to be chance (roughly 95% confidence). Here are three outcomes from the same baseline:
| Scenario | Before | After | z | Verdict |
|---|---|---|---|---|
| A | 46 / 250 = 18.4% | 19 / 210 = 9.0% | 2.87 | Improved |
| B | 46 / 250 = 18.4% | 38 / 240 = 15.8% | 0.75 | No change |
| C | 9 / 48 = 18.8% | 4 / 41 = 9.8% | 1.20 | Not enough data |
Hypothetical numbers for illustration. Scenario C is the dangerous one. The share appears to have halved, exactly like scenario A, and in a meeting that would be celebrated. With fewer than 50 reviews per window, the difference is well within what chance produces. The right call is to keep watching, not to close the issue.
You do not need to be a statistician to use this. The point is discipline: write down the windows and the threshold before you look at the after numbers, so the verdict is not whatever the team hoped for.
Four verdicts
Every fix should end with one explicit verdict, recorded next to the decision that triggered it:
- Improved. The share dropped and the test clears your threshold. Close the issue, and tell the team. Proven wins are rare enough that they deserve to be visible.
- No change. Enough data, no meaningful drop. The fix did not address the cause, or the cause was misdiagnosed. Go back to the reviews and look for the trigger again.
- Worse. The share went up. It happens: a new supplier fixes one thing and breaks another. Catching this in weeks instead of quarters is the main reason to measure at all.
- Not enough data. Extend the window. Set a date to check again rather than letting the question drift.
Watch for confounders
Even with good windows, other things change at the same time as your fix:
- Seasonality. Products used outdoors, or bought as gifts, attract different buyers at different times of year.
- Buyer mix. A deep discount brings in buyers with different expectations. Compare windows with similar pricing when you can.
- Listing changes. A new main image or clearer instructions can reduce complaints on their own. If you changed the listing and the product together, you cannot attribute the result to either one alone.
- Review programs. A batch of Vine reviews or a review request campaign changes who is reviewing. Tag them and check whether the result holds without them.
- New issues masking old ones. Look at the issue you fixed, but also scan for anything newly rising in the after window.
Make it routine
The teams that get value from this do not run it as a special project. They attach a measurement to every fix at the moment they decide on it: the issue, the baseline share, the ship date, the expected date of the verdict. Then they review open verdicts in the same meeting where they review new issues.
That is the loop Reviewly is built around. When you link an improvement to an issue, Reviewly records the before share, waits for the after window to fill with new reviews, and reports a verdict at 30 days, with the reviews from both windows one click away. It works on reviews collected from Amazon, Walmart, Home Depot, Wayfair and Lowe's. If you are new to the underlying analysis, start with how to analyze Amazon reviews. Plans start at $9 a month; see pricing.
FAQ
How long should I wait after a fix before measuring?
Start the after window when the new version is reaching buyers, then give it about 30 days, or longer for slow sellers and for issues that take weeks of use to appear.
What if I can't tell which version a buyer received?
Leave out a transition period after the change, roughly as long as it takes old inventory to sell through. If you can, mark the new version visibly, with a lot code, packaging change or new variant, so future measurements are cleaner.
Can I measure a listing change the same way?
Yes. Listing changes take effect immediately for new buyers, so the old-stock problem disappears, but usage lag still applies. Measure the specific expectation-related issue the change was meant to address.
What threshold should we use?
A z of about 2 (roughly 95% confidence) is a reasonable default. More important than the exact number is agreeing on it before you look at the after data.