How to Analyze Amazon Reviews: A Step-by-Step Method
A practical method for analyzing Amazon reviews: build an issue taxonomy, track issue share instead of star ratings, and turn complaints into decisions.
Most teams "analyze" Amazon reviews by scrolling the one-star tab before a meeting and screenshotting the angriest ones. That produces anecdotes, not decisions. The loudest review becomes the roadmap, and nobody can say whether a complaint is growing, shrinking, or limited to one variant.
This guide walks through a method product and quality teams can run with a spreadsheet on day one, and automate later. The goal is simple: for every product you sell, know which problems buyers report, how often, whether that is changing, and what you are going to do about it.
Key takeaways
- Star rating is a lagging summary. Analyze the text, not the average.
- Build a small issue taxonomy (a "symptom tree") and tag every review against it.
- Track issue share: the percentage of reviews in a period that mention an issue. Raw counts lie when review volume changes.
- Segment by product, variant and time before drawing conclusions.
- End every analysis with a decision and an owner, or it was just reading.
Why the star rating is not enough
A 4.3-star product can be quietly accumulating a defect. Ratings move slowly because they average your entire review history, and a single new complaint pattern barely nudges the number for weeks. By the time the average drops, the problem has usually been shipping for a while.
Ratings also hide what is wrong. Two products at 3.9 stars can have completely different problems: one has a packaging issue that a supplier can fix in a month, the other has a design flaw that needs a new mold. The number is identical; the decision is not.
The text of the reviews carries that information. The work is turning thousands of free-form opinions into a structure you can count.
Step 1: Collect the right reviews
Start with the products that matter commercially: your top sellers, your recent launches, and anything with a rating that dipped in the last quarter. Analyzing your whole catalog at once is how these projects stall.
For each product, collect:
- Review text and title. Titles are often the most compressed version of the complaint.
- Star rating and date. You will need dates to measure change over time.
- Variant or ASIN. Amazon often pools reviews across the variations of a parent listing, so a problem with one color or size can drag down the whole family. You need to know which variant a review is about.
- Verified purchase and program labels. Reviews from the Vine program (free product) or unverified accounts behave differently. Keep them, but be able to filter them.
Cover more than one channel if you sell on more than one. The same product on Amazon, Walmart and Home Depot often attracts different buyers with different expectations, and a complaint that is rare on one channel can dominate another. We cover this in multi-channel review analysis, with channel guides for Walmart, Home Depot and Lowe's and Wayfair.
Step 2: Build a symptom tree
A symptom tree is a short, two-level list of the problems buyers actually describe. The top level is broad (Durability, Noise, Setup, Fit, Packaging, Customer Service). The second level is specific enough to act on (Durability → "seal fails after daily use", "hinge cracks").
Rules that keep the tree useful:
- Name issues the way buyers describe them, not the way engineering does. "Leaks from the bottom" is better than "gasket compression failure" because it is what you will search for in the text.
- Keep it small. Fifteen to thirty second-level issues per product category is plenty. If you have a hundred, nobody will tag consistently.
- Separate product problems from everything else. Delivery damage, late shipping and seller communication are real issues, but they belong to a different owner. Put them in their own branch so they do not inflate product defect numbers.
- Leave room for new phrasings. Buyers invent new ways to describe old problems, and occasionally a genuinely new problem appears. Have a holding bucket for "unclear / new" and review it every cycle.
Build the first version by reading a sample: 50 to 100 recent critical reviews (one to three stars) plus 30 or so four-star reviews. Four-star reviews are underrated. They are full of "great product, but..." sentences, which are early warnings from customers who still like you.
Step 3: Tag every review
Now tag each review with zero, one or several issues from the tree. A review can mention both noise and setup difficulty; tag both. A glowing five-star review with no complaint gets no issue tag, and that is fine. It still counts in the denominator.
If you are doing this by hand, two practical tips:
- Tag in batches by product, not by date. Your judgment stays consistent when you are looking at the same product's vocabulary.
- Have a second person tag a random 10% and compare. Where you disagree, the issue definition is ambiguous. Fix the definition, not the person.
Manual tagging works for a few hundred reviews a month. Past that, it becomes the bottleneck and people quietly stop doing it. That is the point where automated tagging pays for itself, as long as the categories stay yours and every tag can be traced back to the review text.
Step 4: Measure issue share, not counts
This is the step most teams skip, and it is the one that makes the analysis trustworthy.
Issue share = reviews mentioning the issue in a period ÷ all reviews in that period.
Counts are misleading because review volume moves with sales. If you ran a big promotion in March, you got more reviews of every kind, including more complaints. A jump from 20 to 40 "leak" reviews sounds alarming; if total reviews went from 200 to 400, the share did not change at all.
| Month | All reviews | "Leaks" reviews | Issue share |
|---|---|---|---|
| February | 200 | 20 | 10% |
| March (promotion) | 400 | 40 | 10% |
| April | 220 | 35 | 15.9% |
Hypothetical numbers for illustration. March looks like the crisis if you watch counts. April is the real one: fewer reviews overall, but a much larger share of them mention leaking.
Pick a period that gives each product enough reviews to be meaningful. For a high-volume product, weekly works. For a product with ten reviews a month, use a rolling 30 or 90 days, and do not react to a single month.
Step 5: Segment before you conclude
Before you declare that "customers hate the lid", cut the data a few ways:
- By variant. Is the complaint spread evenly, or concentrated in one size or color? Concentration usually points to a supplier, a mold or a specific component.
- By time. Did it start on a particular date? Line that up with production batches, supplier changes and listing edits.
- By channel. If home improvement buyers on Home Depot complain about installation and Amazon buyers do not, the product may be fine and the instructions or expectations are the problem.
- By star rating. An issue that appears mostly in four-star reviews is an irritation. The same issue in one-star reviews is a deal-breaker.
Ten minutes of segmentation prevents the most expensive mistake in review analysis: redesigning a whole product line because of one bad batch.
Step 6: Read the evidence, then decide
Numbers tell you where to look. The reviews tell you what to do. For your top three issues by share, read 15 to 20 actual reviews each. You are looking for:
- The trigger. "After two weeks", "when I fill it past the line", "only on the dishwasher's top rack". Triggers turn a vague defect into a reproducible test.
- The expectation gap. Sometimes the product works as designed and the listing promised something else. That is a copy fix, which is much cheaper than a product fix.
- Comparisons. "My old one from Brand X never did this" tells you what buyers measure you against. That is where competitor review analysis comes in.
Then write the decision down: the issue, the evidence, the change, the owner and the date it ships. A product review meeting that ends without that list has produced reading, not analysis.
Step 7: Close the loop
The last step is the one that turns review analysis into a habit: after the fix ships, check whether the issue share actually dropped. That sounds obvious, but it is harder than it looks because of review lag and old inventory still in the channel. We wrote a separate guide on how to measure whether a product fix worked.
Common mistakes
- Treating the average rating as the KPI. It moves too slowly and says nothing about cause.
- Only reading one-star reviews. You miss the early warnings in three- and four-star reviews.
- Mixing product and logistics complaints. Your defect rate looks worse than it is, and the wrong team gets the blame.
- Changing the taxonomy every month. You lose comparability. Add issues when you must; rename rarely.
- Analysis with no owner. If nobody is accountable for the top issue, the same review will be in next quarter's deck.
Doing this at scale
The method above works in a spreadsheet for a handful of products. It breaks when you have dozens of products across several marketplaces and thousands of new reviews a month. That is the problem Reviewly is built for: it collects reviews from Amazon, Walmart, Home Depot, Wayfair and Lowe's into one catalog, files each review into a symptom tree your team can curate, flags issues whose share is rising, and keeps every chart linked to the underlying review text. When you ship a fix, it watches the issue for 30 days and reports whether complaints actually moved. Plans start at $9 a month; see pricing.
FAQ
How many reviews do I need before the analysis means anything?
For spotting which issues exist, 50 to 100 reviews per product is enough. For measuring change between two periods, you want at least 100 to 200 reviews per period, or you will mistake noise for trends. Low-volume products should use longer windows.
Should I include Vine and unverified reviews?
Include them, but tag them so you can filter. Vine reviewers received the product for free and tend to review early, which makes them useful for spotting issues quickly but less representative of paying buyers.
Can I use AI to tag reviews?
Yes, and past a few hundred reviews a month you probably should. The important part is that the categories are yours, new phrasings are surfaced for a human to confirm rather than silently invented, and every tag links back to the review so anyone can check it.
How often should we run this?
Weekly for your top sellers and recent launches, monthly for the rest of the catalog. The cadence matters less than consistency: the same taxonomy, the same metric, every cycle.