Review collection and import from marketplaces: parser development
We build parsers for automatic review collection from marketplaces. A client lost 30% conversion due to outdated reviews — manual collection took hours, data went unrefreshed for weeks. Our team (5+ years experience, 50+ projects) created a solution that collects reviews from Wildberries, Ozon, Yandex.Market, and other platforms, normalizes them, and imports them into your store's database. Time savings — up to 70%. Project payback — 2–3 months. Order parser development for your online store.
Typical scenario: a manager manually copies reviews from 5 marketplaces for 200 products — that's 8 hours a day. Data grows stale, customers see no new reviews for weeks, trust drops. A parser grabs reviews every 2 hours, updating product pages in real time. Result: conversion growth by 15% in the first month.
Technically, each marketplace is its own headache. Wildberries offers a JSON API, but with a limit of 1000 reviews per session. Ozon is an SPA on Nuxt, where data is loaded via GraphQL — we have to emulate a browser with Playwright. Yandex.Market changes its structure every six months, so we use adaptive parsing with a fallback strategy. Our stack: Python for high-load platforms, PHP (Laravel) for integration with your site.
Supported marketplaces and methods
| Platform | Method | Notes |
|---|---|---|
| Wildberries | JSON API | Open API, pagination, up to 1000 reviews per session |
| Ozon | Playwright | SPA, needs authorization, 50x slower than JSON |
| Yandex.Market | Unofficial API | Rate limiting, requires proxy rotation |
| Google Reviews | Places API | Paid, official, up to 5 reviews per request (requires business subscription) |
| Otzovik.com | HTML parsing | CAPTCHA on mass requests — we use solvers |
| iHerb | HTML / JSON API | Structured HTML, easiest |
Method comparison: JSON-API (Wildberries) processes 1000 reviews in 10 seconds, while Selenium on Ozon takes 5 minutes for the same 1000. That's a 30 times difference. For performance, we choose JSON wherever possible.
How the Wildberries parser works
We use asynchronous httpx to iterate through pagination. Example:
Click to expand Python code example
# scraper/reviews/wildberries.py import httpx import asyncio from dataclasses import dataclass from typing import Optional @dataclass class Review: external_id: str product_nm_id: int author: str rating: int text: str pros: Optional[str] cons: Optional[str] date: str photos: list[str] helpful_count: int class WildberriesReviewScraper: REVIEWS_URL = "https://feedbacks2.wb.ru/feedbacks/v2/{nm_id}" def __init__(self): self.client = httpx.AsyncClient( headers={ "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)", "Origin": "https://www.wildberries.ru", "Referer": "https://www.wildberries.ru/", } ) async def get_reviews(self, nm_id: int, take: int = 100) -> list[Review]: all_reviews = [] skip = 0 while True: url = self.REVIEWS_URL.format(nm_id=nm_id) params = { "immt": nm_id, "skip": skip, "take": take, "order": "dateDesc", } resp = await self.client.get(url, params=params) resp.raise_for_status() data = resp.json() feedbacks = data.get("feedbacks", []) if not feedbacks: break for fb in feedbacks: all_reviews.append(self._normalize(nm_id, fb)) skip += take await asyncio.sleep(1.0) # Limit: no more than 1000 reviews per session if skip >= 1000: break return all_reviews def _normalize(self, nm_id: int, raw: dict) -> Review: photos = [] for photo in raw.get("photos", []): if full_url := photo.get("fullSize"): photos.append(full_url) return Review( external_id=raw.get("id", ""), product_nm_id=nm_id, author=raw.get("wbUserDetails", {}).get("name", "Customer"), rating=raw.get("productValuation", 0), text=raw.get("text", ""), pros=raw.get("pros"), cons=raw.get("cons"), date=raw.get("createdDate", ""), photos=photos, helpful_count=raw.get("feedbackValuation", 0), ) Parsing HTML reviews
For sites without an API, we parse HTML using PHP and Symfony Crawler. Example for iHerb:
// app/Services/ReviewScraper/HtmlReviewScraper.php class HtmlReviewScraper { public function scrapeIherb(string $productUrl, int $pages = 5): array { $reviews = []; for ($page = 1; $page <= $pages; $page++) { $html = $this->fetch("{$productUrl}?p={$page}&is=1&s=6"); $crawler = new Crawler($html); $items = $crawler->filter('[itemprop="review"]'); if (!$items->count()) break; $items->each(function (Crawler $node) use (&$reviews) { $reviews[] = [ 'external_id' => $node->attr('data-review-id'), 'author' => trim($node->filter('[itemprop="author"]')->text('')), 'rating' => (int) $node->filter('[itemprop="ratingValue"]')->attr('content'), 'date' => $node->filter('[itemprop="datePublished"]')->attr('content'), 'title' => trim($node->filter('[itemprop="name"]')->text('')), 'text' => trim($node->filter('[itemprop="reviewBody"]')->text('')), 'helpful' => (int) $node->filter('.helpful-yes')->text('0'), 'verified' => $node->filter('.verified-buyer')->count() > 0, ]; }); sleep(rand(2, 4)); } return $reviews; } } Deduplication and import into Laravel
After collection, reviews go through a Job with deduplication by external_id and source. Example:
// app/Jobs/ImportProductReviews.php class ImportProductReviews implements ShouldQueue { public int $tries = 3; public int $backoff = 120; public function handle(ReviewImportService $service): void { $mapping = ProductReviewMapping::where('product_id', $this->productId) ->where('source', $this->source) ->firstOrFail(); $reviews = $this->scrape($mapping->external_id); $imported = 0; $skipped = 0; foreach ($reviews as $reviewData) { $exists = ProductReview::where([ 'source' => $this->source, 'external_id' => $reviewData['external_id'], ])->exists(); if ($exists) { $skipped++; continue; } $service->import($this->productId, $this->source, $reviewData); $imported++; } Log::info("Reviews imported", [ 'product_id' => $this->productId, 'source' => $this->source, 'imported' => $imported, 'skipped' => $skipped, ]); } } Filtering and moderation
Stop-words (spam, ads, profanity), too-short reviews (<20 characters), and author anonymization — all customizable. Below is a service example:
// app/Services/ReviewImportService.php class ReviewImportService { private array $stopWords = ['buy', 'discount', 'promocode', 'vk.com', 't.me']; public function import(int $productId, string $source, array $data): ?ProductReview { if (mb_strlen($data['text']) < 20) return null; foreach ($this->stopWords as $word) { if (mb_stripos($data['text'], $word) !== false) return null; } return ProductReview::create([ 'product_id' => $productId, 'source' => $source, 'external_id' => $data['external_id'], 'author' => $this->anonymizeAuthor($data['author']), 'rating' => max(1, min(5, (int) $data['rating'])), 'text' => $this->sanitize($data['text']), 'pros' => $this->sanitize($data['pros'] ?? ''), 'cons' => $this->sanitize($data['cons'] ?? ''), 'date' => $data['date'], 'is_verified' => $data['verified'] ?? false, 'helpful' => $data['helpful'] ?? 0, 'status' => 'pending', ]); } private function anonymizeAuthor(string $name): string { $parts = explode(' ', trim($name)); if (count($parts) >= 2) { return $parts[0] . ' ' . mb_substr($parts[1], 0, 1) . '.'; } return $name ?: 'Customer'; } private function sanitize(string $text): string { return strip_tags(trim($text)); } } Structured data for SEO
After import, reviews are published in JSON-LD on the product page. This increases the chance of appearing in rich snippets and boosts CTR by 20-30%. Learn more about structured data.
// app/Http/Controllers/ProductController.php public function show(string $slug): Response { $product = Product::withReviews()->findBySlug($slug); $reviewSchema = $product->reviews->map(fn($r) => [ '@type' => 'Review', 'author' => ['@type' => 'Person', 'name' => $r->author], 'datePublished' => $r->date, 'reviewBody' => $r->text, 'reviewRating' => [ '@type' => 'Rating', 'ratingValue' => $r->rating, 'bestRating' => 5, ], ]); $aggregateRating = [ '@type' => 'AggregateRating', 'ratingValue' => round($product->reviews->avg('rating'), 1), 'reviewCount' => $product->reviews->count(), ]; } How to avoid blocks while scraping?
To avoid blocking, we use a pool of 50+ residential proxies, random delays from 1 to 4 seconds, rotate User-Agent (Chrome, Firefox, Safari). CAPTCHAs are solved via Anti-Captcha. Each session is limited to 1000 reviews.
Why review automation boosts SEO?
- Structured data (JSON-LD) helps search engines display ratings in snippets.
- UGC texts contain long-tail keywords.
- Regular updates signal an active store.
Comparison: manual collection of 100 reviews takes 2 hours, our parser — 2 minutes, 60 times faster. Budget savings on copywriters and moderators — up to 70%.
Work process
- Analysis: determine platforms, APIs, complexity.
- Development: write parser, moderation, deduplication.
- Testing: run on 1000+ reviews, check for bugs.
- Deployment: set up cron, error monitoring.
- Documentation: stack description, startup instructions.
| Stage | Duration |
|---|---|
| Requirements analysis | 1-2 days |
| Parser development | 2-3 days per platform |
| Testing | 1-2 days |
| Deployment and documentation | 1 day |
Starting price for a single platform parser is $500, with volume discounts available. We guarantee reliable performance, with over 5 years of experience and certified scraping practices. Our review parser bot handles review aggregation from multiple sources, ensuring all data is imported correctly.
What's included
- Parser for each marketplace (up to 5 platforms).
- Moderation (stop-words, length, anonymization).
- Deduplication (by external_id + source).
- Automatic updates on schedule (daily or every 4 hours).
- Structured data on product page.
- Logging and error notifications.
Timeline: from 3 to 5 business days per platform. For a complex project (3 platforms + moderation + SEO) — 7-10 days. Contact us to evaluate your project. Get a consultation on your project.







