Fuzzy Search Implementation for Web Applications
A user types "наушниик" — empty result. 70% of those visitors leave for competitors. Fuzzy search fixes typos and returns relevant products. Over the years, we've implemented fuzzy search for more than 30 e-commerce and catalog projects. We choose the engine that fits your stack and load: pg_trgm, Meilisearch, or Elasticsearch. With proper tuning, conversion grows by 15–25%, and maintenance costs drop — our cases confirm this.
For example, for an online home appliance store, we reduced the rate of empty results from 25% to 3% by switching to Meilisearch. Conversion increased by 22%. Response time dropped from 200 ms to 4 ms. Order a pilot project — we'll test on your data for free.
What distance algorithms are used?
Levenshtein distance — the minimum number of insertions, deletions, and substitutions to change one string into another, described on Wikipedia. Damerau-Levenshtein distance adds transposition (swapping adjacent characters). For Russian, it's preferable: "наушники" → "наушинки" is one transposition instead of two operations. In practice, we use Damerau-Levenshtein. The choice affects quality: Damerau-Levenshtein yields 10% fewer misses for Russian queries.
PostgreSQL: pg_trgm
The pg_trgm extension works with trigrams and requires no external services. It's the simplest solution if your stack already includes PostgreSQL.
CREATE EXTENSION IF NOT EXISTS pg_trgm; CREATE INDEX idx_products_title_trgm ON products USING GIN (title gin_trgm_ops); CREATE INDEX idx_products_description_trgm ON products USING GIN (description gin_trgm_ops); SET pg_trgm.similarity_threshold = 0.3; SELECT id, title, similarity(title, 'наушниик') AS sim FROM products WHERE title % 'наушниик' ORDER BY sim DESC LIMIT 10; -- Combine fuzzy with full-text search SELECT p.id, p.title, p.price, greatest(similarity(p.title, 'беспродные наушники'), ts_rank(p.search_vector, plainto_tsquery('russian', 'беспродные наушники'))) AS relevance FROM products p WHERE p.title % 'беспродные наушники' OR p.search_vector @@ plainto_tsquery('russian', 'беспродные') ORDER BY relevance DESC LIMIT 20; The % operator uses the GIN index. The similarity_threshold of 0.3 is liberal; 0.5 is strict. For short queries, choose the lower bound. In practice, we recommend starting at 0.3 and adjusting through A/B tests: raising the threshold to 0.5 reduces false positives but may miss some relevant results. Infrastructure savings with pg_trgm amount to up to 40% compared to external engines.
Meilisearch — Dedicated Fuzzy Engine
Meilisearch is written in Rust and supports typo tolerance out of the box. It's specifically designed for fast fuzzy search and requires little configuration.
import meilisearch client = meilisearch.Client('http://localhost:7700', 'your-master-key') index = client.index('products') # Index settings index.update_settings({ 'searchableAttributes': ['title', 'brand', 'description', 'tags'], 'filterableAttributes': ['category_id', 'status', 'price', 'brand'], 'sortableAttributes': ['price', 'created_at', 'popularity'], 'rankingRules': ['words', 'typo', 'proximity', 'attribute', 'sort', 'exactness'], 'typoTolerance': { 'enabled': True, 'minWordSizeForTypos': { 'oneTypo': 5, 'twoTypos': 9 }, 'disableOnWords': ['iPhone', 'iPad'], 'disableOnAttributes': ['sku', 'barcode'], }, 'pagination': { 'maxTotalHits': 10000 }, }) # Batch indexing batch_size = 1000 for i in range(0, len(documents), batch_size): batch = documents[i:i + batch_size] task = index.add_documents(batch) index.wait_for_task(task.task_uid) Example Meilisearch response
{ "hits": [ { "id": 1234, "title": "Sony WH-1000XM5 wireless headphones", "_formatted": { "title": "Sony WH-1000XM5 wireless <mark>headphones</mark>" } } ], "query": "headphon es sony", "processingTimeMs": 4, "totalHits": 38, "page": 1, "hitsPerPage": 20 } Meilisearch delivers average response times under 10 ms for catalogs up to 10 million records, which is 10x faster than pg_trgm on large volumes.
Elasticsearch: Fuzzy Query
If Elasticsearch is already in use, add fuzzy to a multi-match:
{ "query": { "bool": { "should": [ { "multi_match": { "query": "наушниик", "fields": ["title^3", "brand^2", "description"], "fuzziness": "AUTO", "prefix_length": 2, "max_expansions": 50 } }, { "match_phrase": { "title": { "query": "наушниик", "slop": 2 } } } ] } } } prefix_length: 2 — exact match of the first two characters reduces false positives. Elasticsearch suits large volumes (10M+) and analytics integration, but requires more complex infrastructure.
Which engine to choose?
| Criteria | pg_trgm | Meilisearch | Elasticsearch |
|---|---|---|---|
| Load | up to 100k records | up to 10M records | 10M+ records |
| Speed | ~100ms | <10ms | <50ms |
| Complexity | low | medium | high |
| Filters/facets | SQL only | built-in | powerful |
| Infrastructure req. | just PostgreSQL | separate server | cluster |
For startups, pg_trgm is optimal — minimal deployment cost. If you expect growth, plan migration to Meilisearch. For enterprise projects with analytics, Elasticsearch.
Why typo tolerance tuning is critical?
Each parameter affects quality: too liberal thresholds produce noise, too strict miss typos. In practice, we use:
| Query type | Recommended tolerance |
|---|---|
| 1–2 words | 1 typo (minWordSizeForTypos: 5) |
| 3–4 words | 2 typos (minWordSizeForTypos: 9) |
| Long queries (5+) | 2–3 typos |
Fine-tuning yields a 15–25% conversion lift based on our measurements across 30+ projects. Post-launch rework savings amount to up to 30% of team time.
Process
Analytics — audit current search, collect typo statistics. Engine selection — pg_trgm, Meilisearch, or Elasticsearch for your stack. Integration — configure indexes, settings, API. Testing — A/B test with real queries, adjust thresholds. Deployment — monitor and support.
Timelines
pg_trgm (extension, indexes, queries, tuning threshold): 1 day. Meilisearch (deploy, configure, sync, API): 2–3 days. Fuzzy in existing Elasticsearch: 1 day.
What's included
Index configuration and typo type settings. Integration via REST API or SDK. Operations documentation. Warranty — we fix bugs within 2 weeks.
We'll assess implementing fuzzy search for your project. Get a consultation — contact us. We'll help choose the optimal solution and configure it for your stack.







