Scraping Data from DEX Aggregators: 1inch and Jupiter

Building DeFi market analytics often requires clean routing and price data from DEX aggregators, but rate limits pose a challenge. <cite><a href="https://en.wikipedia.org/wiki/1inch">1inch</a></cite> and <cite><a href="https://en.wikipedia.org/wiki/Jupiter_(blockchain)">Jupiter</a></cite> are key so

Blockchain Development Services

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1309
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1270
  • image_logo-advance_0.webp
    B2B Advance company logo design
    719
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1011
  • image_logo-aider_0.webp
    AIDER company logo development
    954
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1062

Building DeFi market analytics often requires clean routing and price data from DEX aggregators, but rate limits pose a challenge. 1inch and Jupiter are key sources, but their APIs have limits and routes change frequently. We have built a data collector that gathers data from both aggregators, normalizes it, and stores it in TimescaleDB or ClickHouse. With 7+ years of Web3 experience (over 50 aggregation projects) and since 2018 on the market, we have constructed an architecture resilient to outages.

Collecting data from DEX aggregators involves more than scheduled API calls. It means handling different response formats, bypassing limits, normalizing data from 1inch (EVM) and Jupiter (Solana) into a unified schema, and storing it with timestamps for later analytics. In our practice, projects required collecting data for 500+ pairs every 10 seconds — that demanded setting up BullMQ and Redis for request coordination. Our queue handles up to 10,000 requests per minute with 99.9% uptime.

This article covers parsing routes from 1inch and Jupiter, bypassing rate limits, storing data, and building analytics. We also cover MEV data collection, EVM swap parsing, and practical scraping techniques for DEX aggregators.

Collecting Data from 1inch API Without Rate Limits

Swap API vs Fusion API

Swap API (/swap/v6.0/{chain}/swap) is the classic aggregation endpoint. A request returns:

  • tx object with full transaction data
  • protocols — list of protocols in the route with shares
  • toAmount — minimum output amount

For scraping price data without executing a swap, use the /quote endpoint: no slippage parameter, no fromAddress, returns only the quote. This does not load the RPC and requires no permissions.

Fusion API (/fusion/v1.0/{chain}/quote/receive) is a different model: RFQ (request for quote) with market makers. Routing is opaque, protocols are not fully disclosed. For route scraping, it is less useful.

Rate Limits and Bypasses

1inch Public API: 1 request/second, 500k requests/month free. For intensive scraping — 1inch Dev Portal with Pro plan or using your own 1inch router via direct contract calls. Direct contract calls are 2x faster than HTTP requests and bypass rate limits entirely.

Direct call to the 1inch Aggregation Router via eth_call — you get a quote without an HTTP request and without rate limits. Use calldata from SDK for simulation:

const data = routerContract.interface.encodeFunctionData("swap", [ executor, desc, data ]); const result = await provider.call({ to: ROUTER_ADDRESS, data }); 

But this requires understanding the internal format of 1inch — it changes between router versions.

Parsing the Route

The /quote response contains protocols — an array of arrays representing a split route:

"protocols": [ [ [{"name": "UNISWAP_V3", "part": 60, "fromTokenAddress": "...", "toTokenAddress": "..."}], [{"name": "CURVE", "part": 40, ...}] ] ] 

The first level is parallel paths (split by volume). The second level is sequential hops inside each path. For building a liquidity graph: normalize protocol names, aggregate by token pairs, track the dynamics of part over time.

Why Jupiter Price API is Better for Analytics?

V6 Quote API

GET /quote?inputMint=...&outputMint=...&amount=...&slippageBps=50 The response includes routePlan — a detailed route through Solana AMMs:

"routePlan": [ { "swapInfo": { "ammKey": "...", "label": "Orca (Whirlpool)", "inputMint": "...", "outputMint": "...", "inAmount": "1000000", "outAmount": "998432", "feeAmount": "3000", "feeMint": "..." }, "percent": 100 } ] 

For scraping: ammKey is the public key of the pool on Solana. You can directly query the pool state via getAccountInfo. label is a human-readable AMM name.

Price API Jupiter

Jupiter provides a /price?ids=... endpoint — bulk price query for up to 100 tokens per request. It returns the price in USDC with the liquidity source indicated. This is not a quotation (no slippage), but a reference price. It updates every 30 seconds. Jupiter's Price API is 10x more efficient than 1inch's quote endpoint for bulk token prices.

For building price history: poll /price with desired pairs every 30 seconds, store in TimescaleDB or InfluxDB. Per day — ~2880 points per pair.

Jupiter rate limits: public API without key — 600 requests/minute. Jupiter API Pro — higher. For guaranteed uptime in production systems, use your own Jupiter self-hosted or a partner key.

Scraper Architecture

Data Structure

interface RouteSnapshot { timestamp: number; chain: "ethereum" | "solana" | "arbitrum" | ...; inputToken: string; outputToken: string; inputAmount: bigint; outputAmount: bigint; priceImpact: number; // in % protocols: ProtocolHop[]; source: "1inch" | "jupiter"; } interface ProtocolHop { name: string; poolAddress: string; percentOfRoute: number; inputAmount: bigint; outputAmount: bigint; } 

Request Queue and Retries

A collector handling multiple token pairs → parallel requests → quick rate limit hit. The proper architecture: a queue with Bull/BullMQ + Redis, configurable concurrency per source.

Retry with exponential backoff for 429 Too Many Requests: delay = Math.min(base * 2^attempt, maxDelay). For 1inch — base = 1000ms, maxDelay = 30000ms.

Monitor collector health: Prometheus metrics scraper_requests_total{status="success|error"}, scraper_latency_ms. Alert when error rate > 10% over 5 minutes.

Storage and Queries

TimescaleDB (PostgreSQL extension) for time series — optimized for SELECT ... WHERE timestamp BETWEEN ... AND ... queries with aggregation. For high-frequency route scraping — partitioning by day.

ClickHouse as an alternative for very high volumes (>10M rows/day): columnar storage gives 10-100x faster analytical queries over large time ranges. We have stored over 100 million price points for 500+ pairs.

Storage Comparison
Feature TimescaleDB ClickHouse
Optimized for time series, JOINs analytics, aggregations
Max throughput up to 1M rows/day >10M rows/day
Query speed high for point queries high for aggregates
Maintenance auto-partitioning requires tuning
Pair 1inch chains Jupiter pools Frequency
USDC/ETH Ethereum, Arbitrum, Optimism 1 min
SOL/USDC Orca, Raydium 30 sec
BTC/USDC all EVM 5 min

What's Included in the Work

  1. Analysis of your goals and selection of pairs/polling frequency.
  2. Development of the data collector (Node.js/TypeScript) with buffering and retries.
  3. Integration with storage (TimescaleDB / ClickHouse).
  4. Dashboard or API for data access (Grafana, REST).
  5. Documentation of data formats and limits.
  6. Handover of access and training for your team.

Timeline and Budget Estimates

A collector for one source (1inch or Jupiter) with PostgreSQL storage — 1-2 days, budget starts at $1,500. A multi-source collector with data normalization, ClickHouse, and analytics API — 3-5 days. Timeline and cost are discussed individually — contact us for an accurate estimate.

Order development of a data collector for your tasks — get a reliable tool for DeFi analytics. Contact us — we'll choose the optimal solution for your pairs and polling frequency.