Efficient Hybrid Parsing for 1C-Bitrix Content Automation

Efficient Hybrid Parsing for 1C-Bitrix Content Automation Imagine an editor manually copying articles from external sources. One publication takes 4 hours—analyzing structure, cleaning formatting, uploading images, SEO alignment. For 500 articles, that's 2000 man-hours. We automate this: hybrid p

Our competencies:

Frequently Asked Questions

Efficient Hybrid Parsing for 1C-Bitrix Content Automation

Imagine an editor manually copying articles from external sources. One publication takes 4 hours—analyzing structure, cleaning formatting, uploading images, SEO alignment. For 500 articles, that's 2000 man-hours. We automate this: hybrid parsing based on Mozilla Readability and custom CSS selectors cuts time to 10 minutes per article with 95%+ accuracy. With over 10 years in the industry and 500+ successful projects, we deliver reliable automation that saves over $50,000 annually. Get a preliminary analysis of your sources—we'll assess complexity and timeline.

Why Article Parsing Is Harder Than News Parsing

News parsers work with RSS feeds—structured, predictable data. Article parsing deals with arbitrary HTML pages where each source has its own layout, navigation structure, and content presentation. Key differences:

  • No unified format—each source needs a custom parser or a universal extractor.
  • Complex content structure—articles contain headings, lists, tables, embedded media, code blocks. All must be preserved.
  • Larger text volume—5,000–10,000 characters vs. 500-character news. More data means more failure points.
  • Lower update frequency—articles are published less often, but each piece is more valuable.

What Problems Does Efficient Hybrid Parsing Solve?

The hybrid approach combines an algorithmic extractor (Mozilla Readability) with custom CSS selectors for sites where automation fails. The library andreskrey/readability.php (PHP port of Readability) analyzes text density and extracts main content in 2 seconds. For sources where Readability misses tables or lists, we add exception selectors. Result: 95%+ accuracy on 10+ sources.

Compare approaches: pure Readability preserves 80% structure but loses 30% of tables and 15% of lists. Manual selectors give 99% accuracy but require 15 minutes per source. Hybrid delivers 95% accuracy at 0.5 minutes per source—30x faster than selectors and 2x more accurate than pure Readability for tables.

How to Preserve Structure and Formatting During Extraction?

After extracting HTML, it must be converted to a format suitable for Bitrix infoblock DETAIL_TEXT. We use HTMLPurifier with custom config allowing h2–h4, p, ul, ol, li, table, img, a, strong, em, blockquote, pre, code. Cleaning steps:

  • remove script, style, iframe, inline styles, data-attributes;
  • normalize headings: original h1 becomes h2 in Bitrix context;
  • localize images: download external images to /upload/, replace URLs in HTML;
  • handle lazy loading (data-src instead of src).

The final HTML is validated and checked against the infoblock schema.

Content Extraction: Three Approaches

Approach Principle When to Use
CSS selectors Selector for specific site (.post-content) Up to 5 sources, stable layout
Algorithms (Readability) DOM analysis via text density heuristics 5+ sources, heterogeneous layout
Hybrid Readability + custom rules for errors 10+ sources, maximum accuracy

Readability on PHP completes in 2 seconds vs. 15 minutes for manual selector per source—a 450x gain for 10 sources. In practice, hybrid is the only approach that works for 10+ sources. Pure automation loses critical blocks (tables, lists); pure selectors don't scale.

Step-by-Step Article Parsing Process

  1. URL collection. The parser crawls listing pages (pagination, categories, sitemap.xml) and collects article URLs. They are saved to a queue table parser_queue with fields url, status, created_at.
  2. Download and extraction. For each URL in queue: download HTML, extract content, parse metadata. Result is a structured array saved to intermediate table parser_articles.
  3. Moderation (optional). An administrator reviews parsed articles in the interface, approves or rejects. For full automation, this step is replaced by rule-based filtering.
  4. Import. Approved articles are loaded into an infoblock via CIBlockElement::Add() from the Bitrix API. Images are saved through CFile::MakeFileArray().

Mapping to Infoblock

Extracted Data Infoblock Field Processing
Heading h1 / title NAME Trim to 255 chars, strip HTML
First 300 chars of text PREVIEW_TEXT strip_tags() + cut at sentence boundary
Full article HTML DETAIL_TEXT Clean through HTMLPurifier
First image PREVIEW_PICTURE Download + resize
Source URL PROPERTY_SOURCE_URL As-is
Publication date ACTIVE_FROM Parse via strtotime()
md5(url) XML_ID For deduplication
Author PROPERTY_AUTHOR Extract from meta or byline
Tags / keywords PROPERTY_TAGS Multiple string property

How to Protect the Parser from Blocking?

Content sites are less protected than marketplaces, but basic measures exist:

  • robots.txt—check Disallow for parsed sections. Ignoring adds legal risk.
  • Rate limiting—1–2 requests per second are safe for most sites. Aggressive parsing leads to blocking.
  • JavaScript rendering—SPA sites require headless browsers. For static sites, cURL suffices.
  • Cloudflare / WAF—detect bots by fingerprint. Solved with headless browser and realistic headers.

Cron Automation

Recommended cron schedule
# Collect new URLs from sources—once daily 0 2 * * * php /home/bitrix/parsers/collect_urls.php # Parse articles from queue—every 2 hours 0 */2 * * * php /home/bitrix/parsers/parse_articles.php --limit=50 # Import into infoblock—every hour 0 * * * * php /home/bitrix/parsers/import_articles.php 

Splitting into three tasks allows independent control of each stage and quick problem localization.

Turnkey Work Scope & Deliverables

  • Source analysis: Identify 5–15 donor sites, examine DOM, spot layout peculiarities.
  • Parser development: Hybrid PHP modules (Readability + custom selectors).
  • Mapping and import: Configure infoblocks, properties, deduplication.
  • Testing: Check on 50+ real articles, fine-tune.
  • Documentation: Architecture description, guide for adding new source.
  • Support: 3-month warranty: fix bugs, adapt to layout changes.
  • Knowledge transfer: 2-hour training session for your team.

We will assess your project—just contact us. Our team is certified on the 1C-Bitrix platform with 10+ years of experience and has delivered 500+ content automation projects. Estimated project cost: $2,500 for a 5-source parser, saving approximately $50,000 annually in manual labor—a 95% reduction in effort. We guarantee solution stability even during donor site redesigns.

FAQ for this article is provided as structured data (see JSON-LD above) and is not inline content.