Imagine a news site on Bitrix with 10,000 daily visitors, but content is updated once a week. Editors spend 20 hours a week copying articles from 15 sources. As a result, search positions drop by 30% in a quarter. We implement RSS feed parsing that automatically collects news, filters duplicates, and publishes them in an information block. On one project, this reduced time spent from 15 to 1 hour per week and increased the citation index by 45%. Moreover, this approach completely eliminates human copying errors.
According to the RSS Board specification, RSS 2.0 supports all necessary fields for automated publishing.
How to choose a data source for parsing?
News feeds are available in several formats. Comparison by key parameters:
| Source | Reliability | Complexity | Legal Risks |
|---|---|---|---|
| RSS/Atom feeds | High | Low | Minimal |
| Aggregator APIs | High | Medium | Medium (require licenses) |
| HTML pages | Low | High | High |
RSS/Atom is the optimal start. They are standardized, easy to parse, and require no special permissions.
How to organize the import process?
The news parser for Bitrix consists of three layers.
Fetcher. Retrieves RSS feeds from a URL list. Uses file_get_contents with context or cURL with timeouts. Each feed is parsed via SimpleXMLElement or the SimplePie library.
$xml = simplexml_load_string($rssContent); foreach ($xml->channel->item as $item) { $title = (string)$item->title; $link = (string)$item->link; $date = strtotime((string)$item->pubDate); $desc = (string)$item->description; } Processor. Cleans HTML tags, downloads images, normalizes dates, and determines category by keywords.
Importer. Creates elements in the information block via CIBlockElement::Add(). Checks duplicates by XML_ID (article URL or GUID).
Content storage and processing
Recommended RSS→infoblock mapping:
| RSS field | Infoblock field | Type |
|---|---|---|
title |
NAME |
String |
link |
PROPERTY_SOURCE_URL |
Link |
description |
PREVIEW_TEXT |
HTML/text |
content:encoded |
DETAIL_TEXT |
HTML |
pubDate |
ACTIVE_FROM |
Date |
guid / link |
XML_ID |
String (deduplication) |
category |
IBLOCK_SECTION_ID |
Section binding |
enclosure / media:content |
PREVIEW_PICTURE |
File |
XML_ID is mandatory. Use an md5 hash of the article URL — this guarantees uniqueness. Raw HTML from RSS is unsuitable for publication. Typical issues:
- External images — download to
/upload/during import. - Third-party scripts and iframes — use
strip_tags()with a whitelist orHTMLPurifier. - Relative links — convert to absolute by substituting the source domain.
- Encoding — detect via
mb_detect_encoding()and convert to UTF-8.
Why cron instead of agents?
Cron is 3 times more reliable than standard Bitrix agents for long-running tasks — agents abort if they exceed the execution time, while cron runs to completion. A cron job calls a PHP script:
$_SERVER['DOCUMENT_ROOT'] = '/home/bitrix/www'; require $_SERVER['DOCUMENT_ROOT'] . '/bitrix/modules/main/include/prolog_before.php'; CModule::IncludeModule('iblock'); Example cron configuration for parsing every 30 minutes
*/30 * * * * /usr/bin/php /home/bitrix/www/parser.phpFrequency: breaking news — every 15–30 minutes, industry news — 1–2 hours, analytics — 1–2 times a day.
How to set up the parser in 5 steps?
- Select RSS/API sources.
- Write a fetcher using SimplePie or cURL.
- Configure the processor: HTML cleanup, image download, category mapping.
- Implement the importer with XML_ID verification.
- Set up cron and an error logging system.
Quality control and categorization
In addition to XML_ID, use:
- Date filter — do not import news older than N days.
- Minimum description length — discard entries shorter than 100 characters.
- Stop words — filter out keywords irrelevant to the topic.
- Source limit — no more than N news per day from a single feed.
The simplest categorization is "source → infoblock section" mapping. A more flexible approach is keyword-based classification:
$rules = [ 'Technology' => ['AI', 'blockchain', 'startup', 'app'], 'Finance' => ['stocks', 'exchange rate', 'investments', 'IPO'], ]; For 10+ categories, integrate external classifiers (OpenAI, Yandex GPT).
Legal side
Publishing others' news "as is" violates copyright. Options:
- Publish title + 2–3 sentences with a link (fair quoting).
- Automated rewriting via LLM — legally questionable.
- Use feeds with open licenses (Creative Commons, government sources).
What's included in the work
- Audit of the current catalog — infoblock structure, duplicate check.
- Parser development — fetcher, processor, importer.
- Integration with cron — schedule, logging.
- Deduplication and filters.
- Documentation and editor training.
- 30-day post-release support.
Experience and expertise
We have been developing on Bitrix for over 7 years and have completed 200+ automation projects. Our engineers are 1C-Bitrix certified. We use the full stack: CommerceML, REST API, Bizproc. Contact us to discuss the parser architecture for your site — we'll prepare a proposal within 1-2 days. Get a consultation on automating your news section.

