Python Parser for Bitrix: Architecture and Implementation

Python Parser for Bitrix: Architecture and Implementation You're facing a situation where standard CSV import hits performance limits, or a PHP script can't handle a headless browser? For instance, you need to scrape 50,000 products from a competitor's site, but the catalog is served via an SPA,

Our competencies:

Frequently Asked Questions

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1415
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    995
  • image_bitrix-bitrix-24-1c_development_of_an_online_appointment_booking_widget_for_a_medical_center_594_0.webp
    Development based on Bitrix, Bitrix24, 1C for the company Development of an Online Appointment Booking Widget for a Medical Center
    733
  • image_bitrix-bitrix-24-1c_mirsanbel_458_0.webp
    Development based on 1C Enterprise for MIRSANBEL
    863
  • image_crm_dolbimby_434_0.webp
    Website development on CRM Bitrix24 for DOLBIMBY
    772
  • image_crm_technotorgcomplex_453_0.webp
    Development based on Bitrix24 for the company TECHNOTORGKOMPLEKS
    1134

Python Parser for Bitrix: Architecture and Implementation

You're facing a situation where standard CSV import hits performance limits, or a PHP script can't handle a headless browser? For instance, you need to scrape 50,000 products from a competitor's site, but the catalog is served via an SPA, and the partner provides no API. We solve such problems: we design a Python parser that loads data into an intermediate storage, and a PHP importer transfers it to Bitrix infoblocks. With experience from over 40 parsers for Bitrix online stores — from simple RSS aggregators to machine learning systems for content classification — we deliver reliable automation.

With over 5 years of experience and 40+ parsers delivered, we guarantee reliable integration.

Many owners of large Bitrix catalogs face issues updating products: manual entry takes days, Excel import breaks encoding, and partners don't provide APIs. Our approach: Python for collection, Bitrix for storage and delivery. This saves up to 70% of time and ensures full transparency.

Why Python, Not PHP

Specific reasons, not abstract advantages:

  • Asynchrony. asyncio + aiohttp handle 100+ requests in parallel. PHP curl_multi practically achieves 20–50 connections.
  • Headless browser. Playwright for Python works stably with React sites. PHP wrappers for Puppeteer are less reliable.
  • NLP and ML. Text classification, entity extraction — libraries like spaCy and transformers have no equivalent in PHP.
  • Libraries. BeautifulSoup, lxml, Scrapy — proven tools with large communities.

How the Parser Architecture Works

The Python parser runs as a separate service. Data passes through an intermediate storage — tables in a shared database or RabbitMQ queues. Python writes raw data; a PHP agent picks it up and writes into infoblocks using CIBlockElement::Add.

Storage Options

Method Data Volume Key Feature
JSON files up to 1,000 Simple, no dependencies
PostgreSQL/MySQL 1,000–100,000 Indexes, transactions
Bitrix REST API any Direct write, but HTTP overhead
Redis/RabbitMQ streaming Queues, scalability

For most projects, a shared database is optimal: Python writes to a staging table, PHP imports batches every 5–15 minutes via cron.

Comparison: Python vs PHP for Parsing

Criterion Python PHP
Async requests asyncio + aiohttp (100+ parallel) curl_multi (20-50)
Headless browser Playwright (stable) Puppeteer (less reliable)
NLP/ML spaCy, transformers unavailable
Parsing ecosystem Scrapy (full framework) Goutte (limited)

Python is up to 6 times faster than PHP for parsing large catalogs. Scrapy processes 10,000 URLs in 5–10 minutes, while PHP takes 30–40 minutes. Playwright is 3 times more stable than PHP wrappers for Puppeteer due to deeper browser support. Source: internal benchmarks

Example Implementation on Scrapy

Scrapy is a framework that handles URL queues, retries, throttling. A spider for a catalog:

Spider code
import scrapy class CatalogSpider(scrapy.Spider): name = 'catalog' start_urls = ['https://books.toscrape.com/catalogue/page-1.html'] def parse(self, response): for product in response.css('.product_pod'): yield { 'name': product.css('h3 a::attr(title)').get(), 'price': product.css('.price_color::text').get(), 'url': product.css('h3 a::attr(href)').get(), } next_page = response.css('.next a::attr(href)').get() if next_page: yield response.follow(next_page, self.parse) 

Pipeline for writing to the staging table:

import psycopg2 class BitrixPipeline: def open_spider(self, spider): self.conn = psycopg2.connect( host='localhost', port=5433, dbname='bitrix_db', user='bitrix' ) def process_item(self, item, spider): cursor = self.conn.cursor() cursor.execute(""" INSERT INTO parser_staging (name, price, description, image_url, source_url, status) VALUES (%s, %s, %s, %s, %s, 'new') ON CONFLICT (source_url) DO UPDATE SET price = EXCLUDED.price, updated_at = NOW() """, (item['name'], item['price'], item['description'], item['image'], item['url'])) self.conn.commit() return item 

Headless Browser for SPAs

Sites built with React or Vue serve empty HTML. Playwright solves this:

from playwright.async_api import async_playwright async def parse_spa(url): async with async_playwright() as p: browser = await p.chromium.launch(headless=True) page = await browser.new_page() await page.goto(url, wait_until='networkidle') content = await page.content() await browser.close() return content 

Resource consumption: each Chromium instance uses 100–300 MB RAM. For mass parsing, use a pool of 3–5 instances and a task queue.

How to Transfer Data to Bitrix

A PHP script on the Bitrix side fetches data from the staging table:

$rows = $DB->Query("SELECT * FROM parser_staging WHERE status = 'new' LIMIT 100"); while ($row = $rows->Fetch()) { $elementId = (new CIBlockElement())->Add([ 'IBLOCK_ID' => CATALOG_IBLOCK_ID, 'NAME' => $row['name'], 'XML_ID' => md5($row['source_url']), // ... ]); if ($elementId) { $DB->Query("UPDATE parser_staging SET status='imported', bx_id={$elementId} WHERE id={$row['id']}"); } } 

The script runs via cron every 5–15 minutes and processes new records in batches.

Deployment and Monitoring

The Python parser is deployed separately from Bitrix. Use a systemd service or cron for scheduled runs. Virtual environment (venv) isolates dependencies. Logging uses the logging module with rotation. Monitoring — a script checks that the parser ran within the last N hours and sends an alert if it stalls.

Typical crontab:

0 1 * * * cd /opt/parsers && /opt/parsers/venv/bin/scrapy crawl catalog 2>> /var/log/parser.log 0 */4 * * * cd /opt/parsers && /opt/parsers/venv/bin/python news_parser.py 2>> /var/log/parser.log 

We use a custom healthcheck: every 4 hours we verify that the parser completed without errors. If stalled — automatic restart and notification in Telegram. For critical projects, we add alerting based on Prometheus and Grafana.

When to Use a Python Parser Instead of PHP?

If the source is an SPA (React/Vue/Angular), data volume exceeds 10,000 items, content classification is needed, or DDoS protection is required — Python provides a significant advantage. Comparison: Scrapy processes 10,000 URLs in 5–10 minutes, while a PHP solution with curl_multi takes 30–40 minutes. Playwright is 3 times more stable than PHP wrappers for Puppeteer due to deeper browser support.

What's Included in the Work

We provide a full development cycle with the following deliverables:

  • Technical documentation and architecture diagrams.
  • Access to the parser source code and intermediate database.
  • Training for administrators on operation and troubleshooting.
  • Post-launch support for 30 days, including bug fixes and minor adjustments.

How We Develop the Parser

We provide a full development cycle:

  1. Source analysis and architecture agreement.
  2. Development of a spider on Scrapy or an async parser on aiohttp.
  3. Configuration of the Bitrix importer (infoblocks, HL-blocks, SKU offers).
  4. Creation of the intermediate database and synchronization scripts.
  5. Deployment on the server (systemd, cron, monitoring).
  6. Documentation for operation and administrator training.

Estimated Timelines

Development time ranges from 5 to 20 working days, depending on source complexity and data volume. Cost is calculated individually after analyzing your project, typically ranging from $1,000 to $5,000.

We guarantee stable 24/7 parser operation and provide post-launch support. Request a consultation: tell us about your data source, and we'll propose the optimal solution within a day.