Writing a scraper is easy; maintaining it, validating it and delivering reliable data every week is what takes the effort. This guide walks through the real work behind each option and a simple way to choose.
When a data team says "we will do it ourselves with a scraper", they are usually comparing a script that works today with a service that delivers clean data every Monday. Those are different things. A script solves extraction; a managed dataset solves extraction, maintenance, validation, a stable schema, delivery and the licence to use the result. Comparing only the first part gives a false sense of savings.
At Rastriq we work with three delivery models, and it helps to have them clear before deciding: self-service on Apify (you launch a public actor and pay per result), a licensed dataset (you receive an already extracted, normalized and versioned dataset) and a custom project (we design and operate the extraction for your sources). The question "build or buy?" is really "how much of that chain do I want to operate myself?".
A production scraper is not a file, it is a chain with several links that fail for different reasons. It is worth listing them because each one consumes hours from someone with a technical profile:
Building your own scraping is a good decision in several cases. If the source is a single, stable site without serious defences, maintenance cost is low and full control pays off. If the data is so specific to your business that no provider covers it, there is no alternative. And if you have an engineering team with spare capacity and scraping is part of your competitive edge, operating it in-house is strategically sound.
It is also reasonable to build for a proof of concept: write a quick extractor to check that a data point is useful before committing. The trap appears when the prototype stays in production without anyone budgeting it as a service.
A managed dataset pays off when the data is an input rather than the business. If you need prices from several brands, countries or portals, the number of sources multiplies maintenance; with ten sources, something will be broken every week. If the team consuming the data is analytical rather than engineering, it will prefer receiving a file with a stable schema over watching extractors. And if you need history, starting to accumulate it today beats discovering in six months that you never stored it.
Another criterion is risk: when a source changes, the repair depends on who maintains it. With a provider, that response time is part of the service; in-house, it competes with the rest of the team priorities.
The table summarizes how the work is split in each model. It contains no prices because they depend on volume and sources; the aim is to see who is responsible for what.
| Criterion | Own scraping | Self-service on Apify | Licensed dataset | Custom project |
|---|---|---|---|---|
| Time to start | Weeks of development | Minutes: pick an actor and run | Agreed first delivery | Design and testing phase |
| Maintenance when sites change | Your team | Rastriq maintains the actor | Rastriq maintains it | Rastriq maintains it |
| Anti-bot and proxies | Your team | Included in the actor | Included | Included and adapted to your source |
| Common schema across sources | You design it | Actor schema | Normalized | Defined with you |
| History | You accumulate it | You store it | As agreed | As agreed |
| Delivery format | Whatever you program | JSON, CSV, Excel, API | CSV, JSON or bucket | API, S3, CSV or what you need |
| New sources | One per project | Those in the Store | On request | The purpose of the model |
You do not have to pick an extreme. A common pattern is to start with self-service: download the free sample, run an actor with your own filters and check that the data answers your question. If it works and becomes recurring, move to a licensed dataset with scheduled delivery. If a source appears that is not in the catalogue, it is commissioned as a custom project and, once stable, joins the catalogue.
This path reduces risk at each step: you do not pay for a whole project to discover the data is not useful, and you do not get stuck in a tool that does not scale. You can browse the catalogue by vertical on the buy datasets page and see the extraction process under managed web scraping.
If you decide to buy, these questions separate a solid provider from one that only sells extraction:
Manufacturer configurators are a good case to see the whole problem. Each brand publishes prices with different fields, currencies and taxes: BMW delivers gross_list_price, net_list_price and total_taxes; Kia, msrp_before_discount, discount_amount and is_net_price; Toyota, price_list, price_discount, price_net and price_tax. One scraper per brand is manageable; maintaining 28 of them, across several countries, with a common schema, is a service.
That is why we publish one data page per brand with coverage, a downloadable sample and a field dictionary, plus a commercial hub at new car pricing data. If the common schema is what you are missing, the guide to normalizing car configurator pricing explains the method.
Rastriq maintains more than one hundred public actors on the Apify Store, and the same extractors underpin the licensed datasets and the custom projects. That means what you see in a sample is what gets delivered later, not a demo built for the occasion. If you want to see the kind of output before talking, download a sample and review the fields; if you would rather we tell you which source fits your case, the next step is to request a dataset.
There is no universal figure: it depends on the number of sources, their anti-bot defences and how often they change. What is common is that the main cost is not writing the scraper but watching it, repairing it when the source changes and validating data quality every week.
It should not. If the schema is documented and the formats are standard (CSV, JSON, API), you can switch provider or bring extraction in-house later. Always ask for the field dictionary and a real sample before contracting.
Yes. The samples page has real JSON and CSV files of up to 100 rows per source, already sanitized, and the public Apify actors are pay-per-result, so you can run a small extraction before committing to a dataset.
Rastriq extracts public product data and excludes personal fields from its samples, such as phone numbers, emails or private seller details. The data licence spells out what is delivered and what it may be used for.
When your source is not in the catalogue, when you need your own schema or when delivery has specific requirements, such as an S3 bucket or an internal API. It is designed with you and operated as a service.
Download a real sample of up to 100 rows, already sanitized, or tell us which sources you need.