rastriq. Talk to us →
Rastriq / Managed web scraping

Managed web scraping: you define the data, we extract it.

We design, run and maintain web extractions so your team receives structured data, not scripts to fix. From a single new source to a recurring pipeline with a stable schema.

Quick answer: Rastriq designs, runs and maintains custom web extractions on top of 109 public actors on the Apify Store. Data is delivered through Apify, the REST API, CSV or JSON files and S3 buckets, with no personal data and under the Rastriq data licence.

109
Public actors on Apify
4
Delivery channels
0
Personal data

What we already deliver, in real rows

Rastriq maintains 109 public actors on the Apify Store. That codebase and experience is the starting point for every custom project.

SourcePublic actorSample rowsFieldsDownload
Coches.net · car listings in Spaincochesnet-spain5855CSV · JSON
Wallapop · used carswallapop-cars-scraper10023CSV · JSON
Milanuncios · carsmilanuncios-scraper10030CSV · JSON
MachineryZone · used equipment in Europemachineryzone-scraper10020CSV · JSON
Surplex · online industrial auctionssurplex-auction-scraper10044CSV · JSON
BOE · official gazette entriesboe-scraper10016CSV · JSON
Amazon · product reviewsamazon-reviews-scraper2423CSV · JSON
Google Ads Transparency Center · adsgoogle-ads-scraper10015CSV · JSON
TikTok Creative Center · trending hashtagstiktok-trends-scraper5012CSV · JSON
Europages · B2B directoryeuropages-scraper1510CSV · JSON

Data from the current public samples. Each actor is published on the Apify Store and each sample can be downloaded without signing up.

Example of structured output

CountryRankHashtagIndustryPosts
ES1liveiseasyNews & Entertainment10500
ES2octubreNews & Entertainment5600
ES3labolanegraNews & Entertainment7500
ES4lluviaNews & Entertainment3800
ES5octoberNews & Entertainment2000

Five real rows from the TikTok Creative Center (hashtags) sample. It is an example of the structured data an extraction delivers.

Five steps, from idea to pipeline

Every project follows the same path. The first sample arrives before you commit to the full build.

1 · Briefing
Source, fields, markets, volume and frequency. We clarify what you will use the data for and the format you need.
2 · Feasibility and compliance
We review how the information is published, which terms apply and whether the data is public. We rule out anything that requires third-party credentials or involves personal data.
3 · Prototype and sample
We deliver a real sample with the proposed schema so you can validate fields and quality before going to production.
4 · Production
We schedule the extraction, watch for errors and source changes, and apply normalization and QA on every run.
5 · Delivery and maintenance
The data arrives through the agreed channel and we maintain the actor when the source changes its structure.

How we handle anti-bot protections

We do not promise access to every website: we assess each source and choose the lightest technique that works reliably.

Real browsers and consistent sessions
When a source needs it, we use real browsers with consistent fingerprints and stable sessions instead of isolated requests.
Proxies and a conservative pace
We choose the proxy type by country and source, and cap concurrency so we never overload the origin site.
Internal API when available
If the site consumes a public API from its own frontend, we extract from there: it is faster, cheaper and more stable than rendering pages.
Change monitoring
Every run validates record counts and field shapes. If something changes, we catch it before it reaches your table.

Four delivery channels

We deliver where your team already works. If several sources must be unified, we add data normalization to the same pipeline.

Apify
A dataset in your Apify account, with scheduled runs and webhooks. You can rerun or extend the actor whenever you want.
Access: console and API
API
Query results through the Apify REST API with pagination and filters, and plug it into your ETL or your agents.
Format: JSON
CSV / JSON
Files ready to load into your data warehouse, with a field dictionary and the capture date of every row.
Formats: CSV, JSON
S3
Recurring delivery into an S3 bucket you own, with a folder structure by date and source.
Layout: source/date

Compliance: public data, licence and clear limits

Public data only. We extract information any visitor can see without logging in: prices, listings, product sheets, official publications. We do not collect personal data and we exclude phone numbers, emails and private individuals' names.

Clear licence. Data is delivered under Rastriq's data licence, which describes the permitted use. Before we start we review the source terms and explain any limitation.

No touching third-party systems. We do not access private areas or bypass authentication. If a source requires credentials that are not yours, we do not extract it.

Frequently asked questions about managed scraping

How long does a custom extraction take?

It depends on the source. For sites with an internal API or stable HTML, the first sample can be ready within a few days. Sources with anti-bot protection or many structural variants need more prototyping time.

Who maintains the extraction if the website changes?

We do. Every run validates records and fields, and when the source changes its structure we update the actor. Maintenance is part of the recurring service.

Can I keep the actor?

Yes, we can publish it in your Apify account or keep it in ours and deliver only the data. We agree this in the proposal, depending on whether you prefer code ownership or a managed service.

How do I receive the data?

Through the Apify dataset with API access, as CSV or JSON files, or in an S3 bucket you own. We can also integrate webhooks to notify your system when each run finishes.

Do you extract personal data?

No. Rastriq works with public product data and excludes phone numbers, emails, private individuals' names and other personal data. If a project requires them, we do not take it on.

Tell us which source you need

Describe the site and the fields you are after. We reply with a feasibility assessment and a sample of the proposed schema.

  • Feasibility assessment of the source
  • Proposed schema and delivery channel
  • No commitment and no personal data
Could not send. Please use the contact form instead.
Received. Meanwhile, download the real samples at /en/samples/.

Is your source not on the Store?

Tell us what you need and we will assess it.

Keep exploring