Scraping and pricing datasets with a documented schema, downloadable samples with no sign-up and a machine-readable catalog. This page explains how an agent, an assistant or an LLM pipeline can find them, read them and, where needed, launch the extraction.
Rastriq does not run its own API or MCP server: the data is read from static files on this website or requested through the Apify platform, where the actors are published.
Each actor returns the native fields of its source, and every brand or vertical page includes a field dictionary with type and example. An agent can read that dictionary before requesting data and know which columns to expect, which unit the amounts use and which field tells a tax-inclusive price from a tax-exclusive one.
When a delivery combines several sources we apply a common schema: the same columns and types for different brands, portals or gazettes. Each row keeps its source URL and capture date, and quality rules mark doubtful cases in a flags field instead of dropping the row. That leaves the agent to decide what to do with them.
| Field | Type | Meaning |
|---|---|---|
| brand / model | string | Brand and model as the source publishes them |
| market | string | Country or market of the configurator or portal |
| base_price | number | Price published by the source, in the row currency |
| currency | string | ISO 4217 currency code |
| source_url | string | Public URL the row comes from |
| scraped_at | datetime | Capture date and time in ISO 8601 |
| qa_flags | array | Quality warnings: the row is flagged and delivered, not filtered |
Example of the common schema used by normalized datasets; the native field names of each actor are on its source page. More in data normalization and in the OEM configurators hub.
Every source has a public sample of up to 100 rows, sanitized and free of personal fields, in JSON and CSV. It follows the same schema you get in the full delivery, so you can test a parser, a prompt or an agent tool before buying anything.
The datasets come from public sources and contain product and listing attributes, not personal data; samples are sanitized before publication. The data licence allows unlimited internal use, derived products, and training, fine-tuning or evaluating AI models. It does not allow redistributing the raw file or reselling it as an equivalent data product.
The data reflects what each platform published at capture time and is not guaranteed to be complete. The robots.txt file allows the main AI and AI-search bots to crawl the whole site: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended and CCBot.
Yes. An agent can read, without an account, the catalog in catalog.json, the llms.txt index and the JSON or CSV samples in the muestras folder. To run a new extraction it needs an Apify account and must use the Apify API or MCP server, because the actors are published there.
No. Rastriq does not operate its own API or MCP server. The data is served as static files on rastriq.es and as actors on the Apify Store. To automate runs you use the Apify REST API or its MCP server, both configured with your own account and token.
Public samples download as JSON and CSV. Datasets generated by an actor on Apify can also be exported to Excel, XML, HTML and RSS. Recurring deliveries of a licensed dataset come as CSV or JSON, including to your own S3 bucket, with a field dictionary.
Yes. The Rastriq data licence allows using the datasets to train, fine-tune or evaluate machine-learning models, including ones you later commercialise. It does not allow redistributing the raw file or reselling it as an equivalent data product. Derived products are allowed.
No. The datasets are built from public product and listing attributes, not personal data. Samples are sanitized before publication: phone numbers, emails, seller names, authors and chassis numbers are excluded. If a source exposes such fields, they are left out of the samples and the pages.
At rastriq.es/catalog.json, a schema.org DataCatalog with every dataset, its coverage, its fields, the sample link and the licence. Each entry links its sample and its public Apify actor. For a Markdown index written for language models there is rastriq.es/llms.txt, and the full content as clean text is in llms-full.txt.
Tell us sources, markets and frequency and we will tell you what we can deliver, whether as an Apify actor, a licensed dataset or a custom project.