rastriq. Talk to us →
Rastriq / For AI agents

Rastriq for AI agents and LLMs.

Scraping and pricing datasets with a documented schema, downloadable samples with no sign-up and a machine-readable catalog. This page explains how an agent, an assistant or an LLM pipeline can find them, read them and, where needed, launch the extraction.

46
Datasets in the catalog
44
JSON and CSV samples
109
Public actors on Apify

Three ways to query the data

Rastriq does not run its own API or MCP server: the data is read from static files on this website or requested through the Apify platform, where the actors are published.

Static files, no account
An agent can read /catalog.json, /llms.txt and the samples in /muestras/ with a plain GET request. No key or account is needed and nothing is charged. This is the recommended way to discover what data exists and what it looks like.
No token · free
Apify REST API
Every public actor can be run and its dataset read through the Apify API, for example with POST /v2/acts/rastriq~{actor}/run-sync-get-dataset-items and a token from your account. Results come out as JSON, CSV, Excel, XML, HTML or RSS. Pay-per-result actors are charged as stated on their Store page.
Apify token · pay per result
Apify MCP server
Apify provides an MCP server (mcp.apify.com) that lets a compatible client search and run Store actors, including Rastriq's. It is an Apify integration, not a Rastriq-specific connection, and you set it up in your own client with your own account.
Apify account · MCP clients

Documented fields and stable types

Each actor returns the native fields of its source, and every brand or vertical page includes a field dictionary with type and example. An agent can read that dictionary before requesting data and know which columns to expect, which unit the amounts use and which field tells a tax-inclusive price from a tax-exclusive one.

When a delivery combines several sources we apply a common schema: the same columns and types for different brands, portals or gazettes. Each row keeps its source URL and capture date, and quality rules mark doubtful cases in a flags field instead of dropping the row. That leaves the agent to decide what to do with them.

FieldTypeMeaning
brand / modelstringBrand and model as the source publishes them
marketstringCountry or market of the configurator or portal
base_pricenumberPrice published by the source, in the row currency
currencystringISO 4217 currency code
source_urlstringPublic URL the row comes from
scraped_atdatetimeCapture date and time in ISO 8601
qa_flagsarrayQuality warnings: the row is flagged and delivered, not filtered

Example of the common schema used by normalized datasets; the native field names of each actor are on its source page. More in data normalization and in the OEM configurators hub.

Real samples to test an agent

Every source has a public sample of up to 100 rows, sanitized and free of personal fields, in JSON and CSV. It follows the same schema you get in the full delivery, so you can test a parser, a prompt or an agent tool before buying anything.

OEM brands

Volkswagen100 rows
JSONCSV
BMW100 rows
JSONCSV
Mercedes-Benz100 rows
JSONCSV
Audi100 rows
JSONCSV
Toyota100 rows
JSONCSV
Hyundai100 rows
JSONCSV
Kia100 rows
JSONCSV
BYD100 rows
JSONCSV
MINI100 rows
JSONCSV
Stellantis100 rows
JSONCSV
Ford100 rows
JSONCSV
Volvo100 rows
JSONCSV
Peugeot100 rows
JSONCSV
Citroën100 rows
JSONCSV
Dacia & Renault100 rows
JSONCSV
Nissan100 rows
JSONCSV
Honda100 rows
JSONCSV
Mazda100 rows
JSONCSV
Lexus100 rows
JSONCSV
CUPRA100 rows
JSONCSV
Škoda100 rows
JSONCSV
MG100 rows
JSONCSV
Porsche100 rows
JSONCSV
Polestar100 rows
JSONCSV
Smart100 rows
JSONCSV
Lynk & Co76 rows
JSONCSV

Other sources

Coches.net58 rows
JSONCSV
Wallapop (coches)100 rows
JSONCSV
Milanuncios100 rows
JSONCSV
Chileautos100 rows
JSONCSV
AutoTrader UK20 rows
JSONCSV
Car & Classic40 rows
JSONCSV
MachineryTrader100 rows
JSONCSV
MachineryZone100 rows
JSONCSV
Surplex100 rows
JSONCSV
Tiebaobei5 rows
JSONCSV
XCMG Global5 rows
JSONCSV
BOE100 rows
JSONCSV
Amazon reviews24 rows
JSONCSV
Google Ads Transparency Center100 rows
JSONCSV
TikTok Creative Center (hashtags)50 rows
JSONCSV
TikTok search60 rows
JSONCSV
Google Maps (business listings)50 rows
JSONCSV
Europages15 rows
JSONCSV

Catalog and files built for LLMs

catalog.json
schema.org DataCatalog listing every dataset: name, description, source, coverage, fields, sample, Apify actor and licence.
llms.txt
Markdown index of the data products, OEM brands, guides and samples, with absolute URLs.
llms-full.txt
Clean text of the commercial pages, the OEM hub and the guides, concatenated and without menus.
Apify Store
Rastriq public actors, with their input, their output and example tasks.

Licence and responsible data

The datasets come from public sources and contain product and listing attributes, not personal data; samples are sanitized before publication. The data licence allows unlimited internal use, derived products, and training, fine-tuning or evaluating AI models. It does not allow redistributing the raw file or reselling it as an equivalent data product.

The data reflects what each platform published at capture time and is not guaranteed to be complete. The robots.txt file allows the main AI and AI-search bots to crawl the whole site: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended and CCBot.

Questions from agents and the people who build them

Can an AI agent query Rastriq data?

Yes. An agent can read, without an account, the catalog in catalog.json, the llms.txt index and the JSON or CSV samples in the muestras folder. To run a new extraction it needs an Apify account and must use the Apify API or MCP server, because the actors are published there.

Does Rastriq have its own API or MCP server?

No. Rastriq does not operate its own API or MCP server. The data is served as static files on rastriq.es and as actors on the Apify Store. To automate runs you use the Apify REST API or its MCP server, both configured with your own account and token.

Which formats is the data delivered in?

Public samples download as JSON and CSV. Datasets generated by an actor on Apify can also be exported to Excel, XML, HTML and RSS. Recurring deliveries of a licensed dataset come as CSV or JSON, including to your own S3 bucket, with a field dictionary.

Can I use the data to train or evaluate an AI model?

Yes. The Rastriq data licence allows using the datasets to train, fine-tune or evaluate machine-learning models, including ones you later commercialise. It does not allow redistributing the raw file or reselling it as an equivalent data product. Derived products are allowed.

Do the datasets contain personal data?

No. The datasets are built from public product and listing attributes, not personal data. Samples are sanitized before publication: phone numbers, emails, seller names, authors and chassis numbers are excluded. If a source exposes such fields, they are left out of the samples and the pages.

Where is the machine-readable list of datasets?

At rastriq.es/catalog.json, a schema.org DataCatalog with every dataset, its coverage, its fields, the sample link and the licence. Each entry links its sample and its public Apify actor. For a Markdown index written for language models there is rastriq.es/llms.txt, and the full content as clean text is in llms-full.txt.

Need a specific dataset for your agent?

Tell us sources, markets and frequency and we will tell you what we can deliver, whether as an Apify actor, a licensed dataset or a custom project.

Explore the data