ExtractionTemplate.
The scraping framework's unit: what a page kind looks like, how far to walk, how to render and navigate, what to extract, how each field is normalised, what makes a row the same row, which source object and mapping the rows become, under which policies, proved by recorded samples
Figure 1ExtractionTemplatespecified
- kind
- definition
- scope
- follows
- key
- its name and version, a tenant’s or the platform’s
- store
- postgres
- family
- sources
Fields#
| field | type | required | note |
|---|---|---|---|
name | text | yes | |
page_kind | one of listing | detail | feed | table | search results | infinite scroll | document | API response | yes | |
scope | CrawlScope | yes | maps onto spider-rs’s configuration |
render | one of static | auto | rendered | yes | |
navigation | BrowsingStep | yes | run on CloakBrowser before extraction: open, click, type, scroll, wait, paginate |
fields | FieldSpec | yes | |
schema | json | no | for the model-assisted reader, every value with its evidence span |
normalisation | NormalisationRule | yes | |
dedup_key | [text] | yes | the fields that make one row the same row across reads |
source_object | SourceObject | yes | |
mapping | Mapping | no | |
filter | ResourceFilter | yes | |
identity_needed | text | no | |
proxy_policy | ProxyPolicy | no | |
freshness | duration | yes | how long a read stays good; a re-read before it is a read of bronze |
attempt_caps | {text:count} | yes | by failure class |
budget | {text:count} | yes | pages, bytes, money |
copied_from | id | no | the platform original a tenant copied |
samples_hash | hash | no | the recorded pages and their expected rows |
state | Readiness | yes | ready only when the samples check passes |
kind_of_entity | KindOfEntity | no | the kind of entity a page of this template shows first |
Routes that use it#
/v1/ask/propose_templatereturns it/v1/act/scrape_templatereturns it