ScrapeRequest.
Arbitrary scraping as data: a scope, an extraction of fields or a schema for the model-assisted reader, the source object the rows become, a resource filter, a held identity and a proxy policy
Figure 1ScrapeRequestspecified
- kind
- value
- scope
- none
- key
- none
- store
- none
- family
- sources
Fields#
| field | type | required | note |
|---|---|---|---|
template | id | no | an extraction template; any of its parts may be overridden below |
scope | CrawlScope | no | required when no template |
fields | FieldSpec | no | declared selectors |
schema | json | no | for the model-assisted reader, every value with its evidence span or refused |
source_object | SourceObject | yes | declared inline for one scrape, or a registered row |
mapping | Mapping | no | |
filter | ResourceFilter | yes | |
as_identity | id | no | |
proxy_policy | ProxyPolicy | no | |
credit_cells | [id] | no | |
asked_by | text | no | who asked for the run, as the caller’s kind, or the crawler for an address a walk found; the beat takes what a person or the agent asked before the crawler’s own discoveries, oldest first within each |
Routes that use it#
/v1/ask/scrapetakes it as input/v1/act/crawltakes it as input