ContentRequest.
Get_content: one address read as the person or the platform, with what to return and what may be downloaded
Figure 1ContentRequestspecified
- kind
- value
- scope
- none
- key
- none
- store
- none
- family
- sources
Fields#
| field | type | required | note |
|---|---|---|---|
address | text | yes | |
returns | [one of article | html | markdown | text | links | metadata | stylesheets | scripts | images | screenshot | pdf | network events] | yes | article is the main content with its title, author, published time and language, as CrawlKit’s extractor returns it |
filter | ResourceFilter | yes | |
as_identity | id | no | a held identity, lent by a grant: its cookie jar, session token, access token or API key |
proxy_policy | ProxyPolicy | no | |
wait_for | text | no | a selector or a time |
cache_for | duration | no | |
credit_cells | [id] | no |
Routes that use it#
/v1/ask/get_contenttakes it as input