CrawlScope.
What a scrape may walk: the seeds, the address patterns allowed and denied, the greatest depth and count of pages, whether the sitemap drives discovery, and the pagination rule
Figure 1CrawlScopespecified
- kind
- value
- scope
- none
- key
- none
- store
- none
- family
- sources
Fields#
| field | type | required | note |
|---|---|---|---|
seeds | [text] | yes | |
allow_patterns | [text] | yes | |
deny_patterns | [text] | yes | |
max_depth | count | yes | |
max_pages | count | yes | |
use_sitemap | bool | yes | |
pagination | text | no | a selector or a pattern for the next page |
Routes that use it#
No route takes or returns CrawlScope directly. The record holds it, and the objects that point at it reach it.