ExtractedArticle.
The main content of a page as the carried extractor returns it: title, author, published time with its basis, language, word count, the extractor used and its quality score
Figure 1ExtractedArticlespecified
- kind
- value
- scope
- none
- key
- none
- store
- none
- family
- sources
Fields#
| field | type | required | note |
|---|---|---|---|
title | text | yes | |
author | text | no | |
published_at | time | no | |
basis_of_time | one of published | scheduled | first seen | yes | |
content_markdown | text | yes | |
language | text | yes | |
word_count | count | yes | |
extractor | text | yes | which extractor produced it |
quality | share | yes | prose quality in [0, 1]; 0 when unknown |
content_hash | hash | yes |
Routes that use it#
No route takes or returns ExtractedArticle directly. The record holds it, and the objects that point at it reach it.