New:Introducing Socket Scanning for VS Code Marketplace Extensions.Learn more →
Get Started

webreader

Package Overview
Dependencies
Maintainers
1
Versions
1
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

webreader

Tree-based HTML reader for Python.

pipPyPI
Version
2.3.8
Weekly downloads
110
Maintainers
1
Weekly downloads
 
Created

Python Dependencies Tests

Features

  • Tree-based parsing into an editable document tree
  • CSS selectors: combinators, attributes, pseudo-classes, selector lists
  • Extractors: metadata, links, tables, text, images, forms, headings, scripts, stylesheets
  • Sanitization with a removal report
  • Minify, pretty-print, Markdown conversion
  • JSON / CSV / Markdown export
  • CLI tool
  • Standard library only at runtime

Installation

pip install webreader

Quickstart

import webreader

html = '<h1 class="title">Hello</h1><a href="/x">link</a>'
doc = webreader.parse(html)

webreader.select(doc, 'h1.title')   # [{'tag': ..., 'attrs': ..., 'text': ..., 'html': ...}]
doc.find('a').get_attr('href')       # '/x'

webreader.minify(html)              # collapsed whitespace
webreader.pretty(html)              # re-indented, readable markup
webreader.to_markdown(html)         # convert to Markdown

clean, report = webreader.sanitize(html)

Bytes input is decoded using detected encoding: BOM sniffing first, then <meta charset> in the first 2 KB, then the configured encoding fallback.

In non-strict mode, malformed HTML does not raise; the warnings are available on the document:

doc = webreader.parse('<div><p>x</p>')   # unclosed <div>
doc.errors                                 # ['unclosed tag(s) at end of input: <div>']

Extraction

webreader.extract_metadata(html)    # title, OG/Twitter, canonical, charset, JSON-LD
webreader.extract_links(html)       # links and asset references
webreader.extract_tables(html)      # matrices with colspan/rowspan expanded
webreader.extract_text(html)        # readable text without boilerplate
webreader.extract_images(html)      # src, alt, dimensions, srcset candidates
webreader.extract_forms(html)       # form fields with types and defaults
webreader.extract_headings(html)    # h1-h6 outline + word count + reading time
webreader.extract_scripts(html)     # inline and external script inventory
webreader.extract_stylesheets(html) # link[rel~=stylesheet] and inline styles

CSS selectors

Supported syntax: type selectors, *, #id, .class, attribute selectors ([attr], [attr=value] and the ^=, $=, *=, ~=, |= operators), the pseudo-classes :first-child, :last-child, :nth-child(an+b), :not(compound) and :contains(text), compounds such as div.content#main[href="x"], the combinators whitespace (descendant), > (child), + (adjacent sibling) and ~ (general sibling), and comma-separated selector lists.

webreader.select(doc, 'a[href^="https://"]')
webreader.select(doc, 'ul > li:nth-child(2n+1):not(.skip)')
webreader.select(doc, 'h2 + p, blockquote p:contains("note")')

Invalid or unsupported selectors raise SelectorSyntaxError.

Tree editing

Nodes support a mutation API, and the edited tree can be re-serialized:

doc = webreader.parse('<div><a href="/x">link</a></div>')
link = doc.find('a')
link.set_attr('rel', 'noopener').remove_attr('class')

new = webreader.Node('p')
new.children.append('added text')
doc.find('div').append_child(new)

link.unwrap()        # replace the <a> with its inner text
doc.to_html()        # '<div>link<p>added text</p></div>'

Node also provides remove() and replace_with(*nodes).

Markdown conversion

to_markdown renders headings, paragraphs, links, images, emphasis/strong/strikethrough, inline and fenced code (with language detection from class="language-..."), nested lists, blockquotes, horizontal rules and tables.

webreader.to_markdown('<h1>T</h1><p>Hi <b>bold</b></p>')
# '# T\n\nHi **bold**'

Sanitization

sanitize() returns (clean_html, report):

  • tags outside the whitelist are unwrapped, inner content kept
  • script, style, noscript, template are dropped entirely
  • attributes filtered through a safe-list, on* handlers removed
  • javascript:, vbscript: and non-image data: URLs blocked
  • the report lists removed tags/attributes, blocked URLs and detected threats

CLI

python -m webreader parse -f input.html
python -m webreader select -f input.html -s 'a[href]'
python -m webreader links -f input.html --base-url https://example.com
python -m webreader meta -f input.html
python -m webreader tables -f input.html --format csv
python -m webreader text -f input.html
python -m webreader images -f input.html
python -m webreader forms -f input.html
python -m webreader minify -f input.html --out min.html
python -m webreader pretty -f input.html --out pretty.html
python -m webreader md -f input.html --out page.md
python -m webreader sanitize -f input.html --out clean.html --allowed p,a,ul

Use -f - to read from stdin. Every output command accepts --out to write to a file instead of stdout.

Default settings

KeyDefaultPurpose
strip_whitespaceTruecollapse whitespace in text nodes
encodingutf-8fallback when bytes input has no detectable encoding
max_depth100maximum element nesting
preserve_commentsFalsekeep comments in the tree
strict_modeFalseraise on malformed HTML
remove_commentsTruestrip comments when minifying
allowed_tagsp, br, b, i, a, ul, ol, li, div, spansanitize whitelist
max_file_size_mb10reject oversized input
cache_size50parse cache capacity

Exceptions

HTMLPaferError (base), ConfigLoadError, SanitizationError, SelectorSyntaxError.

FAQs

Related posts