
Security News
arXiv Is Rate Limiting Authors Following a Flood of AI Slop Submissions
arXiv now limits authors to two submissions a month as AI slop overwhelms moderators, delays good papers, and sparks debate over applying the limit to everyone.
webreader
Advanced tools
pip install webreader
import webreader
html = '<h1 class="title">Hello</h1><a href="/x">link</a>'
doc = webreader.parse(html)
webreader.select(doc, 'h1.title') # [{'tag': ..., 'attrs': ..., 'text': ..., 'html': ...}]
doc.find('a').get_attr('href') # '/x'
webreader.minify(html) # collapsed whitespace
webreader.pretty(html) # re-indented, readable markup
webreader.to_markdown(html) # convert to Markdown
clean, report = webreader.sanitize(html)
Bytes input is decoded using detected encoding: BOM sniffing first,
then <meta charset> in the first 2 KB, then the configured
encoding fallback.
In non-strict mode, malformed HTML does not raise; the warnings are available on the document:
doc = webreader.parse('<div><p>x</p>') # unclosed <div>
doc.errors # ['unclosed tag(s) at end of input: <div>']
webreader.extract_metadata(html) # title, OG/Twitter, canonical, charset, JSON-LD
webreader.extract_links(html) # links and asset references
webreader.extract_tables(html) # matrices with colspan/rowspan expanded
webreader.extract_text(html) # readable text without boilerplate
webreader.extract_images(html) # src, alt, dimensions, srcset candidates
webreader.extract_forms(html) # form fields with types and defaults
webreader.extract_headings(html) # h1-h6 outline + word count + reading time
webreader.extract_scripts(html) # inline and external script inventory
webreader.extract_stylesheets(html) # link[rel~=stylesheet] and inline styles
Supported syntax: type selectors, *, #id, .class, attribute
selectors ([attr], [attr=value] and the ^=, $=, *=, ~=,
|= operators), the pseudo-classes :first-child, :last-child,
:nth-child(an+b), :not(compound) and :contains(text), compounds
such as div.content#main[href="x"], the combinators whitespace
(descendant), > (child), + (adjacent sibling) and ~ (general
sibling), and comma-separated selector lists.
webreader.select(doc, 'a[href^="https://"]')
webreader.select(doc, 'ul > li:nth-child(2n+1):not(.skip)')
webreader.select(doc, 'h2 + p, blockquote p:contains("note")')
Invalid or unsupported selectors raise SelectorSyntaxError.
Nodes support a mutation API, and the edited tree can be re-serialized:
doc = webreader.parse('<div><a href="/x">link</a></div>')
link = doc.find('a')
link.set_attr('rel', 'noopener').remove_attr('class')
new = webreader.Node('p')
new.children.append('added text')
doc.find('div').append_child(new)
link.unwrap() # replace the <a> with its inner text
doc.to_html() # '<div>link<p>added text</p></div>'
Node also provides remove() and replace_with(*nodes).
to_markdown renders headings, paragraphs, links, images,
emphasis/strong/strikethrough, inline and fenced code (with language
detection from class="language-..."), nested lists, blockquotes,
horizontal rules and tables.
webreader.to_markdown('<h1>T</h1><p>Hi <b>bold</b></p>')
# '# T\n\nHi **bold**'
sanitize() returns (clean_html, report):
script, style, noscript, template are dropped entirelyon* handlers removedjavascript:, vbscript: and non-image data: URLs blockedpython -m webreader parse -f input.html
python -m webreader select -f input.html -s 'a[href]'
python -m webreader links -f input.html --base-url https://example.com
python -m webreader meta -f input.html
python -m webreader tables -f input.html --format csv
python -m webreader text -f input.html
python -m webreader images -f input.html
python -m webreader forms -f input.html
python -m webreader minify -f input.html --out min.html
python -m webreader pretty -f input.html --out pretty.html
python -m webreader md -f input.html --out page.md
python -m webreader sanitize -f input.html --out clean.html --allowed p,a,ul
Use -f - to read from stdin. Every output command accepts --out
to write to a file instead of stdout.
| Key | Default | Purpose |
|---|---|---|
strip_whitespace | True | collapse whitespace in text nodes |
encoding | utf-8 | fallback when bytes input has no detectable encoding |
max_depth | 100 | maximum element nesting |
preserve_comments | False | keep comments in the tree |
strict_mode | False | raise on malformed HTML |
remove_comments | True | strip comments when minifying |
allowed_tags | p, br, b, i, a, ul, ol, li, div, span | sanitize whitelist |
max_file_size_mb | 10 | reject oversized input |
cache_size | 50 | parse cache capacity |
HTMLPaferError (base), ConfigLoadError, SanitizationError,
SelectorSyntaxError.
FAQs
Tree-based HTML reader for Python.
The pypi package webreader receives a total of 111 weekly downloads. As such, webreader popularity was classified as not popular.
We found that webreader demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Security News
arXiv now limits authors to two submissions a month as AI slop overwhelms moderators, delays good papers, and sparks debate over applying the limit to everyone.

Research
/Security News
A new GhostAction wave hits hundreds of GitHub repos, expanding CI/CD secret theft to cloud and AI credentials in source code and git history.

Research
/Security News
Tensorlake npm SDK version 0.5.144 was compromised in a ChainDrop / Shai-Hulud attack, delivering credential-stealing malware.