Make broken HTML useful.
Parse real-world markup into a traversable tree. Query what matters, clean unsafe content, and extract an article without a browser or native extension.
Incoming markup
<main>
<h1>Field notes</h1>
<p class="lead">A story worth keeping
<img src="https://remote.test/photo.png">
<script>track()</script>
</main>Safe document tree
documenthtmlbodymainh1 Field notesp.lead A story worth keepingimg remote source blocked
Jabbah::Sanitize.clean(document, profile: :feed)1 remote image blocked
A small toolkit for documents.
The parser tolerates imperfect input without turning into a browser engine.
Parse
Documents and fragments with implicit element closing and charset decoding.
Select
Find nodes by tag, id, class, attributes, and common structural selectors.
Extract
Find article-like content and a readable title, byline, and excerpt.
Serialize
Clone a document and write safe, escaped HTML.
Clean before display.
Docs, feed, and mail profiles drop active elements and unsafe URLs. Remote images are blocked by default; cid: images survive for mail readers.
Sanitization returns a clone. The original parsed document remains available for inspection.
require "jabbah"
document = Jabbah.parse(html)
safe = Jabbah::Sanitize.clean(document, profile: :feed)
puts safe.to_html
puts safe.blocked_count