Jabbah

Make broken HTML useful.

Parse real-world markup into a traversable tree. Query what matters, clean unsafe content, and extract an article without a browser or native extension.

document.htmlparse / select / clean

Incoming markup

<main>
  <h1>Field notes</h1>
  <p class="lead">A story worth keeping
  <img src="https://remote.test/photo.png">
  <script>track()</script>
</main>

Safe document tree

documenthtmlbodymainh1 Field notesp.lead A story worth keepingimg remote source blocked
Jabbah::Sanitize.clean(document, profile: :feed)1 remote image blocked

A small toolkit for documents.

The parser tolerates imperfect input without turning into a browser engine.

Parse

Documents and fragments with implicit element closing and charset decoding.

Select

Find nodes by tag, id, class, attributes, and common structural selectors.

Extract

Find article-like content and a readable title, byline, and excerpt.

Serialize

Clone a document and write safe, escaped HTML.

Clean before display.

Docs, feed, and mail profiles drop active elements and unsafe URLs. Remote images are blocked by default; cid: images survive for mail readers.

Sanitization returns a clone. The original parsed document remains available for inspection.

require "jabbah"
document = Jabbah.parse(html)
safe = Jabbah::Sanitize.clean(document, profile: :feed)
puts safe.to_html
puts safe.blocked_count