# htmlgrabr

> A Node.js library to grab and clean HTML content.

Latest version **1.1.1** (published 2021-01-02) · MIT license · 0 weekly downloads

## Install

```sh
npm install htmlgrabr
pnpm add htmlgrabr
yarn add htmlgrabr
bun add htmlgrabr
```

## Health

**Score 30/100 (F)** — status: abandoned.

Positive: has types; esm support; no vulnerabilities; high quality score.

Warnings: low downloads.

Negative: abandoned; low maintenance score.

## Facts

| | |
|---|---|
| Version | 1.1.1 |
| Published | 2021-01-02 |
| First published | 2018-11-04 |
| Weekly downloads | 0 |
| License | MIT |
| TypeScript types | bundled |
| Module format | ESM + CommonJS |
| Node | >=12.0.0 |
| Dependencies | 13 |
| Unpacked size | 73.9 KB |
| Known vulnerabilities | 0 (+17 in 2 direct dependencies) |
| Install scripts | no |
| GitHub stars | 5 |
| Author | Nicolas Carlier |
| Maintainers | ncarlier |

## Links

- npm: https://www.npmjs.com/package/htmlgrabr
- Repository: https://github.com/ncarlier/htmlgrabr
- Homepage: https://github.com/ncarlier/htmlgrabr#readme
- Issues: https://github.com/ncarlier/htmlgrabr/issues
- npm.io page: https://npm.io/package/htmlgrabr

## Dependencies (13)

- [jsdom](https://npm.io/package/jsdom.md) ^16.4.0
- [parse5](https://npm.io/package/parse5.md) ^6.0.1
- [pretty](https://npm.io/package/pretty.md) ^2.0.0
- [dompurify](https://npm.io/package/dompurify.md) ^2.2.6
- [mime-types](https://npm.io/package/mime-types.md) ^2.1.28
- [node-fetch](https://npm.io/package/node-fetch.md) ^2.6.1
- [@types/jsdom](https://npm.io/package/@types/jsdom.md) ^16.2.5
- [html2plaintext](https://npm.io/package/html2plaintext.md) ^2.1.2
- [@types/dompurify](https://npm.io/package/@types/dompurify.md) ^2.1.0
- [@types/mime-types](https://npm.io/package/@types/mime-types.md) ^2.1.0
- [@types/node-fetch](https://npm.io/package/@types/node-fetch.md) ^2.5.7
- [@mozilla/readability](https://npm.io/package/@mozilla/readability.md) ^0.4.0
- [@types/mozilla-readability](https://npm.io/package/@types/mozilla-readability.md) ^0.2.0

## Recent versions

- 1.1.1 (latest) — 2021-01-02
- 1.1.0 — 2021-01-02
- 1.0.4 — 2020-05-09
- 1.0.3 — 2020-05-09
- 1.0.1 — 2018-11-05
- 1.0.0 — 2018-11-04

## README

# HTMLGrabr library

[![Travis](https://img.shields.io/travis/ncarlier/htmlgrabr.svg)](https://travis-ci.org/ncarlier/htmlgrabr)
[![Coverage Status](https://coveralls.io/repos/github/ncarlier/htmlgrabr/badge.svg?branch=master)](https://coveralls.io/github/ncarlier/htmlgrabr?branch=master)
[![Donate](https://img.shields.io/badge/donate-paypal-blue.svg)](https://paypal.me/nunux)

A Node.js library to grab and clean HTML content.

### Features

- Extract page content from an URL (`HTMLGrabr.grabURL(url: URL): GrabbedPage`)
- Extract page content from a string (`HTMLGrabr.grab(s: string): GrabbedPage`)
- Extract Open Graph properties
- Clean the page content:
  - Extract main HTML content using [mozilla-readability](https://github.com/mozilla/readability)
  - Sanitize HTML content using [DOMPurify](https://github.com/cure53/DOMPurify), with some extras:
    - Remove unwanted links or images
    - Remove pixel tracker
    - Remove unwanted attributes (such as `style`, `class`, `id`, ...)
    - And more

### Usage

```bash
npm install --save htmlgrabr
```

The in your code:

```javascript
const HTMLGrabr = require('htmlgrabr').HTMLGrabr
const { URL } = require('url')

const grabber = new HTMLGrabr()

grabber.grabUrl(new URL('https://about.readflow.app'))
  .then(page => {
    console.log(page)
  }, err => {
    console.error(err)
  })
```

### API

Create new instance:

```js
const HTMLGrabr = require('htmlgrabr').HTMLGrabr
const grabber = new HTMLGrabr(config)
```

Configuration object:

```typescript
interface GrabberConfig {
  debug?: boolean                     // Print debug logs if true
  pretty?: boolean                    // Beautify HTML content if true
  isBlockedHost?: BlockedHostCtrlFunc // Function used to detect unwanted URLs
  rewriteURL?: URLRewriterFunc        // Function used to rewrite HTML src attributes
  rules?: Map<string, Rule>           // Rule definitions (see below)
  headers?: Headers                   // HTTP headers to set
}
```

Rule definition:

```typescript
export interface Rule {
  selector: string             // HTML query selector
  type: 'redirect' | 'content' // Rule type:
  // - 'redirect' will use 'src' or 'href' attribute to redirect content extraction
  // - 'content' to specify content to extract
}
```

Grab a page:

```js
const result = grabber.grabUrl(new URL('https://...'))
```

Result object:

```typescript
interface GrabbedPage {
  title: string        // Page title
  url: string | null   // Source URL
  image: string | null // Page illustration
  html: string         // HTML content
  text: string         // Text content (from HTML)
  excerpt: string      // Excerpt (from meta data or HTML)
  length: number       // Read length
  images: ImageMeta[]  // Embedded image URLs
}
```

---

---
_Source: https://npm.io/package/htmlgrabr · Machine-readable twin of the npm.io package page. Health data is recomputed on every publish._
