# squid-crawler

> [![npm Package](https://img.shields.io/npm/v/squid-crawler.svg?style=flat-square)](https://www.npmjs.com/package/squid-crawler)

Latest version **1.0.2** (published 2017-07-20) · GPL-3.0 license · 0 weekly downloads

## Install

```sh
npm install squid-crawler
pnpm add squid-crawler
yarn add squid-crawler
bun add squid-crawler
```

## Health

**Score 15/100 (F)** — status: abandoned.

Positive: no vulnerabilities.

Warnings: low downloads; no types; no esm support.

Negative: abandoned; low maintenance score.

## Facts

| | |
|---|---|
| Version | 1.0.2 |
| Published | 2017-07-20 |
| First published | 2017-07-20 |
| Weekly downloads | 0 |
| License | GPL-3.0 |
| TypeScript types | none |
| Module format | CommonJS |
| Dependencies | 19 |
| Known vulnerabilities | 0 (+5 in 3 direct dependencies) |
| Install scripts | no |
| Maintainers | jberlin |

## Links

- npm: https://www.npmjs.com/package/squid-crawler
- npm.io page: https://npm.io/package/squid-crawler

## Dependencies (19)

- [uuid](https://npm.io/package/uuid.md) ^3.1.0
- [chalk](https://npm.io/package/chalk.md) ^2.0.1
- [ramda](https://npm.io/package/ramda.md) ^0.24.1
- [yargs](https://npm.io/package/yargs.md) ^8.0.2
- [lodash](https://npm.io/package/lodash.md) ^4.17.4
- [moment](https://npm.io/package/moment.md) ^2.18.1
- [string](https://npm.io/package/string.md) ^3.3.3
- [bluebird](https://npm.io/package/bluebird.md) ^3.5.0
- [fs-extra](https://npm.io/package/fs-extra.md) ^4.0.0
- [commander](https://npm.io/package/commander.md) ^2.11.0
- [pretty-ms](https://npm.io/package/pretty-ms.md) ^3.0.0
- [validator](https://npm.io/package/validator.md) ^8.0.0
- [remote-dom](https://npm.io/package/remote-dom.md) ^0.0.3
- [parse-domain](https://npm.io/package/parse-domain.md) ^1.1.0
- [pretty-error](https://npm.io/package/pretty-error.md) ^2.1.1
- [eventemitter3](https://npm.io/package/eventemitter3.md) ^2.0.3
- [normalize-url](https://npm.io/package/normalize-url.md) ^1.9.1
- [filenamify-url](https://npm.io/package/filenamify-url.md) ^1.0.0
- [chrome-remote-interface](https://npm.io/package/chrome-remote-interface.md) ^0.24.2

## Recent versions

- 1.0.2 (latest) — 2017-07-20
- 1.0.0 — 2017-07-20

## README

# Squid
[![npm Package](https://img.shields.io/npm/v/squid-crawler.svg?style=flat-square)](https://www.npmjs.com/package/squid-crawler)

Squid is a high fidelity archival crawler that uses Chrome or Chrome Headless.

`Squid` aims to address the need for a high fidelity crawler akin to Heritrix while still
easy enough for the personal archivist to setup and use.

Squid does not seek (at the moment) to dethrone Heritrix as the king of wide archival crawls rather
seeks to address Heritrix's short comings namely
- No JavaScript execution
- Everything is plain text
- Requiring configuration to known how to preserve the web
- Setup time and technical knowledge required of its users

For more information about this see
- [Adapting the Hypercube Model to Archive Deferred Representations and Their Descendants](https://arxiv.org/abs/1601.05142)
- [2012-10-10: Zombies in the Archives](http://ws-dl.blogspot.ca/2012/10/2012-10-10-zombies-in-archives.html)
- [2013-11-28: Replaying the SOPA Protest](http://ws-dl.blogspot.ca/2013/11/2013-11-28-replaying-sopa-protest.html)
- [2015-06-26: PhantomJS+VisualEvent or Selenium for Web Archiving?](http://ws-dl.blogspot.ca/2015/06/2015-06-26-phantomjsvisualevent-or.html)

`Squid` is built using Node.js and [chrome-remote-interface](https://github.com/cyrus-and/chrome-remote-interface).

Can't install Node on your system
then `Squid` highly recommends [WARCreate](http://warcreate.com/) or [WAIL](https://github.com/N0taN3rd/wail/releases).   
WARCreate did this first and if it had not `Squid` would not exist :two_hearts:

If recording the web is what you seek `Squid` highly recommends [Webrecorder](https://webrecorder.io/).


# Out Of The Box Crawls
### Page Only
Preserve the page such that there is no difference when replaying the page from viewing the page in a web browser at preservation time

### Page + Same Domain Links
Page Only option plus preserve all links found on the page that are on the same domain as the page

### Page + All internal and external links
Page + Same Domain Link option plus all links from other domains

#### Crawls Operate In Terms Of A Composite memento
A Memento is an archived copy of web resource [RFC 7089](http://www.rfc-editor.org/info/rfc7089)  The datetime when the copy was archived is called its Memento-Datetime.  A composite memento is a root resource such as an HTML web page and all of the embedded resources (images, CSS, etc.) required for a complete presentation.

More information about this terminology can be found via [ws-dl.blogspot.com](http://ws-dl.blogspot.com/search?q=composite)

# Usage

There are two shell scripts provided to help you use the project at the current stage.

### run-chrome.sh   
You can change the variable `chromeBinary` to point to the Chrome command to use,
that is to launch Chrome via.

The value for `chromeBinary` currently is `google-chrome-beta`

The `remoteDebugPort` variable is used for `--remote-debugging-port=<port>`

Chrome v59 (stable) or v60 (beta) are actively tested on Ubuntu 16.04.

v60 is currently used and known to work well :+1:  

Chrome < v59 will not work.  

No testing is done on canary or google-chrome-unstable so your millage may vary
if you use these versions. Windows sorry your not supported yet.

Takes one argument `headless` if you wish to use Chrome headless otherwise runs Chrome with a head :grinning:

For more information see [Google web dev updates](https://developers.google.com/web/updates/2017/04/headless-chrome).

### run-crawler.sh
Once Chrome has been started you can use `run-crawler.sh`  passing it `-c <path-to-config.json>`

More information can be retrieved by using `-h` or `--help`

The `config.json` file example below is provided beside the two shell scripts without annotations as the annotations (comments) are not valid `json`

```js
{
 // supports page-only, page-same-domain, page-all-links
// crawl only the page, crawl the page and all same domain links,
// and crawl page and all links. In terms of a composite memento
  "mode": "page-same-domain",
 // an array of seeds or a single seed
  "seeds": [
    "http://acid.matkelly.com"
  ],
  "warc": {
    "naming": "url", // currently this is the only option supported do not change.....
    "output": "path" // where do you want the WARCs to be placed. optional defaults to cwd
  },
 // Chrome instance we are to connect to is running on host, port.  
// must match --remote-debugging-port=<port> set when launching chrome.
// localhost is default host when only setting --remote-debugging-port
  "connect": {
    "host": "localhost",
    "port": 9222
  },
// time is in milliseconds
  "timeouts": {
   // wait at maxium 8s for Chrome to navigate to a page
    "navigationTimeout": 8000,
 // wait 7 seconds after page load
    "waitAfterLoad": 7000
  },
// optional auto scrolling of the page. same feature as webrecorders auto-scroll page
// time is in milliseconds and indicates the duration of the scroll
// in proportion to page size. Higher values means longer smooth scrolling, shorter values means faster smooth scroll
 "scroll": 4000
}
```
[![JavaScript Style Guide](https://cdn.rawgit.com/feross/standard/master/badge.svg)](https://github.com/feross/standard)

---
_Source: https://npm.io/package/squid-crawler · Machine-readable twin of the npm.io package page. Health data is recomputed on every publish._
