# data-kraken

> A command line tool that fetches info about users, commits, repositories, Docker images and npm dependencies from GitHub

Latest version **2.1.0** (published 2022-12-19) · MIT license · 0 weekly downloads

## Install

```sh
npm install data-kraken
pnpm add data-kraken
yarn add data-kraken
bun add data-kraken
```

Provides the commands `0`, `1`, `2`, `3`, `4`, `5`, `6`, `7`, `8`, `9`, `10`, `11`, `12`, `13`, `14`, `15`, `16`, `17`.

## Health

**Score 15/100 (F)** — status: abandoned.

Positive: no vulnerabilities.

Warnings: low downloads; no types; no esm support.

Negative: abandoned; low maintenance score.

## Facts

| | |
|---|---|
| Version | 2.1.0 |
| Published | 2022-12-19 |
| First published | 2022-11-29 |
| Weekly downloads | 0 |
| License | MIT |
| TypeScript types | none |
| Module format | CommonJS |
| Dependencies | 8 |
| Unpacked size | 2 MB |
| Known vulnerabilities | 0 |
| Install scripts | yes |
| Maintainers | pahund |

## Links

- npm: https://www.npmjs.com/package/data-kraken
- npm.io page: https://npm.io/package/data-kraken

## Dependencies (8)

- [yaml](https://npm.io/package/yaml.md) ^2.1.3
- [debug](https://npm.io/package/debug.md) ^4.3.4
- [semver](https://npm.io/package/semver.md) ^7.3.8
- [commander](https://npm.io/package/commander.md) ^9.4.1
- [@ladjs/env](https://npm.io/package/@ladjs/env.md) ^4.0.0
- [supports-color](https://npm.io/package/supports-color.md) ^9.2.3
- [@fast-csv/format](https://npm.io/package/@fast-csv/format.md) ^4.3.5
- [node-fetch-cache](https://npm.io/package/node-fetch-cache.md) ^3.0.5

## Recent versions

- 2.1.0 (latest) — 2022-12-19
- 2.1.0-beta.4 (beta) — 2022-12-19
- 2.1.0-beta.3 — 2022-12-19
- 2.1.0-beta.2 — 2022-12-16
- 2.1.0-beta.1 — 2022-12-15
- 2.1.0-beta.0 — 2022-12-15
- 2.0.1 — 2022-12-13
- 2.0.1-beta.1 — 2022-12-13
- 2.0.1-beta.0 — 2022-12-13
- 2.0.0 — 2022-12-13
- 2.0.0-beta.15 — 2022-12-13
- 2.0.0-beta.14 — 2022-12-13
- 2.0.0-beta.13 — 2022-12-13
- 2.0.0-beta.12 — 2022-12-13
- 2.0.0-beta.11 — 2022-12-13
- … 15 more at https://npm.io/package/data-kraken/versions

## README

<img alt="Data Kraken Logo" src="docs/images/data-kraken-logo.png" style="float: right">

# Data Kraken

A command line tool that fetches info about users, commits, repositories, Docker images and npm dependencies from
GitHub.

## Prerequisites

This tool runs with [Node.js](https://nodejs.org/). Make sure you have an up-to-date version installed.

## Preparation

You need to have a
[personal access token](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/creating-a-personal-access-token)
for the orgs and repositories on GitHub that you want to examine.

<details>
  <summary>Getting a token – click here for a step by step guide</summary>

- Go to your [GitHub token settings page](../../../../settings/tokens):

![Screenshot: GitHub personal access tokens configuration page](docs/images/screenshot-github-personal-access-token01.png)

- Click on “Generate new token (classic)”

![Screenshot: Adding a new personal access token](docs/images/screenshot-github-personal-access-token02.png)

- Enter any name you find suitable
- Tick the checkboxes for access to “repo” and “user”
- Confirm with “Generate token”
- Copy the personal access token to a safe place for later use

Refer to the
[instructions](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/creating-a-personal-access-token)
on GitHub for further information on this.

</details>

## Installation

Install the _data-kraken_ as a global command line tool using npm like so:

```
npm install -g data-kraken
```

### Create a configuration file

Before you can use _data-kraken_, you have to set it up so that it knows yours the GitHub personal access token
you've created (see chapter [Preparation](#preparation)).

For this, you create a configuration file with in your home directory, like so:

```
echo DK_ACCESS_TOKEN=personal-access-token123 > $HOME/.data-kraken
```

### GitHub Enterprise users

By default, _data-kraken_ uses the API of the public GitHub, [github.com](https://github.com/). If your company
is hosting its own GitHub Enterprise instance, like we do at [Adevinta](https://www.adevinta.com/), add the
`DK_BASE_URL` option to your _.data-kraken_ config file, for example:

```
DK_BASE_URL=github.es.ecg.tools
```

## How to run

With the access token in place as described in the previous chapters, run the _data-kraken_ on the command line
like so:

```
data-kraken
```

(this will display a help message to get you on your way)

## Getting the latest version

By default, once installed, _data-kraken_ will always run the locally installed version. If you don't want to
miss the latest features, run this every once in a while to get the latest version:

```
npm update -g data-kraken
```

## Usage examples

### Tech Debt

Shows a technical debt score for one or more GitHub repositories.

```
data-kraken tech-debt --org mobile-de --repo consumer-fe
```

The repository parameter is optional:

- if repository is provided, the output will show the score along with some improvement hints
- if repository is omitted, the output will show a list of all the repos in the org, ranked according to their
  tech debt score

#### More info

Run `--help` for more info on the `tech-debt` command:

```
data-kraken help tech-debt
```

### Guilds

Shows the guilds associated with repos – backend, frontend, data, qa, android, ios or devops.

```
data-kraken guild --org mobile-de --repo consumer-fe
```

The repository parameter is optional:

- if repository is provided, the output will show the repo's associated guilds
- if repository is omitted, the output will show a list of all guilds found in the org, along with the repos
  associated with them

Many commands, including the `guild` command, have a `--guilds` option that lets you specify guilds. The
generated output will then only show data from repos associated with the specified guilds.

#### More info

Run `--help` for more info on the `guild` command:

```
data-kraken help guild
```

### Inactive

Shows the level of inactivity of GitHub repository. The inactivity score is a value from 0 to 100, 0 being a repo
that currently gets updated every day and 100 being a repo that has not been updated in a very long time.

```
data-kraken inactive --org mobile-de --repo consumer-fe
```

The repository parameter is optional:

- if repository is provided, the output will show the repo's inactivity score along with additional info on how
  the score is componsed
- if repository is omitted, the output will show a list of all the repos in the org, ranked according to their
  level of inactivity

#### More info

Run `--help` for more info on the `inactive` command:

```
data-kraken help inactive
```

### Docker images

Shows the Docker images used in the specified GitHub org and repository, found in a search of all the Dockerfiles
in each repo.

```
data-kraken docker-images --org mobile-de --repo consumer-fe
```

Repository is optional, if omitted, the whole org is searched.

#### Search expressions

You can pass a regular expression to match the images against. In the simplest usage example, the expression can
just be a search term:

```
data-kraken docker-images --org mobile-de node
```

This will give you a list of repositories that use a Node.js image.

More advanced example:

```
data-kraken docker-images --org mobile-de "^.+/shared/node1[46].+$"
```

This will list all the repos that use _dock.es.ecg.tools/shared/node14_ or _dock.es.ecg.tools/shared/node16_, but
not _dock.es.ecg.tools/shared/node12_.

- See also: [Regular expressions](#regular-expressions)

#### More info

Run `--help` for more info on the `docker-images` command:

```
data-kraken help docker-images
```

### Npm packages

Shows the npm packages that repositories are dependent on according to their _package.json_ files.

```
data-kraken npm-packages --org mobile-de --repo consumer-fe
```

Repository is optional, if omitted, the whole org is searched.

#### Search expressions

You can pass one or two regular expressions to match the package names or versions against. In the simplest usage
example, the expression can just be a search term:

```
data-kraken npm-packages --org mobile-de react
```

…gives you results for packages that have “react” in them (e.g. _react_, _react-dom_, _react-router_, etc.).

More advanced example:

```
data-kraken npm-packages --org mobile-de ^react$ "^[~^]*1[68]{1}"
```

…gives you results for precise package “react” with major versions 16 or 18.

- See also: [Regular expressions](#regular-expressions)

#### More info

Run `--help` for more info on the `npm-packages` command:

```
data-kraken help npm-packages
```

### Repos

Shows info about the repositories a user contributed to in “pretty print” on the console:

```
data-kraken repos patrick-hund
```

You can use the _--org_ option to constrain output to a specific GitHub org:

```
data-kraken repos --org mobile-de patrick-hund
```

You can specify multiple users:

```
data-kraken repos patrick-hund daniel-korger uwe-loydl
```

- Caveat: [Data time range](#data-time-range)

#### More info

Run `--help` for more info on the `repos` command:

```
data-kraken help repos
```

### Files

Shows info about what kinds of files the user modified (frontend or backend):

```
data-kraken files patrick-hund
```

As with the [repos](#repos) command, you can specify multiple users and a GitHub org. In addition, you can also
constrain output to a specific repository:

```
data-kraken files --org mobile-de --repo consumer-fe nina-maass
```

- Caveat: [Data time range](#data-time-range)

#### More info

Run `--help` for more info on the `files` command:

```
data-kraken help files
```

## Options

### CSV output

To facilitate importing the output into a Google Sheet, you can specify CSV format:

```
data-kraken repos --format csv patrick-hund
```

…or…

```
data-kraken files --format csv patrick-hund
```

This is particularly useful when using multiple users. You can pipe a list of usernames into _data-kraken_ using
_xargs_ and store the output in a CSV file, like this:

```
cat users.txt | xargs data-kraken files --format csv > files.csv
```

You can then upload and import the CSV file into Google Sheets.

- See also: [CSV date format](#csv-date-format)

### JSON output

You can also have _data-kraken_ deliver its output in JSON format, for example:

```
data-kraken npm-packages --org mobile-de --format json
```

### Verbose output

All commands support a flag for getting more verbose output:

`-v` or `--verbose`

The effect of using verbose mode is different depending on the command and the format type.

## Caching

When executing a command, _data-kraken_ does **a lot** of requests to the GitHub API, which can take a long time.
Be patient when executing a command that you haven't used before!

For subsequent command executions, _data-kraken_ uses cached data from previous API calls to speed things up.

The time to live of the caching can be configured through the environment variable `DK_FETCH_CACHE_TTL`. You can
set it in the [.data-kraken](#preparation) config file in your home directory. In
[.data-kraken.defaults](.data-kraken.defaults), this is set to 86400000 milliseconds, which is one day.

## Additional notes and caveats

### Data time range

For commands related to users (e.g. [repos](#repos), [files](#files)), _data-kraken_ fetches commit data of the
users.

We fetch data from GitHub as far back as it is allowed to by constraints of the GitHub API. This is usually data
for around two weeks, depending on how active the user was (less activity – data ranges further back in time).

### Regular expressions

Some hints on how to use regular expression with commands that support them (e.g. [docker-images](#docker-images)
, [npm-packages](#npm-packages)):

- Specify regular expressions without enclosing forward slashes
- Providing regular expression flags (g, i, u, etc.) is not supported
- The search is always case-insensitive
- Complex regular expressions need to be quoted, otherwise your shell will complain because it tries to evaluate
  the expression

### CSV date format

For commands that create CSV data with times in them (e.g. [repos](#repos), [files](#files)), importing the CSV
file in Google Sheets works best if you set the `DK_LOCALE` and `DK_TIME_ZONE` options in the _.data-kraken_ file
in your home directory to the locale and time zone your Google Sheets is set to. Then dates and times will be
imported properly as dates you can calculate with rather than mere strings.

If your Google Workspace is in German, for example, you want to specify `DK_LOCALE=de-DE`. If you are located in
Toronto, you want to specify `DK_TIME_ZONE=EST`.

Default locale is English / Great Britain (`en-GB`) and Barcelona / Berlin / Amsterdam time (`CET`).

## Contributing

You are most welcome to fork this repository and create a pull request. The following will hopefully get you on
your way.

### How to install for development

- Check out the source code
- Use correct Node.js version:

```
nvm use
```

- Install dependencies:

```
yarn install
```

- Create _.data-kraken_ config file:

```
cp .data-kraken.example .data-kraken
```

- Uncomment the line with `DK_ACCESS_TOKEN` in the .data-kraken file, replace the value with your personal GitHub
  access token
  ([instructions](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/creating-a-personal-access-token))

### Running the script

You can run the script with `node src/dataKraken.mjs`. For your convenience, there is also an npm script that
does this, with [debugging](#debugging) already enabled.

#### Configuration

To determine the tech debt score, the program analyses the Dockerfiles and package.json files of the repositories
and assigns tech debt scores for dependencies that are outdated or banned. The algorithm uses a YAML config file
to do this:

- [config/tech-debt-evaluator.yaml](.data-kraken-config/tech-debt-evaluator.yaml)

### Tests

This package uses [Jest](https://jestjs.io/) for automated testing.

#### Running tests

To run unit test:

```
yarn test
```

#### Style considerations

Write unit tests mostly for low-level functions that have lots of different input to make sure that they return
the expected result. Use `test` and `test.each` instead of `describe` and `it`.

- Code example: [getTechDebtScore.test.js](src/commands/techDebt/evaluation/getTechDebtScore.test.mjs)

### Terminating with error

Whenever the program encounters a situation where it can't continue, e.g. network errors from API request
attempts, it should terminate with an error code. Use the function `die` in these cases, supplying an error
message:

```javascript
import die from "./utils/die.js";

die("Failed to execute command");
```

### Using the GitHub API

The codebase provides a package with utility functions for fetching data from the GitHub API.

#### Main API functions

The main function for fetching data are:

- [fetchResult](src/api/fetchResult.mjs) – given a REST API path and an optional result page, fetches the result
  from that path
- [fetchSearchResult](src/api/fetchSearchResult.mjs) – given a search query and an optional result page, fetches
  search results

#### Additional API utilities

This program includes numerous ways to reduce the number of requests to the GitHub API while making it resilient
against connection problems and improving performance.

If you implement additional commands that fetch data from GitHub, you need to use these the same way the existing
commands do:

- [inBatches](src/api/inBatches.mjs) – executes fetch commands in batches rather than executing them all at once
- [withPagination](src/api/withPagination.mjs) – fetches paged results one page after another
- [withRetry](src/api/withRetry.mjs) – retries API requests if they fail
- [fetchWithCache](src/api/fetchWithCache.mjs) – caches fetch results using the local file system; **note:** this
  is already built-in into _fetchData_, so you'll only need this when implementing your own fetch function ( see
  [Caching](#caching))

### Debugging

You can turn on a debug logger through the environment variable `DEBUG`, example:

```
DEBUG=* yarn data-kraken docker-images --org mobile-de
```

This will print log statements to the console that are created through the log function.

The asterisk argument in the above example means show all log statements; you can only show specific log
statements by specifying a logger name.

The logger name is the relative path to the logging JavaScript module, prefixed with _data-kraken:_, with forward
slashes replaced by colons and without the file extension.

For example, the logger name for module `src/commands/dockerImages/run.js` is
`data-kraken:commands:dockerImages:run`, and you can show only log statements from this module with this command:

```
DEBUG=data-kraken:commands:dockerImages:run yarn data-kraken docker-images --org mobile-de
```

- See also: [debug lib documentation](https://github.com/debug-js/debug) on GitHub

#### Object logging depth

Objects are logged only up to a certain depth. You can increase this depth with the environment variable
`DEBUG_DEPTH`.

#### Adding log statements in the code

You can add log statements to any module using debug, like this:

```javascript
import createLogFunction from "./utils/createLogFunction.js";

const log = createLogFunction();

log("I'm a happy camper");
```

The logger name will be set to “data-kraken” automatically. You can override this behaviour by providing a name
as a string argument to `createLogFunction` (recommended!):

```javascript
const log = createLogFunction("my:awesome:logger");
```

In this case, the logger name you provide is prefixed with `data-kraken:`, i.e. the resulting logger name will be
`data-kraken:my:awesome:logger`.

If you intend to leave the log statements in the code, please use sensible names according to the
[conventions of the debug library](https://github.com/debug-js/debug#conventions). Recommended is the path to the
logging JavaScript module, with slashes replaced by colons, without file extension.

Example:

If your module's path is `src/command/myCommand/doSomething.js`, initialize a logger with this statement:

```javascript
const log = createLogFunction("command:myCommand:doSomething");
```

## Publishing a new package version

### Prerequisites

To be able to publish, you need to have the permission on [npmjs.org](https://www.npmjs.org/). Ask one of the
maintainers to grant you the access rights.

### Versioning

This project uses [semantic versioning](https://semver.org/), a.k.a. SemVer. If you're not familiar with the
concept, [please read up on it](https://semver.org/).

In a nutshell:

- If your new release contains only bugfixes, publish a **patch** version (e.g. old version 1.0.0 → new version
  1.0.**1**)
- If your new release contains new features that are compatible with all existing features, publish a **minor**
  version (e.g. 1.0.0 → 1.**1**.0)
- If your new release contains new features that are _not_ compatible with all existing features (also known as
  “breaking changes”), publish a **major** version (e.g. 1.0.0 → **1**.0.0)

Beta versions are suffixed with `-beta.x`, where `x` is a number starting at zero that is incremented with every
beta release.

### Beta versions

Before you publish a final version of the package, make sure you test everything with a beta release.

1. Make sure tests pass: `yarn test`
1. Bump the version number in [package.json](package.json) – example: `"version": "2.0.0-beta.0"`
1. Bump the version number in [src/dataKraken.mjs](src/dataKraken.mjs) – example: `.version("2.0.0-beta.0")`
1. Build the _*bin*_ file (in _dist_ directory*)*: `yarn build`
1. Run the publish command: `yarn npm publish --tag beta`
1. Verify that it worked: `npx data-kraken@beta --version`

### Final versions

When you are confident your new version is ready for the public at large, follow the same steps as
[above](#beta-versions), but this time, without the `beta` parts:

1. Make sure tests pass: `yarn test`
1. Bump the version number in [package.json](package.json) – example: `"version": "2.0.0"`
1. Bump the version number in [src/dataKraken.mjs](src/dataKraken.mjs) – example: `.version("2.0.0")`
1. Build the _*bin*_ file (in _dist_ directory*)*: `yarn build`
1. Run the publish command: `yarn npm publish`
1. Verify that it worked: `npx data-kraken@latest --version`

## License

[MIT license](LICENSE.md) – copyright 2022 mobile.de GmbH

---
_Source: https://npm.io/package/data-kraken · Machine-readable twin of the npm.io package page. Health data is recomputed on every publish._
