# ollama

> Ollama Javascript library

Latest version **0.6.4** (published 2026-09-29) · MIT license · 0 weekly downloads

## Install

```sh
npm install ollama
pnpm add ollama
yarn add ollama
bun add ollama
```

## Health

**Score 70/100 (B)** — status: active.

Positive: has types; esm support; no vulnerabilities; recently updated; high maintenance score; high quality score.

Warnings: low downloads; pre 1.0.

## Facts

| | |
|---|---|
| Version | 0.6.4 |
| Published | 2026-09-29 |
| First published | 2023-09-14 |
| Weekly downloads | 0 |
| License | MIT |
| TypeScript types | bundled |
| Module format | ESM + CommonJS |
| Dependencies | 1 |
| Unpacked size | 134.7 KB |
| Known vulnerabilities | 0 |
| Install scripts | no |
| Author | Saul Boyd |
| Maintainers | jmorgan, mchiang0610, drifkin, bruce-ollama, parthollama, saulb |

## Links

- npm: https://www.npmjs.com/package/ollama
- Repository: https://github.com/ollama/ollama-js
- Issues: https://github.com/ollama/ollama-js/issues
- npm.io page: https://npm.io/package/ollama

## Dependencies (1)

- [whatwg-fetch](https://npm.io/package/whatwg-fetch.md) ^3.6.20

## Recent versions

- 0.6.4 (latest) — 2026-09-29
- 0.6.3 — 2025-11-13
- 0.6.2 — 2025-10-30
- 0.6.1 — 2025-10-30
- 0.6.0 — 2025-09-25
- 0.5.18 — 2025-09-16
- 0.5.17 — 2025-08-07
- 0.5.16 — 2025-05-30
- 0.5.15 — 2025-04-16
- 0.5.14 — 2025-02-24
- 0.5.13 — 2025-02-10
- 0.5.12 — 2025-01-14
- 0.5.11 — 2024-12-06
- 0.5.10 — 2024-11-12
- 0.5.9 — 2024-09-09
- … 22 more at https://npm.io/package/ollama/versions

## README

# Ollama JavaScript Library

The Ollama JavaScript library provides the easiest way to integrate your JavaScript project with [Ollama](https://github.com/jmorganca/ollama).

## Getting Started

```
npm i ollama
```

## Usage

```javascript
import ollama from 'ollama'

const response = await ollama.chat({
  model: 'llama3.1',
  messages: [{ role: 'user', content: 'Why is the sky blue?' }],
})
console.log(response.message.content)
```

### Browser Usage

To use the library without node, import the browser module.

```javascript
import ollama from 'ollama/browser'
```

## Streaming responses

Response streaming can be enabled by setting `stream: true`, modifying function calls to return an `AsyncGenerator` where each part is an object in the stream.

```javascript
import ollama from 'ollama'

const message = { role: 'user', content: 'Why is the sky blue?' }
const response = await ollama.chat({
  model: 'llama3.1',
  messages: [message],
  stream: true,
})
for await (const part of response) {
  process.stdout.write(part.message.content)
}
```

## Cloud Models

Run larger models by offloading to Ollama’s cloud while keeping your local workflow.

[You can see models currently available on Ollama's cloud here.](https://ollama.com/search?c=cloud)

### Run via local Ollama

1) Sign in (one-time):

```
ollama signin
```

2) Pull a cloud model:

```
ollama pull gpt-oss:120b-cloud
```

3) Use as usual (offloads automatically):

```javascript
import { Ollama } from 'ollama'

const ollama = new Ollama()
const response = await ollama.chat({
  model: 'gpt-oss:120b-cloud',
  messages: [{ role: 'user', content: 'Explain quantum computing' }],
  stream: true,
})
for await (const part of response) {
  process.stdout.write(part.message.content)
}
```

### Cloud API (ollama.com)

Access cloud models directly by pointing the client at `https://ollama.com`.

1) Create an [API key](https://ollama.com/settings/keys), then set the `OLLAMA_API_KEY` environment variable:

```
export OLLAMA_API_KEY=your_api_key
```

2) Generate a response via the cloud API:

```javascript
import { Ollama } from 'ollama'

const ollama = new Ollama({
  host: 'https://ollama.com',
  headers: { Authorization: 'Bearer ' + process.env.OLLAMA_API_KEY },
})

const response = await ollama.chat({
  model: 'gpt-oss:120b',
  messages: [{ role: 'user', content: 'Explain quantum computing' }],
  stream: true,
})

for await (const part of response) {
  process.stdout.write(part.message.content)
}
```

## API

The Ollama JavaScript library's API is designed around the [Ollama REST API](https://github.com/jmorganca/ollama/blob/main/docs/api.md)

### chat

```javascript
ollama.chat(request)
```

- `request` `<Object>`: The request object containing chat parameters.

  - `model` `<string>` The name of the model to use for the chat.
  - `messages` `<Message[]>`: Array of message objects representing the chat history.
    - `role` `<string>`: The role of the message sender ('user', 'system', or 'assistant').
    - `content` `<string>`: The content of the message.
    - `images` `<Uint8Array[] | string[]>`: (Optional) Images to be included in the message, either as Uint8Array or base64 encoded strings.
    - `tool_name` `<string>`: (Optional) Add the name of the tool that was executed to inform the model of the result 
  - `format` `<string>`: (Optional) Set the expected format of the response (`json`).
  - `stream` `<boolean>`: (Optional) When true an `AsyncGenerator` is returned.
  - `think` `<boolean | "high" | "medium" | "low">`: (Optional) Enable model thinking. Use `true`/`false` or specify a level. Requires model support.
  - `logprobs` `<boolean>`: (Optional) Return log probabilities for tokens. Requires model support.
  - `top_logprobs` `<number>`: (Optional) Number of top log probabilities to return per token when `logprobs` is enabled.
  - `keep_alive` `<string | number>`: (Optional) How long to keep the model loaded. A number (seconds) or a string with a duration unit suffix ("300ms", "1.5h", "2h45m", etc.)
  - `tools` `<Tool[]>`: (Optional) A list of tool calls the model may make.
  - `options` `<Options>`: (Optional) Options to configure the runtime.

- Returns: `<ChatResponse>`

### generate

```javascript
ollama.generate(request)
```

- `request` `<Object>`: The request object containing generate parameters.
  - `model` `<string>` The name of the model to use for the chat.
  - `prompt` `<string>`: The prompt to send to the model.
  - `suffix` `<string>`: (Optional) Suffix is the text that comes after the inserted text.
  - `system` `<string>`: (Optional) Override the model system prompt.
  - `template` `<string>`: (Optional) Override the model template.
  - `raw` `<boolean>`: (Optional) Bypass the prompt template and pass the prompt directly to the model.
  - `images` `<Uint8Array[] | string[]>`: (Optional) Images to be included, either as Uint8Array or base64 encoded strings.
  - `format` `<string>`: (Optional) Set the expected format of the response (`json`).
  - `stream` `<boolean>`: (Optional) When true an `AsyncGenerator` is returned.
  - `think` `<boolean | "high" | "medium" | "low">`: (Optional) Enable model thinking. Use `true`/`false` or specify a level. Requires model support.
  - `logprobs` `<boolean>`: (Optional) Return log probabilities for tokens. Requires model support.
  - `top_logprobs` `<number>`: (Optional) Number of top log probabilities to return per token when `logprobs` is enabled.
  - `keep_alive` `<string | number>`: (Optional) How long to keep the model loaded. A number (seconds) or a string with a duration unit suffix ("300ms", "1.5h", "2h45m", etc.)
  - `width` `<number>`: (Optional, Experimental) Width of the generated image in pixels. For image generation models only.
  - `height` `<number>`: (Optional, Experimental) Height of the generated image in pixels. For image generation models only.
  - `steps` `<number>`: (Optional, Experimental) Number of diffusion steps. For image generation models only.
  - `options` `<Options>`: (Optional) Options to configure the runtime.
- Returns: `<GenerateResponse>`

### pull

```javascript
ollama.pull(request)
```

- `request` `<Object>`: The request object containing pull parameters.
  - `model` `<string>` The name of the model to pull.
  - `insecure` `<boolean>`: (Optional) Pull from servers whose identity cannot be verified.
  - `stream` `<boolean>`: (Optional) When true an `AsyncGenerator` is returned.
- Returns: `<ProgressResponse>`

### push

```javascript
ollama.push(request)
```

- `request` `<Object>`: The request object containing push parameters.
  - `model` `<string>` The name of the model to push.
  - `insecure` `<boolean>`: (Optional) Push to servers whose identity cannot be verified.
  - `stream` `<boolean>`: (Optional) When true an `AsyncGenerator` is returned.
- Returns: `<ProgressResponse>`

### create

```javascript
ollama.create(request)
```

- `request` `<Object>`: The request object containing create parameters.
  - `model` `<string>` The name of the model to create.
  - `from` `<string>`: The base model to derive from.
  - `stream` `<boolean>`: (Optional) When true an `AsyncGenerator` is returned.
  - `quantize` `<string>`: Quanization precision level (`q8_0`, `q4_K_M`, etc.).
  - `template` `<string>`: (Optional) The prompt template to use with the model.
  - `license` `<string|string[]>`: (Optional) The license(s) associated with the model.
  - `system` `<string>`: (Optional) The system prompt for the model.
  - `parameters` `<Record<string, unknown>>`: (Optional) Additional model parameters as key-value pairs.
  - `messages` `<Message[]>`: (Optional) Initial chat messages for the model.
  - `adapters` `<Record<string, string>>`: (Optional) A key-value map of LoRA adapter configurations.
- Returns: `<ProgressResponse>`

Note: The `files` parameter is not currently supported in `ollama-js`.

### delete

```javascript
ollama.delete(request)
```

- `request` `<Object>`: The request object containing delete parameters.
  - `model` `<string>` The name of the model to delete.
- Returns: `<StatusResponse>`

### copy

```javascript
ollama.copy(request)
```

- `request` `<Object>`: The request object containing copy parameters.
  - `source` `<string>` The name of the model to copy from.
  - `destination` `<string>` The name of the model to copy to.
- Returns: `<StatusResponse>`

### list

```javascript
ollama.list()
```

- Returns: `<ListResponse>`

### show

```javascript
ollama.show(request)
```

- `request` `<Object>`: The request object containing show parameters.
  - `model` `<string>` The name of the model to show.
  - `system` `<string>`: (Optional) Override the model system prompt returned.
  - `template` `<string>`: (Optional) Override the model template returned.
  - `options` `<Options>`: (Optional) Options to configure the runtime.
- Returns: `<ShowResponse>`

### embed

```javascript
ollama.embed(request)
```

- `request` `<Object>`: The request object containing embedding parameters.
  - `model` `<string>` The name of the model used to generate the embeddings.
  - `input` `<string> | <string[]>`: The input used to generate the embeddings.
  - `truncate` `<boolean>`: (Optional) Truncate the input to fit the maximum context length supported by the model.
  - `keep_alive` `<string | number>`: (Optional) How long to keep the model loaded. A number (seconds) or a string with a duration unit suffix ("300ms", "1.5h", "2h45m", etc.)
  - `options` `<Options>`: (Optional) Options to configure the runtime.
- Returns: `<EmbedResponse>`

### web search
- Web search capability requires an Ollama account. [Sign up on ollama.com](https://ollama.com/signup) 
- Create an API key by visiting https://ollama.com/settings/keys
```javascript
ollama.webSearch(request)
```

- `request` `<Object>`: The search request parameters.
  - `query` `<string>`: The search query string.
  - `max_results` `<number>`: (Optional) Maximum results to return (default 5, max 10).
- Returns: `<SearchResponse>`

### web fetch

```javascript
ollama.webFetch(request)
```

- `request` `<Object>`: The fetch request parameters.
  - `url` `<string>`: The URL to fetch.
- Returns: `<FetchResponse>`

### ps

```javascript
ollama.ps()
```

- Returns: `<ListResponse>`

### version

```javascript
ollama.version()
```

- Returns: `<VersionResponse>`

### abort

```javascript
ollama.abort()
```

This method will abort **all** streamed generations currently running with the client instance.
If there is a need to manage streams with timeouts, it is recommended to have one Ollama client per stream.

All asynchronous threads listening to streams (typically the `for await (const part of response)`) will throw an `AbortError` exception. See [examples/abort/abort-all-requests.ts](examples/abort/abort-all-requests.ts) for an example.

### systemone

```typescript
import ollama from 'ollama'

const response = await ollama.systemone({
  model: 'nimble',
  state: 'Our checkout has returned 500 errors since 9am.',
  questions: {
    team: {
      type: 'choice',
      instructions: 'Which team should handle this ticket?',
      criteria: { billing: 'Payments and refunds', technical: 'Software errors' },
    },
  },
})
console.log(response.answers.team)
```

System One uses `POST /v1/systemone` and requires Ollama v0.35.0 or later with a
compatible local model such as `nimble`. It returns one JSON response; streaming
and cloud models are not supported.

`state` and question `instructions` accept text, JSON objects, or arrays. Questions
are evaluated in their supplied order:

- `choice`: 2–26 option keys mapped to descriptions; `null` uses the key as its description.
- `noul`: probability of true, with optional `{"false": "No", "true": "Yes"}` descriptions.
- `score`: 2–26 descriptions ordered lowest to highest; returns a potentially fractional, zero-based score.

Responses contain `model`, `answers`, and `usage.input_tokens` / `usage.output_tokens`.
Confidence measures probability concentration, not calibrated correctness.
Token usage comes from the server; output tokens are not necessarily zero.
Optional `keep_alive` accepts seconds or a duration string. Requests use the client's
existing host, headers, and HTTP error handling. The server validates its body and
model context limits without truncating input.

Available on the default client and `new Ollama()` from both `ollama` and
`ollama/browser`. Both entry points export `SystemOneRequest`, `SystemOneResponse`,
and the question/answer types. Browser requests follow the server's normal CORS rules.
See [the combined question example](examples/systemone.mjs).

## Custom client

A custom client can be created with the following fields:

- `host` `<string>`: (Optional) The Ollama host address. Default: `"http://127.0.0.1:11434"`.
- `fetch` `<Object>`: (Optional) The fetch library used to make requests to the Ollama host.
- `headers` `<Object>`: (Optional) Custom headers to include with every request.

```javascript
import { Ollama } from 'ollama'

const ollama = new Ollama({ host: 'http://127.0.0.1:11434' })
const response = await ollama.chat({
  model: 'llama3.1',
  messages: [{ role: 'user', content: 'Why is the sky blue?' }],
})
```

## Custom Headers

You can set custom headers that will be included with every request:

```javascript
import { Ollama } from 'ollama'

const ollama = new Ollama({
  host: 'http://127.0.0.1:11434',
  headers: {
    Authorization: 'Bearer <api key>',
    'X-Custom-Header': 'custom-value',
    'User-Agent': 'MyApp/1.0',
  },
})
```

## Building

To build the project files run:

```sh
npm run build
```

---
_Source: https://npm.io/package/ollama · Machine-readable twin of the npm.io package page. Health data is recomputed on every publish._
