npm.io
5.13.0 • Published 6d ago

chat-about-video

Licence
Apache-2.0
Version
5.13.0
Deps
3
Size
281 kB
Vulns
0
Weekly
0

chat-about-video

Chat about zero or one or more video clip(s) or audio file(s) using the powerful OpenAI ChatGPT (hosted in OpenAI or Microsoft Azure) or Google Gemini (hosted in Google Could). It provides a standardized interface for interacting with OpenAI ChatGPT (OpenAI or Azure) and Google Gemini,

Version Downloads/week License

chat-about-video is a powerful Unified Abstraction Layer designed to accelerate the development of conversational AI applications. It provides a standardized interface for interacting with OpenAI ChatGPT (OpenAI or Azure) and Google Gemini, allowing you to switch between providers with zero or minimal changes to your application logic.

Why use chat-about-video?

  • Provider Agnostic: Write your code once and swap between ChatGPT and Gemini via configuration. This future-proofs your application against model changes or pricing shifts.
  • Unified Video Handling: Seamlessly handles the complexities of frame extraction and cloud storage uploading (for ChatGPT) or direct ingestion (for Gemini) through a single API.
  • Simplified Tool Calling: A standardized way to define and handle tool/function calls across different model providers.
  • Production Ready: Built-in retries for throttling, server errors, and connectivity issues.

Key features

  • Switch providers effortlessly: Change from ChatGPT to Gemini (or vice-versa) without rewritten your conversation logic.
  • Multi-Cloud Support: Supports models hosted in Azure OpenAI, OpenAI, NVIDIA NIM (OpenAI compatible), and Google Cloud.
  • Flexible Media Input: Extract frames automatically via FFmpeg, supply your own images, or provide audio files.
  • Rich Conversations: Supports multiple videos, image groups, and audio files in a single chat.
  • Conversation Branching & Rewind: Fork conversations (fork()) into independent branches and rewind prompt history (rewind()) by turn checkpoints.
  • Mandated Output: Force JSON responses with or without schemas.
  • Resilient: Automatic backoff and retries for 429, 5xx, and network errors.
  • Usage Tracking: Built-in token usage metadata collection.

Usage

Installation (quick start)

To use chat-about-video in your Node.js application, add it as a dependency along with other necessary packages based on your usage scenario. Below are examples for typical setups:

# ChatGPT on OpenAI or Azure with Azure Blob Storage
npm i chat-about-video openai @ffmpeg-installer/ffmpeg @azure/storage-blob
# Gemini in Google Cloud
npm i chat-about-video @google/generative-ai @ffmpeg-installer/ffmpeg
# ChatGPT on OpenAI or Azure with AWS S3
npm i chat-about-video openai @ffmpeg-installer/ffmpeg @handy-common-utils/aws-utils @aws-sdk/s3-request-presigner @aws-sdk/client-s3

If ffmpeg binary is already available, you don't need to add dependency @ffmpeg-installer/ffmpeg.

Optional dependencies

ChatGPT

To use ChatGPT hosted on OpenAI or Azure:

npm i openai

Gemini

To use Gemini hosted on Google Cloud:

npm i @google/generative-ai

ffmpeg

If you need ffmpeg for extracting video frame images, ensure it is installed. You can use a system package manager or an NPM package:

sudo apt install ffmpeg
# or
npm i @ffmpeg-installer/ffmpeg

Azure Blob Storage

To use Azure Blob Storage for frame images (not needed for Gemini):

npm i @azure/storage-blob

AWS S3

To use AWS S3 for frame images (not needed for Gemini):

npm i @handy-common-utils/aws-utils @aws-sdk/s3-request-presigner @aws-sdk/client-s3

How the video is provided to ChatGPT or Gemini

ChatGPT

chat-about-video supports uploading video frames into cloud storage and making them available to ChatGPT.

  • Integrate ChatGPT from Microsoft Azure or OpenAI effortlessly.
  • Utilize ffmpeg integration provided by this package for frame image extraction or opt for a DIY approach.
  • Store frame images with ease, supporting Azure Blob Storage and AWS S3.
  • Models hosted in Azure seems to allow less number of images per request than models hosted in OpenAI.
Gemini

chat-about-video supports sending video frames directly to Google's API without requiring cloud storage.

  • Utilize ffmpeg integration provided by this package for frame image extraction or opt for a DIY approach.
  • The number of frame images is only limited by the Gemini API in Google Cloud.

Concrete types and low level clients

ChatAboutVideo and Conversation are generic classes. Use them without concrete generic type parameters when you want the flexibility to easily switch between ChatGPT and Gemini.

Otherwise, you may want to use concrete type. Below are some examples:

// cast to a concrete type
const castToChatGpt = chat as ChatAboutVideoWithChatGpt;

// you can also just leave the ChatAboutVideo instance generic, but narrow down the conversation type
const conversationWithGemini = (await chat.startConversation(...)) as ConversationWithGemini;
const conversationWithChatGpt = await (chat as ChatAboutVideoWithChatGpt).startConversation(...);

To access the underlying API wrapper, use the getApi() function on the ChatAboutVideo instance. To get the raw API client, use the getClient() function on the awaited object returned from getApi().

Cleaning up

Intermediate files, such as extracted frame images, can be saved locally or in the cloud. To remove these files when they are no longer needed, remember to call the end() function on the Conversation instance when the conversion finishes.

Switching between configurations

You can define multiple configurations and switch between them using the activeSupportedChatApiOptions function. This is useful when you want to easily switch between different environments (e.g. dev, prod) or different models. Note that nested objects are deeply merged, while arrays are replaced rather than concatenated.

import { activeSupportedChatApiOptions, ChatAboutVideo } from 'chat-about-video';

const options = {
  active: process.env.ACTIVE_CONFIG || 'dev',
  base: {
    storage: {
      azureStorageConnectionString: process.env.AZURE_STORAGE_CONNECTION_STRING!,
    },
  },
  dev: {
    credential: { key: process.env.DEV_KEY! },
    completionOptions: { model: 'gpt-4o' },
  },
  prod: {
    credential: { key: process.env.PROD_KEY! },
    completionOptions: { model: 'gpt-4' },
  },
};

const chat = new ChatAboutVideo(activeSupportedChatApiOptions(options));

Mandating JSON response

JSON response can be guaranteed either with a JSON Schema or without. Below example code works for both ChatGPT and Gemini:

// Without specifying a JSON schema
const explanation = await conversation.say(
  'Explain your answer. The response should be in JSON like this: {"referencedFrames": [1, 5], "why": "Reason for giving this response."}',
  { jsonResponse: true },
);
console.log(chalk.grey("\nAI's Explanation: " + JSON.stringify(JSON.parse(explanation!), null, 2)));

// With a JSON schema
const detailedExplanation = await conversation.say('Explain your answer in detail. The response should be in JSON.', {
  jsonResponse: {
    name: 'DetailedExplanation',
    schema: {
      type: 'object',
      properties: {
        referencedFrames: {
          type: 'array',
          items: { type: 'integer' },
        },
        understandingOfTheQuestion: { type: 'string' },
        reasoningSteps: { type: 'array', items: { type: 'string' } },
      },
      required: ['referencedFrames', 'understandingOfTheQuestion', 'reasoningSteps'],
    },
  },
});
console.log(chalk.grey("\nAI's detailed explanation: " + JSON.stringify(JSON.parse(detailedExplanation!), null, 2)));

Tool Calling (Function Calling)

chat-about-video supports tool calling for both ChatGPT and Gemini. This allows the AI to request information by calling functions you've defined.

1. Define Tools

Pass your tool definitions in the completion options. The structure follows the underlying API (OpenAI or Gemini). You can also use the ChatGPT style structure for Gemini providers, as the package will automatically convert it for Gemini if needed:

const tools = [
  {
    type: 'function',
    function: {
      name: 'get_weather',
      description: 'Get the current weather',
      parameters: {
        type: 'object',
        properties: {
          location: { type: 'string' },
        },
        required: ['location'],
      },
    },
  },
];

const answer = await conversation.say<ConversationResponse>("What's the weather like in Melbourne?", { tools });
2. Handle Tool Calls

The say and submitToolCallResults methods will return an object containing toolCalls if the AI wants to call tools. You are responsible for executing the tools and submitting the results back.

import { ConversationResponse, ToolCallResult } from 'chat-about-video';

let response = await conversation.say<ConversationResponse>('What is the weather in Melbourne?', { tools });

// Loop to handle potential multiple rounds of tool calling
while (typeof response !== 'string' && response?.toolCalls) {
  if (response.responseText) {
    console.log(`AI: ${response.responseText}`);
  }
  const toolResults: ToolCallResult[] = [];
  for (const call of response.toolCalls) {
    console.log(`AI requests tool: ${call.name}(${JSON.stringify(call.arguments)})`);

    // Execute your tool logic
    const result = await myWeatherFunction(call.arguments.location);

    toolResults.push({
      name: call.name,
      result: { temperature: result.temp, unit: 'C' },
      toolCallId: call.id, // Required for OpenAI
    });
  }
  // Submit results back to the AI
  response = await conversation.submitToolCallResults<ConversationResponse>(toolResults);
}

// Final text response
console.log('AI Answer:', response);

Forking and Rewinding Conversations (fork and rewind)

chat-about-video provides state management methods on Conversation instances to allow branching conversations and step rewinding without losing accumulated token usage history.

1. Branching a Conversation (fork)

Use fork() to create an independent clone of an existing conversation at its current point (or at an earlier turn).

  • Isolated Prompt State: The forked conversation gets a deep copy of the prompt history. Further turns or rewinds on either conversation do not affect the other.
  • Reference-Counted Cleanup: Shared resources (e.g., extracted video frame files) are preserved until every forked conversation in the family has called end().
  • Usage Independence: Token usage on the fork is tracked separately starting from zero.
// Start a conversation
const conversation = await chat.startConversation('/path/to/video.mp4');
await conversation.say('Analyze the video content.');

// Fork the conversation into two separate branches
const branchA = conversation.fork();
const branchB = conversation.fork();

// Branch A explores one topic
await branchA.say('What color is the car in the video?');

// Branch B explores another topic independently
await branchB.say('Describe the background music.');

// Remember to end all forked conversations when finished
await conversation.end();
await branchA.end();
await branchB.end();
2. Rewinding History (rewind)

Use rewind(steps) to remove the last $N$ successful turns (say or submitToolCallResults) from the conversation prompt history.

  • Preserved Token Usage: Token usage already recorded on the conversation instance is preserved.
  • Checkpoint Rewind: Rewinds the prompt back to the state after the specified turn. Passing a step count larger than the number of completed turns restores the prompt back to its initial state.
await conversation.say('First question'); // Turn 1
await conversation.say('Second question'); // Turn 2

// Rewind the last turn (drops 'Second question' and AI response)
conversation.rewind(1);

// Continue conversation from Turn 1 state
await conversation.say('Alternative second question');

You can also combine fork and rewind in a single call to fork from an earlier checkpoint:

// Fork at 1 turn prior to the current state
const earlierFork = conversation.fork(1);

Customisation

Frame extraction

If you would like to customise how frame images are extracted and stored, consider these:

  • In the options object passed to the constructor of ChatAboutVideo, there's a property extractVideoFrames. This property allows you to customise how frame images are extracted.
    • format, interval, limit, width, height - These allows you to specify your expectation on the extraction.
    • deleteFilesWhenConversationEnds - This flag allows you to specify whether you want extracted frame images to be deleted from the local file system when the conversation ends, or not.
    • framesDirectoryResolver - You can supply a function for determining where extracted frame image files should be stored locally.
    • extractor - You can supply a function for doing the extraction.
  • In the options object passed to the constructor of ChatAboutVideo, there's a property storage. For ChatGPT, storing frame images in the cloud is recommended. You can use this property to customise how frame images are stored in the cloud.
    • azureStorageConnectionString - If you would like to use Azure Blob Storage, you need to put the connection string in this property. If this property does not have a value, ChatAboutVideo would assume that you'd like to use AWS S3, and default AWS identity/credential will be picked up from the OS.
    • storageContainerName, storagePathPrefix - They allows you to specify where those images should be stored.
    • downloadUrlExpirationSeconds - For images stored in the cloud, presigned download URLs with expiration are generated for ChatGPT to access. This property allows you to control the expiration time.
    • deleteFilesWhenConversationEnds - This flag allows you to specify whether you want extracted frame images to be deleted from the cloud when the conversation ends, or not.
    • uploader - You can supply a function for uploading images into the cloud.
Settings of the underlying model

In the options object passed to the constructor of ChatAboutVideo, there's a property clientSettings, and there's another property completionSettings. Settings of the underlying model can be configured through those two properties.

You can also override settings using the last parameter of startConversation(...) function on ChatAboutVideo, or the last parameter of say(...) function on Conversation.

Code examples

The following integration test files demonstrate various features and providers:

File AI Provider Features
chatgpt-openai-azure-storage.ts ChatGPT (OpenAI) Basic usage with Azure Storage
chatgpt-openai-azure-storage-multi-video.ts ChatGPT (OpenAI) Multiple videos in one conversation
chatgpt-azure-azure-storage-json.ts ChatGPT (Azure) JSON response mode
gemini-json.ts Google Gemini JSON response mode
chatgpt-manual-frames.ts ChatGPT Manual frame extraction using FFmpeg
chatgpt-azure-azure-storage-tools.ts ChatGPT (Azure) Tool/Function calling
gemini-tools.ts Google Gemini Tool/Function calling
gemini-chatgpt-style-tools.ts Google Gemini ChatGPT-style tool calling
gemini-audio.ts Google Gemini Audio file support
nvidia-nim-tools.ts NVIDIA NIM OpenAI-compatible tools usage
Example 1: Using ChatGPT hosted in OpenAI with Azure Blob Storage

Source: test/integration/chatgpt-openai-azure-storage.ts

// This is a demo utilising ChatGPT hosted in OpenAI.
// Video frame images are uploaded to Azure Blob Storage and then made available to GPT from there.
//
// This script can be executed with a command line like this from the project root directory:
// export OPENAI_API_KEY=...
// export AZURE_STORAGE_CONNECTION_STRING=...
// export OPENAI_MODEL_NAME=...
// export AZURE_STORAGE_CONTAINER_NAME=...
// ENABLE_DEBUG=true DEMO_VIDEO=~/Downloads/test1.mp4 npx ts-node test/integration/chatgpt-openai-azure-storage.ts
//

import { consoleWithColour } from '@handy-common-utils/misc-utils';
import chalk from 'chalk';
import readline from 'node:readline';

import { ChatAboutVideo, ConversationWithChatGpt } from '../src';

async function demo() {
  const chat = new ChatAboutVideo(
    {
      credential: {
        key: process.env.OPENAI_API_KEY!,
      },
      storage: {
        azureStorageConnectionString: process.env.AZURE_STORAGE_CONNECTION_STRING!,
        storageContainerName: process.env.AZURE_STORAGE_CONTAINER_NAME || 'vision-experiment-input',
        storagePathPrefix: 'video-frames/',
      },
      completionOptions: {
        // model is required by OpenAI
        model: process.env.OPENAI_MODEL_NAME || 'gpt-4o', // 'gpt-4-vision-preview', // or gpt-4o
      },
      extractVideoFrames: {
        limit: 100,
        interval: 2,
      },
    },
    consoleWithColour({ debug: process.env.ENABLE_DEBUG === 'true' }, chalk),
  );

  const conversation = (await chat.startConversation(process.env.DEMO_VIDEO!)) as ConversationWithChatGpt;

  const rl = readline.createInterface({ input: process.stdin, output: process.stdout });
  const prompt = (question: string) => new Promise<string>((resolve) => rl.question(question, resolve));
  while (true) {
    const question = await prompt(chalk.red('\nUser: '));
    if (!question) {
      continue;
    }
    if (['exit', 'quit', 'q', 'end'].includes(question)) {
      await conversation.end();
      break;
    }
    const answer = await conversation.say(question, { max_tokens: 2000 });
    console.log(chalk.blue('\nAI:' + answer));
  }
  console.log('Demo finished');
  rl.close();
}

demo().catch((error) => console.log(chalk.red(JSON.stringify(error, null, 2))));
Example 2: Multiple videos using ChatGPT hosted in OpenAI with Azure Blob Storage

Source: test/integration/chatgpt-openai-azure-storage-multi-video.ts

async function demo() {

  ...

  const conversation = (await chat.startConversation([
    { videoFile: process.env.DEMO_VIDEO_1!, promptText: 'This is the first video:' },
    { videoFile: process.env.DEMO_VIDEO_2!, promptText: 'This is the second video:' },
    { videoFile: process.env.DEMO_VIDEO_1!, promptText: 'This is the third video:' },
  ])) as ConversationWithChatGpt;

  ...

}
Example 3: Using ChatGPT hosted in Azure with Azure Blob Storage

Source: test/integration/chatgpt-azure-azure-storage-json.ts

// This is a demo utilising ChatGPT hosted in Azure.
// Video frame images are uploaded to Azure Blob Storage and then made available to GPT from there.
//
// This script can be executed with a command line like this from the project root directory:
// export AZURE_OPENAI_API_ENDPOINT=..
// export AZURE_OPENAI_API_KEY=...
// export AZURE_OPENAI_DEPLOYMENT_NAME=...
// export AZURE_STORAGE_CONNECTION_STRING=...
// export AZURE_STORAGE_CONTAINER_NAME=...
// ENABLE_DEBUG=true DEMO_VIDEO=~/Downloads/test1.mp4 npx ts-node test/integration/chatgpt-azure-azure-storage-json.ts

import { consoleWithColour } from '@handy-common-utils/misc-utils';
import chalk from 'chalk';
import readline from 'node:readline';

import { ChatAboutVideo, ConversationWithChatGpt } from '../src';

async function demo() {
  const chat = new ChatAboutVideo(
    {
      endpoint: process.env.AZURE_OPENAI_API_ENDPOINT!,
      credential: {
        key: process.env.AZURE_OPENAI_API_KEY!,
      },
      storage: {
        azureStorageConnectionString: process.env.AZURE_STORAGE_CONNECTION_STRING!,
        storageContainerName: process.env.AZURE_STORAGE_CONTAINER_NAME || 'vision-experiment-input',
        storagePathPrefix: 'video-frames/',
      },
      clientSettings: {
        // deployment is required by Azure
        deployment: process.env.AZURE_OPENAI_DEPLOYMENT_NAME || 'gpt4vision',
        // apiVersion is required by Azure
        apiVersion: '2024-10-21',
      },
    },
    consoleWithColour({ debug: process.env.ENABLE_DEBUG === 'true' }, chalk),
  );

  const conversation = (await chat.startConversation(process.env.DEMO_VIDEO!)) as ConversationWithChatGpt;

  const rl = readline.createInterface({ input: process.stdin, output: process.stdout });
  const prompt = (question: string) => new Promise<string>((resolve) => rl.question(question, resolve));
  while (true) {
    const question = await prompt(chalk.red('\nUser: '));
    if (!question) {
      continue;
    }
    if (['exit', 'quit', 'q', 'end'].includes(question)) {
      await conversation.end();
      break;
    }
    const answer = await conversation.say(question, { max_tokens: 2000 });
    console.log(chalk.blue('\nAI:' + answer));
  }
  console.log('Demo finished');
  rl.close();
}

demo().catch((error) => console.log(chalk.red(JSON.stringify(error, null, 2))));
Example 4: Using Gemini hosted in Google Cloud

Source: test/integration/gemini-json.ts

// This is a demo utilising Google Gemini through Google Generative Language API.
// Google Gemini allows many frame images to be supplied because of its huge context length.
// Video frame images are sent through Google Generative Language API directly.
//
// This script can be executed with a command line like this from the project root directory:
// export GEMINI_API_KEY=...
// ENABLE_DEBUG=true DEMO_VIDEO=~/Downloads/test1.mp4 npx ts-node test/integration/gemini-json.ts

import { consoleWithColour } from '@handy-common-utils/misc-utils';
import chalk from 'chalk';
import readline from 'node:readline';

import { HarmBlockThreshold, HarmCategory } from '@google/generative-ai';

import { ChatAboutVideo, ConversationWithGemini } from '../src';

async function demo() {
  const chat = new ChatAboutVideo(
    {
      credential: {
        key: process.env.GEMINI_API_KEY!,
      },
      clientSettings: {
        modelParams: {
          model: 'gemini-2.5-flash',
        },
      },
      extractVideoFrames: {
        limit: 100,
        interval: 0.5,
      },
      completionOptions: {
        safetySettings: [
          {
            category: 'HARM_CATEGORY_HATE_SPEECH' as any,
            threshold: 'BLOCK_NONE' as any,
          },
        ],
      },
    },
    consoleWithColour({ debug: process.env.ENABLE_DEBUG === 'true' }, chalk),
  );

  const conversation = (await chat.startConversation(process.env.DEMO_VIDEO!)) as ConversationWithGemini;

  const rl = readline.createInterface({ input: process.stdin, output: process.stdout });
  const prompt = (question: string) => new Promise<string>((resolve) => rl.question(question, resolve));
  while (true) {
    const question = await prompt(chalk.red('\nUser: '));
    if (!question) {
      continue;
    }
    if (['exit', 'quit', 'q', 'end'].includes(question)) {
      await conversation.end();
      break;
    }
    const answer = await conversation.say(question, {
      safetySettings: [{ category: HarmCategory.HARM_CATEGORY_SEXUALLY_EXPLICIT, threshold: HarmBlockThreshold.BLOCK_NONE }],
    });
    console.log(chalk.blue('\nAI:' + answer));
  }
  console.log('Demo finished');
  rl.close();
}

demo().catch((error) => console.log(chalk.red(JSON.stringify(error, null, 2)), error));
Example 5: Multiple groups of extracted frame images using ChatGPT hosted in Azure with Azure Blob Storage

Source: test/integration/chatgpt-manual-frames.ts

async function demo() {
  const tmpDir = os.tmpdir();
  const video1 = process.env.DEMO_VIDEO_1!;
  const video2 = process.env.DEMO_VIDEO_2!;
  const outputDir1 = path.join(tmpDir, 'video1-frames');
  const outputDir2 = path.join(tmpDir, 'video2-frames');

  console.log(chalk.green('Extracting frames from the first video...'));
  const { relativePaths: frames1, cleanup: cleanupFrames1 } = await extractVideoFramesWithFfmpeg(video1, outputDir1, 1, 'jpg', 200);

  console.log(chalk.green('Extracting frames from the second video...'));
  const { relativePaths: frames2, cleanup: cleanupFrames2 } = await extractVideoFramesWithFfmpeg(video2, outputDir2, 3, 'jpg', 200);

  const chat = new ChatAboutVideo(
    {
      credential: {
        key: process.env.OPENAI_API_KEY!,
      },
      storage: {
        azureStorageConnectionString: process.env.AZURE_STORAGE_CONNECTION_STRING!,
        storageContainerName: process.env.AZURE_STORAGE_CONTAINER_NAME || 'vision-experiment-input',
        storagePathPrefix: 'video-frames/',
      },
      completionOptions: {
        model: process.env.OPENAI_MODEL_NAME || 'gpt-4o',
      },
    },
    consoleWithColour({ debug: process.env.ENABLE_DEBUG === 'true' }, chalk),
  );

  const conversation = (await chat.startConversation([
    {
      promptText: 'Frame images from sample 1:',
      images: frames1.map((frame, i) => ({ imageFile: path.join(outputDir1, frame), promptText: `Frame CodeRed-${i + 1}` })),
    },
    {
      promptText: 'Frame images from sample 2, also known as the "good example":',
      images: frames2.map((frame) => ({ imageFile: path.join(outputDir2, frame) })),
    },
  ])) as ConversationWithChatGpt;

  ...

}
Example 6: Using NVIDIA NIM (OpenAI-compatible)

Source: test/integration/nvidia-nim-tools.ts

// This is a demo utilizing NVIDIA NIM via its OpenAI-compatible API.
//
// This script can be executed with a command line like this from the project root directory:
// export NVIDIA_NIM_API_KEY=...
// ENABLE_DEBUG=true npx ts-node test/integration/nvidia-nim-tools.ts

import { consoleWithColour, consoleWithoutColour } from '@handy-common-utils/misc-utils';
import chalk from 'chalk';
import readline from 'node:readline';

import { ChatAboutVideo, ConversationWithChatGpt, ToolCallResult } from '../src';

async function demo() {
  const chat = new ChatAboutVideo(
    {
      endpoint: process.env.NVIDIA_NIM_API_ENDPOINT || 'https://integrate.api.nvidia.com/v1',
      credential: {
        key: process.env.NVIDIA_NIM_API_KEY!,
      },
      completionOptions: {
        model: process.env.NVIDIA_NIM_MODEL || 'qwen/qwen3.5-397b-a17b',
      },
    },
    consoleWithColour({ debug: process.env.ENABLE_DEBUG === 'true' }, chalk),
  );

  const conversation = (await chat.startConversation(consoleWithoutColour({ debug: false, quiet: false }))) as ConversationWithChatGpt;

  const tools: any[] = [
    {
      type: 'function',
      function: {
        name: 'get_current_time',
        description: 'Get the current local time',
        parameters: {
          type: 'object',
          properties: {},
        },
      },
    },
  ];

  // ... handling tool calls as shown in other examples ...
}
Example 7: Using audio files with Gemini

Source: test/integration/gemini-audio.ts

import { consoleWithColour } from '@handy-common-utils/misc-utils';
import chalk from 'chalk';
import path from 'node:path';
import readline from 'node:readline';

import { ChatAboutVideo, ConversationWithGemini } from '../src';

const sampleAudioFile = path.resolve(__dirname, '../sample-media-files/engine-start.h264.aac.mp4'); // Or a real audio file like an mp3

async function demo() {
  const chat = new ChatAboutVideo(
    {
      credential: {
        key: process.env.GEMINI_API_KEY!,
      },
      clientSettings: {
        modelParams: {
          model: 'gemini-2.5-flash',
        },
      },
    },
    consoleWithColour({ debug: process.env.ENABLE_DEBUG === 'true' }, chalk),
  );

  const conversation = (await chat.startConversation([{ audioFile: sampleAudioFile }])) as ConversationWithGemini;

  const rl = readline.createInterface({ input: process.stdin, output: process.stdout });
  const prompt = (question: string) => new Promise<string>((resolve) => rl.question(question, resolve));

  while (true) {
    const question = await prompt(chalk.red('\nUser: '));
    if (!question) continue;
    if (['exit', 'quit', 'q', 'end'].includes(question)) {
      await conversation.end();
      break;
    }
    const answer = await conversation.say(question);
    console.log(chalk.blue('\nAI: ' + answer));
  }
  rl.close();
}

demo().catch((error) => console.log(chalk.red(JSON.stringify(error, null, 2))));

API

chat-about-video

Modules

Classes

Class: ChatAboutVideo<CLIENT, OPTIONS, PROMPT, RESPONSE>

chat.ChatAboutVideo

Type parameters
Name Type
CLIENT any
OPTIONS extends AdditionalCompletionOptions = any
PROMPT any
RESPONSE any
Constructors
constructor

• new ChatAboutVideo<CLIENT, OPTIONS, PROMPT, RESPONSE>(options, log?)

Type parameters
Name Type
CLIENT any
OPTIONS extends AdditionalCompletionOptions = any
PROMPT any
RESPONSE any
Parameters
Name Type
options SupportedChatApiOptions
log undefined | LineLogger<(message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void>
Properties
Property Description
Protected apiPromise: Promise<ChatApi<CLIENT, OPTIONS, PROMPT, RESPONSE>>
Protected log: undefined | LineLogger<(message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void>
Protected options: SupportedChatApiOptions
Methods
getApi

▸ getApi(): Promise<ChatApi<CLIENT, OPTIONS, PROMPT, RESPONSE>>

Get the underlying API instance.

Returns

Promise<ChatApi<CLIENT, OPTIONS, PROMPT, RESPONSE>>

The underlying API instance.


startConversation

▸ startConversation(log?): Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

Start a conversation without a video

Parameters
Name Type Description
log? LineLogger<(message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void> Optional logger for this conversation, if not provided, the logger of ChatAboutVideo instance will be used.
Returns

Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

The conversation.

▸ startConversation(options?, log?): Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

Start a conversation without a video

Parameters
Name Type Description
options? OPTIONS Overriding options for this conversation
log? LineLogger<(message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void> Optional logger for this conversation, if not provided, the logger of ChatAboutVideo instance will be used.
Returns

Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

The conversation.

▸ startConversation(videoFile, log?): Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

Start a conversation about a video.

Parameters
Name Type Description
videoFile string Path to a video file in local file system.
log? LineLogger<(message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void> Optional logger for this conversation, if not provided, the logger of ChatAboutVideo instance will be used.
Returns

Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

The conversation.

▸ startConversation(videoFile, options?, log?): Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

Start a conversation about a video.

Parameters
Name Type Description
videoFile string Path to a video file in local file system.
options? OPTIONS Overriding options for this conversation
log? LineLogger<(message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void> Optional logger for this conversation, if not provided, the logger of ChatAboutVideo instance will be used.
Returns

Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

The conversation.

▸ startConversation(videos, log?): Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

Start a conversation about a video.

Parameters
Name Type Description
videos (VideoInput | ImagesInput | AudioInput)[] Array of videos, images, or audios to be used in the conversation. For each video/audio, the file path and the prompt before it should be provided. For each group of images, the image file paths and the prompt before the image group should be provided.
log? LineLogger<(message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void> Optional logger for this conversation, if not provided, the logger of ChatAboutVideo instance will be used.
Returns

Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

The conversation.

▸ startConversation(videos, options?, log?): Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

Start a conversation about a video.

Parameters
Name Type Description
videos (VideoInput | ImagesInput | AudioInput)[] Array of videos, images, or audios to be used in the conversation. For each video/audio, the file path and the prompt before it should be provided. For each group of images, the image file paths and the prompt before the image group should be provided.
options? OPTIONS Overriding options for this conversation
log? LineLogger<(message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void> Optional logger for this conversation, if not provided, the logger of ChatAboutVideo instance will be used.
Returns

Promise<Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>>

The conversation.

Class: Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>

chat.Conversation

Type parameters
Name Type
CLIENT any
OPTIONS extends AdditionalCompletionOptions = any
PROMPT any
RESPONSE any
Constructors
constructor

• new Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>(conversationId, api, prompt, options, cleanup?, log?)

Type parameters
Name Type
CLIENT any
OPTIONS extends AdditionalCompletionOptions = any
PROMPT any
RESPONSE any
Parameters
Name Type
conversationId string
api ChatApi<CLIENT, OPTIONS, PROMPT, RESPONSE>
prompt undefined | PROMPT
options OPTIONS
cleanup? () => Promise<any>
log undefined | LineLogger<(message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void>
Properties
Property Description
Protected api: ChatApi<CLIENT, OPTIONS, PROMPT, RESPONSE>
Protected checkpoints: number[] = [] Prompt length after each successful say or submitToolCallResults.

Note on PROMPT type assumption:
The turn-checkpoint and rewind mechanisms assume that PROMPT is an Array of message objects (as implemented
by standard providers such as Gemini and ChatGPT). Array lengths are recorded as checkpoint markers. If PROMPT is not an
Array (or is undefined), checkpoint recording and restoring safely degrade to no-ops.
Protected conversationId: string
Protected ended: boolean = false
Protected initialPromptLength: number Prompt length when this conversation was constructed, before any successful turn.
Rewind past every checkpoint restores the prompt to this length.
Protected log: undefined | LineLogger<(message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void, (message?: any, ...optionalParams: any[]) => void>
Protected options: OPTIONS
Protected prompt: undefined | PROMPT
Protected sharedCleanup: SharedCleanup
Protected usage: undefined | UsageMetadata
Methods
end

▸ end(): Promise<void>

End this conversation. Shared resources are deleted when this is the last living conversation in the fork family. Calling end again on the same instance does nothing.

Returns

Promise<void>

nothing


fork

▸ fork(steps?): Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>

Create a new conversation with a deep copy of this conversation's prompt and checkpoints. The fork starts with no usage of its own. Later turns and rewind calls on either conversation do not affect the other.

Pass steps to fork from an earlier checkpoint. That is a fork followed by rewind on the new conversation only.

Cleanup of shared resources (extracted frames, uploaded images) runs only after every conversation in the family has called end.

Parameters
Name Type Description
steps? number Optional number of successful turns to remove from the fork, with the same meaning as rewind. Omitted, zero, and negative values fork at the current prompt.
Returns

Conversation<CLIENT, OPTIONS, PROMPT, RESPONSE>

The forked conversation.

Throws

When this conversation or its family resources have already been ended / cleaned up.


getApi

▸ getApi(): ChatApi<CLIENT, OPTIONS, PROMPT, RESPONSE>

Get the underlying API instance.

Returns

ChatApi<CLIENT, OPTIONS, PROMPT, RESPONSE>

The underlying API instance.


getPrompt

▸ getPrompt(): undefined | PROMPT

Get the prompt for the current conversation. The prompt is the accumulated messages in the conversation so far.

Returns

undefined | PROMPT

The prompt which is the accumulated messages in the conversation so far.


getUsage

▸ getUsage(): undefined | UsageMetadata

Get usage statistics of the conversation. Please note that the usage statistics would be undefined before the first say call. It could also be undefined if the underlying API does not support usage statistics. The usage statistics may not cover those failed requests due to content filtering or other reasons. Therefore, it could be less than the billable usage.

Returns

undefined | UsageMetadata

The usage statistics of the conversation. Or undefined if not available.


progressConversation

▸ Protected progressConversation(updatedPrompt, effectiveOptions): Promise<undefined | string | ConversationResponse>

Parameters
Name Type
updatedPrompt PROMPT
effectiveOptions OPTIONS
Returns

Promise<undefined | string | ConversationResponse>


promptLength

▸ Protected promptLength(): number

Returns

number


recordCheckpoint

▸ Protected recordCheckpoint(): void

Record a checkpoint marker of the current prompt length after a successful turn. Assumes this.prompt is an Array of message objects. If this.prompt is not an Array (or is undefined), recording is safely skipped (no-op).

Returns

void

nothing


restorePromptLength

▸ Protected restorePromptLength(length): void

Shrink the prompt back to length. Assumes this.prompt is an Array of message objects (as used by Gemini and ChatGPT APIs). If this.prompt is not an Array (or is undefined), shrinking is safely skipped (no-op), ensuring a failed restore cannot mask the error that caused it. Note: If PROMPT is not an Array and an API call fails mid-turn, automatic prompt restoration on error will be a no-op, leaving this.prompt in its partially-appended state.

Parameters
Name Type Description
length number Prompt length to restore.
Returns

void

nothing


rewind

▸ rewind(steps): void

Remove the last successful turns from the prompt. One turn is one successful say or submitToolCallResults. Usage already recorded on this conversation is left as it is.

Note: Rewind assumes this.prompt is an Array of message objects (as used by Gemini and ChatGPT APIs). If PROMPT is not an Array (or is undefined), rewind becomes a no-op because no checkpoints are recorded.

Parameters
Name Type Description
steps number Number of successful turns to remove. Values past the number of successful turns remove every turn and restore the prompt to the length it had when this conversation was created. Zero and negative values do nothing.
Returns

void

nothing


say

▸ say<RT>(message, options?): Promise<RT>

Say something in the conversation, and get the response from AI

Type parameters
Name Type Description
RT extends string | ConversationResponse = string The type of the response. It can be a string | undefined, or ConversationResponse, or the combination of them. You need to choose the correct type based on whether tool call could be returned.
Parameters
Name Type Description
message string The message to say in the conversation.
options? Partial<OPTIONS> Options for fine control.
Returns

Promise<RT>

The response text if there's no tool call, or a ConversationResponse object if there's tool call.


submitToolCallResults

▸ submitToolCallResults<RT>(toolResults, options?): Promise<RT>

Submit tool call results to the conversation, and get the response from AI.

Type parameters
Name Type Description
RT extends string | ConversationResponse = string The type of the response. It can be a string or ConversationResponse, or the combination of them. You need to choose the correct type based on whether tool call could be returned.
Parameters
Name Type Description
toolResults ToolCallResult[] Array of tool call results.
options? Partial<OPTIONS> Options for fine control
Returns

Promise<RT>

The response text if there's no further tool call, or a ConversationResponse object if there's further tool call.

▸ submitToolCallResults<RT>(toolResults, additionalMessage?, options?): Promise<RT>

Submit tool call results to the conversation, and get the response from AI.

Type parameters
Name Type Description
RT extends string | ConversationResponse = string The type of the response. It can be a string or ConversationResponse, or the combination of them. You need to choose the correct type based on whether tool call could be returned.
Parameters
Name Type Description
toolResults ToolCallResult[] Array of tool call results.
additionalMessage? string Optional message to append to the prompt
options? Partial<OPTIONS> Options for fine control
Returns

Promise<RT>

The response text if there's no further tool call, or a ConversationResponse object if there's further tool call.

Class: ChatGptApi

chat-gpt.ChatGptApi

Implements
Constructors
constructor

• new ChatGptApi(options)

Parameters
Name Type
options ChatGptOptions
Properties
Property Description
Protected client: ChatGptClient
Protected Optional extractVideoFrames: EffectiveExtractVideoFramesOptions
Protected options: ChatGptOptions
Protected Optional storage: Required<Pick<StorageOptions, "uploader">> & StorageOptions
Protected tmpDir: string
Methods
appendToPrompt

▸ appendToPrompt(newPromptOrResponse, prompt?): Promise<ChatCompletionMessageParam[]>

Append a new prompt or response to the form a full prompt. This function is useful to build a prompt that contains conversation history.

Parameters
Name Type Description
newPromptOrResponse ChatCompletionMessageParam[] | ChatCompletion A new prompt to be appended, or previous response to be appended.
prompt? ChatCompletionMessageParam[] The conversation history which is a prompt containing previous prompts and responses. If it is not provided, the conversation history returned will contain only what is in newPromptOrResponse.
Returns

Promise<ChatCompletionMessageParam[]>

The full prompt which is effectively the conversation history.

Implementation of

ChatApi.appendToPrompt


buildAudioPrompt

▸ buildAudioPrompt(audioFile, _conversationId?): Promise<BuildPromptOutput<ChatCompletionMessageParam[], ChatGptCompletionOptions>>

Build prompt for sending audio content to AI. Sometimes, to include audio in the conversation, additional options and/or clean up is needed. In such case, options to be passed to generateContent function and/or a clean up callback function can be returned from this function.

Parameters
Name Type Description
audioFile string Path to the audio file.
_conversationId? string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<ChatCompletionMessageParam[], ChatGptCompletionOptions>>

An object containing the prompt, optional options, and an optional cleanup function.

Implementation of

ChatApi.buildAudioPrompt


buildImagesPrompt

▸ buildImagesPrompt(imageInputs, conversationId?): Promise<BuildPromptOutput<ChatCompletionMessageParam[], ChatGptCompletionOptions>>

Build prompt for sending images content to AI. Sometimes, to include images in the conversation, additional options and/or clean up is needed. In such case, options to be passed to generateContent function and/or a clean up callback function can be returned from this function.

Parameters
Name Type Description
imageInputs ImageInput[] Array of image inputs.
conversationId string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<ChatCompletionMessageParam[], ChatGptCompletionOptions>>

An object containing the prompt, optional options, and an optional cleanup function.

Implementation of

ChatApi.buildImagesPrompt


buildTextPrompt

▸ buildTextPrompt(text, _conversationId?): Promise<{ prompt: ChatCompletionMessageParam[] }>

Build prompt for sending text content to AI

Parameters
Name Type Description
text string The text content to be sent.
_conversationId? string Unique identifier of the conversation.
Returns

Promise<{ prompt: ChatCompletionMessageParam[] }>

An object containing the prompt.

Implementation of

ChatApi.buildTextPrompt


buildToolCallResultsPrompt

▸ buildToolCallResultsPrompt(toolResults, _conversationId?): Promise<BuildPromptOutput<ChatCompletionMessageParam[], ChatGptCompletionOptions>>

Build prompt for tool results.

Parameters
Name Type Description
toolResults ToolCallResult[] Array of tool call results.
_conversationId? string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<ChatCompletionMessageParam[], ChatGptCompletionOptions>>

An object containing the prompt.

Implementation of

ChatApi.buildToolCallResultsPrompt


buildVideoPrompt

▸ buildVideoPrompt(videoFile, conversationId?): Promise<BuildPromptOutput<ChatCompletionMessageParam[], ChatGptCompletionOptions>>

Build prompt for sending video content to AI. Sometimes, to include video in the conversation, additional options and/or clean up is needed. In such case, options to be passed to generateContent function and/or a clean up callback function can be returned from this function.

Parameters
Name Type Description
videoFile string Path to the video file.
conversationId string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<ChatCompletionMessageParam[], ChatGptCompletionOptions>>

An object containing the prompt, optional options, and an optional cleanup function.

Implementation of

ChatApi.buildVideoPrompt


generateContent

▸ generateContent(prompt, options): Promise<ChatCompletion>

Generate content based on the given prompt and options.

Parameters
Name Type Description
prompt ChatCompletionMessageParam[] The full prompt to generate content.
options ChatGptCompletionOptions Optional options to control the content generation.
Returns

Promise<ChatCompletion>

The generated content.

Implementation of

ChatApi.generateContent


getClient

▸ getClient(): Promise<ChatGptClient>

Get the raw client. This function could be useful for advanced use cases.

Returns

Promise<ChatGptClient>

The raw client.

Implementation of

ChatApi.getClient


getResponseText

▸ getResponseText(result): Promise<string>

Get the text from the response object

Parameters
Name Type Description
result ChatCompletion the response object
Returns

Promise<string>

Implementation of

ChatApi.getResponseText


getToolCalls

▸ getToolCalls(result): Promise<undefined | ToolCall[]>

Extract tool calls from the response object.

Parameters
Name Type Description
result ChatCompletion the response object
Returns

Promise<undefined | ToolCall[]>

Array of tool calls if tool calling is requested by AI, or undefined otherwise.

Implementation of

ChatApi.getToolCalls


getUsageMetadata

▸ getUsageMetadata(result): Promise<undefined | UsageMetadata>

Extract usage metadata from the response object.

Parameters
Name Type Description
result ChatCompletion the response object
Returns

Promise<undefined | UsageMetadata>

Usage metadata from the response, if available. If the response does not contain usage metadata, it returns undefined.

Implementation of

ChatApi.getUsageMetadata


isConnectivityError

▸ isConnectivityError(error): boolean

Check if the error is a connectivity error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a connectivity error, false otherwise.

Implementation of

ChatApi.isConnectivityError


isDownloadError

▸ isDownloadError(error): boolean

Check if the error is a temporary download error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a temporary connectivity error, false otherwise.

Implementation of

ChatApi.isDownloadError


isServerError

▸ isServerError(error): boolean

Check if the error is a server error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a server error, false otherwise.

Implementation of

ChatApi.isServerError


isThrottlingError

▸ isThrottlingError(error): boolean

Check if the error is a throttling error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a throttling error, false otherwise.

Implementation of

ChatApi.isThrottlingError

Class: GeminiApi

gemini.GeminiApi

Implements
Constructors
constructor

• new GeminiApi(options)

Parameters
Name Type
options GeminiOptions
Properties
Property Description
Protected client: GenerativeModel
Protected extractVideoFrames: EffectiveExtractVideoFramesOptions
Protected options: GeminiOptions
Protected tmpDir: string
Methods
appendToPrompt

▸ appendToPrompt(newPromptOrResponse, prompt?): Promise<Content[]>

Append a new prompt or response to the form a full prompt. This function is useful to build a prompt that contains conversation history.

Parameters
Name Type Description
newPromptOrResponse Content[] | GenerateContentResult A new prompt to be appended, or previous response to be appended.
prompt? Content[] The conversation history which is a prompt containing previous prompts and responses. If it is not provided, the conversation history returned will contain only what is in newPromptOrResponse.
Returns

Promise<Content[]>

The full prompt which is effectively the conversation history.

Implementation of

ChatApi.appendToPrompt


buildAudioPrompt

▸ buildAudioPrompt(audioFile, _conversationId?): Promise<BuildPromptOutput<Content[], GeminiCompletionOptions>>

Build prompt for sending audio content to AI. Sometimes, to include audio in the conversation, additional options and/or clean up is needed. In such case, options to be passed to generateContent function and/or a clean up callback function can be returned from this function.

Parameters
Name Type Description
audioFile string Path to the audio file.
_conversationId? string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<Content[], GeminiCompletionOptions>>

An object containing the prompt, optional options, and an optional cleanup function.

Implementation of

ChatApi.buildAudioPrompt


buildImagesPrompt

▸ buildImagesPrompt(imageInputs, _conversationId): Promise<BuildPromptOutput<Content[], GeminiCompletionOptions>>

Build prompt for sending images content to AI. Sometimes, to include images in the conversation, additional options and/or clean up is needed. In such case, options to be passed to generateContent function and/or a clean up callback function can be returned from this function.

Parameters
Name Type Description
imageInputs ImageInput[] Array of image inputs.
_conversationId string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<Content[], GeminiCompletionOptions>>

An object containing the prompt, optional options, and an optional cleanup function.

Implementation of

ChatApi.buildImagesPrompt


buildTextPrompt

▸ buildTextPrompt(text, _conversationId?): Promise<{ prompt: Content[] }>

Build prompt for sending text content to AI

Parameters
Name Type Description
text string The text content to be sent.
_conversationId? string Unique identifier of the conversation.
Returns

Promise<{ prompt: Content[] }>

An object containing the prompt.

Implementation of

ChatApi.buildTextPrompt


buildToolCallResultsPrompt

▸ buildToolCallResultsPrompt(toolResults, _conversationId?): Promise<BuildPromptOutput<Content[], GeminiCompletionOptions>>

Build prompt for tool results.

Parameters
Name Type Description
toolResults ToolCallResult[] Array of tool call results.
_conversationId? string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<Content[], GeminiCompletionOptions>>

An object containing the prompt.

Implementation of

ChatApi.buildToolCallResultsPrompt


buildVideoPrompt

▸ buildVideoPrompt(videoFile, conversationId?): Promise<BuildPromptOutput<Content[], GeminiCompletionOptions>>

Build prompt for sending video content to AI. Sometimes, to include video in the conversation, additional options and/or clean up is needed. In such case, options to be passed to generateContent function and/or a clean up callback function can be returned from this function.

Parameters
Name Type Description
videoFile string Path to the video file.
conversationId string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<Content[], GeminiCompletionOptions>>

An object containing the prompt, optional options, and an optional cleanup function.

Implementation of

ChatApi.buildVideoPrompt


generateContent

▸ generateContent(prompt, options): Promise<GenerateContentResult>

Generate content based on the given prompt and options.

Parameters
Name Type Description
prompt Content[] The full prompt to generate content.
options GeminiCompletionOptions Optional options to control the content generation.
Returns

Promise<GenerateContentResult>

The generated content.

Implementation of

ChatApi.generateContent


getClient

▸ getClient(): Promise<GenerativeModel>

Get the raw client. This function could be useful for advanced use cases.

Returns

Promise<GenerativeModel>

The raw client.

Implementation of

ChatApi.getClient


getResponseText

▸ getResponseText(result): Promise<string>

Get the text from the response object

Parameters
Name Type Description
result GenerateContentResult the response object
Returns

Promise<string>

Implementation of

ChatApi.getResponseText


getToolCalls

▸ getToolCalls(result): Promise<undefined | ToolCall[]>

Extract tool calls from the response object.

Parameters
Name Type Description
result GenerateContentResult the response object
Returns

Promise<undefined | ToolCall[]>

Array of tool calls if tool calling is requested by AI, or undefined otherwise.

Implementation of

ChatApi.getToolCalls


getUsageMetadata

▸ getUsageMetadata(result): Promise<undefined | UsageMetadata>

Extract usage metadata from the response object.

Parameters
Name Type Description
result GenerateContentResult the response object
Returns

Promise<undefined | UsageMetadata>

Usage metadata from the response, if available. If the response does not contain usage metadata, it returns undefined.

Implementation of

ChatApi.getUsageMetadata


isConnectivityError

▸ isConnectivityError(error): boolean

Check if the error is a connectivity error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a connectivity error, false otherwise.

Implementation of

ChatApi.isConnectivityError


isDownloadError

▸ isDownloadError(_error): boolean

Check if the error is a temporary download error.

Parameters
Name Type Description
_error any any error object
Returns

boolean

true if the error is a temporary connectivity error, false otherwise.

Implementation of

ChatApi.isDownloadError


isServerError

▸ isServerError(error): boolean

Check if the error is a server error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a server error, false otherwise.

Implementation of

ChatApi.isServerError


isThrottlingError

▸ isThrottlingError(error): boolean

Check if the error is a throttling error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a throttling error, false otherwise.

Implementation of

ChatApi.isThrottlingError

Interfaces

Interface: AdditionalCompletionOptions

types.AdditionalCompletionOptions

Properties
Property Description
Optional backoffOnConnectivityError: number[] Array of retry backoff periods (unit: milliseconds) for situations that the network connection couldn't be established or lost or request/response timeout.
Optional backoffOnDownloadError: number[] Array of retry backoff periods (unit: milliseconds) for situations that the AI temporarily fails to download a file.
This kind of situation has a chance to happen when many image URLs are passed to OpenAI at the same time.
Optional backoffOnServerError: number[] Array of retry backoff periods (unit: milliseconds) for situations that the server returns 5xx response
Optional backoffOnThrottling: number[] Array of retry backoff periods (unit: milliseconds) for situations that the server returns 429 response
Optional jsonResponse: boolean | JSONSchema | { schema: Schema } & Partial<Omit<JSONSchema, "schema">>
Optional startPromptText: string The user prompt that will be sent before the video content.
If not provided, nothing will be sent before the video content.
Optional systemPromptText: string System prompt text. If not provided, a default prompt will be used.

Interface: AudioInput

types.AudioInput

Properties
Property Description
audioFile: string Path to an audio file in local file system.
promptText: string The prompt before the audio.

Interface: BuildPromptOutput<PROMPT, OPTIONS>

types.BuildPromptOutput

Type parameters
Name
PROMPT
OPTIONS
Properties
Property Description
Optional cleanup: () => Promise<any>
Optional options: Partial<OPTIONS>
prompt: PROMPT

Interface: ChatApi<CLIENT, OPTIONS, PROMPT, RESPONSE>

types.ChatApi

Type parameters
Name Type
CLIENT CLIENT
OPTIONS extends AdditionalCompletionOptions
PROMPT PROMPT
RESPONSE RESPONSE
Implemented by
Methods
appendToPrompt

▸ appendToPrompt(newPromptOrResponse, prompt?): Promise<PROMPT>

Append a new prompt or response to the form a full prompt. This function is useful to build a prompt that contains conversation history.

Parameters
Name Type Description
newPromptOrResponse PROMPT | RESPONSE A new prompt to be appended, or previous response to be appended.
prompt? PROMPT The conversation history which is a prompt containing previous prompts and responses. If it is not provided, the conversation history returned will contain only what is in newPromptOrResponse.
Returns

Promise<PROMPT>

The full prompt which is effectively the conversation history.


buildAudioPrompt

▸ buildAudioPrompt(audioFile, conversationId?): Promise<BuildPromptOutput<PROMPT, OPTIONS>>

Build prompt for sending audio content to AI. Sometimes, to include audio in the conversation, additional options and/or clean up is needed. In such case, options to be passed to generateContent function and/or a clean up callback function can be returned from this function.

Parameters
Name Type Description
audioFile string Path to the audio file.
conversationId? string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<PROMPT, OPTIONS>>

An object containing the prompt, optional options, and an optional cleanup function.


buildImagesPrompt

▸ buildImagesPrompt(imageInputs, conversationId?): Promise<BuildPromptOutput<PROMPT, OPTIONS>>

Build prompt for sending images content to AI. Sometimes, to include images in the conversation, additional options and/or clean up is needed. In such case, options to be passed to generateContent function and/or a clean up callback function can be returned from this function.

Parameters
Name Type Description
imageInputs ImageInput[] Array of image inputs.
conversationId? string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<PROMPT, OPTIONS>>

An object containing the prompt, optional options, and an optional cleanup function.


buildTextPrompt

▸ buildTextPrompt(text, conversationId?): Promise<{ prompt: PROMPT }>

Build prompt for sending text content to AI

Parameters
Name Type Description
text string The text content to be sent.
conversationId? string Unique identifier of the conversation.
Returns

Promise<{ prompt: PROMPT }>

An object containing the prompt.


buildToolCallResultsPrompt

▸ buildToolCallResultsPrompt(toolResults, conversationId?): Promise<BuildPromptOutput<PROMPT, OPTIONS>>

Build prompt for tool results.

Parameters
Name Type Description
toolResults ToolCallResult[] Array of tool call results.
conversationId? string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<PROMPT, OPTIONS>>

An object containing the prompt.


buildVideoPrompt

▸ buildVideoPrompt(videoFile, conversationId?): Promise<BuildPromptOutput<PROMPT, OPTIONS>>

Build prompt for sending video content to AI. Sometimes, to include video in the conversation, additional options and/or clean up is needed. In such case, options to be passed to generateContent function and/or a clean up callback function can be returned from this function.

Parameters
Name Type Description
videoFile string Path to the video file.
conversationId? string Unique identifier of the conversation.
Returns

Promise<BuildPromptOutput<PROMPT, OPTIONS>>

An object containing the prompt, optional options, and an optional cleanup function.


generateContent

▸ generateContent(prompt, options?): Promise<RESPONSE>

Generate content based on the given prompt and options.

Parameters
Name Type Description
prompt PROMPT The full prompt to generate content.
options? OPTIONS Optional options to control the content generation.
Returns

Promise<RESPONSE>

The generated content.


getClient

▸ getClient(): Promise<CLIENT>

Get the raw client. This function could be useful for advanced use cases.

Returns

Promise<CLIENT>

The raw client.


getResponseText

▸ getResponseText(response): Promise<string>

Get the text from the response object

Parameters
Name Type Description
response RESPONSE the response object
Returns

Promise<string>


getToolCalls

▸ getToolCalls(response): Promise<undefined | ToolCall[]>

Extract tool calls from the response object.

Parameters
Name Type Description
response RESPONSE the response object
Returns

Promise<undefined | ToolCall[]>

Array of tool calls if tool calling is requested by AI, or undefined otherwise.


getUsageMetadata

▸ getUsageMetadata(response): Promise<undefined | UsageMetadata>

Extract usage metadata from the response object.

Parameters
Name Type Description
response RESPONSE the response object
Returns

Promise<undefined | UsageMetadata>

Usage metadata from the response, if available. If the response does not contain usage metadata, it returns undefined.


isConnectivityError

▸ isConnectivityError(error): boolean

Check if the error is a connectivity error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a connectivity error, false otherwise.


isDownloadError

▸ isDownloadError(error): boolean

Check if the error is a temporary download error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a temporary connectivity error, false otherwise.


isServerError

▸ isServerError(error): boolean

Check if the error is a server error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a server error, false otherwise.


isThrottlingError

▸ isThrottlingError(error): boolean

Check if the error is a throttling error.

Parameters
Name Type Description
error any any error object
Returns

boolean

true if the error is a throttling error, false otherwise.

Interface: ChatApiOptions<CS, CO>

types.ChatApiOptions

Type parameters
Name
CS
CO
Properties

| Property | Description | | ----------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | ---- | ---- | ---- | ------- | ------- | ---- | ----- | -------- | --- | | Optional clientSettings: CS | | | Optional completionOptions: AdditionalCompletionOptions & CO | | | credential: Object | Type declaration

| Name | Type |
| :------ | :------ |
| key | string | | | Optional endpoint: string | | | Optional tmpDir: string | Temporary directory for storing temporary files.
If not specified, then the temporary directory of the OS will be used. |

Interface: ConversationResponse

types.ConversationResponse

Properties
Property Description
Optional responseText: string Response text from AI.
Optional toolCalls: ToolCall[] Array of tool calls if tool calling is requested by AI.

Interface: ExtractVideoFramesOptions

types.ExtractVideoFramesOptions

Properties
Property Description
Optional deleteFilesWhenConversationEnds: boolean Whether files should be deleted when the conversation ends.
Optional extractor: VideoFramesExtractor Function for extracting frames from the video.
If not specified, a default function using ffmpeg will be used.
Optional format: string Image format of the extracted frames.
Default value is 'jpg'.
Optional framesDirectoryResolver: (inputFile: string, tmpDir: string, conversationId: string) => string Function for determining the directory location for storing extracted frames.
If not specified, a default function will be used.
The function takes three arguments:
Optional height: number Video frame height, default is undefined which means the scaling
will be determined by the videoFrameWidth option.
If both videoFrameWidth and videoFrameHeight are not specified,
then the frames will not be resized/scaled.
Optional interval: number Intervals between frames to be extracted. The unit is second.
Default value is 5.
Optional limit: number Maximum number of frames to be extracted.
Default value is 10 which is the current per-request limitation of ChatGPT Vision.
Optional width: number Video frame width, default is 200.
If both videoFrameWidth and videoFrameHeight are not specified,
then the frames will not be resized/scaled.

Interface: ImageInput

types.ImageInput

Properties
Property Description
imageFile: string Path to an image file in local file system.
Optional promptText: string The prompt text before the image.
This is optional, and could be used to provide the timestamp or other information about the image.

Interface: ImagesInput

types.ImagesInput

Properties
Property Description
images: ImageInput[]
promptText: string The prompt before the images.

Interface: StorageOptions

types.StorageOptions

Properties
Property Description
Optional azureStorageConnectionString: string
Optional deleteFilesWhenConversationEnds: boolean Whether files should be deleted when the conversation ends.
Optional downloadUrlExpirationSeconds: number Expiration time for the download URL of the frame images in seconds. Default is 3600 seconds.
Optional storageContainerName: string Storage container for storing frame images of the video.
Optional storagePathPrefix: string Path prefix to be prepended for storing frame images of the video.
Default is empty.
Optional uploader: FileBatchUploader Function for uploading files

Interface: ToolCall

types.ToolCall

Properties
Property Description
arguments: Record<string, any> The arguments to the function call, already parsed into an object.
Optional id: string Unique identifier for the tool call.
OpenAI always provides this, Gemini does not.
name: string The name of the function to be called.

Interface: ToolCallResult

types.ToolCallResult

Properties
Property Description
name: string The name of the function being responded to.
Required by Gemini.
result: Record<string, any> The result of the function call, as an object.
Optional toolCallId: string Unique identifier for the tool call being responded to.
Required by OpenAI.

Interface: UsageMetadata

types.UsageMetadata

Properties
Property Description
Optional completionTokens: number
Optional promptTokens: number
totalTokens: number

Interface: VideoInput

types.VideoInput

Properties
Property Description
promptText: string The prompt before the video.
videoFile: string Path to a video file in local file system.

## Modules

Module: aws
Functions
createAwsS3FileBatchUploader

▸ createAwsS3FileBatchUploader(s3Client, expirationSeconds, parallelism?): FileBatchUploader

Parameters
Name Type Default value
s3Client S3Client undefined
expirationSeconds number undefined
parallelism number 3
Returns

FileBatchUploader

Module: azure
Functions
createAzureBlobStorageFileBatchUploader

▸ createAzureBlobStorageFileBatchUploader(blobServiceClient, expirationSeconds, parallelism?): FileBatchUploader

Parameters
Name Type Default value
blobServiceClient BlobServiceClient undefined
expirationSeconds number undefined
parallelism number 3
Returns

FileBatchUploader

Module: chat
Classes
Type Aliases
ChatAboutVideoWith

Ƭ ChatAboutVideoWith<T>: ChatAboutVideo<ClientOfChatApi<T>, OptionsOfChatApi<T>, PromptOfChatApi<T>, ResponseOfChatApi<T>>

Type parameters
Name
T

ChatAboutVideoWithChatGpt

Ƭ ChatAboutVideoWithChatGpt: ChatAboutVideoWith<ChatGptApi>


ChatAboutVideoWithGemini

Ƭ ChatAboutVideoWithGemini: ChatAboutVideoWith<GeminiApi>


ConversationWith

Ƭ ConversationWith<T>: Conversation<ClientOfChatApi<T>, OptionsOfChatApi<T>, PromptOfChatApi<T>, ResponseOfChatApi<T>>

Type parameters
Name
T

ConversationWithChatGpt

Ƭ ConversationWithChatGpt: ConversationWith<ChatGptApi>


ConversationWithGemini

Ƭ ConversationWithGemini: ConversationWith<GeminiApi>


MultipleSupportedChatApiOptions

Ƭ MultipleSupportedChatApiOptions: { active: string ; base?: Partial<SupportedChatApiOptions> | null } & Record<string, Partial<SupportedChatApiOptions> | string | null | undefined>

Options for multiple supported chat APIs. Its "base" property is the base options to be used for merging with the active options. Its "active" property specifies the name of the active options. The active options will be merged with the base options, with the active options taking precedence.


SupportedChatApiOptions

Ƭ SupportedChatApiOptions: ChatGptOptions | GeminiOptions

Functions
accumulateUsage

▸ accumulateUsage(totalUsage, incrementalUsage): undefined | UsageMetadata

Add up usage.

Parameters
Name Type Description
totalUsage UsageMetadata Existing usage that will be updated.
incrementalUsage undefined | UsageMetadata New usage to add. If it is undefined, then there will be no change to totalUsage.
Returns

undefined | UsageMetadata

nothing, the totalUsage is updated in place.


activeSupportedChatApiOptions

▸ activeSupportedChatApiOptions(options): SupportedChatApiOptions

Get the active options from the multiple options. It first finds the active options using the active key, and then merges the base options with the active options.

Parameters
Name Type Description
options MultipleSupportedChatApiOptions The multiple options. It will not be mutated by this function.
Returns

SupportedChatApiOptions

The active options which can be passed into the constructor of ChatAboutVideo


buildImagesPromptFromVideo

▸ buildImagesPromptFromVideo<CLIENT, OPTIONS, PROMPT, RESPONSE>(api, extractVideoFrames, tmpDir, videoFile, conversationId?): Promise<BuildPromptOutput<PROMPT, OPTIONS>>

Build prompt for sending frame images of a video content to AI. This function is usually used for implementing the buildVideoPrompt function of ChatApi by utilising already implemented buildImagesPrompt function. It extracts frame images from the video and builds a prompt containing those images for the conversation.

Type parameters
Name Type
CLIENT CLIENT
OPTIONS extends AdditionalCompletionOptions
PROMPT PROMPT
RESPONSE RESPONSE
Parameters
Name Type Description
api ChatApi<CLIENT, OPTIONS, PROMPT, RESPONSE> The API instance.
extractVideoFrames EffectiveExtractVideoFramesOptions The options for extracting video frames.
tmpDir string The temporary directory to store the extracted frames.
videoFile string Path to a video file in local file system.
conversationId string The conversation ID.
Returns

Promise<BuildPromptOutput<PROMPT, OPTIONS>>

The prompt and options for the conversation.


generateTempConversationId

▸ generateTempConversationId(): string

Convenient function to generate a temporary conversation ID.

Returns

string

A temporary conversation ID.

Module: chat-gpt
Classes
Type Aliases
ChatGptClient

Ƭ ChatGptClient: AzureOpenAI | OpenAI


ChatGptCompletionOptions

Ƭ ChatGptCompletionOptions: AdditionalCompletionOptions & Omit<OpenAI.ChatCompletionCreateParamsNonStreaming, "messages" | "stream">


ChatGptOptions

Ƭ ChatGptOptions: { extractVideoFrames?: ExtractVideoFramesOptions ; storage?: StorageOptions } & ChatApiOptions<AzureClientOptions, ChatGptCompletionOptions>


ChatGptPrompt

Ƭ ChatGptPrompt: OpenAI.ChatCompletionCreateParamsNonStreaming["messages"]


ChatGptResponse

Ƭ ChatGptResponse: OpenAI.ChatCompletion

Module: gemini
Classes
Type Aliases
GeminiClient

Ƭ GeminiClient: GenerativeModel


GeminiClientOptions

Ƭ GeminiClientOptions: Object

Type declaration
Name Type
modelParams ModelParams
requestOptions? RequestOptions

GeminiCompletionOptions

Ƭ GeminiCompletionOptions: AdditionalCompletionOptions & Omit<GenerateContentRequest, "contents">


GeminiOptions

Ƭ GeminiOptions: { clientSettings: GeminiClientOptions ; extractVideoFrames?: ExtractVideoFramesOptions } & ChatApiOptions<GeminiClientOptions, GeminiCompletionOptions>


GeminiPrompt

Ƭ GeminiPrompt: GenerateContentRequest["contents"]


GeminiResponse

Ƭ GeminiResponse: GenerateContentResult

Module: index
References
AdditionalCompletionOptions

Re-exports AdditionalCompletionOptions


AudioInput

Re-exports AudioInput


BuildPromptOutput

Re-exports BuildPromptOutput


ChatAboutVideo

Re-exports ChatAboutVideo


ChatAboutVideoWith

Re-exports ChatAboutVideoWith


ChatAboutVideoWithChatGpt

Re-exports ChatAboutVideoWithChatGpt


ChatAboutVideoWithGemini

Re-exports ChatAboutVideoWithGemini


ChatApi

Re-exports ChatApi


ChatApiOptions

Re-exports ChatApiOptions


ClientOfChatApi

Re-exports ClientOfChatApi


Conversation

Re-exports Conversation


ConversationResponse

Re-exports ConversationResponse


ConversationWith

Re-exports ConversationWith


ConversationWithChatGpt

Re-exports ConversationWithChatGpt


ConversationWithGemini

Re-exports ConversationWithGemini


EffectiveExtractVideoFramesOptions

Re-exports EffectiveExtractVideoFramesOptions


ExtractVideoFramesOptions

Re-exports ExtractVideoFramesOptions


FileBatchUploader

Re-exports FileBatchUploader


ImageInput

Re-exports ImageInput


ImagesInput

Re-exports ImagesInput


MultipleSupportedChatApiOptions

Re-exports MultipleSupportedChatApiOptions


OptionsOfChatApi

Re-exports OptionsOfChatApi


PromptOfChatApi

Re-exports PromptOfChatApi


ResponseOfChatApi

Re-exports ResponseOfChatApi


StorageOptions

Re-exports StorageOptions


SupportedChatApiOptions

Re-exports SupportedChatApiOptions


ToolCall

Re-exports ToolCall


ToolCallResult

Re-exports ToolCallResult


UsageMetadata

Re-exports UsageMetadata


VideoFramesExtractor

Re-exports VideoFramesExtractor


VideoInput

Re-exports VideoInput


accumulateUsage

Re-exports accumulateUsage


activeSupportedChatApiOptions

Re-exports activeSupportedChatApiOptions


buildImagesPromptFromVideo

Re-exports buildImagesPromptFromVideo


extractVideoFramesWithFfmpeg

Re-exports extractVideoFramesWithFfmpeg


generateTempConversationId

Re-exports generateTempConversationId


lazyCreatedFileBatchUploader

Re-exports lazyCreatedFileBatchUploader


lazyCreatedVideoFramesExtractor

Re-exports lazyCreatedVideoFramesExtractor

Module: storage
References
FileBatchUploader

Re-exports FileBatchUploader

Functions
lazyCreatedFileBatchUploader

▸ lazyCreatedFileBatchUploader(creator): FileBatchUploader

Parameters
Name Type
creator Promise<FileBatchUploader>
Returns

FileBatchUploader

Module: storage/types
Type Aliases
FileBatchUploader

Ƭ FileBatchUploader: (dir: string, relativePaths: string[], containerName: string, blobPathPrefix: string) => Promise<{ cleanup: () => Promise<any> ; downloadUrls: string[] }>

Type declaration

▸ (dir, relativePaths, containerName, blobPathPrefix): Promise<{ cleanup: () => Promise<any> ; downloadUrls: string[] }>

Function that uploads files to the cloud storage.

####### Parameters

Name Type Description
dir string The directory path where the files are located.
relativePaths string[] An array of relative paths of the files to be uploaded.
containerName string The name of the container where the files will be uploaded.
blobPathPrefix string The prefix for the blob paths (file paths) in the container.

####### Returns

Promise<{ cleanup: () => Promise<any> ; downloadUrls: string[] }>

A Promise that resolves with an object containing an array of download URLs for the uploaded files and a cleanup function to remove the uploaded files from the container.

Module: types
Interfaces
Type Aliases
ClientOfChatApi

Ƭ ClientOfChatApi<T>: T extends ChatApi<infer CLIENT, any, any, any> ? CLIENT : never

Type parameters
Name
T

EffectiveExtractVideoFramesOptions

Ƭ EffectiveExtractVideoFramesOptions: Pick<ExtractVideoFramesOptions, "height"> & Required<Omit<ExtractVideoFramesOptions, "height">>


OptionsOfChatApi

Ƭ OptionsOfChatApi<T>: T extends ChatApi<any, infer OPTIONS, any, any> ? OPTIONS : never

Type parameters
Name
T

PromptOfChatApi

Ƭ PromptOfChatApi<T>: T extends ChatApi<any, any, infer PROMPT, any> ? PROMPT : never

Type parameters
Name
T

ResponseOfChatApi

Ƭ ResponseOfChatApi<T>: T extends ChatApi<any, any, any, infer RESPONSE> ? RESPONSE : never

Type parameters
Name
T

Module: utils
Functions
effectiveExtractVideoFramesOptions

▸ effectiveExtractVideoFramesOptions(options?): EffectiveExtractVideoFramesOptions

Calculate the effective values for ExtractVideoFramesOptions by combining the default values and the values provided

Parameters
Name Type Description
options? ExtractVideoFramesOptions the options containing the values provided
Returns

EffectiveExtractVideoFramesOptions

The effective values for ExtractVideoFramesOptions


effectiveStorageOptions

▸ effectiveStorageOptions(options): Required<Pick<StorageOptions, "uploader">> & StorageOptions

Calculate the effective values for StorageOptions by combining the default values and the values provided

Parameters
Name Type Description
options StorageOptions the options containing the values provided
Returns

Required<Pick<StorageOptions, "uploader">> & StorageOptions

The effective values for StorageOptions


findCommonParentPath

▸ findCommonParentPath(paths): Object

Find the common parent path of the given paths. If there is no common parent path, then the root path of the current process will be returned.

Parameters
Name Type Description
paths string[] Input paths to find the common parent path for. It can be absolute or relative paths.
Returns

Object

The common parent path and the relative paths from the common parent.

Name Type
commonParent string
relativePaths string[]

Module: video
References
VideoFramesExtractor

Re-exports VideoFramesExtractor


extractVideoFramesWithFfmpeg

Re-exports extractVideoFramesWithFfmpeg

Functions
lazyCreatedVideoFramesExtractor

▸ lazyCreatedVideoFramesExtractor(creator): VideoFramesExtractor

Parameters
Name Type
creator Promise<VideoFramesExtractor>
Returns

VideoFramesExtractor

Module: video/ffmpeg
Functions
extractVideoFramesWithFfmpeg

▸ extractVideoFramesWithFfmpeg(inputFile, outputDir, intervalSec, format?, width?, height?, startSec?, endSec?, limit?): Promise<{ cleanup: () => Promise<any> ; relativePaths: string[] }>

Function that extracts frame images from a video file.

Parameters
Name Type Description
inputFile string Path to the input video file.
outputDir string Path to the output directory where frame images will be saved.
intervalSec number Interval in seconds between each frame extraction.
format? string Format of the output frame images (e.g., 'jpg', 'png').
width? number Width of the output frame images in pixels.
height? number Height of the output frame images in pixels.
startSec? number Start time of the video segment to extract in seconds, inclusive.
endSec? number End time of the video segment to extract in seconds, exclusive.
limit? number Maximum number of frames to extract.
Returns

Promise<{ cleanup: () => Promise<any> ; relativePaths: string[] }>

An object containing an array of relative paths to the extracted frame images and a cleanup function for deleting those files.

Module: video/types
Type Aliases
VideoFramesExtractor

Ƭ VideoFramesExtractor: (inputFile: string, outputDir: string, intervalSec: number, format?: string, width?: number, height?: number, startSec?: number, endSec?: number, limit?: number) => Promise<{ cleanup: () => Promise<any> ; relativePaths: string[] }>

Type declaration

▸ (inputFile, outputDir, intervalSec, format?, width?, height?, startSec?, endSec?, limit?): Promise<{ cleanup: () => Promise<any> ; relativePaths: string[] }>

Function that extracts frame images from a video file.

####### Parameters

Name Type Description
inputFile string Path to the input video file.
outputDir string Path to the output directory where frame images will be saved.
intervalSec number Interval in seconds between each frame extraction.
format? string Format of the output frame images (e.g., 'jpg', 'png').
width? number Width of the output frame images in pixels.
height? number Height of the output frame images in pixels.
startSec? number Start time of the video segment to extract in seconds, inclusive.
endSec? number End time of the video segment to extract in seconds, exclusive.
limit? number Maximum number of frames to extract.

####### Returns

Promise<{ cleanup: () => Promise<any> ; relativePaths: string[] }>

An object containing an array of relative paths to the extracted frame images and a cleanup function for deleting those files.

Keywords