open-wa
DocsAI Agent Patterns

AI Agent Patterns

Architecture patterns for building AI agents that interact with WhatsApp through open-wa.

AI Agent Patterns

This guide covers common architecture patterns for connecting LLMs to WhatsApp through open-wa. Snippets marked as runnable include setup or local stubs. Snippets marked as patterns are meant to show structure, not a complete application.

Application-Provided AI Helpers

The examples do not assume an official LLM, vision, or speech-to-text SDK. In your app, provide helpers with the same shape, then replace the stub bodies with your own provider calls.

type ChatTurn = {
  role: 'user' | 'assistant';
  content: string;
};

async function callLLM(input: string | ChatTurn[] | unknown): Promise<string> {
  void input;
  return 'Thanks, I received your message.';
}

async function callVisionLLM(imageDataUrlOrBase64: string, mimeType?: string): Promise<string> {
  void imageDataUrlOrBase64;
  return `Image received${mimeType ? ` as ${mimeType}` : ''}.`;
}

async function callSTT(audioDataUrlOrBase64: string): Promise<string> {
  void audioDataUrlOrBase64;
  return 'Transcribed voice note text.';
}

Expected behavior with the stubs: a message comes in, the bot sends a deterministic placeholder response, and no external AI service is called. Add your own logging around each helper before replacing the stubs so you can see input, response, and provider errors.

The Conversation Loop Pattern

The simplest AI agent receives messages, sends them to an LLM helper, and replies with the response.

import { Client, createClient } from '@open-wa/wa-automate';
import { PuppeteerDriver } from '@open-wa/driver-puppeteer';

type ChatTurn = {
  role: 'user' | 'assistant';
  content: string;
};

async function callLLM(input: string | ChatTurn[]): Promise<string> {
  void input;
  return 'Thanks, I received your message.';
}

const runtime = await createClient({
  sessionId: 'ai-bot',
  driver: new PuppeteerDriver(),
  headless: false,
});

const client = new Client({ client: runtime, transport: runtime.getTransport() });
await client.start();

client.onMessage(async (message) => {
  if (message.type !== 'chat' || message.fromMe) return;

  const response = await callLLM(message.body);
  await client.sendText(message.from, response);
});

Expected behavior: when a user sends a text message, the bot replies with the helper response. Messages from the bot itself and non-text messages are ignored.

Key Decisions

  • When to respond: Filter by message type, sender, group membership, or keywords
  • What to skip: System messages, media, status updates, your own messages
  • How to reply: sendText for text, sendImage for generated images, reply for quoted responses

Context Management

LLMs need conversation context to maintain coherent multi-turn conversations.

Sliding Window

Pattern snippet: this assumes client and callLLM come from your application setup.

type ChatTurn = {
  role: 'user' | 'assistant';
  content: string;
};

const contextWindow = new Map<string, ChatTurn[]>();

client.onMessage(async (message) => {
  const chatId = message.from;
  const history = contextWindow.get(chatId) ?? [];

  history.push({ role: 'user', content: message.body });

  if (history.length > 20) {
    history.splice(0, history.length - 20);
  }

  const response = await callLLM(history);
  history.push({ role: 'assistant', content: response });
  contextWindow.set(chatId, history);

  await client.sendText(chatId, response);
});

Expected behavior: each chat keeps its own rolling context. After more than 10 user and assistant exchanges, older turns are dropped before the next LLM call.

Use getMessagesForLLM

Pattern snippet: if this generated Easy API method is available in your version, it can provide WhatsApp-native context. Its second argument is last, not count; provide your own callLLM helper for the returned message shape. Check the method reference before using it on an embedded client.

client.onMessage(async (message) => {
  const recentMessages = await client.getMessagesForLLM(message.chatId, 20);

  const response = await callLLM(recentMessages);
  await client.sendText(message.from, response);
});

Expected behavior: the bot asks open-wa for recent chat context, sends that context to your helper, then replies to the same chat.

Rate Limits and Safety

Implementing Rate Limiting

Install the limiter in your application. This documentation example does not add dependencies to the open-wa repo.

npm install bottleneck

Runnable snippet, assuming client and callLLM are defined as shown above:

import Bottleneck from 'bottleneck';

const limiter = new Bottleneck({
  minTime: 15000,
  maxConcurrent: 1,
});

client.onMessage(async (message) => {
  if (message.type !== 'chat' || message.fromMe) return;

  try {
    const response = await callLLM(message.body);
    await limiter.schedule(() => {
      console.log('ai-agent: sending rate-limited reply', { chatId: message.from });
      return client.sendText(message.from, response);
    });
  } catch (error) {
    console.error('ai-agent: queued reply failed', { chatId: message.from, error });
  }
});

Expected behavior: model calls run before the global outbound limiter, and outgoing replies start at least 15 seconds apart within this process. The catch block handles both provider rejection and queue failure; this schedule is not a WhatsApp safety guarantee and does not preserve per-chat ordering by itself.

Ban Risk Profile

Provider restrictions can reflect volume, repeated content, recipient consent, account context, or signals that are not documented by open-wa. Keep consent, opt-out state, queue depth, and retry ownership in your application, and pause the workflow when the provider rejects sends. Random timing and an account's age are observations or assumptions, not controls this guide can promise.

Media in LLM Pipelines

Processing Incoming Images

Pattern snippet: callVisionLLM is an application-provided helper. Keep it as a stub until you wire your chosen vision provider.

client.onMessage(async (message) => {
  if (message.type !== 'image') return;

  const mediaData = await client.decryptMedia(message);
  const base64 = mediaData.split(',')[1] ?? mediaData;
  const description = await callVisionLLM(base64, message.mimetype);

  await client.reply(message.from, `Image: ${description}`, message.id);
});

Expected behavior: The bot receives and decrypts an image message. The vision helper returns text. The bot replies with the description.

Processing Voice Notes

Pattern snippet: callSTT and callLLM are application-provided helpers. Keep them as stubs until you wire your chosen speech and LLM providers.

client.onMessage(async (message) => {
  if (message.type !== 'audio' && !message.mimetype?.includes('ogg')) return;

  const mediaData = await client.decryptMedia(message);
  const transcription = await callSTT(mediaData);
  const response = await callLLM(transcription);

  await client.sendText(message.from, response);
});

Expected behavior: The bot receives and decrypts a voice note. The speech helper returns text. The LLM helper creates a response, and the bot sends it.

Sending Generated Media

Pattern snippet: sendImage on the embedded client uses sendImage(to, dataUrlOrBase64, filename, caption?). The filename is necessary before the optional caption.

const imageDataUrl = 'data:image/png;base64,...';
const imageFilename = 'ai-generated.png';
const imageCaption = 'Generated image';

await client.sendImage(message.from, imageDataUrl, imageFilename, imageCaption);

Expected behavior: the image is sent to the incoming chat with ai-generated.png as the attachment filename and Generated image as the caption. sendPtt is a generated Easy API method; check that method's runtime support before adding voice-note output to an embedded client.

Group Chats vs Direct Messages

Detecting the Chat Type

Use this pattern in your message handler when group behavior must differ from direct-message behavior.

client.onMessage(async (message) => {
  const isGroup = message.from.includes('@g.us');
  const isDirect = message.from.includes('@c.us');

  if (isGroup) {
    if (!message.body.includes('@botname')) return;
  }

  if (isDirect) {
    console.log('ai-agent: direct message received', { chatId: message.from });
  }
});

Expected behavior: group messages only continue when the bot is mentioned. Direct messages continue through the direct-message branch and can be logged or handled separately.

Group-Specific Considerations

  • Respond only when @botname detection finds a mention
  • Respect group admin rules
  • Be aware of group size and message volume
  • Use group-specific context windows when necessary

Interactive Messages

Buttons and Lists

Use these sends when you want your agent to offer fixed choices instead of free-form text. These are generated Easy API methods: sendButtons and sendListMessage are deprecated and require an Insiders license, and the current embedded Client does not declare them. The alias sendList refers to sendListMessage on surfaces that expose aliases.

await client.sendButtons(message.from, 'Choose an option:', [
  { id: 'opt1', body: 'Option 1' },
  { id: 'opt2', body: 'Option 2' },
  { id: 'opt3', body: 'Option 3' },
]);

await client.sendList(message.from, 'Select from the list:', 'Choose', [
  { title: 'Section 1', rows: [{ id: 'r1', title: 'Row 1' }] },
]);

Expected behavior: the chat receives either buttons or a list. Your handler must still process the response message type.

Handling Button and List Responses

Pattern snippet: branch on the response type and pass the selected ID into your own business logic.

client.onMessage(async (message) => {
  if (message.type === 'buttons_response') {
    const selectedId = message.selectedButtonId;
    console.log('ai-agent: button selected', { selectedId });
  }

  if (message.type === 'list_response') {
    const selectedId = message.listResponse?.selectedRowId;
    console.log('ai-agent: list row selected', { selectedId });
  }
});

Expected behavior: button and list replies are logged with the selected ID. Replace the logs with your own handler once you know the menu shape.

Concurrent Message Handling

When multiple messages arrive simultaneously, handle them without overwhelming your LLM API or WhatsApp.

Install the queue in your application:

npm install p-queue

Runnable snippet, assuming client and callLLM are defined as shown above:

import PQueue from 'p-queue';

const queue = new PQueue({ concurrency: 3 });

client.onMessage(async (message) => {
  if (message.type !== 'chat' || message.fromMe) return;

  void queue.add(async () => {
    console.log('ai-agent: processing queued message', { chatId: message.from });
    const response = await callLLM(message.body);
    await client.sendText(message.from, response);
  }).catch((error) => {
    console.error('ai-agent: queued message failed', { chatId: message.from, error });
  });
});

Expected behavior: incoming messages enter the queue immediately, but only three combined LLM-and-send jobs run at once. The rejection handler prevents an unobserved queued failure; this compact example does not provide per-chat ordering or a separate outbound schedule.

Queue Configuration

  • concurrency: Maximum parallel queued jobs in this compact example; use a separate LLM queue when send pacing must not occupy a worker
  • timeout: Maximum time per message, for example 30 seconds
  • retry: Retry failed calls with exponential backoff in your own queue wrapper

Complete Runnable Skeleton

Install the runtime, driver, and queue package in your application before you use this skeleton:

npm install @open-wa/wa-automate@5.1.0 @open-wa/driver-puppeteer@5.1.0 p-queue

Save the following JavaScript as ai-agent.mjs. The local callLLM stub returns a placeholder response, so the example does not call an AI provider.

import { Client, createClient } from '@open-wa/wa-automate';
import { PuppeteerDriver } from '@open-wa/driver-puppeteer';
import PQueue from 'p-queue';

async function callLLM(input) {
  void input;
  return 'Thanks, I received your message.';
}

// At most two model calls run at once across different chats.
const llmQueue = new PQueue({ concurrency: 2 });

// One process-local queue starts at most one outbound task every 15 seconds.
// This controls local scheduling, not provider acceptance or account status.
const sendQueue = new PQueue({ concurrency: 1, intervalCap: 1, interval: 15000, strict: true });

const contextWindow = new Map();
const chatTails = new Map();

function enqueueForChat(chatId, task) {
  const previous = chatTails.get(chatId) ?? Promise.resolve();
  const next = previous.catch(() => undefined).then(task);
  chatTails.set(chatId, next);

  const cleanup = () => {
    if (chatTails.get(chatId) === next) chatTails.delete(chatId);
  };
  void next.then(cleanup, cleanup);

  return next;
}

const runtime = await createClient({
  sessionId: 'ai-bot',
  driver: new PuppeteerDriver(),
  headless: true,
});
const client = new Client({ client: runtime, transport: runtime.getTransport() });
await client.start();

client.onMessage(async (message) => {
  if (message.type !== 'chat' || message.fromMe) return;

  const isGroup = message.from.includes('@g.us');
  if (isGroup && !message.body.includes('@bot')) return;

  const job = enqueueForChat(message.from, async () => {
    const history = [
      ...(contextWindow.get(message.from) ?? []),
      { role: 'user', content: message.body },
    ];

    if (history.length > 20) {
      history.splice(0, history.length - 20);
    }

    const response = await llmQueue.add(() => callLLM([...history]));

    const sendResult = await sendQueue.add(() => {
      console.log('ai-agent: sending response', { chatId: message.from });
      return client.sendText(message.from, response);
    });
    if (sendResult === false) {
      throw new Error('WhatsApp did not accept the send request');
    }

    history.push({ role: 'assistant', content: response });
    contextWindow.set(message.from, history);
  });

  void job.catch((error) => {
    console.error('ai-agent: message job failed', { chatId: message.from, error });
  });
});

Run it with node ai-agent.mjs and complete the normal client authentication when the browser opens.

Expected behavior: The bot receives text messages and ignores group messages that do not mention @bot. Two llmQueue workers can call the model at once across different chats. sendQueue runs one send task at a time and starts at most one every 15 seconds in this process. chatTails waits for the previous model-and-send job for the same chat, preserving the order in which that chat's callbacks enter this process and keeping its context updates sequential. A rejected model or send reaches job.catch, where the application logs it; the next queued job for that chat still runs. A local interval is not a delivery or account-safety guarantee.

Was this helpful?

Your answer includes the page path and docs version.

On this page