Quick overview
For agencies and freelancers who run one AI agent for several paying clients. Each client has a monthly budget, a pause switch and its own instructions in an n8n Data table. A run starts only if the client can afford it, and the tokens used are booked to that client.
How it works
- Your app POSTs client_id and prompt to the webhook. Unknown or paused clients get 403, and clients without budget left get 429.
- A worst-case reservation is written as the run's own ledger row. The ledger is read again, and the run backs off if parallel runs used the budget first.
- The AI Agent runs with Max Iterations and a token limit from Settings.
- The prompt and completion tokens that n8n recorded for each model call are read through the n8n API and booked. The answer is returned with a usage summary.
- You get an email at your alert threshold and when a client can't afford another run.
- Hourly, stuck reservations expire. Monthly, you get totals per client and old rows are deleted. A GET endpoint returns usage per client.
- If a Data table can't be read or written, the run is refused with 503.
Setup
- Create the two Data tables listed in the sticky note: ai_clients and ai_usage_ledger.
- Add credentials: OpenAI on Chat model, n8n API on Read token usage, Header Auth on both webhooks and SMTP on the email nodes.
- Fill in the three Settings nodes and keep "Save execution progress" on.
- Add a client row to ai_clients.
- Publish the workflow and run the curl example in the "Try it" note.
Requirements
- n8n with Data tables (tested on 2.41.6)
- An OpenAI API key, or another provider with an OpenAI-compatible Chat Completions endpoint
- An n8n API key (the n8n API isn't available during the free trial)
- SMTP credentials
Customization
- Set the two cost rates to count in your currency instead of tokens.
- Swap the Calculator for your own tools, or the email nodes for Slack or Telegram.
Additional info
The reservation is an estimate, and Data tables have no atomic increment. The "Limits" note explains what that means for parallel runs and high volume. Without the n8n API credential, each run is charged its full reservation. We're building Keelstamp, which isn't released yet. This template doesn't depend on it.
Prepared with AI assistance. Editorial responsibility: PowerQuant ApS.