nexa
ExploreServicesAboutTeamFAQBlogContact
GitHubStart a project
Blog · Engineering

Adding AI to an app without a surprise bill

Most of the cost of an AI feature is decided before the first request. Model routing, caching, context size and hard limits, in order of impact.

Ethan

Tech Lead · Mobile & Web

October 7, 20267 min read

On this page

  1. Where the money goes
  2. The request pipeline
  3. Five techniques, in order of impact
  4. Limits that actually stop spending
  5. Degrade, do not break
  6. Measure the right number
  7. The checklist

Adding an AI feature to an app takes an afternoon. Keeping it fast, correct and affordable once real users arrive takes a design. Most of the cost is decided before the first request is sent: what you send, which model answers, and what you do when the answer is already known. This is the checklist we use when a client asks for "a chatbot" or "AI search", written so you can apply it to your own product.

TL;DR
  • You pay per token, in and out. The prompt you send is usually the bigger half of the bill.
  • Route each request to the smallest model that can handle it. Escalate only when needed.
  • Cache what repeats: system prompts, shared documents, identical questions.
  • Put hard limits in code: per user, per day, per request. Budgets in a spreadsheet do not stop a loop.
  • Measure cost per successful answer, not cost per request.
Stacked bars showing one AI request: a locked system prompt bar, a wide documents bar, a chat history bar, a thin pink question bar, and an answer bar below an arrow
One request, top to bottom: system prompt (cached), retrieved documents, chat history, the user's question in pink, and the answer below the arrow. The question is the thinnest layer.

Where the money goes

Language models are billed by the token, roughly three quarters of an English word. Every request is charged twice: once for everything you send (input) and once for everything the model writes back (output). Output tokens cost more per token, but input is usually where the volume is.

A typical support assistant request looks like this:

Part of the requestTokensChanges per request?
System prompt and rules~1,500No, identical every time
Retrieved help articles~3,000Sometimes
Chat history~1,000Grows every turn
User's question~50Yes
Model's answer (output)~300Yes

The question is about 1% of the input. That is the key insight behind every technique below: shrink, reuse or skip the parts that do not change.

Work out the cost before you build

Multiply expected requests per day by tokens per request, using your provider's current price sheet. If the number makes you uncomfortable on paper, it will be worse in production, because real users ask follow-up questions.

The request pipeline

We put every AI call behind one server-side function. The app never talks to the model provider directly. That single choke point is what makes the rest possible: caching, routing, limits and logging all live in one place.

  1. Check limitsUser quota and daily budget
  2. Check cacheSeen this exact question?
  3. RoutePick the smallest capable model
  4. Call + streamWith a token cap on the answer
  5. LogTokens, cost, latency, outcome
Every request passes the same five steps. Most savings happen before the model is called.
export async function ask(userId: string, question: string) {
  await assertWithinQuota(userId);           // 1. hard limits first

  const cached = await answerCache.get(question);
  if (cached) return cached;                 // 2. free answer

  const model = pickModel(question);         // 3. small by default
  const answer = await llm.complete({
    model,
    system: SYSTEM_PROMPT,                   // cached by the provider
    messages: await buildContext(userId, question),
    maxTokens: 500,                          // 4. cap the output
  });

  await usageLog.record(userId, model, answer.usage); // 5. measure
  await answerCache.set(question, answer.text);
  return answer.text;
}

Five techniques, in order of impact

  1. Route to the smallest model that works

    Providers sell several sizes of model. The small ones are many times cheaper and faster, and they handle classification, extraction, short rewrites and simple FAQ answers well. Send everything to a small model first; escalate to a large one only for long reasoning, code, or when the small model reports low confidence.

    A simple router is often a few rules (question length, detected intent, whether documents were retrieved). You do not need a model to choose the model.

  2. Use prompt caching for the parts that repeat

    Major providers can cache a fixed prefix of your prompt, such as system instructions and shared reference documents, and bill cached reads at a fraction of the normal input price. The rule is simple: put stable content first and variable content last, so the prefix stays identical between requests.

  3. Send less context

    Retrieval (RAG) should return the three most relevant passages, not the whole manual. Summarise old chat turns instead of resending them in full. Strip HTML, boilerplate and duplicate text before it reaches the prompt. Every token removed here is saved on every single request.

  4. Cache whole answers

    Many apps receive the same handful of questions over and over: opening hours, pricing, how to reset a password. Store the answer keyed by a normalised version of the question and return it instantly. It costs nothing, and it is faster than any model.

  5. Cap the output

    Set a maximum answer length and ask for concise answers in the system prompt. Long answers are expensive and, on a phone screen, usually worse anyway.

Many requests flowing into a router connected to a cache; a thick pink path leads to a fast model marked with a lightning bolt, a thin gray path to a larger model
Routing in practice: the router checks the cache first, sends most traffic to the small, fast model, and passes only the hard cases to the large one.

Limits that actually stop spending

A monthly budget alert tells you about a problem after it has happened. Limits enforced in code prevent it. We set three layers:

01

Per request

Maximum input and output tokens. Stops one huge document or a runaway answer.

02

Per user

Requests per minute and per day. Stops abuse, scripts and accidental loops in the client.

03

Per app

A daily spending ceiling. When reached, degrade gracefully instead of failing silently.

Never call the model from the client

An API key shipped inside a mobile app or a web bundle can be extracted in minutes. Keep keys on the server, behind your own authenticated endpoint, so every limit above can be enforced.

Degrade, do not break

When a limit is reached or the provider is slow, the feature should still be useful. Good fallbacks, from best to worst:

  • Serve a cached or pre-written answer for common questions.
  • Fall back to the smaller model with a shorter answer.
  • Show search results from your own help content without generation.
  • Explain clearly that the assistant is busy and offer a contact option.

Do

  • Stream answers so users see progress immediately
  • Log tokens and cost for every request from day one
  • Version your prompts like code

Don't

  • Resend the full chat history on every turn
  • Use the largest model "to be safe"
  • Rely only on the provider's billing dashboard

Measure the right number

Cost per request is easy to track and slightly misleading. A cheap answer that is wrong leads to a follow-up question, a support ticket, or a user who leaves. The metric we report to clients is cost per successful answer: total AI spend divided by the answers users rated helpful or did not need to rephrase.

A dashboard with four KPI tiles, a daily spend line staying under a dashed budget line, and a chart of helpful versus unhelpful answers
The numbers worth watching: volume, cache hit rate, spend against the budget line, and helpful (green) versus unhelpful (pink) answers.

Add a thumbs up and down to every answer. It costs one button, and it turns "the AI feels okay" into data you can improve against: which questions fail, which prompts need work, and which requests never needed the large model.

The checklist

  • All model calls go through one server-side function
  • Small model by default, large model on escalation
  • Stable prompt content first, so provider caching works
  • Retrieval returns a few passages, not whole documents
  • Common questions answered from a cache
  • Output length capped
  • Limits per request, per user and per app, enforced in code
  • A fallback for when limits or the provider fail
  • Tokens, cost and user feedback logged for every requestYou cannot reduce what you do not measure.
The cheapest model call is the one you never make. The second cheapest is the one that sends only what the model needs.

If you are planning an AI feature and want a second opinion on the architecture or the expected cost, talk to us. We are happy to review a design before any code is written.

Work with us

Need something like this built?

Our services

Keep reading

All posts
  • AI can design an app now. Are designers redundant?
    Read post

    Design

    AI can design an app now. Are designers redundant?

    AI makes the first draft of a screen cheap. That shifts design work to what matters most: understanding users and choosing what is right for them.

    Tina· Oct 10, 20266 min read
  • AI in software development: turn change into an advantage
    Read post

    Engineering

    AI in software development: turn change into an advantage

    AI does not replace technical thinking. With the right process, it frees teams to focus on the decisions that make products better.

    Ethan· Oct 9, 20264 min read
  • Why we wrote our own file-system wrapper for React Native
    Read post

    Engineering

    Why we wrote our own file-system wrapper for React Native

    Three of our apps needed to save, read and export files. Each one did it slightly differently, so we pulled the shared bits into a small library.

    Ethan· Sep 18, 20263 min read
nexa
ExploreServicesAboutTeamFAQBlogContact

© 2026 Nexa Tech. All rights reserved.

Web · Mobile · Open source