Back

17/09/2026

AI Tokens

A token is the unit an AI model reads and bills in, roughly four characters of English. Tokens cover both the question and every document passage retrieved to answer it, which is why cost tracks source material rather than message count.

Tokens are not always complete words. Depending on the model and its tokenizer, one token may be a character, part of a word, a whole word, or punctuation, while the same text can produce a different count across models and languages.

What Counts Toward Token Use in Customer Support

A customer may type one short question, yet the request sent to the model can contain far more text than the customer sees. The full input can include system instructions, earlier messages, retrieved help-centre passages, tool descriptions, and formatting added by the support platform.

The model then creates output tokens for its reply, so providers usually separate usage into two groups.

  • Input tokens cover the text and other supported content supplied to the model, including the customer’s question and any retrieved context.
  • Output tokens cover the content generated by the model, with some reasoning models also counting internal reasoning that does not appear in the visible answer.

A short message is therefore not the same as a small request. If an agent retrieves three long articles to answer “Can I change my plan?”, those passages may account for most of the token use.

Illustration of What Counts Toward Token Use in Customer Support

Token-based cost tracks retrieved context and generated output, not the number of customer messages, while flat-plan cost follows the agreed monthly charge.

How Tokens Set the Model’s Context Limit

A context window is the maximum token budget a model can work with during one request, although providers may define separate limits for input, output, or both together. When the request exceeds that limit, the application must remove, shorten, summarize, or split some of the content.

A larger context window allows the model to consider more material, but loading every available document can raise cost and crowd the request with unrelated details. Strong retrieval selects the passages most likely to answer the question, which can use the token budget more carefully than sending an entire knowledge base.

How Token Billing Changes the Support Bill

Model providers commonly price input and output tokens at separate rates, which makes the cost of each request depend on its actual token use. A basic estimate uses the following calculation.

Request cost = (input tokens ÷ 1,000,000 × input rate) + (output tokens ÷ 1,000,000 × output rate)

The final amount can change with the model, response length, conversation history, retrieved content, language, cached input, and tool calls. Two support teams handling the same number of messages can therefore pay different amounts when their requests contain different amounts of context.

Token billing is only one pricing model. Per-resolution pricing charges for conversations counted as resolved under the vendor’s method, while flat pricing sets a recurring charge with stated plan allowances. Buyers should compare what triggers a charge, which usage limits apply, and whether failed or unresolved requests still consume billable units.

Wonderchat prices on flat monthly plans rather than per token or per resolution. Flat pricing trades a possible saving for a more predictable bill, because a team with light usage may pay less under metered pricing while a team with heavier or longer requests may value a set monthly cost.

Frequently Asked Questions

What is a token in AI?

A token is a piece of text that an AI model reads or generates. It may be a character, part of a word, a whole word, or punctuation.

How many tokens is a page of text?

There is no fixed number. Using the rough estimate of three-quarters of an English word per token, a 500-word page contains about 667 tokens before added instructions, formatting, or retrieved context.

Why do AI support tools charge per token?

Token counts reflect how much content a model processes and generates, so some providers use them to calculate usage. Long context, retrieved documents, conversation history, and detailed replies can increase the cost even when the customer sends a short message.

Related Terms

  • Retrieval
  • Context Window
  • Tokenization
  • Resolution-Based Pricing