HowDoesAIDoThat?

Glossary / plain-language definition

What is AI inference?

Inference is running a trained AI model on an input to produce an output. It is the computation that happens when you ask a model a question or generate an image.

By HowDoesAIDoThatSources checked 11 October 2026
Illustrative example

You submit “Explain this invoice in three bullet points”. A language model processes the request and generates its reply.

Using a model after it has been trained

Inference is the use stage of an AI model. A model has learned numerical parameters, usually called weights, during training. Inference applies those weights to a new input. The output might be a predicted category, a transcription, an image or a written answer. It does not have to be a chatbot. IBM explains training and inference.

  1. 01 / inputSupply a prompt

    Text is represented as tokens, ready for the language model.

  2. 02 / processingUse trained weights

    The model processes the context and predicts subsequent tokens.

  3. 03 / outputGenerate an answer

    Chosen tokens form the response. Useful does not always mean correct.

Illustrative text-generation pipeline. Other types of inference produce classifications, images, audio or other outputs.

For a practical example, imagine asking a local language model: “Explain this invoice in three bullet points”, followed by the invoice text. Below is a simplified account of a typical autoregressive language model. The example is illustrative; it is not a recorded run, and other model architectures can work differently.

1. Your prompt becomes the input

The application prepares your request. It may add conversation history, role markers, instructions or retrieved documents. This matters because the visible text you type can be only part of what the model receives. A chat application needs the format expected by its model; Hugging Face documents chat templates.

If an app retrieves invoice records first, that retrieval is a separate step around the model. If it calls a calculator, the calculator does the arithmetic. “An AI answered” does not tell you whether the answer came entirely from model generation or also used other software.

2. The text is tokenised

A tokenizer converts text into units the model can process, then represents those units with numerical IDs. Tokens can correspond to pieces of words, punctuation or other units. The exact split depends on the tokenizer. This is why a word count is not an exact token count. Hugging Face's tokenization guide describes the stages involved.

You do not normally need to handle these IDs yourself. What matters is that long invoices, previous messages and the requested answer all create processing work. A short question attached to a large document is still a large input.

3. Prefill processes the supplied context

In the prefill stage, the model processes input tokens and builds intermediate state used for generation. A typical Transformer stores attention information in a key/value, or KV, cache. The model can reuse that information instead of recalculating all previous work for each new token. NVIDIA explains prefill and decode.

This helps explain a pause before the first visible text. The application may also be loading the model, waiting for a server or doing retrieval. A pause alone does not identify the cause. Distinguish time until the first token from the time to finish the whole answer.

4. Decode generates the answer

The model produces scores for possible next tokens. Generation settings determine how a token is selected. That token becomes part of the growing sequence, and generation continues until a stop condition or output limit is reached. Hugging Face describes generation strategies.

The tokenizer converts the resulting IDs back into readable text. The application displays the answer, sometimes as a stream. A fluent explanation of the invoice still needs checking against the original amounts. Token generation is not a guarantee that a claim is correct, a calculation was executed, or a source was consulted.

Training changes weights; ordinary inference uses them

Training and fine-tuning deliberately update model parameters. Ordinary inference does not update the model's weights with every request. Supplying another invoice can change the current answer through its context without permanently teaching the base model. Hugging Face's fine-tuning guide covers the separate training process.

An application may save chat history, maintain a memory database or use interactions in a later training process. Those are application or provider behaviours; they should not be confused with the forward computation that generates a reply. Check the particular application's documentation before assuming a conversation is forgotten or used for training.

Local and cloud inference both need resources

With local inference, computation runs on your hardware. With cloud inference, a remote service supplies the hardware. A browser interface can be used for either, so a webpage does not prove the model is running on your PC. IBM Research explains inference in deployed systems.

Local use involves memory, storage, electricity and setup time. Cloud use can involve subscriptions or usage charges, depending on the service. Longer inputs, longer outputs and larger models can change the work required. Choose a route using the actual task and its documented requirements; check the particular model and runtime rather than assuming a universal PC requirement.

Sources

Primary references checked 11 October 2026. This explanation is not a benchmark of a particular model or computer.