Choose Context Compaction Thresholds for Local AI Models
To choose a context compaction threshold for a local AI model, start with the context window your runtime actually allocates. Reserve space for the next response, incoming messages or tool results, and counting uncertainty. Set the trigger below what remains, then check whether compaction preserves the information your task needs.
There is no universally safe 80% setting. A short writing conversation and a coding agent that reads large files can need different thresholds even when they use the same model. My recommendation is to calculate a starting budget, test it in a disposable conversation, and adjust one setting at a time.

The method below applies to local chat applications and agents that summarize older context. Ollama and Open WebUI provide a concrete example, but the calculation depends on your application's counting and compaction behavior.
Separate the context window from the compaction threshold
The context window is the working space available to a model. The compaction threshold tells an application when to replace older conversation content with a summary. Changing the threshold does not enlarge the model's window.
Keep these controls separate when configuring your setup:
- Allocated context: the window available to the running model.
- Compaction trigger: the token count or percentage that starts summarization.
- Response budget: room for the model's generated output.
- Retention: how much recent conversation remains verbatim.
- Summary budget: how much output the summarizer can produce.
Also check what your application means by compaction. Summarizing older messages, dropping them, and shifting a runtime's context are different operations. For example, llama.cpp exposes context-shifting controls separately from context size. A runtime's ability to continue generating does not establish that it created a useful summary of discarded material.
1. Find the context window your setup actually uses
A model advertised with a large maximum window may be loaded with a smaller allocation. Base your calculation on the effective limit for the current request, within the model's supported range. Account for a smaller application or serving limit when one applies.
Check Ollama while the model is loaded
Send a short request through your usual application, then run this on the machine serving Ollama:
ollama psRead the CONTEXT column and check PROCESSOR for the CPU/GPU allocation. Record the context value before touching the compaction threshold.
Ollama supports the server default OLLAMA_CONTEXT_LENGTH and the request option num_ctx. In Open WebUI, an explicit num_ctx under the model's parameters or Chat Controls overrides that server default. Check the affected chat if its behavior differs from a fresh conversation.
Increasing the allocation requires more memory. I would first choose a window the machine can serve at an acceptable speed, then tune compaction around it. For the hardware implications, this explanation of Ollama memory requirements provides useful background on GPU memory and offloading.
Check the serving configuration in other runtimes
For llama.cpp, inspect --ctx-size and the server's reported context limits. Its configuration also includes parallel slots and shared cache controls, so do not assume a headline allocation describes every request's usable budget. Record the limit applicable to your session.
If you switch models, change concurrency, or reload with a different allocation, recalculate. A threshold that worked with one loaded configuration is not automatically appropriate for another.
2. Calculate a threshold with enough headroom
For a setup whose threshold counter covers the complete input context, use this planning rule:
Starting threshold <= C - R - G - M- C: the effective context window.
- R: the output space you want to reserve.
- G: additional input that can arrive before compaction protects the next request.
- M: a margin for counting uncertainty and variable overhead.
This is a budgeting method, not a documented universal algorithm or a guarantee against overflow. The timing of the threshold check matters. An application that checks before adding a tool result needs enough room for that result; one that checks afterward may already be trying to recover from an oversized input.
Check what the counter includes. If it measures only conversation history, also subtract the system instructions, tool definitions, retrieved passages, or other material added outside that measurement. Do not subtract those components twice when the reported count already includes them.
A worked example for a 16K window
Suppose the effective window is 16,384 tokens and your counter includes the complete input. You reserve 2,048 tokens for output, 1,024 for incoming content before the next effective check, and 1,024 for uncertainty:
16,384 - 2,048 - 1,024 - 1,024 = 12,288 tokensThat gives a starting threshold of 12,288 tokens, or 75% of the window. The percentage is the result of the assumptions. If your tools can add another 4,000 tokens in one step, those assumptions no longer fit.
The following examples show how different reserves change the starting point. These are illustrative calculations, not measured performance results or preset recommendations.
| Window C | Output R | Growth + margin | Threshold |
|---|---|---|---|
| 8,192 | 1,024 | 1,536 | 5,632 |
| 16,384 | 2,048 | 2,048 | 12,288 |
| 32,768 | 4,096 | 8,192 | 20,480 |
All values are tokens. Each row assumes that fixed input content is already included in the threshold counter. If your interface accepts percentages, round down when translating a calculated token ceiling into its supported setting.
Choose reserves from the work you actually do
For short text conversations, estimate growth from representative user messages and replies. For coding, include file reads, search results, error logs, and tool responses. For document questions, account for the passages injected into each request.
I would inspect the largest plausible turn rather than the average turn. An average of 300 tokens offers little protection when one routine command can return thousands. Limit unnecessarily broad tool output or retrieve a relevant document section before increasing the window.
Reserve generation space for the behavior you enable, including reasoning when it consumes the backend's generation or context budget. A short visible answer does not necessarily mean the model generated only a short output.
3. Make sure the summary request also fits
Compaction requires the summarizer to read input and generate a replacement. A threshold can leave room for an ordinary reply while still being unsuitable for the summarization request.
Check the summarizer's own loaded window, prompt overhead, output allowance, and the amount of history it receives. Some implementations include recent messages as background even when they summarize only the older portion. If you select a separate model, its context allocation may be smaller than the chat model's.
My first choice would be a summarizer that can accept the expected input comfortably and preserve the task's details. A smaller model is worth considering only after it passes that check. For a fully local workflow, explicitly verify the selected summary model and endpoint as well as the main chat model.
A summary that ends halfway through a requirement needs a different fix from a summary request rejected for excessive input. Inspect the failure before changing the trigger: the first may need a better output allowance, while the second may require earlier compaction or a smaller input.
4. Apply the threshold in Open WebUI
For installations with built-in Context Compaction, open Settings > Admin > Interface. Record existing settings, enable the feature, and enter the calculated Token Threshold. The documented default is 80,000 tokens, which is already larger than an 8K, 16K, or 32K window.
The global setting is CONTEXT_COMPACTION_TOKEN_THRESHOLD. The advanced parameter compact_token_threshold provides an override, bounded by CONTEXT_COMPACTION_TOKEN_CAP. When unset, the cap follows the global threshold. A cap limits threshold overrides; it does not enforce a maximum final prompt size.
For multiple local models, use appropriate model-level values instead of treating one threshold as suitable for every allocation. Save and reopen the controls to verify the effective settings. Match the options to your installed release.
If a reasonable threshold still produces no compaction, follow the checks for Open WebUI compaction. A disabled feature, an unsuitable turn boundary, and a failed summary request require different fixes.
5. Tune retention by the space it actually leaves
Open WebUI's Retained Messages setting defaults to 40% and is constrained to 10–50%. It describes message count, not token count. One large recent message can occupy more space than many older exchanges.
Evaluate the complete input after compaction:
Fixed instructions + summary + retained messages + injected contentThat total should leave room for new work and output. It should also sit comfortably below the next trigger. Otherwise, the application may spend much of the session repeatedly summarizing nearly the same material.
For example, suppose a 12,288-token trigger leaves an input of about 11,500 tokens after compaction. A modest new turn could cross the threshold again. I would inspect the largest retained messages and the summary length before raising the trigger and consuming the remaining headroom.
Reducing retention is a trade-off: it gives more room but moves more information into a lossy summary. Keep recent exchanges intact when exact phrasing, code, or immediate troubleshooting details matter. Remove irrelevant bulk at its source when that is the real space problem.
6. Validate both capacity and remembered details
A successful answer after compaction is only part of the check. The model must also continue with the correct requirements. I would use a disposable conversation that resembles the real workload and contains several separate exchanges.
- Record the configuration. Note the application version, model, allocated window, trigger, retention, summary model, and output settings.
- Plant verifiable requirements. Include an exact filename, a decision that replaced an earlier choice, an unresolved question, and a constraint on the next action.
- Build representative history. Use realistic message sizes and tool results. Cross the threshold through normal work rather than one enormous paste.
- Verify that compaction occurred. Use the application's event indicators or server logs. The assistant saying it summarized the conversation is insufficient evidence.
- Check the continuation. Ask for a next-step plan that must apply the planted requirements, then compare it with your notes.
- Continue through another cycle. Check whether earlier constraints and later corrections survive repeated summarization.
Where available, compare input and output counts around the event. Ollama's chat response exposes prompt_eval_count for prompt tokens and eval_count for generated tokens, alongside timing fields. Compare equivalent requests and distinguish model loading time from prompt processing and generation.
Record compaction frequency, waiting time, request failures, and missing requirements. A lower threshold may reduce input pressure while creating more summary interruptions. A higher threshold may reduce interruptions but leave too little space for a large turn. Choose the setting that meets your workload's needs.
What to change when the first setting fails
| Observed problem | First adjustment to investigate |
|---|---|
| Context errors before compaction | Confirm the effective window, then lower the trigger or reduce incoming content. |
| Compaction on nearly every turn | Measure the remaining input; inspect bulky retained messages and summary size. |
| Important details disappear | Inspect summary quality and retention; keep exact requirements in an authoritative source. |
| Summary requests fail | Check the summary model's allocation, connection, and generation settings. |
| A larger window becomes too slow | Check memory allocation and offloading, then reassess the workload's necessary context. |
Change one variable and repeat the same comparison. If it makes the outcome worse, restore the recorded value before trying another explanation.
Keep durable requirements outside an evolving summary when possible. For coding workflows, the guide to preserving project instructions explains how to separate permanent rules from temporary task state. Confirm that your application actually loads the relevant instructions.
Frequently asked questions
Should every local model compact at 80%?
No. Eighty percent leaves only 1,638 tokens in an 8,192-token window. That may be insufficient for the next input and response. Calculate the reserve first, then derive a percentage if your interface requires one.
Does compaction guarantee that the next prompt will fit?
No. Large new content can still exceed the window. Open WebUI can also skip compaction when no suitable split exists, or continue with uncompacted history after an error. Treat the trigger as a management mechanism rather than a hard request-size limit.
Does lowering the threshold reduce allocated VRAM?
Do not assume it will. A shorter active prompt and a smaller runtime allocation are different things. Check actual memory usage; if the allocation itself is too large, review the runtime's context setting separately.
When is a fresh chat preferable?
Start a fresh chat when the objective changes substantially or old decisions are no longer useful. Carry over a checked handoff with the current goal, constraints, source locations, and next action. Preserve the original conversation for reference.
Choose your first threshold
Record the loaded window and estimate the largest normal input burst. Subtract that allowance, the response budget, and a counting margin. Use the resulting value as a starting trigger, then verify two compaction cycles with representative work. Keep the setting only if both the requests fit and the important details survive.