Open WebUI Context Compaction Not Triggering: What to Check

If Open WebUI context compaction is not triggering, first check whether it is enabled, whether the effective token threshold fits your backend's actual context window, and whether the conversation contains enough separate turns to compact. The default threshold is 80,000 tokens, which is unsuitable for a backend serving a much smaller window.

Also distinguish a missing trigger from a failed summarization request. Changing the threshold cannot repair an authentication error in the model generating the summary. Conversely, replacing that model will not help if compaction is disabled.

Open WebUI Context Compaction Not Triggering

Work through the checks below in order. Keep your original conversation, record existing settings, and use a disposable chat for experiments.

1. Confirm that your installed version includes the feature

Open WebUI introduced built-in automatic context compaction in v0.10.0. Older installations may instead use community filters or custom pipelines, whose controls and behavior differ.

This guide describes the built-in feature using v0.11.4-era behavior and current documentation. Check your running version before looking for every setting shown here, particularly if you use a fork or an older pinned container image.

If the feature is absent, review the release notes and upgrade procedure for your deployment. Back up the persistent data first. After an upgrade, refresh the browser with Ctrl+F5 on Windows or Cmd+Shift+R on macOS to clear stale frontend assets.

For a shared deployment, confirm that all application instances run the intended version. A database migration can make mixed-version deployments unsafe; follow the documented migration procedure rather than updating replicas independently without checking compatibility.

2. Check the saved administrator settings

Sign in as an administrator and open Settings > Admin > Interface. Locate the Context Compaction controls and verify these values:

SettingWhat to check
Context CompactionEnabled. Its default is off.
Token ThresholdA trigger below the backend's usable window; the default is 80,000.
Token CapThe ceiling on threshold overrides, not a limit on the final prompt.
Retained MessagesThe default keeps 40% of recent messages. Allowed values are clamped to 10–50%.
Context Compaction ModelThe intended model for producing summaries.

Save your changes, reopen the settings, and confirm they remain in place.

These compaction settings use persistent configuration. On an existing installation, changing ENABLE_CONTEXT_COMPACTION or another environment variable and restarting may leave an already saved database value in effect. Editing the setting through the administrator interface is the simplest first check.

Avoid deleting the data volume to reset one setting. Disabling persistent configuration globally also changes how other settings are saved, so it is a deployment decision rather than a routine compaction fix.

3. Compare the effective threshold with the real context window

A model's advertised maximum and the context allocated by your serving backend are different things. For Ollama, run this on the machine serving the model while it is loaded:

ollama ps

Check the CONTEXT column. Increasing context allocation consumes additional memory, so raising it blindly can create a separate resource problem. If you are balancing model size against a small machine's memory, the local LLM setup guide provides broader background.

Illustrative example: suppose the backend is actually serving an 8,192-token window. An 80,000-token compaction threshold is already above that entire window. Fix the mismatch before investigating more subtle behavior.

Leave room for the next user message, generated output, and other request content. There is no single threshold that is safe for every model and workload.

Check overrides in the affected chat

Open WebUI supports advanced parameters at the model, account, and chat levels. A chat-specific value can override an account or model default. Inspect the affected conversation's Chat Controls as well as the model configuration.

The compaction override is compact_token_threshold. In the current implementation, its positive value replaces the global threshold but remains bounded by CONTEXT_COMPACTION_TOKEN_CAP. If the cap is unset, it follows the global threshold.

For example, a 30,000-token model override cannot take effect as 30,000 when the cap is 20,000. Equally, a saved chat override can explain why changing a model default appears ineffective. Record the values at each level before changing them.

4. Make sure the conversation can be split

Exceeding the threshold is not sufficient by itself. In v0.11.4, automatic compaction skips an active conversation list containing three or fewer messages. It also needs a suitable user-message boundary that leaves both an older section to summarize and a recent section to retain.

A huge first message, one oversized attachment, or a long sequence of tool activity within a single exchange can therefore be a poor trigger test. Thousands of tokens do not necessarily provide several usable conversational turns.

Use several ordinary user-and-assistant exchanges in a separate test chat. Keep individual messages modest enough for the backend to accept them. If that conversation compacts but your original one does not, inspect the original conversation's structure and unusually large recent messages.

Retained Messages is a percentage of message count, not a promise to retain that percentage of tokens. One recent message may contain more text than many older ones combined.

5. Check for evidence of compaction

Successful compaction does not remove old messages from the visible conversation. Open WebUI preserves the stored history while changing the context used for subsequent model requests.

Watch for the interface notifications Compacting context, followed by Context compacted or Context compaction failed. A missing notification alone is weaker evidence than a server-side check.

Administrators can inspect the backend logs around the timestamp of a test request. For a Docker container named open-webui:

docker logs --since 10m open-webui

Substitute your actual container name. Check for compaction events and nearby model or provider errors. Open WebUI logs backend activity to standard output at the INFO level by default. If more detail is necessary, temporarily use GLOBAL_LOG_LEVEL=DEBUG and restore the previous level afterward; debug logs can contain full request payloads.

Compare the same test before and after one configuration change. A normal answer, a shorter answer, or the assistant claiming that it summarized the chat does not independently establish that the compaction mechanism ran.

6. Verify the model that generates the summary

Compaction requires a separate generation request. The main chat can work while this background request fails.

In v0.11.4 and the current Task Models documentation, compaction has its own Context Compaction Model selector. Current Model uses the chat model; the Local and External Task Model selectors do not choose the compaction model. Some reference text describes older fallback behavior, so match advice to your installed release.

For diagnosis, select a known-working model explicitly or use Current Model. Confirm that the relevant user can access it and that its connection accepts requests. A successful short chat is only a connectivity check: summarizing a long conversation requires a much larger input.

Look for the actual failure before changing settings: authentication, rate limits, input length, unavailable models, and rejected generation parameters need different fixes.

If compaction runs but the summary is incomplete

Inspect TASK_MODEL_PARAMS. These parameters apply to background tasks, including compaction summaries, and can affect other tasks such as title generation too.

With no explicit task parameters, the current behavior uses the summary model's configured max_tokens when available, otherwise 1,000 output tokens. Reasoning can consume that allowance before a useful summary is finished. An explicit output budget may help, but verify that the selected provider supports the parameter and avoid changing unrelated task settings.

That is a summary-quality investigation. Do not confuse it with a trigger that never ran.

7. Check what still occupies the request after compaction

Compaction is not a guaranteed ceiling on the request sent to a model. Older conversation content can shrink while fresh instructions, retrieved material, attachments, or tool definitions still consume substantial space.

Compare a plain text test chat with the failing setup. Reintroduce attached knowledge, tools, and custom filters individually. Treat the result as a diagnostic comparison: if adding one component recreates the failure, inspect that component's input size and configuration.

For document-heavy chats, check the retrieval mode. Focused Retrieval uses relevant passages, whereas Full Context can inject the complete file. For chat-attached files, Full Context content is injected on each message. Model-attached knowledge also depends on the function-calling mode.

Use retrieval when the question needs selected passages from a large collection. Keep full-document context for cases that actually require the entire text. This changes which material the model receives, so verify answers against the original document after switching.

If the assistant forgets a rule after a successful compaction, investigate where the rule was supplied. The guide to preserving project instructions covers the separate problem of instruction discovery and task handoffs.

For background on the broader approach, this coverage of longer Claude conversations explains how another assistant uses summarization to continue lengthy chats. Its product behavior should not be treated as Open WebUI configuration guidance.

Run a controlled test before changing your main setup

Use this procedure to isolate the failure while keeping the test easy to repeat:

  1. Record the starting configuration. Note the app version, backend, allocated context, threshold, cap, and summary model.
  2. Create a disposable chat. Use one known-working model with no optional attachments or tools.
  3. Temporarily lower the threshold for that test. Keep it comfortably below the allocated window, leaving space for output and the summary request. Prefer an available per-chat override to a global change.
  4. Build several separate exchanges. Cross the test threshold through normal turns, not one enormous pasted message.
  5. Check notifications and server logs. Record whether compaction was skipped, succeeded, or attempted and failed.
  6. Restore the settings. Then reintroduce the original workload a component at a time.

If the minimal test succeeds, concentrate on the original chat's overrides, structure, and additional content. If it fails, keep the investigation focused on configuration, summary-model requests, and the installed version.

Frequently asked questions

Does telling the assistant to summarize trigger built-in compaction?

A conversational summary is not proof that Open WebUI created a compaction checkpoint. Verify the application event or backend behavior instead.

Why does a fresh chat work while the old one fails?

That comparison points toward accumulated context or chat-specific configuration. It does not identify which component is responsible. Compare overrides, large recent inputs, and attached resources before replacing the model.

Should I increase the threshold to fix a context-length error?

Usually that moves the trigger in the wrong direction. First establish the backend's real limit and why the existing request exceeds it. A higher threshold delays compaction.

What should I include in a bug report?

Include the Open WebUI version, deployment method, backend and model, effective settings, a small reproduction, and relevant redacted log errors. Distinguish a skipped trigger from a failed summary request, and state whether the plain text test worked.

Your next check

Start with the saved enable switch and the effective threshold, then run one small, observable test. That gives you a concrete result to investigate without sacrificing the original conversation or resetting the installation.