Skip to content

Runtime & UI

These settings control the technical limits and the default behavior of the agents when interacting with tools and the long-term memory.

Determines the maximum time an agent may spend in a single “loop” calling tools.

  • Range: 60 to 3600 seconds.
  • Default: 600 seconds (10 minutes).
  • Purpose: Prevents agents from getting into infinite tool call loops or consuming excessive resources when they cannot find a solution.

Defines how many relevant fragments from the vector memory are passed to the LLM per request.

  • Range: 1 to 50 entries.
  • Default: 5 entries.
  • Note: Higher values provide more context but consume more tokens and can confuse the model (“Lost in the Middle”).

Controls the agents’ default write access to the memory.

  • Allow write access: If active, agents can automatically store important information from the conversation in the long-term memory.
  • Effect: Applies as the default for all new agents/tasks, but can be overridden by specific policies (see Memory documentation).

A global rate limiting for outgoing API calls to AI providers (OpenAI, Anthropic etc.).

  • Range: 1 to 500 requests.
  • Default: 10 requests per minute.
  • Purpose: Protection against unexpected costs and avoidance of “429 Too Many Requests” errors at the providers.

Determines the local time for the entire Ontheia host.

  • Format: IANA timezone string (e.g., Europe/Berlin, UTC).
  • Default: Europe/Berlin (or value from APP_TIMEZONE).
  • Effect:
    • Chat Titles: Automatically generated titles use this timezone for dates.
    • Logs (Trace): Events are converted to this local time for display.
    • Cron Jobs: Schedules are executed based on this timezone.
    • Agent Context: The “Current Time” injected into the agent follows this setting.

A global switch that enables or disables token-by-token streaming of LLM responses into the chat.

  • Default: enabled.
  • Effect: When enabled, agent responses appear in the chat while the model generates them. When disabled, the full response appears as one block after generation completes.

Scope: The switch applies to the Anthropic API path and all OpenAI-compatible providers (OpenAI, xAI, Google, Ollama, etc.). CLI providers (e.g. Claude CLI) always deliver the response as a block by nature. Individual OpenAI-compatible providers whose endpoint does not support SSE can be excluded via provider or model metadata with "stream": false — they then keep responding as a block regardless of the global switch.

When to disable? Only when something misbehaves — for example if a provider rejects streaming requests, or a reverse proxy in front of Ontheia buffers SSE responses so streaming never reaches the browser anyway.

A global switch that enables or disables prompt caching on the Anthropic API path.

  • Default: enabled.
  • Effect: When enabled, Ontheia places cache_control markers on the stable prefix (tools + system prompt) and the growing chat history. Recurring requests then read that prefix at the heavily reduced cache price (~0.1× input).

Why Anthropic only? Anthropic is the provider where caching can cost more than it saves without any way for us to opt out per model: writing the cache (cache_creation) is billed at ~1.25× the normal input price. If the cached prefix is not re-read within the 5-minute TTL — e.g. for sporadic single-shot runs (cron tasks with a single LLM call, long thinking pauses in chat) — you pay the write premium without ever collecting the cheap read: +25% instead of savings. With Anthropic, we decide whether cache_control is sent — which is why an on/off switch works here technically.

OpenAI (except gpt-5.6), xAI and Mistral cache automatically and without a write premium (there is only “Input” and “Cached input”, no write column). Caching can never be more expensive there than no caching — this switch does not affect them.

Exception since gpt-5.6 (July 2026): the gpt-5.6 models (sol/terra/luna) now also bill a 1.25× cache-write premium — the exact same mechanic as Anthropic (pricing table example: gpt-5.6-sol $5.00 input → $6.25 cache write). Unlike Anthropic, though, this caching is implicit on OpenAI’s side (mode auto, the documented default) — Ontheia sends no parameter of its own that this switch could toggle. This switch therefore does not cover gpt-5.6. OpenAI instead offers an explicit mode where you mark exactly which prefixes to cache — that isn’t a simple on/off setting but a structural change to how the request is built, and it isn’t implemented yet (see roadmap: Responses API expansion).

When to disable? If your Anthropic usage consists mostly of sporadic single-shot runs and you see no cache savings (the ⚡ symbol with cache-read tokens) in the response token usage. For dense tool loops and fast chat sequences the switch should stay enabled — the cheap read clearly dominates there.

Observe gpt-5.6 instead of switching it off: the ⚡ badge in the trace panel now also surfaces OpenAI’s cache-write tokens (as cache creation, mirroring Anthropic). Use it to judge whether gpt-5.6 is actually costing you extra in your usage pattern (sporadic runs) or whether repeated reads make it a net saving (dense tool loops, fast chat sequences).