feat(cli): persist token count in graph state across sessions#2323
Merged
Conversation
Merging this PR will not alter performance
Comparing Footnotes
|
Mason Daugherty (mdrxy)
enabled auto-merge (squash)
March 30, 2026 20:25
james8814
pushed a commit
to james8814/deepagents
that referenced
this pull request
Apr 2, 2026
…ain-ai#2323) The CLI's token count (`/tokens`, status bar) was stored only in memory via `TextualTokenTracker`, so resuming a thread with `-r` always showed "No token usage yet" — even though `/offload` independently counted ~291 tokens from messages. This persists the count in the LangGraph graph state so it survives across sessions and both commands agree. - Store the token count in the graph's checkpointed state so it survives across sessions, using a private field the model never sees - Replace the dedicated `TextualTokenTracker` class with a simple integer on the app plus three callbacks for updating the UI - After each LLM response, persist the token count to graph state - When resuming a thread, load the saved token count alongside the message history so the display is immediately correct - Add `_context_tokens` as a `PrivateStateAttr` field on a new `TokenTrackingState` schema, registered via a no-op `TokenStateMiddleware` in `create_cli_agent` — checkpointed but hidden from the model - Remove the `TextualTokenTracker` class from `app.py`; replace with a plain `_context_tokens: int` cache on the app and three typed callbacks (`_on_tokens_update`, `_on_tokens_hide`, `_on_tokens_show`) wired to the `TextualUIAdapter` - After each LLM response, `_persist_context_tokens` writes the count to graph state via `aupdate_state`; failures log at `warning` level rather than silently disappearing - On thread resume, `_fetch_thread_history_data` now returns a `_ThreadHistoryPayload` that bundles messages + the persisted token count; `_load_thread_history` seeds the local cache from it - `/tokens`, `/clear`, `/offload`, and thread-switch all read/write `self._context_tokens` instead of the removed tracker - `_get_thread_state_values` fallback (server mode) now recovers `_context_tokens` from the SQLite checkpointer alongside `_summarization_event`
Mason Daugherty (mdrxy)
added a commit
that referenced
this pull request
Apr 7, 2026
> [!CAUTION] > Merging this PR will automatically publish to **PyPI** and create a **GitHub release**. For the full release process, see [`.github/RELEASING.md`](https://github.com/langchain-ai/deepagents/blob/main/.github/RELEASING.md). --- _Everything below this line will be the GitHub release body._ --- ## [0.0.35](deepagents-cli==0.0.34...deepagents-cli==0.0.35) (2026-04-07) ### Highlights * **Skills:** Invoke SDK skills directly from the CLI with `/skill:name` or at startup via `--skill`. Skills are a composable extension point for domain-specific agent behaviors. * **Themes & configuration:** New theme system with color overrides, global dotenv at `~/.deepagents/.env`, and a `DEEPAGENTS_CLI_` env var prefix for conflict-free configuration. * **Auto-updates:** Full update lifecycle — `/update` to upgrade in-place, `/auto-update` to toggle background checks, and a refreshed install script UX. * **Headless workflows:** Agent-friendly UX for scripted/headless invocations, AgentCore Code Interpreter sandbox provider, and improved non-interactive tracing and guidance. * **Performance:** Sub-250ms first paint via aggressive import deferral (pydantic, adapters, heavy SDK modules), markdown stack prewarming, and reduced health-poll intervals. ### Features * Load `~/.deepagents/.env` as global dotenv ([#1909](#1909)) ([5a21d0a](5a21d0a)) * Skill invocation via `/skill:name` ([#2037](#2037)) ([cc8cce7](cc8cce7)) * `--skill` startup invocation ([#2477](#2477)) ([5f0f1d4](5f0f1d4)) * Auto-update lifecycle, `/update` command, install script ux ([#2095](#2095)) ([fd92f6e](fd92f6e)) * `/auto-update` to toggle auto-updates ([#2276](#2276)) ([ad70bde](ad70bde)) * Themes ([#2134](#2134)) ([db67af0](db67af0)) * Allow color overrides on built-in themes, default dark to false ([#2275](#2275)) ([8f71865](8f71865)) * Add `DEEPAGENTS_CLI_` env var prefix and fix dotenv load order ([#2303](#2303)) ([29647bb](29647bb)) * Add `ls_integration` metadata to langsmith traces ([#2272](#2272)) ([5dd8098](5dd8098)) * Add animated spinner to non-interactive verbose mode ([#2001](#2001)) ([153f465](153f465)) * Add async backend support to local context middleware ([#2118](#2118)) ([a0d623c](a0d623c)) * Agent-friendly ux for scripted/headless workflows ([#2271](#2271)) ([386438f](386438f)) * AgentCore Code Interpreter sandbox provider ([#2120](#2120)) ([92556c7](92556c7)) * Context-aware connecting banner for resume and local server ([#2092](#2092)) ([18b385b](18b385b)) * Default langsmith project to `'deepagents-cli'` ([#2277](#2277)) ([7178b87](7178b87)) * Enhance tool-call UI, add `Ctrl+U` shortcut for chat input ([#1757](#1757)) ([800c552](800c552)) * Persist token count in graph state across sessions ([#2323](#2323)) ([5be352d](5be352d)) * Pop queued messages individually on `esc` instead of clearing all ([#2089](#2089)) ([c76d855](c76d855)) * Render ask-user questions as markdown ([#2339](#2339)) ([5fbb14a](5fbb14a)) * Show platform-specific ripgrep install command in missing-tool warning ([#1997](#1997)) ([f000ce5](f000ce5)) * Support root/MDM installs ([#2346](#2346)) ([f618acc](f618acc)) * Surface unsupported input modalities in system prompt ([#2327](#2327)) ([95620e7](95620e7)) ### Bug Fixes * Align headless todo guidance with non-interactive mode ([#2459](#2459)) ([281899b](281899b)) * Disable markup parsing for blocked-link notifications ([#2170](#2170)) ([15867bf](15867bf)) * Dismiss slash command autocomplete on space ([#2478](#2478)) ([02e46bc](02e46bc)) * Eliminate autocomplete popup flicker ([#2020](#2020)) ([4b2db1e](4b2db1e)) * Eliminate trace fragmentation in non-interactive mode ([#2136](#2136)) ([9bddc52](9bddc52)) * Enforce approval toggle when launched with `-y` ([#2278](#2278)) ([28a32b7](28a32b7)) * Escape exception text in rich markup error output ([#2307](#2307)) ([42bccca](42bccca)) * Escape markup in toast notifications ([#2139](#2139)) ([90ccc28](90ccc28)) * Exit app on `ctrl+d` when thread list is empty ([#2270](#2270)) ([e859077](e859077)) * Guard against textual cursor/document desync crash ([#2494](#2494)) ([c14a748](c14a748)) * Human-readable duration and consistent dim styling on teardown screen ([#1995](#1995)) ([901a0a4](901a0a4)) * Mark token count as approximate after interrupted generation ([#2353](#2353)) ([cb9a0c7](cb9a0c7)) * Resolve misleading "missing package" error when provider import fails ([#1960](#1960)) ([b90fbad](b90fbad)) * Open trace in browser immediately when busy ([#2305](#2305)) ([b452032](b452032)) * Patch model identity in system prompt on `/model` swap ([#2024](#2024)) ([36aecbf](36aecbf)) * Harden MCP pre-flight health checks ([#2019](#2019)) ([2b27055](2b27055)) * Pre-flight health checks for MCP servers ([#2008](#2008)) ([30d60e3](30d60e3)) * Prevent premature thinking state with parallel subtasks ([#1858](#1858)) ([189104c](189104c)) * Prevent session stats loss on mid-turn exit ([#2238](#2238)) ([b1807aa](b1807aa)) * Rebind toggle tool output to `ctrl+o` to unblock `cmd+right` ([#2088](#2088)) ([b486fe5](b486fe5)) * Remove duplicate server failure notification ([#2141](#2141)) ([c1cfe72](c1cfe72)) * Remove keybinding overrides that shadow textual built-ins ([#2084](#2084)) ([08fc5d0](08fc5d0)) * Resolve conflicting langsmith env var precedence ([#2455](#2455)) ([b6997d8](b6997d8)) * Show server startup error instead of generic agent message ([#2397](#2397)) ([a3e1e93](a3e1e93)) * Certain slash commands should not require server connection ([#1974](#1974)) ([32bd814](32bd814)) * Sort MCP tools deterministically for prompt-cache stability ([#2497](#2497)) ([39b43cf](39b43cf)) * Squash loading widget timer leak ([#2396](#2396)) ([01d3d86](01d3d86)) * Support pre-release versions in update checker ([#2164](#2164)) ([e18e9dc](e18e9dc)) * Surface clear error for missing sandbox provider deps ([#1999](#1999)) ([939f56a](939f56a)) * Use relative paths in langgraph config for Windows compat ([#2244](#2244)) ([d10dfbd](d10dfbd)) * Warn agent that local filesystem is inaccessible in sandbox mode ([#2274](#2274)) ([a3b61e5](a3b61e5)) * Wire `enable_ask_user` flag to remove tool in non-interactive mode ([#2105](#2105)) ([2399747](2399747)) ### Performance Improvements * Sub 250ms first paint ([#2027](#2027)) ([e42e05c](e42e05c)) * Defer `/model` selector data loading off event loop ([#2259](#2259)) ([a32ce7f](a32ce7f)) * Defer heavy imports from startup path ([#2022](#2022)) ([b7f5a99](b7f5a99)) * Defer pydantic and adapter imports out of startup hot path ([#2269](#2269)) ([0a410b4](0a410b4)) * Prewarm markdown stack and cache skill body render ([#2236](#2236)) ([0a3ba47](0a3ba47)) * Reduce health poll interval for local langgraph dev server ([#2283](#2283)) ([7f5c3de](7f5c3de)) --- _Everything above this line will be the GitHub release body._ --- > [!NOTE] > A **New Contributors** section is appended to the GitHub release notes automatically at publish time (see [Release Pipeline](https://github.com/langchain-ai/deepagents/blob/main/.github/RELEASING.md#release-pipeline), step 2). --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Mason Daugherty <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The CLI's token count (
/tokens, status bar) was stored only in memory viaTextualTokenTracker, so resuming a thread with-ralways showed "No token usage yet" — even though/offloadindependently counted ~291 tokens from messages. This persists the count in the LangGraph graph state so it survives across sessions and both commands agree.Changes (TL;DR)
TextualTokenTrackerclass with a simple integer on the app plus three callbacks for updating the UIChanges (detailed)
_context_tokensas aPrivateStateAttrfield on a newTokenTrackingStateschema, registered via a no-opTokenStateMiddlewareincreate_cli_agent— checkpointed but hidden from the modelTextualTokenTrackerclass fromapp.py; replace with a plain_context_tokens: intcache on the app and three typed callbacks (_on_tokens_update,_on_tokens_hide,_on_tokens_show) wired to theTextualUIAdapter_persist_context_tokenswrites the count to graph state viaaupdate_state; failures log atwarninglevel rather than silently disappearing_fetch_thread_history_datanow returns a_ThreadHistoryPayloadthat bundles messages + the persisted token count;_load_thread_historyseeds the local cache from it/tokens,/clear,/offload, and thread-switch all read/writeself._context_tokensinstead of the removed tracker_get_thread_state_valuesfallback (server mode) now recovers_context_tokensfrom the SQLite checkpointer alongside_summarization_event