Skip to content

Report real Grok token usage and list-price cost from CLI logs - #3135

Open
olddonkey wants to merge 19 commits into
steipete:mainfrom
olddonkey:feat/grok-real-token-usage
Open

Report real Grok token usage and list-price cost from CLI logs#3135
olddonkey wants to merge 19 commits into
steipete:mainfrom
olddonkey:feat/grok-real-token-usage

Conversation

@olddonkey

@olddonkey olddonkey commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

Problem

#3085 enabled Grok token cost, but the local session projection was not measuring actual consumption:

  • signals.json exposes ending context-window occupancy, not per-turn token usage. On a real corpus it reported 653K tokens where the completed turns contained 54.1M.
  • Grok cost was always nil, and the models.dev resolver did not include the xAI catalog.
  • Expanding the scan from small metadata files to growing updates.jsonl files could not safely remain on @MainActor.

What this changes

Read the completed-turn usage that the Grok CLI actually records

GrokLocalSessionScanner reads turn_completed events from each session's updates.jsonl, matches both session/update and _x.ai/session/update, and buckets each line by its own timestamp. It preserves raw aggregate token totals and uses the recorded modelCalls count only to approximate per-call tiered list pricing in O(1).

The parser now streams through the shared chunked JSONL reader instead of loading whole files. Production bounds are explicit: a 64 MiB tail per file, 1 MiB per record, 20,000 retained turns per file, and global scan budgets of 256 recent sessions, 256 MiB, and 100,000 turns. The process-wide LRU cache retains at most 64 files or 50,000 turns. Cancellation is checked between chunks, I/O/cancelled results are not cached, and any truncation marks history incomplete rather than presenting a partial total as complete.

Price the Grok models and preserve provenance

The models.dev xAI catalog is now eligible for cost pricing. Responses-API names such as grok-4.6-build resolve to the base grok-4.6 SKU after exact lookup, while real independent names such as grok-build-0.1 remain untouched. Cost is published as .listPriceEstimate; Grok's internal costUsdTicks is not presented as billed spend.

On a stale catalog, the scan prices immediately and refreshes in the background. On a fresh install with no catalog artifact, the first scan now awaits the initial best-effort refresh attempt before creating the snapshot, so a successful first refresh is visible in the first publication. A refresh failure still degrades to token-only data rather than failing the local scan.

Keep scanning and publication off the main actor

The provider projection consumes the async probe's snapshot. Remaining fallback paths scan on a detached utility task with a single scan in flight. A maximum-window snapshot is narrowed by each consumer through CostUsageTokenSnapshot.narrowed(toHistoryDays:calendar:), including the spend dashboard's 365-day request.

Keep OpenCodex xAI history out of the Grok subscription row

OpenCodex usage.jsonl records provider/model usage but does not retain the credential mode used for each request. Reading today's config.json cannot distinguish older API-key traffic from older OAuth traffic after a configuration switch.

For that reason, xAI entries now remain .tokenOnly and are not merged into the Grok subscription row until the producer records request-time credential provenance. The current-config reader and its routing parameter were removed, with dispatcher/fan-out regressions and updated documentation. Other existing OpenCodex subscription routes are unchanged.

Review findings addressed

  • P1 — preserve record-time xAI credential attribution: resolved by pausing OpenCodex xAI-to-Grok attribution when record-time evidence is unavailable. Historical traffic can no longer move between provider rows when the current auth config changes.

  • P2 — publish pricing after the first catalog refresh: resolved by awaiting the initial refresh only when no cached catalog exists, with a regression proving that a successful refresh prices the first returned summary. Stale catalogs still use the non-blocking refresh path.

  • P2 — bound growing session logs: resolved at bdfc10e0d with chunked tail reads, per-file and global byte/turn/session budgets, bounded LRU retention, cancellation checks, and incomplete-history propagation. Regressions cover tail truncation, global scan budgets, and cache eviction limits.

  • P3 — correct the documented session source: the earlier signals.json/30-day description now documents updates.jsonl, the requested window up to 365 days, list-price provenance, and every production bound. signals.json is metadata-only and also byte-limited.

  • The earlier repeated-probe failure was also fixed: an existing Grok local publication survives consecutive remote refresh failures instead of falling through to the generic clear path.

  • P1 — restore the published fallback for Grok menu consumers: resolved on exact head 2f80ae834 with a live-consumer-only Grok projection fallback. Menu cards and cost history now consume the current-config local publication when the remote snapshot is absent, while override cards retain the original isolated projection and cannot inherit provider-level data. Regressions cover both the visible fallback and override isolation.

  • P1 — rescan after an empty Grok publication: resolved on exact head c85067dda. A current-config publication with no snapshot is now treated as no usable fallback, so a later remote failure rescans newly written local turns. Once a nonempty snapshot is published, subsequent failures reuse it without redundant scans. The empty-to-nonempty transition is covered by the scanner regression.

  • P1 — keep bare OpenCodex xAI history token-only: standalone aggregation no longer assigns list-price dollars when request-time credential provenance is unavailable.

  • P1 — label populated Grok cost surfaces: menu details, charts, dashboard rows, and spend views identify the amount as a public xAI list-price estimate rather than a bill.

  • P1 — keep failed-refresh local totals advancing: every failed Grok remote billing refresh schedules the bounded local scan; the existing single-flight task coalesces concurrent scans.

  • P1 — isolate xAI pricing invalidation: xAI now has a separate pricing fingerprint, so xAI-only catalog changes do not invalidate Codex caches.

  • P2 — bound discovery I/O: Grok session discovery stops at 4,096 tree entries and marks history incomplete when that bound is reached.

  • P2 — refresh after preservable network failures: resolved on exact head c4c4000e0. A timeout can retain the last remote provider snapshot while still rescanning local sessions; live Grok consumers compare timestamps and select the newer current-config local publication. Empty local scans leave the retained remote projection in place, and account/context override cards remain isolated from provider-level live data.

  • P1 — keep Usage & Spend on the freshest Grok data: resolved on exact head fcdffd107. Dashboard capture and retained-publication paths now use the same timestamp selector as live menu consumers; a newer 100-token local publication wins over a preserved 77-token remote snapshot.

  • P3 — leave release notes to the release flow: the normal-PR CHANGELOG.md entry was removed; the release-note context remains in this PR body.

  • Latest-main rebase integration: rebased onto main 41d904fd6, retained main’s off-main Grok scanning path, migrated its projection regression from signals.json metadata to completed-turn JSONL, and adopted the previous main parser hash without rebuilding compatible Codex rows.

Real-session evidence

The opt-in proof was rerun on exact head fcdffd1077147dc538ddd1412f37fac7336acacf against the native Grok CLI session corpus through the shipped bounded scanner, snapshot projection, models.dev xAI pricing, and spend-catalog path:

CODEXBAR_LIVE_GROK_CATALOG_PROOF=1 swift test --filter GrokXAISpendCatalogTests

catalog_source=grok
today_tokens=0
last_30_days_tokens=2739923
today_cost_usd=nil
window_cost_usd=2.181282
cost_provenance=listPriceEstimate
history_days=365
priced_days=1
token_days=1
daily_buckets=1
available_sources=grok

All 3 tests in the suite passed. This opt-in route reads the real local ~/.grok/sessions/**/updates.jsonl corpus and the local models.dev catalog rather than temporary session fixtures; its transcript prints aggregate fields only. The current corpus has no completed-turn tokens today and one priced day in the last 30 days, so these values naturally change as local logs age. The populated-surface assertion in the same suite also verifies the visible Public xAI list-price estimate · not a bill. disclosure.

The same exact tree was then packaged with ./Scripts/package_app.sh. The production build, widget packaging, code-sign validation, helper-resource smoke, symlinked-helper smoke, and app-launch smoke passed. The packaged helper reports CodexBar 0.55.0, config validate returns Config: OK, and the running process is /Users/olddonkey/Documents/CodexBar-pr3135/CodexBar.app/Contents/MacOS/CodexBar.

Evidence boundary: the public packaged usage JSON intentionally does not encode UsageSnapshot.costUsage, and a forced --provider grok --source cli call in this environment returned a provider error. I am therefore not presenting that command as spend proof. The package result establishes exact-tree production build/sign/launch health; the redacted transcript above establishes final-head native-session token/pricing behavior. No credential values, account identity, session contents, or interactive Keychain/browser-cookie reads were emitted.

Exact-head deterministic behavior proof

Exact head fcdffd1077147dc538ddd1412f37fac7336acacf drives the production bounded JSONL scanner, publication path, failure preservation, timestamp selection, menu consumers, and spend-dashboard capture through temporary on-disk Grok session files:

initial_completed_turn_tokens=77
retained_remote_tokens_after_timeout=77
local_completed_turn_tokens_after_append=100
selected_live_consumer_tokens=100
selected_source=newer_current_config_local_publication
second_timeout_local_scan_count=2
dashboard_selected_tokens=100

The regression first writes a completed 77-token turn and installs its projection as the retained remote snapshot. It then appends a second completed 23-token turn, injects URLError.timedOut, verifies the remote snapshot is preserved at 77, verifies the local publication advances to 100, and verifies the live consumer selects the newer 100-token snapshot. A second failed refresh proves that local logs are rescanned again. The status-menu regression separately proves that newer local tokens beat stale remote tokens in both the visible card and cost-history submenu, while the override-card isolation regression remains green. The dashboard regression independently installs a one-minute-older 77-token remote snapshot and a current 100-token local publication, then verifies capture-only Usage & Spend selects the local 100-token, $1 list-price snapshot and its newer timestamp.

This is exact-head deterministic production-path evidence using real JSONL files in a temporary Grok home. The separate native-session aggregate above was rerun on the same exact final head against the installed Grok CLI corpus; the two proofs are intentionally distinguished.

Testing

  • Exact head fcdffd1077147dc538ddd1412f37fac7336acacf: make check passed — parser hash 678fd59821eccb04, provider/package/documentation gates, SwiftFormat (0/2002), and SwiftLint (0 violations in 2001 files).
  • Exact head fcdffd1077147dc538ddd1412f37fac7336acacf: make test passed — 933 selections, 78/78 groups successful on the first pass, 0 failed groups, 0 retries, and 0 timeouts (648.5 seconds).
  • Exact-head focused validation includes SpendDashboardGrokFreshnessTests (1/1), SpendDashboardControllerTests (25/25), and ProviderArchitectureGatekeeperTests (38/38); the preservable-timeout and stale-remote menu regressions also remain in the fully green suite.
  • Pricing tests now use isolated models.dev cache roots, preventing user-level catalog artifacts refreshed by Grok tests from changing built-in GPT-5.6 expectations.
  • The branch is rebased on upstream main 41d904fd6d884b7df73ac180a913ccbf04a9d7d6, 0 behind / 19 ahead, and git diff --check is clean. GitHub reports MERGEABLE; all exact-head CI checks passed.
  • Earlier native-session proof head 3b642530a: focused OpenCodex/Grok suites (35 tests), production bundle/signing/launch smoke, packaged CodexBarCLI config validate, packaged Grok usage through grok-cli-proxy, and the authorized two-test live scanner/pricing/catalog proof passed.

Normal automated tests use temporary homes, injected catalogs/transports, and no live credentials or interactive Keychain reads. The earlier explicitly authorized live checks used the configured Grok provider and local session logs; no credential values were printed.

Maintainer decision

The owner-requested code findings are addressed on exact head fcdffd107, including preserved-timeout freshness in both menu and Usage & Spend consumers and removal of the release-owned changelog entry. The branch is rebased and locally green. All exact-head CI checks passed; ClawSweeper re-review is in progress, and the final-head native-session proof is recorded above. Owner approval of the default estimate semantics and maintainer re-review remain the external gates. The xAI amount remains clearly labeled as a non-billed public list-price estimate, distinct from SuperGrok subscription credits.

Release-note context is retained in this PR body; the normal-PR changelog entry was removed per repository release ownership.

@clawsweeper

clawsweeper Bot commented Aug 22, 2026

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

olddonkey added a commit to olddonkey/CodexBar that referenced this pull request Aug 22, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 09cf7edb0f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread Sources/CodexBar/UsageStore+Refresh.swift Outdated
Comment thread Sources/CodexBarCore/Providers/Grok/GrokLocalSessionScanner.swift
@clawsweeper clawsweeper Bot added merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. P2 Normal priority bug or improvement with limited blast radius. rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. labels Aug 22, 2026
@clawsweeper

clawsweeper Bot commented Aug 22, 2026

Copy link
Copy Markdown

Codex review: needs maintainer review before merge. Reviewed August 25, 2026, 6:40 PM ET / 22:40 UTC.

ClawSweeper review

What this changes

This PR reads bounded completed-turn Grok CLI logs for usage, estimates xAI list-price cost, and publishes the freshest local result to menu and Usage & Spend views.

Merge readiness

⚠️ Ready for maintainer review - 3 items remain

The patch is coherent and evidence-backed, but defaulting Grok to a visible non-billed list-price amount changes established display semantics and needs owner approval before merge.

Priority: P2
Reviewed head: fcdffd1077147dc538ddd1412f37fac7336acacf
Owner decision: Required. See Decision needed.

Review scores

Measure Result What it means
Overall readiness 🐚 platinum hermit (4/6) Strong exact-head proof and targeted regression coverage support the patch; only product acceptance of the default estimate semantics remains.
Proof confidence 🦞 diamond lobster (5/6) Sufficient (live_output): The PR supplies redacted exact-head native-session output, deterministic production-path JSONL evidence, and packaged-app smoke results.
Patch quality 🐚 platinum hermit (4/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Verified Sufficient (live_output): The PR supplies redacted exact-head native-session output, deterministic production-path JSONL evidence, and packaged-app smoke results.
Evidence reviewed 5 items Bounded scanner and estimate provenance: The scanner reads completed-turn JSONL records with explicit file, record, session, discovery, byte, and turn limits; generated snapshots are marked as list-price estimates.
Live consumer freshness selection: Grok’s live menu and dashboard consumers select a newer current-config local publication over an older remote projection while keeping override cards isolated.
Rebase provenance: The PR head integrates the current-main off-menu-thread Grok scanning work from the merged #3198 change.
Findings None None.
Security None None.

Live Verification

Command: swift run CodexBarCLI --help

Result: FAIL (failed) — execution before step 1 run: sh -lc pnpm install --ignore-scripts --frozen-lockfile failed: ! Corepack is about to download https://registry.npmjs.org/pnpm/-/pnpm-11.24.0.tgz

sh -lc pnpm install --ignore-scripts --frozen-lockfile failed: ! Corepack is about to download https://registry.npmjs.org/pnpm/-/pnpm-11.24.0.tgz

Assertions:

  • FAIL expect_output: Usage:

How this fits together

CodexBar gathers remote provider data and local CLI logs, then projects usage snapshots to its menu and Usage & Spend dashboard. This changes Grok’s local-log input, pricing calculation, and snapshot selection.

flowchart LR
A[Grok CLI session logs] --> B[Bounded local scanner]
C[xAI price catalog] --> D[List-price calculation]
B --> D
D --> E[Usage snapshot publication]
F[Remote billing probe] --> E
E --> G[Menu and Usage and Spend views]
Loading

Decision needed

Question Recommendation
Should CodexBar show public xAI list-price estimates by default for locally recorded Grok CLI usage? Accept disclosed estimates: Merge the default estimate with the existing visible “not a bill” disclosure.

Why: This PR changes the user-visible meaning of Grok cost from unavailable to a non-billed USD estimate, which cannot be resolved by implementation correctness alone.

Before merge

  • Resolve merge risk (P1) - Existing Grok users will see materially different historical token totals and a new USD amount; the disclosure avoids calling it a bill, but the default estimate semantics still require explicit acceptance.
  • Complete next step (P2) - A human must choose the default semantics for the newly displayed Grok list-price estimate; no mechanical repair is identified.
Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Production versus test delta production +1,581/-339, tests +2,154/-100 The large scanner and projection change is accompanied by more test coverage than production growth.
Changed surface 34 files affected The update spans parsing, snapshot publication, native UI disclosure, documentation, and regressions.

Merge-risk options

Maintainer options:

  1. Confirm default estimate semantics (recommended)
    Accept the compatibility change only if a clearly disclosed public list-price estimate is the intended default for existing Grok users.
  2. Pause the estimate surface
    Keep the scanner work out of the default display until a token-only or opt-in direction is selected.

Technical review

Best possible solution:

Merge the bounded completed-turn implementation only after the owner explicitly accepts disclosed public xAI list-price estimates as the default Grok cost presentation.

Do we have a high-confidence way to reproduce the issue?

Not applicable: this is a provider-reporting enhancement rather than a single current-main failure report; the PR supplies exact-head runtime evidence for its new behavior.

Is this the best way to solve the issue?

Unclear: the implementation and disclosure are well-tested, but an owner must decide whether this estimate belongs in the default product surface.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning high; reviewed against 41d904fd6d88.

Labels

Label justifications:

  • P2: This is a substantial provider-usage improvement with a normal-priority owner decision about default behavior.
  • merge-risk: 🚨 compatibility: The merge changes historical Grok totals and introduces a new default USD estimate for existing users.
  • rating: 🐚 platinum hermit: Overall readiness is 🐚 platinum hermit; proof is 🦞 diamond lobster and patch quality is 🐚 platinum hermit.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Sufficient (live_output): The PR supplies redacted exact-head native-session output, deterministic production-path JSONL evidence, and packaged-app smoke results.
  • proof: sufficient: Contributor real behavior proof is sufficient. The PR supplies redacted exact-head native-session output, deterministic production-path JSONL evidence, and packaged-app smoke results.

Evidence

What I checked:

Likely related people:

  • Peter Steinberger: Merged the current-main Grok scan scheduling work and has extensive history in the shared usage and spend paths. (role: recent area contributor and merger; confidence: high; commits: 41d904fd6d88, 61fbe9fac507; files: Sources/CodexBarCore/Providers/Grok/GrokLocalSessionScanner.swift, Sources/CodexBar/UsageStore+TokenCost.swift)
  • Alec Gutman, Chip: Introduced the merged Grok Usage & Spend source that this PR refines. (role: introduced Grok Usage & Spend surface; confidence: high; commits: 3bfbffdcea58; files: Sources/CodexBar/SpendDashboardController.swift, Sources/CodexBarCore/Providers/Grok/GrokLocalSessionScanner.swift)

Rank-up moves

Optional improvements that raise the rating; they are not merge blockers.

  • Obtain owner confirmation that disclosed public xAI list-price estimates are the intended default Grok cost behavior.

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (23 earlier review cycles; latest 8 shown)
  • reviewed 2026-08-24T11:17:19.095Z sha c85067d :: needs maintainer review before merge. :: none
  • reviewed 2026-08-24T11:48:16.086Z sha c85067d :: needs maintainer review before merge. :: none
  • reviewed 2026-08-25T02:35:57.554Z sha 1846517 :: needs maintainer review before merge. :: none
  • reviewed 2026-08-25T04:15:09.806Z sha 584b03a :: needs real behavior proof before merge. :: [P2] Refresh local Grok usage after preservable probe failures
  • reviewed 2026-08-25T05:13:02.943Z sha c4c4000 :: needs real behavior proof before merge. :: [P1] Use the fresher local snapshot in Usage & Spend | [P3] Leave the changelog entry to the release flow
  • reviewed 2026-08-25T06:15:15.014Z sha ff85e9a :: needs real behavior proof before merge. :: none
  • reviewed 2026-08-25T06:52:20.055Z sha ff85e9a :: needs real behavior proof before merge. :: none
  • reviewed 2026-08-25T08:55:43.508Z sha ff85e9a :: needs maintainer review before merge. :: none

@olddonkey
olddonkey force-pushed the feat/grok-real-token-usage branch from 08360b5 to e3cd3b9 Compare August 22, 2026 06:21
@olddonkey olddonkey changed the title Report real Grok token usage and list-price cost from CLI session logs Report real Grok token usage and list-price cost, from the CLI logs and OpenCodex alike Aug 22, 2026
@clawsweeper clawsweeper Bot added merge-risk: 🚨 auth-provider 🚨 Merging this PR could break OAuth, tokens, provider routing, model choice, or credentials. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. and removed rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. labels Aug 22, 2026
@olddonkey

Copy link
Copy Markdown
Contributor Author

Both automated findings are addressed, plus the review's other checklist items. The inline comments were left against 09cf7edb0, which no longer exists — the branch has since been rebased onto 27c7f334e and the head is now 03e5a25dc, so I'm summarising here rather than replying in a stale diff.

P1 — Preserve the Grok fallback on repeated probe failures

Fixed in 923193ec0, Sources/CodexBar/UsageStore+Refresh.swift. The guard had been hoisted into the if provider == .grok, publication == nil condition, so a Grok failure with a publication fell through to the generic else if tokenCostRequiresProviderSnapshot { clearTokenSnapshot } branch. Grok now owns its branch outright and can never reach the clear:

if provider == .grok {
    if self.tokenSnapshotPublicationForCurrentProviderConfig(for: provider) == nil {
        Task { @MainActor [weak self] in
            await self?.scanAndPublishGrokLocalTokenSnapshot(...)
        }
    }
} else if Self.tokenCostRequiresProviderSnapshot(provider) {
    self.clearTokenSnapshot(for: provider)
}

Regression coverage is in missing remote snapshot scans and publishes local tokens then clears empty data. Per the review's request it now drives two consecutive failing refreshes (03e5a25dc) rather than one — which matters here, because the first failure is what publishes through the fallback scan and only the second arrives with a publication in place, i.e. the failure that used to wipe the row. Both iterations assert the row still reads 77 tokens and that no redundant rescan ran.

P2 — Refresh pricing before scanning Grok sessions

Correct, and thank you — this was a genuine gap and not one the local tests would have surfaced. refreshPricingIfAllowed is gated to Codex and Claude, and Grok never reaches it at all because its snapshot comes from the provider probe rather than CostUsageFetcher.loadTokenSnapshot. On a machine with Codex or Claude also enabled the shared cache is already populated, so the failure is invisible there; enable only Grok and the catalog never appears and the Cost row shows tokens with no money, permanently.

Fixed in 744677e68. The Grok scan paths now request ModelsDevPricingPipeline.refreshIfNeeded through a summarizeRequestingPricingRefresh wrapper, called from all four scan sites (GrokStatusProbe, both branches in GrokProviderDescriptor, and UsageStore.scanAndPublishGrokLocalTokenSnapshot). It is detached rather than awaited, matching how the Codex and Claude paths already treat it — pricing availability must not delay or fail a local scan — and it is safe to call repeatedly, since it returns immediately unless the cache is stale and serialises through its own coordinator. summarize itself stays synchronous and side-effect free.

Note the inline comment still points at GrokLocalSessionScanner.swift:662; that line is the unchanged pricing lookup, and the fix is upstream of it in the new wrapper, so the anchor looks live even though it is addressed.

Coverage: absent models dev cache requests a background refresh, stale models dev cache requests a background refresh, and fresh models dev cache skips the background refresh. All three assert whether a refresh was requested through an injected transport — no test touches the network.

Real-session evidence

CODEXBAR_LIVE_GROK_CATALOG_PROOF=1 swift test --filter GrokXAISpendCatalogTests, against real local Grok CLI sessions, through the shipped code path:

catalog_source=grok
today_tokens=5043749
last_30_days_tokens=52696354
today_cost_usd=3.3471699999999993
window_cost_usd=49.353424
cost_provenance=listPriceEstimate
history_days=365
priced_days=4
token_days=4
daily_buckets=4
available_sources=grok

The same corpus on main reports 653K tokens and no cost. history_days=365 shows the requested window is honoured (it was pinned to 30). priced_days == token_days shows no day was silently left unpriced. The gated proof was extended in e3cd3b9ce to print cost, provenance and priced-day coverage, since tokens alone cannot evidence the half of this change that is about money.

Those figures were cross-checked against an independent reimplementation of the pricing formula over the same logs; the two agree to the cent.

Merge risk / branch state

Rebased onto current main (27c7f334e); the branch reports clean. Full suite on the head: 77/77 groups, 922 selections, 0 failures. swiftformat --lint and swiftlint --strict clean. Upstream CI green on the previous head including all three Linux builds.

One thing deliberately left undone: no CHANGELOG.md entry. 0.54.1 was finalized and there is no open Unreleased section, so I did not invent a version heading — happy to add one wherever you prefer.

@clawsweeper clawsweeper Bot added rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. and removed rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. labels Aug 22, 2026
@olddonkey

olddonkey commented Aug 22, 2026

Copy link
Copy Markdown
Contributor Author

Both new findings addressed at 44d79a95a.

P1 — Do not map every xAI log record to the Grok subscription

Agreed, and taken as specified rather than argued down. Routing on the prefix alone is right for the case that motivated this — traffic authenticated with the user's Grok account, which is what makes it consume SuperGrok quota — but it silently folds an API-key user's pay-as-you-go xAI spend into the subscription row. CodexBar already models the developer platform as its own xai provider precisely to keep those apart, so the old behaviour crossed a boundary the app deliberately maintains.

The usage log carries no per-record credential evidence; I checked every field emitted for xai rows (requestId, timestamp, provider, model, requestedModel, resolvedModel, usage, usageStatus, status, routeDecision, …) and there is nothing about auth, account or key. The signal that does exist is ~/.opencodex/config.json, which records authMode per provider.

So attribution now requires positive OAuth evidence:

  • xai routes to .subscription(.grok) only when its configured authMode is OAuth. Anything else returns .tokenOnly — the spend is real, it just belongs to no tracked subscription — rather than .unknown, which would read as "unrecognised provider".
  • Fail closed. A missing or malformed config, a providers block without xai, or an entry without authMode all count as no evidence and keep the records off the Grok row.
  • OpenCodexRouteDispatcher stays a pure function. The set of OAuth-backed provider ids is threaded in from the caller (OpenCodexUsageFanOutSpendDashboardSource), so the routing site never touches the filesystem and every existing caller and test that does not care about auth keeps working.
  • The gate applies only to xai. openai, kimi-coding, deepseek and opencode-go are untouched — changing them would be an unreviewed behaviour change for other providers — and a test pins that they ignore xAI auth state entirely.

Coverage: xai OAuth config routes to Grok, xai non OAuth config stays token only (parameterised over several non-OAuth values), xai routing fails closed without readable complete OAuth config, non xai subscription routes ignore xai auth state, plus fan-out cases proving the same entries land on the Grok row under an OAuth config and are absent under an API-key one. No test reads the developer's real ~/.opencodex; the home directory is injected.

docs/grok.md no longer claims this path cannot distinguish OAuth from API-key traffic, because it now can.

P2 — Republish the Grok snapshot after a missing catalog refreshes

I looked at this closely and am deliberately not adding a republish path. Reasoning, so you can overrule it if you disagree:

The refresh is fire-and-forget, so the scan that requests it returns whatever the cache currently holds — that part is accurate. But the parse cache stores parsed turns, not prices, so aggregation and pricing re-run on every summarize. The next Grok scan therefore prices against the refreshed catalog with no extra machinery, bounding the unpriced window to a single refresh cycle. That is the same behaviour Codex and Claude already have: refreshPricingIfAllowed dispatches into Task.detached and their current scan does not wait for it either.

The alternative — plumbing a completion signal back across the actor boundary into the @MainActor publication path — buys one refresh cycle of latency on first run, at the cost of a new cross-actor completion path in code that publishes user-visible spend. That trade looked disproportionate, and inconsistent with how the two established providers behave. I have recorded the reasoning as a comment at the call site rather than leaving it implicit, so the next reader does not have to re-derive it.

Happy to build it if you would rather have it.

Evidence

The attribution itself only becomes visible in the app: SpendDashboardSource.mergingOpenCodexInputs is what merges the fan-out into provider rows, and the CLI's cost command reports OpenCodex as its own source rather than routing it, so terminal output cannot show this path. The figures below are read off the freshly packaged build running against real local data, on a machine whose ~/.opencodex/config.json has "xai": { "authMode": "oauth" }; screenshots of both panes follow.

The two halves stay distinguishable in the UI, which makes the attribution legible rather than something you have to take on trust: the CLI goes through the responses API so its SKU is grok-4.6-build, while OpenCodex's records resolve to the bare grok-4.6 / grok-4.5 / grok-4.3. Both sit under the Grok provider.

model row source shown
grok-4.6-build Grok CLI session logs $50.52 · 54M
grok-4.6 OpenCodex $161.21 · 182M
grok-4.5 OpenCodex $4.54 · 5.5M
grok-4.3 OpenCodex $0.50 · 201K

Independently recomputing the same corpus agrees to the cent on both halves: 54,121,501 tokens / $50.52 for the CLI logs, and $166.25 across 1,520 OpenCodex xai records. The CLI half is reproducible by anyone on their own machine through the gated proof test (CODEXBAR_LIVE_GROK_CATALOG_PROOF=1), whose output is in the PR body.

The negative direction — API-key traffic staying off the Grok row — is covered by tests rather than a screenshot, since demonstrating it live would mean rewriting the machine's OpenCodex config.

State

Full suite on 44d79a95a: 77/77 groups, 922 selections, 0 failures. swiftformat --lint and swiftlint --strict clean. Rebased on 27c7f334e.

image image

@clawsweeper clawsweeper Bot added proof: 📸 screenshot Contributor real behavior proof includes screenshot evidence. proof: sufficient Contributor real behavior proof is sufficient. status: ⏳ waiting on author ClawSweeper has contributor-facing work open and is waiting for author action. and removed status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. labels Aug 22, 2026
@olddonkey olddonkey changed the title Report real Grok token usage and list-price cost, from the CLI logs and OpenCodex alike Report real Grok token usage and list-price cost from CLI logs Aug 23, 2026
@olddonkey

Copy link
Copy Markdown
Contributor Author

Addressed both current findings in 211e1977d and resolved the two review threads.

  • P1 / historical xAI attribution: removed current-config-based xAI → Grok routing. usage.jsonl has no request-time credential provenance, so xAI records now remain token-only until the producer can persist that evidence. Removed the config reader/plumbing and added dispatcher/fan-out regressions.
  • P2 / first pricing publication: when no models.dev artifact exists, the first Grok scan now awaits the initial best-effort refresh attempt before summarizing. A successful refresh prices the first returned snapshot; stale catalogs still price immediately and refresh in the background. Added a regression that writes the catalog during refresh and asserts the first summary is priced.
  • Updated the PR title/body and docs/grok.md so they no longer claim OpenCodex xAI traffic is merged into the Grok subscription row.

Validation on the exact pushed head:

  • focused Grok/OpenCodex suites: 32 tests passed
  • make check: passed
  • make test: 922 selections, 77/77 groups, 0 failed groups, 0 retries
  • branch is based on current main (27c7f334e) and the merge-tree is clean

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Aug 23, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@clawsweeper clawsweeper Bot added merge-risk: 🚨 availability 🚨 Merging this PR could cause crashes, hangs, restart loops, stalls, or process outages. and removed merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. proof: sufficient Contributor real behavior proof is sufficient. proof: 📸 screenshot Contributor real behavior proof includes screenshot evidence. labels Aug 23, 2026
@clawsweeper clawsweeper Bot added rating: 🦪 silver shellfish Thin PR readiness signal; proof, validation, or implementation needs work. and removed rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. labels Aug 25, 2026
@olddonkey

Copy link
Copy Markdown
Contributor Author

Addressed the two latest findings on exact head ff85e9a6439199fd7e171856fc840468c8960af8.

  • P1 — Usage & Spend freshness: both dashboard capture and retained-publication paths now use the shared live Grok timestamp selector. A regression installs a one-minute-older 77-token remote snapshot and a current 100-token local publication, then verifies capture-only Usage & Spend selects the local 100-token / $1 list-price snapshot and newer timestamp.
  • P3 — release-owned changelog: removed the normal-PR CHANGELOG.md entry and retained the release-note context in the PR body.
  • Updated the provider-architecture gate exact anchors/fingerprint after the dashboard branch changed; the full 38-test gate passes.

Exact-head validation:

  • SpendDashboardGrokFreshnessTests: 1/1 passed
  • SpendDashboardControllerTests: 25/25 passed
  • ProviderArchitectureGatekeeperTests: 38/38 passed
  • make check: passed; parser hash 78aab1f789a2f6cc, SwiftFormat 0/1996, SwiftLint 0 violations in 1995 files
  • make test: 931/931 selections, 78/78 groups successful on the first pass, 0 failures, 0 retries, 0 timeouts (648.3s)
  • Rebased on upstream main b6a4ce968; 0 behind / 18 ahead; git diff --check clean

The PR body now contains the exact-head deterministic dashboard proof and still distinguishes the earlier native-session proof from the final head. No credentialed live-account probe was run for this repair. Exact-head CI is running.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Aug 25, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@clawsweeper clawsweeper Bot added rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. and removed rating: 🦪 silver shellfish Thin PR readiness signal; proof, validation, or implementation needs work. labels Aug 25, 2026
@olddonkey

Copy link
Copy Markdown
Contributor Author

Final exact-head status for ff85e9a6439199fd7e171856fc840468c8960af8:

  • All GitHub checks passed: changes, lint, Linux x64/ARM64/musl, macOS 0/2 and 1/2, aggregate lint-build-test, and GitGuardian.
  • GitHub reports MERGEABLE / CLEAN; the branch remains rebased on b6a4ce968, 0 behind / 18 ahead.
  • ClawSweeper reviewed this exact SHA with Findings: None; both review threads remain resolved.
  • Local exact-head gates remain green: make check and make test (931/931 selections, 78/78 first-pass groups, 0 retries/timeouts).

No further code comments are outstanding. Remaining external gates are the requested exact-head native-session proof, owner approval of the default non-billed list-price estimate semantics, and maintainer re-review of the older CHANGES_REQUESTED review. I did not run a credentialed live-account probe or merge the PR.

@olddonkey

Copy link
Copy Markdown
Contributor Author

Added final-head native-session evidence for exact SHA ff85e9a6439199fd7e171856fc840468c8960af8 and updated the PR body.

Redacted real-corpus transcript from the opt-in production-path proof:

CODEXBAR_LIVE_GROK_CATALOG_PROOF=1 swift test --filter GrokXAISpendCatalogTests

catalog_source=grok
today_tokens=0
last_30_days_tokens=2739923
today_cost_usd=nil
window_cost_usd=2.181282
cost_provenance=listPriceEstimate
history_days=365
priced_days=1
token_days=1
daily_buckets=1
available_sources=grok

All 3 tests passed. This route reads the installed native Grok session corpus and local models.dev catalog through the shipped bounded scanner/snapshot/spend path; it does not inject temporary session fixtures. The same suite verifies the populated-surface disclosure Public xAI list-price estimate · not a bill.

I also packaged the exact tree with ./Scripts/package_app.sh: production build, widget packaging, code-sign validation, helper-resource smoke, symlinked-helper smoke, and app-launch smoke passed. The packaged helper reports CodexBar 0.55.0, config validate is OK, and the exact worktree app is running.

Evidence boundary: packaged usage JSON intentionally omits UsageSnapshot.costUsage, and the forced Grok CLI fetch returned a provider error in this environment, so I did not represent it as spend proof. No credentials, identity, session contents, Keychain reads, or browser-cookie imports were emitted.

The remaining owner decision about default estimate semantics is unchanged.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Aug 25, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@clawsweeper clawsweeper Bot added the proof: sufficient Contributor real behavior proof is sufficient. label Aug 25, 2026
…on logs

two ways and expensive in a third.

Wrong tokens: the scanner summed `contextTokensUsed` from `signals.json`, which
is the session's ENDING context-window occupancy, not what it consumed. On a
real machine that reported 653K where actual consumption was 48.0M. Read the
sibling `updates.jsonl` instead, where every `turn_completed` event carries the
turn's real usage, and bucket by the per-line timestamp so a session crossing
local midnight lands in both days.

No cost: `toCostUsageTokenSnapshot` hardcoded nil dollars, and nothing could
have priced a Grok model anyway because `codexModelsDevProviderIDs` had no
`xai`. Add it, and resolve `grok-<version>-build` onto its base catalog model —
the `-build` suffix is an artifact of the responses-API surface, not a separate
SKU. `grok-build-0.1` is a real model and is never rewritten. Cost is the public
xAI card via models.dev, provenance `.listPriceEstimate`, so Grok stays
comparable with Claude and Codex. grok's own `costUsdTicks` is deliberately not
used for display.

A turn's `usage` is the aggregate of `modelCalls` API calls, so tiering on the
turn total would push nearly every multi-call turn into the >=200k bracket.
Price on the per-call average instead, in closed form over the two synthetic
call groups. This under-tiers slightly when context grows within a turn
(measured ~4% below the vendor's own accounting on a 27-turn sample, against
~+10% for aggregate tiering); the trade is documented at the call site and
pinned by a test.

Main-actor cost: the scan ran synchronously inside `@MainActor UsageStore` on
every menu-card build, refresh and dashboard load. It now reads the projection
the async probe already produced, and the remaining fallback scans on a
detached task with one scan in flight at a time. The probe projects the maximum
window and consumers narrow it, so `costUsageHistoryDays` and the dashboard's
365-day request are both honoured.

Hardening: `modelCalls` comes from a file, so it is validated before it can size
any work; parsing is cached per (path, size, mtime) with entries evicted when a
file is no longer visited; the cache lock is not held across file reads.

Note for upgraders: adding `xai` to `codexModelsDevProviderIDs` changes the
Codex pricing-cache key, so the first launch after this re-prices existing Codex
history once. Same one-time cost as when kimi and deepseek were added.
OpenCodex sends inference straight to api.x.ai using the Grok account's OAuth
credentials, so it burns the same SuperGrok subscription the Grok provider
reports on. It only spawns the `grok` binary to refresh tokens, so those
requests never reach ~/.grok/sessions and the local session scanner cannot see
them — 1,435 requests on one real machine that CodexBar attributed to nothing.

Route the `xai` provider prefix to the Grok subscription, the same way `openai`
already routes to Codex. Like that mapping, this routes on the prefix and does
not distinguish OAuth from API-key traffic. The `-build` suffix seen in the data
is a responses-API protocol artifact, not a separate billing pool, so traffic is
not split by it.

Routing alone would have produced tokens with no dollars. The aggregator priced
the bare `entry.model`, and a name without a route prefix is resolved against
the `openai` provider — which is why `gpt-5.6-sol` prices today and `grok-4.6`
resolved to `openai/grok-4.6` and missed. Qualify an unprefixed model with its
provider before pricing. Codex rows are unaffected (the qualified name resolves
to the same target), and providers outside the supported set keep returning nil.
Grok resolved list prices straight out of the cached models.dev catalog, but
nothing in its path ever fetched that catalog. The only fetch trigger is
CostUsageFetcher.refreshPricingIfAllowed, which is gated to Codex and Claude —
and Grok never reaches it at all, because its snapshot comes from the provider
probe rather than the shared token-cost pipeline.

On a machine where Codex or Claude is also enabled the cache is already there,
so this is invisible. Enable only Grok and the file never appears: every price
lookup returns nil and the Cost row shows tokens with no money, permanently.

Request ModelsDevPricingPipeline.refreshIfNeeded from the Grok scan paths. It is
safe to call repeatedly — it returns immediately unless the cache is stale and
serialises through its own coordinator — and it is detached rather than awaited,
matching how the Codex and Claude paths already treat it: pricing availability
must never delay or fail a local scan, and the next refresh fills in the value.

`summarize` stays synchronous and side-effect free; the refresh lives in a
wrapper so the parse-cache behaviour and existing tests are untouched.

Reported as P2 by the automated review on the pull request.
… proof

The opt-in live proof scanned real sessions but printed tokens only, which
cannot evidence the half of this change that is about money. It now also
reports today's and the window's list-price cost, the provenance, the window
actually used, and how many days carried a price versus tokens — so an
all-unpriced result is visible in the output instead of reading as zero.

Still skipped unless CODEXBAR_LIVE_GROK_CATALOG_PROOF=1.
The regression guard drove a single failing refresh after a local publication
existed. The defect it covers is specifically about the *second* failure: the
first one publishes through the fallback scan, and only the next one arrives
with a publication already in place — which is what used to hit the generic
clear branch. Drive the failure twice and assert the row and the scan count
both hold.
Routing every OpenCodex `xai` record to the Grok subscription is right for the
case that motivated it — traffic authenticated with the user's Grok account,
which is what makes it burn the SuperGrok quota. It is wrong for anyone using an
xAI API key: their pay-as-you-go developer-platform spend gets folded into the
subscription row, silently inflating it. CodexBar models that platform as its
own xAI provider precisely to keep the two apart.

The usage log carries no per-record credential evidence, so the decision has to
come from the OpenCodex provider config, which records `authMode` per provider.
Read it, and attribute to Grok only when that mode is OAuth; anything else is
token-only spend that belongs to no tracked subscription.

Fail closed: a missing or malformed config, no `xai` entry, or an absent
`authMode` all count as no OAuth evidence and keep the records off the Grok row.
The dispatcher stays a pure function — the set of OAuth-backed provider ids is
threaded in from the caller rather than read at the routing site — and the gate
applies only to `xai`, leaving the other routes exactly as they were.

Also records why the Grok pricing refresh stays fire-and-forget: the parse cache
holds parsed turns rather than prices, so the next scan reprices against the
refreshed catalog, and plumbing completion back to republish was judged
disproportionate to a delay Codex and Claude already share.

Raised as P1 by the automated review; the owner chose verifiable attribution
over prefix-only routing.
@olddonkey

Copy link
Copy Markdown
Contributor Author

Rebased PR #3135 onto current upstream main (41d904fd6) and force-pushed exact head fcdffd1077147dc538ddd1412f37fac7336acacf.

  • Preserved the new main-thread safety work from Keep Grok session scans off the menu thread #3198: Grok local session I/O remains on the detached utility scan path, while failed remote refreshes still schedule the single-flight completed-turn refresh.
  • Migrated the latest-main Grok projection regression from metadata-only signals.json input to completed-turn updates.jsonl fixtures.
  • Regenerated the Codex parser hash (678fd59821eccb04) and added the previous main hash as a compatible predecessor; the adoption regression passes.
  • Scoped Grok parse-cache decode assertions to their own fixture tree so the full suite is not order-dependent.

Exact-head validation:

  • make check: passed; SwiftFormat 0/2002, SwiftLint 0 violations in 2001 files.
  • make test: 933/933 selections, 78/78 groups passed on the first attempt, 0 failures/retries/timeouts (648.5s).
  • Live native-session proof: 3/3 passed; 2,739,923 last-30-day tokens and $2.181282 list-price estimate, with listPriceEstimate provenance and the non-bill disclosure.
  • ./Scripts/package_app.sh: production build, widget packaging, signing validation, resource probes, and launch smoke passed; the exact worktree app remained running.
  • GitHub exact-head CI: all checks passed, including both macOS shards, Linux x64/ARM64/musl, lint, aggregate, and GitGuardian.

The PR body now carries the rebased exact-head evidence. ClawSweeper is already reviewing fcdffd107; the owner product decision on default disclosed estimate semantics remains unchanged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. P2 Normal priority bug or improvement with limited blast radius. proof: sufficient Contributor real behavior proof is sufficient. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants