Skip to content

feat(acceptance): validate the public entry points against live infrastructure - #34

Merged
AlexMikhalev merged 1 commit into
mainfrom
feat/public-release-acceptance
Sep 27, 2026
Merged

AlexMikhalev merged 1 commit into
mainfrom
feat/public-release-acceptance

Conversation

@AlexMikhalev

Copy link
Copy Markdown
Contributor

Purpose

Adds a repeatable acceptance run for the public release entry points, and
records what an attempt to publish the dependency family actually proved.

scripts/acceptance-public-release.py

Read-only, and reports pass, fail or not-executed per entry point, so an
environment limitation is never counted as success. Checks:

  • the documented installer, run against a scratch directory: the installed
    binary must report the channel version and the run must report checksum
    verification
  • every advertised archive against its manifest size and SHA-256
  • the documented Homebrew formula resolves in the tap
  • the crates.io versions a cargo install would serve

Exit codes: 0 all executed checks passed, 1 a check failed, 2 the
installer check failed, 3 nothing was reachable.

Current live result:

Checks
  SKIP  documented installer: disabled
  PASS  channel manifests: 3 binaries at 1.21.16, 20 archives byte-verified
  SKIP  homebrew formula: brew not available on this host
  PASS  crates.io cli family: 3 crates at None

The installer check reports not-executed rather than a false pass until
terraphim/terraphim-ai#964 is merged, because the published script is still the
pre-fix one. The Homebrew check reports the same on a host without brew.

What the family publish proved

Publishing the dependency family was attempted against live infrastructure. A
dry run of the full client family on main
(run 36314289075)
reaches terraphim_config 1.20.4 and fails to compile:

error[E0432]: unresolved import `terraphim_automata::parse_markdown_directives_dir`
 --> terraphim_config-1.20.4/src/lib.rs:24:21
  note: the item is gated behind the `fs-traversal` feature
error: failed to verify package tarball

The chain, verified end to end:

  1. cargo install on a client crate resolves against crates.io
  2. crates.io carries terraphim_config and terraphim_persistence at 1.20.4,
    while the client crates require 1.20.2 (pinned exactly, because 1.20.4 is
    yanked on the Gitea registry)
  3. crates.io's terraphim_automata is 1.21.1, which gates
    parse_markdown_directives_dir behind fs-traversal, a feature
    terraphim_config 1.20.4 does not enable
  4. The build of terraphim_config 1.20.4 fails, so no client crate is
    publishable

This is the failure the [patch.crates-io] block already documents, reached
through the publish pipeline itself. The repair lives in terraphim-core and
terraphim-config-persistence, not here.

The same run confirmed two smaller facts: publish-crates.yml refuses
terraphim_agent before mutating any manifest (#95), as designed; and
terraphim_update 1.20.2 and terraphim_command_runtime 0.1.0 already exist on
crates.io, so they are no-ops in a future crate_list.

Contents

  • scripts/acceptance-public-release.py (new)
  • docs/plans/ — the research and design documents for the public-release
    items, including the addendum recording the above

Testing

python3 scripts/acceptance-public-release.py runs against live infrastructure;
the run above is its actual output. The script is also exercised with
--skip-installer and --skip-crates.

@AlexMikhalev
AlexMikhalev force-pushed the feat/public-release-acceptance branch 2 times, most recently from 741d0fd to caa7791 Compare September 27, 2026 11:36
…structure

Adds the repeatable acceptance run and records what the family publish proved.

scripts/acceptance-public-release.py is read-only and reports pass, fail or
not-executed per entry point, so an environment limitation is never counted as
success. It runs the documented installer against a scratch directory and
checks the installed binary reports the channel version, verifies every
advertised archive against the manifest size and SHA-256, resolves the
Homebrew formula, and reports the crates.io versions a cargo install would
serve.

The plans under docs/plans/ gain the measured result of the family publish.
A dry run of the full client crate family on main (run 36314289075) reaches
terraphim_config 1.20.4 and fails to compile:

    unresolved import terraphim_automata::parse_markdown_directives_dir
    the item is gated behind the `fs-traversal` feature

crates.io's terraphim_automata is 1.21.1, which gates that symbol, while
crates.io's terraphim_config is 1.20.4, which does not enable the feature, and
the client crates require config 1.20.2. No client crate is publishable until
that is repaired, and the repair lives in terraphim-core and
terraphim-config-persistence, not here.

Run 36314289075 also confirmed that publish-crates.yml refuses terraphim_agent
before mutating any manifest (#95), and that terraphim_update 1.20.2 and
terraphim_command_runtime 0.1.0 already exist on crates.io.

Verified: acceptance script compiles; record manifest check passes (3 binaries
at 1.21.16, 20 archives byte-verified); the failing dry run is cited by id.
@AlexMikhalev
AlexMikhalev force-pushed the feat/public-release-acceptance branch from caa7791 to ea6d880 Compare September 27, 2026 13:12
@AlexMikhalev
AlexMikhalev merged commit 299f0b5 into main Sep 27, 2026
2 checks passed
AlexMikhalev added a commit that referenced this pull request Oct 4, 2026
* fix(ci): add --features enrichment to Gitea native CI Refs #2171

Mirror the GitHub CI enrichment steps in the Gitea native-ci workflow:
- cargo clippy -p terraphim_sessions --features enrichment -- -D warnings
- cargo test -p terraphim_sessions --features enrichment --lib --no-fail-fast

All 67 tests pass locally including the 3 enrichment tokio tests and
2 concept unit tests that were silently skipped before.

* test(grep): add default-feature zero-chunk smoke guard

Adds an explicit, named CI guard for the silent zero-chunk regression
documented in terraphim/terraphim-ai#3025 / #4325. A default-feature
build of terraphim_grep must return non-zero chunks for a matching
query; the test fails loudly if `code-search` is ever removed from the
`default` feature set (which would compile search_code() to a no-op
stub returning Ok(vec![]) -- success-with-zero-items).

- crates/terraphim_grep/tests/default_feature_smoke.rs: new integration
  test. Distinct from no_thesaurus_cli.rs (KG-absent fallback) -- this
  test's single purpose is the default-feature contract.
- .github/workflows/ci.yml: named, visible smoke-test step after the
  workspace test run (functional gating already existed via
  `cargo test --workspace`; this makes it explicit and documented).

Verified bi-directionally this session (real binaries, no mocks):
- default features (code-search on): chunks_returned=1 -> PASS
- code-search removed from default: chunks_returned=0 -> FAIL with the
  named regression message -> revert -> PASS again.

Gates: fmt clean; clippy -p terraphim_grep --all-targets -D warnings 0;
full crate test suite green.

Refs terraphim/terraphim-ai#4325 (cross-repo: terraphim_grep lives in
terraphim-clients, extracted via #1910).

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(quality): capture verification + validation reports for PRs #44, #45, #49, #51, #52, #59, #60

- PR #44: Insufficient KG propagation (#2721)
- PR #45: rust-engineer shortname fix (#2723)
- PR #49: CI enrichment feature (#2171)
- PR #51: thesaurus NotFound ERROR suppress (#48)
- PR #52: Gitea CI enrichment feature (#2171)
- PR #59: grep default-feature smoke test (#4325)
- PR #60: terraphim_grep crates.io publishable metadata (#58)

Also refreshes Cargo.lock to register env_logger as a terraphim_agent
runtime dependency introduced by PR #51 (Cargo.lock had drifted from
Cargo.toml since the merge).

Refs #108

* feat(memory): scaffold terraphim-agent memory CLI namespace with 10 subcommands Refs #1899

Add  top-level command wrapping the eight-stage agentic memory
lifecycle behind a single discoverable CLI surface.

Subcommands:
- capture (routes to learn hook)       - validate (rubric scorer stub)
- distill (routes to learn compile)    - retire (learned-rules stub)
- scope (role/project boundaries)      - rubric (6-dimension diagnostic)
- provenance (routes to sessions)      - second-run (token delta)
- retrieve (routes to search)
- apply (routes to terraphim_hooks)

Seven subcommands are routing stubs delegating to existing handlers.
Three (validate, rubric, second-run) are placeholder stubs for net-new
code to be wired in subsequent steps.

Implemented in crates/terraphim_agent/src/main.rs following existing
Command/Subcommand enum pattern (Memory variant in Command enum,
MemorySub enum with clap derive, run_memory_command async handler).

Tests: cargo check passes, 246 lib tests pass, all 10 subcommands
respond to --help and execute their handler stubs.

* feat(memory): wire terraphim_agent_evolution into CLI with capture/list/show/export/scope Refs #1899

Add terraphim_agent_evolution v1.20.2 from terraphim registry as dependency.
Implement real memory lifecycle operations backed by AgentEvolutionSystem:

- capture: write MemoryItem to evolution store with provenance tags
- list: enumerate memory items with optional item_type filter
- show: display full details of a memory item or lesson by ID
- export: dump memory items + lessons as JSON or markdown
- scope: show role/project KG boundaries from ~/.config/terraphim/kg

capture/scope are now real implementations instead of routing stubs.
Three new subcommands (list, show, export) added for evolution inspection.

Tests: 246 lib tests pass. cargo check clean with 0 warnings.

* feat(memory): implement rubric scorer, second-run signal, policy doc, README Refs #1899

Implement all remaining memory lifecycle components:

Rubric subcommand:
- 6-dimension rule-based scorer: faithfulness, scope, provenance,
  actionability, decay, risk
- Composite score with weighted average (faithfulness 0.30,
  actionability 0.25, scope 0.15, provenance 0.10, decay 0.10,
  risk 0.10 inverted)
- Markdown readout with dimension table, top-3 offenders,
  recommended retirements
- RubricScore struct + score_memory_item/scoring helpers

Validate subcommand:
- Scores all, specific, or most recent memory items
- Per-item scores with composite average

Retire subcommand:
- Writes retirement proposals to ~/.config/terraphim/
  (learned-rules-retirements.md) with CTO approval flag

Second-run subcommand:
- Reads ADF artefact directory (~/.cache/terraphim/adf-artefacts/)
- Finds JSON run metrics files per Gitea issue
- Computes token delta, retry delta, wall-time delta
- Emits structured JSON with interpretation

MEMORY_POLICY.md:
- Public commons vs permissioned memory boundary
- Storage location table
- Enforcement via memory scope --check

README:
- Full memory command reference (13 subcommands documented)
- Link to MEMORY_POLICY.md

Tests: 246 lib tests pass. cargo check clean.

* feat(memory): add cross-invocation persistence via JSON file store Refs #1899

Add load_evolution()/save_evolution() helpers that persist
AgentEvolutionSystem state as JSON to:
  ~/.config/terraphim/evolution/cli-agent.json

- MemoryState and LessonsState are serialized/deserialized
- Captured items survive CLI invocations
- List/show/export/rubric consume persisted state
- save_evolution() called after every capture mutation

Verified: capture → list shows item, capture → rubric scores,
capture → export produces markdown. All 246 tests pass.

* fix(memory): address PR review findings -- P1 capture semantics, P2 dedup scoring Refs #1899

P1: Capture now stores provenance_tag as a tag, not as content.
     Content gets a descriptive string. Tags field receives
     'provenance:<tag>' entries. Help text updated accordingly.

P2: Extract compute_decay() and compute_risk() shared functions
     to eliminate duplicate scoring logic between score_memory_item
     and the rubric retirement filter. Add TERRAPHIM_ADF_ARTEFACTS_DIR
     env var override for second-run artefact path.

* fix(grep): resolve project thesaurus by role shortname

* fix(grep): rank KG matches above substring metadata

* feat(grep): add update commands via shared updater

* fix(update): rotate zipsign release verifier key

* fix(update): preserve release asset name for verification

* fix(agent): use clients release assets for updates

* fix(rebase): repair post-rebase fallout in PR #61

The rebase of task/1899-memory-lifecycle-cli onto current main
d54f28f left several files in a non-compiling state due to upstream
renames (terraphim_update, terraphim_automata 1.21.0 borrows &Thesaurus,
thesaurus registry pins from Refs #112).

- crates/terraphim_grep/Cargo.toml: drop the path-dep duplicate
  left by the rebase; keep the Refs #112 registry-bearing entry.
- crates/terraphim_grep/src/hybrid_searcher.rs:
  - bind the kg_concepts result of search_kg (was dropped by an
    orphan semicolon);
  - pass &thesaurus to find_matches (1.21.0 borrows it);
  - iterate over all search_paths, apply boost_chunks_with_kg
    before truncating, so KG-ranked chunks survive the candidate
    cut.
- crates/terraphim_grep/src/main.rs: restore the function
  signatures that the conflict resolution collapsed (push_unique_candidate,
  discover_project_thesaurus) and add the missing #[test] attributes
  on the two post-resolution tests.
- crates/terraphim_update/src/signature.rs: remove an orphan closing
  brace left after the conflict on get_embedded_public_keys.

Verified locally: cargo check --workspace --all-features, cargo clippy
--workspace --all-features --all-targets -- -D warnings, and cargo test
on the affected crates (terraphim_grep, terraphim_update, terraphim_sessions)
all pass. Pre-existing Refs #113 failures in cross_mode_consistency_test,
mcp_server integration tests, and terraphim_update::test_resolve_asset_url_against_local_manifest
are unchanged.

Refs #1899

* fix(update): revert redundant with_repo to unbreak packaged install gate

PR #61 added UpdaterConfig::with_repo and called it from agent + grep
with the constructor's own default values (terraphim/terraphim-clients).
The packaged_install_graph_regression test (cargo package + cargo install
--path) resolves terraphim_update from the terraphim registry, where the
published 1.20.2 lacks with_repo, so the install failed to compile.

The design's own rollback plan says to revert with_repo if no consumer
uses it functionally. Here every caller passed the same value the
constructor already sets, so the calls were no-ops; the method had no
real consumer and would have required a registry publish to keep CI
green. Reverting restores install-graph correctness and aligns the
branch with main's updater API.

Refs #1899

* docs(quality): add PR #61 verification report

Captures the 8 native-ci steps, gitea native-ci status check, traceability
matrix for Refs #1899 / #95 / #62 / #112, and the defect register (the
with_repo rollback plus two rebase-fallout defects resolved before
merge). Refs #1899

* docs(quality): add PR #61 validation report

Acceptance criteria trace from Refs #1899 / #95 / #62 / #112, performance
and security review notes, defect register (the three closed defects from
the verification report plus follow-ups). Refs #1899

* docs(quality): PR #84 verification + validation -- documents 57 pre-existing failures exposed by --all-targets

Refs #84, closes part of #108 (campaign summary)

Findings (full report in .quality/pr-84-{verification,validation}.md):
- 57 unique test failures surface on Linux CI when --all-targets replaces --lib
- 37 are pre-existing (also fail on macOS main baseline); 20 are Linux-only
- Pre-existing failures are tracked by #113 (PR #135 in progress) and
  separately by the missing docs/src/kg fixture path bug
- PR's empirical claim 'does not destabilise the pipeline' does not hold on
  the Gitea runner because is_ci_environment() does not recognise the runner
  (no CI=true / GITHUB_ACTIONS / dockerenv / root)
- Recommendation: do not merge PR #84 as-is; merge PR #135 first or extend
  PR #84 with env: CI: true plus minimal #[ignore] for the MCP server-binary
  tests, plus fix the wrong <workspace>/docs/src/kg path in replace_feature_tests

* test(fixtures): sync terraphim_server test fixtures from terraphim-ai v1.21.3 (Refs #113)

The integration tests under crates/terraphim_agent/tests/ that depend
on a real terraphim_server binary reference fixtures under
terraphim_server/{default,fixtures}/ that lived only in the
terraphim-ai repo. Sync the minimal set required by the 7 #113 tests
(2 in cross_mode_consistency_test, 2 in integration_tests, 3 in
kg_ranking_integration_test) so the tests can run from terraphim-clients
without fetching from terraphim-ai at test time.

Files synced (verbatim copy from terraphim-ai v1.21.3):
- terraphim_server/default/terraphim_engineer_config.json
- terraphim_server/fixtures/cross_mode_test_config.json
- terraphim_server/fixtures/term_to_id.json
- terraphim_server/fixtures/thesaurus_Default.json
- terraphim_server/fixtures/haystack/Engineer_thesaurus.json
- terraphim_server/fixtures/haystack/System Operator_thesaurus.json
- terraphim_server/fixtures/haystack/*.md (16 files)

Source: https://git.terraphim.cloud/terraphim/terraphim-ai @ v1.21.3

Refs #113

* fix(agent): make PreToolUse hook rewriting opt-in, default-on guard

The PreToolUse hook silently rewrote destructive commands when their
substring matched a thesaurus entry (e.g. `rm -rf /tmp/foo` could
become `rm -Readiness Feedback /tmp/foo`). This is dangerous because:

  * Substitution was always-on, never warned the user.
  * A typo'd KG entry or stray synonym could mutate an unrelated
    destructive command, blowing away files instead of running the
    user-intended action.

This commit flips the default to safe-by-construction:

  * New `--rewrite` flag (default false) opts in to KG substitution.
    When false, the hook still probes for replacements but emits a
    `warnings` array entry describing the suppressed match.
  * Guard check defaults to ON for `pre-tool-use` (still off for
    `post-tool-use`, `pre-commit`, `prepare-commit-msg` because they
    fire after execution or on text inputs).
  * `--no-with-guard` flag explicitly disables the guard (clap does
    not auto-derive `--no-with-guard`; explicit override required).
  * `with_guard` and `no_with_guard` are mutually exclusive via
    `conflicts_with`.

Adds `tests/hook_safety.rs` with five regression tests covering:

  * `rm -rf /tmp/foo` passes through with a warning (the original bug).
  * `rm -rf /` is denied by the default guard.
  * `--rewrite` enables substitution as before.
  * `--no-with-guard` suppresses the guard envelope.
  * `/tmp/` allowlist still allows `rm -rf /tmp/foo`.

Refs #126

* feat(agent): --explain for guard, README robot-mode fix, priority docs

Refs #127, #129

#127 (README robot-mode example)

  The example `terraphim-agent search "retry policy" --robot --format json`
  was wrong. `--robot` and `--format` are global flags on the top-level
  Cli struct and must precede the subcommand. Placed them after the
  subcommand clap fails with `error: unexpected argument '--robot' found`.
  Updated the README with both the corrected example and a parenthetical
  explaining the failure mode.

#129 (guard --explain + priority docs)

  * New `--explain` flag on `terraphim-agent guard` that prints the
    per-stage evaluation trace (allowlist > destructive > suspicious >
    default). With `--json` it emits a structured `GuardTrace`; without
    `--json` the trace goes to stderr in a human-readable form.
  * New `CommandGuard::check_with_trace` returns a `GuardTrace`
    containing the final `GuardResult` plus per-stage outcomes
    (allow/block/sandbox/continue/no_match) and matched terms.
  * README now documents the priority order with a worked example
    showing the allowlist short-circuiting destructive.
  * Added `tests/guard_priority.rs` with 5 regression tests:
      - allowlist_short_circuits_before_destructive
      - destructive_short_circuits_before_suspicious
      - default_allow_path_emits_default_stage
      - explain_exits_one_on_block
      - explain_text_output_is_human_readable

Both tests (`hook_safety` and `guard_priority`) and the existing
`learn_no_service_tests` pass. Clippy clean with -D warnings.

* test(grep): pin --search-only flag and LLM-client skip

Refs #128

The installed `~/.cargo/bin/terraphim-grep` was 1.21.11 while the
workspace source is 1.21.13. The 1.21.11 binary is missing the
`--search-only` flag (Refs terraphim-clients#81) and several other
recent fixes (chunks_returned counter, etc.). Rebuilt and installed
the 1.21.13 release binary.

Adds `tests/search_only_flag.rs` with three regression tests so CI
catches a future drift between source and installed binary:

  * search_only_flag_is_accepted -- the flag is parsed.
  * search_only_skips_llm_client_with_openrouter_key_present --
    even with a stray API key, --search-only must skip the LLM client
    build (asserted by checking stderr for the
    "skipping LLM client setup" debug log).
  * help_documents_search_only_flag -- defensive: --help must mention
    --search-only so users can discover it.

Clippy clean with -D warnings.

* feat(robot): mark REPL-only commands in schema output

Refs #131

`terraphim-agent robot schemas` previously listed every REPL command
(name, description, args, flags, examples, response_schema) without
flagging which ones are actually available as top-level CLI
subcommands. Consumers parsing the schema expected parity with
`terraphim-agent --help`, which doesn't list REPL-only entries like
`vm` or `chat`. The mismatch was confusing.

Adds a `repl_only: bool` field to `CommandDoc` (defaults to `false`
via serde). Marks the two REPL-only commands called out in #131:

  * `vm`  -- always REPL-only (firecracker-gated, no top-level CLI).
  * `chat` -- REPL-only, feature-gated behind `repl-chat`.

All other 13 commands keep `repl_only: false`. Comments next to the
literals document why each is REPL-only.

Adds `tests/robot_schemas.rs` with four regression tests:

  * every_command_has_repl_only_field
  * top_level_cli_commands_are_not_repl_only
  * vm_is_marked_repl_only
  * chat_is_marked_repl_only (tolerates missing chat when the
    `repl-chat` feature is off in the test binary)

All 4 tests pass with `--features repl-chat`. Clippy clean with
-D warnings.

* docs: ADRs for guard priority and pretool rewrite, reference, blog posts

Refs #130, #132, #133

#130 (reference doc)

  * docs/agent-reference.md — 271-line reference enumerating all 22
    top-level `terraphim-agent` subcommands. The README's "Key
    Commands" table only listed 8; the remaining 14
    (roles, config, kg, extract, replace, validate, suggest,
    interactive, repl, setup, check-update, update, listen, cache)
    now have one worked example each.
  * Quick-reference family grouping at the bottom.
  * Cross-links to the two ADRs and the design doc.
  * Cross-links to the five new blog posts.

#132 (five blog posts)

  * docs/blog/terraphim-agent-sessions.md — Claude Code / Cursor /
    Aider import flow, bounded walker, JSON output schema.
  * docs/blog/terraphim-agent-setup.md — onboarding wizard and the
    10 role templates.
  * docs/blog/terraphim-agent-robot-mode.md — JSON output, exit
    codes (0..7), self-describing schemas, global flag caveat.
  * docs/blog/terraphim-agent-shared-learning.md — markdown-backed
    BM25-deduped learning store with trust levels.
  * docs/blog/terraphim-update-r2-backend.md — R2 manifest format,
    verification path, fallback chain, backend overrides.
  * README "Further reading" section links to all five plus the
    reference doc.

#133 (two ADRs)

  * adr/ADR-002-guard-priority-order.md — rationale for the
    allowlist > destructive > suspicious > default priority with
    fail-open per stage. Includes worked examples and the four
    alternatives that were rejected (destructive-first, voting,
    most-specific-wins, etc.).
  * adr/ADR-003-pretool-hook-rewrite.md — rationale for making
    thesaurus substitution opt-in via `--rewrite`. Documents the
    two modes (default warn-only vs opt-in substitution), the
    residual risk of warnings being dropped by future agent
    runtimes, and the four alternatives that were rejected
    (always-off, always-on, sentinel-in-command, default-on).

Both ADRs follow the same Context / Decision / Consequences /
Alternatives / References format as ADR-001.

* feat(robot): add CLI chat schema, mark summarize REPL-only

The structural-pr-review for #134 flagged two cross-file
inconsistencies in [
  {
    "aliases": [
      "q",
      "query",
      "find"
    ],
    "arguments": [
      {
        "description": "Search query text",
        "name": "query",
        "required": true,
        "type": "string"
      }
    ],
    "description": "Search documents using semantic and keyword matching",
    "examples": [
      {
        "command": "/search async error handling",
        "description": "Basic search"
      },
      {
        "command": "/search database migration --role DevOps --limit 5",
        "description": "Search with role and limit"
      }
    ],
    "flags": [
      {
        "default": "current",
        "description": "Role context for search",
        "name": "--role",
        "short": "-r",
        "type": "string"
      },
      {
        "default": "10",
        "description": "Maximum results to return",
        "name": "--limit",
        "short": "-l",
        "type": "integer"
      },
      {
        "default": "false",
        "description": "Enable semantic search",
        "name": "--semantic",
        "type": "boolean"
      },
      {
        "default": "false",
        "description": "Include concept matches",
        "name": "--concepts",
        "type": "boolean"
      }
    ],
    "name": "search",
    "response_schema": {
      "properties": {
        "concepts_matched": {
          "items": {
            "type": "string"
          },
          "type": "array"
        },
        "results": {
          "items": {
            "properties": {
              "id": {
                "type": "string"
              },
              "preview": {
                "type": "string"
              },
              "rank": {
                "type": "integer"
              },
              "score": {
                "type": "number"
              },
              "title": {
                "type": "string"
              },
              "url": {
                "type": "string"
              }
            },
            "type": "object"
          },
          "type": "array"
        },
        "total_matches": {
          "type": "integer"
        }
      },
      "type": "object"
    }
  },
  {
    "aliases": [
      "c",
      "cfg"
    ],
    "arguments": [
      {
        "description": "Subcommand: show, set",
        "name": "subcommand",
        "required": true,
        "type": "string"
      }
    ],
    "description": "View and modify configuration",
    "examples": [
      {
        "command": "/config show",
        "description": "Show current configuration"
      },
      {
        "command": "/config set selected_role Engineer",
        "description": "Set configuration value"
      }
    ],
    "flags": [],
    "name": "config",
    "response_schema": {
      "properties": {
        "config": {
          "type": "object"
        }
      },
      "type": "object"
    }
  },
  {
    "aliases": [
      "r"
    ],
    "arguments": [
      {
        "description": "Subcommand: list, select",
        "name": "subcommand",
        "required": true,
        "type": "string"
      }
    ],
    "description": "Manage roles",
    "examples": [
      {
        "command": "/role list",
        "description": "List available roles"
      },
      {
        "command": "/role select Engineer",
        "description": "Select a role"
      }
    ],
    "flags": [],
    "name": "role",
    "response_schema": {
      "properties": {
        "current_role": {
          "type": "string"
        },
        "roles": {
          "items": {
            "type": "string"
          },
          "type": "array"
        }
      },
      "type": "object"
    }
  },
  {
    "aliases": [
      "g",
      "kg"
    ],
    "arguments": [],
    "description": "Display knowledge graph concepts",
    "examples": [
      {
        "command": "/graph",
        "description": "Show top concepts"
      },
      {
        "command": "/graph --top-k 20",
        "description": "Show top 20 concepts"
      }
    ],
    "flags": [
      {
        "default": "10",
        "description": "Number of top concepts to show",
        "name": "--top-k",
        "short": "-k",
        "type": "integer"
      }
    ],
    "name": "graph",
    "response_schema": {
      "properties": {
        "concepts": {
          "items": {
            "properties": {
              "count": {
                "type": "integer"
              },
              "term": {
                "type": "string"
              }
            },
            "type": "object"
          },
          "type": "array"
        }
      },
      "type": "object"
    }
  },
  {
    "aliases": [],
    "arguments": [
      {
        "description": "Subcommand: list, pool, status, metrics, execute, agent, tasks, allocate, release, monitor",
        "name": "subcommand",
        "required": true,
        "type": "string"
      }
    ],
    "description": "Manage Firecracker VMs",
    "examples": [
      {
        "command": "/vm list",
        "description": "List VMs"
      },
      {
        "command": "/vm execute python print('hello')",
        "description": "Execute code in VM"
      }
    ],
    "flags": [
      {
        "description": "VM identifier",
        "name": "--vm-id",
        "type": "string"
      }
    ],
    "name": "vm",
    "response_schema": {
      "properties": {
        "status": {
          "type": "string"
        },
        "vms": {
          "type": "array"
        }
      },
      "type": "object"
    }
  },
  {
    "aliases": [
      "h",
      "?"
    ],
    "arguments": [
      {
        "description": "Command to get help for",
        "name": "command",
        "required": false,
        "type": "string"
      }
    ],
    "description": "Show help information",
    "examples": [
      {
        "command": "/help",
        "description": "Show all commands"
      },
      {
        "command": "/help search",
        "description": "Get help for search"
      }
    ],
    "flags": [],
    "name": "help",
    "response_schema": {
      "properties": {
        "commands": {
          "type": "array"
        },
        "help_text": {
          "type": "string"
        }
      },
      "type": "object"
    }
  },
  {
    "aliases": [],
    "arguments": [
      {
        "description": "Subcommand: capabilities, schemas, examples",
        "name": "subcommand",
        "required": true,
        "type": "string"
      }
    ],
    "description": "Robot mode commands for AI agents",
    "examples": [
      {
        "command": "/robot capabilities",
        "description": "Get capabilities"
      },
      {
        "command": "/robot schemas search",
        "description": "Get schema for search"
      }
    ],
    "flags": [
      {
        "default": "json",
        "description": "Output format: json, jsonl, minimal, table",
        "name": "--format",
        "short": "-f",
        "type": "string"
      }
    ],
    "name": "robot",
    "response_schema": {
      "type": "object"
    }
  }
]:

P1: `Command::Chat` in main.rs is a top-level CLI subcommand
    gated by `--features llm` (default-on), but the only
    `chat` schema entry was the REPL `chat` (gated by
    `--features repl-chat`) marked `repl_only: true`.
    A downstream agent introspecting via `robot schemas` would
    conclude there is no top-level `chat` subcommand when there
    actually is one. Added a separate `CommandDoc` for the CLI
    `chat` (gated by `#[cfg(feature = "llm")]`) with
    `repl_only: false`, the required `prompt` argument, and
    `--role` / `--model` flags. The REPL entry is unchanged.

P2: `summarize` was marked `repl_only: false` but it is
    REPL-only (registered in `repl::commands`, no top-level
    CLI subcommand). Flipped to `repl_only: true`.

Both fixes are now visible in `robot schemas` output. The
`chat` descriptions are also disambiguated (CLI one-shot vs
REPL interactive) so a downstream consumer can pick the right
entry without reading the `repl_only` field first.

Refs terraphim-clients#134 P1, P2 (summarize)
Refs docs/plans/design-terraphim-grep-agent-fixes-2026-08-30.md

* test(agent): pin chat and summarize repl_only correctness

Pins the cross-file consistency for #134:

- `top_level_cli_commands_are_not_repl_only`: now includes
  `chat` and filters by `repl_only == false` so the assertion
  picks the CLI entry (and is not double-counted by the REPL
  one in `repl-chat` builds).
- `repl_chat_is_marked_repl_only` (renamed from
  `chat_is_marked_repl_only`): filters by `repl_only == true`
  so it survives a `repl-chat` build that also contains the
  new CLI `chat` entry.
- `cli_chat_is_marked_not_repl_only` (new): asserts the CLI
  `chat` is present with `repl_only: false` and that the
  `prompt` argument is required. Catches a future edit that
  silently drops the required argument or flips the flag.
- `summarize_is_marked_repl_only` (new): asserts `summarize`
  is `repl_only: true` when present (gated by `repl-chat`).

All four pass in default builds and in `--features repl-chat`
builds (6/6 tests in each).

Refs terraphim-clients#134 P1, P2 (summarize)

* refactor(agent): extract GuardTrace::print, make check delegate to check_with_trace

Two related simplifications in `terraphim-agent guard`:

P2.2: The `--explain` trace formatting was duplicated verbatim
      across `run_offline_command` (lines ~2018-2048) and
      `run_server_command` (lines ~4928-4952) — about 30 lines
      each, differing only in the `*fail_open` deref. Drift risk.
      Extracted `GuardTrace::print(json: bool)` in
      `guard_patterns.rs`; the two call sites collapse to
      `trace.print(*json)?`.

P2.3: `CommandGuard::check` and `CommandGuard::check_with_trace`
      had ~95% identical bodies (`check` returns `GuardResult`,
      `check_with_trace` returns `GuardTrace`). Made `check`
      a one-line wrapper:
      `pub fn check(&self, command: &str) -> GuardResult {
          self.check_with_trace(command).result
      }`
      The trace is cheap to build (`Vec<4>` plus three
      Aho-Corasick matches), so centralising the pipeline
      eliminates ~70 lines of duplication.

Verified: 5/5 `guard_priority` tests, 5/5 `hook_safety` tests,
6/6 `robot_schemas` tests, 3/3 `search_only_flag` tests pass.
Clippy clean with `-D warnings` on
`--all-targets --features repl-chat`.

Refs terraphim-clients#134 P2.2, P2.3

* docs: clarify REPL vs CLI chat in robot-mode blog and reference

The robot-mode blog post and the 271-line agent reference both
documented `chat` as a single REPL-only command behind
`--features repl-chat`, omitting the top-level CLI `chat`
subcommand (gated by `--features llm`, default-on).

Update both to:

- Mention both flavours and the `repl_only` flag that
  distinguishes them.
- Show both `chat` entries in the `robot schemas` example
  output (with `false` and `true` rows).
- Note that `--features repl-chat` transitively enables `llm`,
  so a `repl-chat` build carries two `chat` schemas.
- Add a `chat` (CLI) section to `docs/agent-reference.md`
  before the existing `chat` (REPL-only) section, with the
  one-shot `prompt` argument and `--role`/`--model` flags.

Refs terraphim-clients#134 P1

* ci(native-ci): build terraphim_server from terraphim-ai and run #113 integration tests (Refs #113)

terraphim_server is not a workspace member of terraphim-clients; it
lives in the private terraphim-ai repo. Build it from the v1.21.3 git
tag with cargo install --git so the runner has a real binary on disk,
then export TERRAPHIM_SERVER_BIN and run the three integration test
targets whose ensure_server_binary() / server_binary_path() helpers
look up the binary by that env var first.

Three landmines in 'cargo install --git', in order:

1. Multiple-binary repo. terraphim-ai has firecracker, dsm,
   gitea_runner, merge_coordinator, server, eval_check -- refuses --bin
   unless a positional <PACKAGE> selects one. Fix: trailing 'terraphim_server'.

2. Isolated context. cargo install does NOT inherit the workspace's
   .cargo/config.toml, but terraphim-ai v1.21.3 has [patch.crates-io]
   entries with registry = 'terraphim'. Without the registry declared,
   parsing the cloned manifest fails: 'registry index was not found
   in any configuration: terraphim'. Fix: --config 'registries
   .terraphim.index=...' and --config 'registry.global-credential
   -providers=[cargo:token]'.

3. Patch resolution. The v1.21.3 patches use caret ranges ('version
   = "1.20.2"'), and the registry now publishes both 1.20.2 and
   1.21.0. cargo refuses: 'patch for terraphim_automata resolved to
   more than one candidate: 1.20.2, 1.21.0'. Fix: --locked honours
   the v1.21.3 Cargo.lock, which pins terraphim_automata and
   terraphim_types to exactly 1.20.2.

Also patch kg_ranking_integration_test.rs::ensure_server_binary() to
honour TERRAPHIM_SERVER_BIN (mirrors the cross_mode_consistency_test
helper; the kg_ranking helper had the right error message but never
actually read the env var).

This makes the server-binary-dependent integration tests run on every
push instead of silently failing fast because the binary is missing:
- crates/terraphim_agent/tests/cross_mode_consistency_test.rs (2)
- crates/terraphim_agent/tests/integration_tests.rs (5)
- crates/terraphim_agent/tests/kg_ranking_integration_test.rs (3)

Verified locally: install succeeds (4m08s cold), 2/2 + 5/5 + 1/3 pass
on the install side. The remaining 2 kg_ranking failures are pre-existing
test-data setup issues (docs/src haystack does not exist; tracked in
#84's 57 pre-existing failures).

Refs #113

* style(robot,grep,agent): apply rustfmt to cherry-picked commits (Refs #135b)

Run 278 / job 61817 failed cargo fmt --all -- --check on the
scope-creep commits cherry-picked from PR #135. The diffs are
all line-wrapping cleanups that rustfmt wants:

- crates/terraphim_agent/src/main.rs (GuardTrace::print refactor)
- crates/terraphim_agent/tests/guard_priority.rs
- crates/terraphim_agent/tests/hook_safety.rs
- crates/terraphim_agent/tests/robot_schemas.rs
- crates/terraphim_grep/tests/search_only_flag.rs

cargo fmt --all brings the tree back into compliance; no
semantic changes.

Refs #135b

* test(kg-ranking): point haystack at synced fixtures, replace missing 'python' term

Refs #113

The test_config.json haystacks referenced docs/src/, which exists in
terraphim-ai but not in this repo (terraphim-clients). After syncing
fixtures from terraphim-ai v1.21.3, the engineer's markdown content
lives at terraphim_server/fixtures/haystack/. Pointing both the
"Test Engineer" and "Default" Ripgrep haystacks at that path lets
the Ripgrep scorer pick up actual .md content instead of returning
zero results.

test_term_specific_boosting also searched for the term 'python', which
is not present in the terraphim-ai v1.21.3 fixture corpus (only
'rust', 'machine_learning', and 'neural_networks' markdown files are
synced). Swapping 'python' for 'neural networks' keeps the same test
intent (three varied terms) while using terms that genuinely exist
in the haystack.

Verified locally with TERRAPHIM_SERVER_BIN pointing at the cached
cargo install binary:
- test_role_switching: ok
- test_term_specific_boosting: ok (3/3 terms returned results)
- test_knowledge_graph_ranking_impact: ok

* test(cross-mode): un-ignore test_role_consistency_across_modes via per-role pre-warm (Refs #113b)

Removes the last `#[ignore]` in the integration suite. The cold-cache
slowness is paid exactly once per role, so warming up each role before
the timing-critical loop makes the 30s default ApiClient timeout
sufficient on CI. Results from the warm-up queries are discarded; only
the count-consistency assertions in the real loop are checked.

The warm-up uses `limit: Some(1)` so each call only pays the cold-cache
cost, not full pagination. If even a warm-up query times out, the test
fails loudly with the underlying transport error (no silent reliance on
a longer timeout).

Verified locally:
- test_mode_specific_verification ... ok (unchanged)
- test_cross_mode_consistency ... ok (unchanged)
- test_role_consistency_across_modes ... ok (un-ignored)
- 3 passed; 0 failed; 0 ignored; total 25.7s

* test(cross-mode): drop unnecessary explicit deref flagged by clippy

Removes the `*` deref of `warm_role` in the warm-up SearchQuery
construction. `warm_role: &&str` already auto-derefs to `&str` at
the `RoleName::new` call site; clippy::explicit_auto_deref (implied
by -D warnings) refused the explicit deref.

* fix(terraphim_agent): filter sub-word matches in extract (Refs #46)

Re-applies 31baedf8 onto current main. extract reported phantom term
labels (e.g. 'learning path' for the two-letter abbreviation 'lp'
matching inside 'aLPha') and started paragraphs at mid-word byte
offsets, because short thesaurus terms match as substrings via
Aho-Corasick with no word-boundary check.

extract_paragraphs now drops matches not flanked by word boundaries,
and labels each surviving paragraph with the surface form actually
present (Matched.term) rather than the concept expansion. Also
corrects the REPL/MCP extract path which shares this method.

Adds unit tests in service.rs covering: standalone word matches,
sub-word rejection, text-edge cases, and the full phantom-label /
mid-word-offset regression which is the headline defect in #46.

Note: the regression test borrows the thesaurus (`&thesaurus`) because
terraphim_automata 1.21.0 takes `&Thesaurus`. The original commit
passed an owned Thesaurus, which fails to compile against 1.21.0.

* test: provision docs/src/kg fixture + extend manifest fixture (Refs #84) (#141)

* test(terraphim_update): derive manifest fixture + document KG format (Refs #141) (#145)

* ci(terraphim-clients): run all workspace targets in test gate

Second per-repo remediation for terraphim/terraphim-agents#91.
Replace the library-only test gate with all-targets testing so
binaries, examples, and integration tests in tests/*.rs are
exercised. Regression context: 2026-07-31 (cargo test --lib
silently excluded the integration suite). ADF was not restarted;
no orchestrator configuration changed.

Scope: workflow yml only. Mirrors terraphim/terraphim-ai#3159.

* Fix #142: repoint find_files KG-scorer fixture at mcp_server (#12)

Repoint KG-scorer fixture at mcp_server.

find_files_with_kg_scorer_boosts_matching_paths built a thesaurus
containing the term 'automata' and asserted that a path under
crates/terraphim_automata/ would be boosted to the top of the
results. The terraphim_automata crate does not exist in this workspace,
so the assertion always failed.

Repoint the fixture at 'mcp_server' (which has a real
crates/terraphim_mcp_server/ directory) and update the assertion to
look for the matching path segment. The thesaurus still exercises the
KG-scorer boosting path; only the keyword and assertion substring
change.

Refs #142.

(cherry picked from commit b1c8247e0f59d339e6aa2d95307aa47647e568b7)

* Fix #143: hermetic MCP stdio tests (#13)

Make test_tools_list and test_all_mcp_tools deterministic and hermetic:

- Spawn the server with cwd set to a unique temp dir so
  terraphim_config::project::discover() cannot walk up to a host
  .terraphim/.
- Drain stderr on a background thread so the OS pipe buffer never
  fills and SIGPIPEs the server mid-test.
- Drop the leading '--' separator before --verbose (clap Args::parse
  rejects '--').
- Send the notifications/initialized frame between initialize and
  tools/list.
- Switch test_all_mcp_tools to lightweight tools (json_decode,
  find_files, grep_files) instead of the KG-backed tools whose
  ensure_thesaurus_loaded walk hangs in CI when the path is empty.
- Move shared helpers into tests/support/mod.rs and silence the
  per-binary dead-code warnings.

Refs #143.

(cherry picked from commit 5f54243b7f629a2806af99d369e7778dff721716)

* Fix #144: hermetic user_prompt_submit tests via TERRAPHIM_DEFAULT_DATA_PATH (#14)

The user-prompt-submit hook path uses LearningCaptureConfig::default() which
resolves global_dir via dirs::data_dir(). On macOS/Windows that ignores
XDG_DATA_HOME and returns $HOME/Library/Application Support, so the test
that set HOME and XDG_DATA_HOME never found the file it expected.

Production:
- Honour TERRAPHIM_DEFAULT_DATA_PATH in Default::default().
- storage_location() short-circuits to global_dir when the env var is set.

Test:
- Rewrite user_prompt_submit_tests to use support::cli_test_env helpers.
- No mocks, no #[ignore], no timeout increases.

All 4 tests pass on macOS.

Refs #144.

(cherry picked from commit 2ceda189769ae95f31a1a5f00e45c4aca3eaeb2a)

* test: resolve workspace binaries via CARGO_BIN_EXE, honour TERRAPHIM_SERVER_BIN

The terraphim_mcp_server and terraphim_agent integration tests located
their binaries by guessing <workspace>/target/debug/<name>. That guess is
wrong whenever CARGO_TARGET_DIR is set, which is how the native runner
builds (~/.cargo/build/by-runner/...), so under the all-targets gate every
stdio-driven MCP test and the extract-validation suite failed with
"binary not found" even though cargo had just built the binary.

Use env!("CARGO_BIN_EXE_<name>") instead: cargo sets it for a package's
own integration tests and guarantees the binary is built first, under any
target dir. TERRAPHIM_MCP_SERVER_BIN / TERRAPHIM_AGENT_BIN remain as
explicit overrides.

server_binary() in terraphim_agent's test support now honours
TERRAPHIM_SERVER_BIN, matching the other server-dependent tests, so the
CI-installed terraphim_server (not a workspace member, Refs #113) is used
by extract_functionality_validation too.

Refs #91

* ci: provision terraphim_server before the all-targets test gate

Run the #113 cargo install step before the workspace test step and export
TERRAPHIM_SERVER_BIN on it. With --all-targets the workspace step now
executes server_mode_tests, cross_mode_consistency_test, integration_tests
and kg_ranking_integration_test, all of which need the prebuilt server;
with --lib they were never compiled, so the ordering did not matter.

The focused #113 re-runs are kept for fast failure attribution.

Refs #91

* test: make hook_safety and search_only_flag hermetic

Both suites passed on developer machines and failed on the Gitea runner
(run 29586) because they read the host environment:

- hook_safety spawned terraphim-agent with the host's HOME, so the
  PreToolUse thesaurus came from ~/.config/terraphim. On one machine a
  personal KG term happened to match "-rf" (rewriting it to "Readiness
  Feedback"); on the runner nothing matched, so the "KG-replaceable"
  warning and the --rewrite substitution never fired. The tests now run
  under support::cli_test_env::apply_hermetic_env (fixture role config,
  KG at tests/test_kg) and a fixture concept tests/test_kg/trash.md maps
  "rm -rf" to "trash" so the replacement path is genuinely exercised.

- search_only_flag asserts on a debug!-level line. terraphim-grep's
  tracing filter defers to RUST_LOG when set, and the runner exports
  RUST_LOG=info, which hid the line. The tests now pin
  RUST_LOG=info,terraphim_grep=debug on the spawned process.

Verified with the full workflow step list under RUST_LOG=info and an
empty HOME (CARGO_HOME/RUSTUP_HOME kept): 91 binaries, 2231 passed, 0 failed, 3 ignored.

Refs #91

* docs(plans): add cass-parity session-search test plan + traceability matrix

Research artefact from the disciplined research pass (2026-09-03):
- 100-capability cass v0.6.11 catalog (C01-C100) benchmarked against
  terraphim session search; verdicts 0 FULL/34 PARTIAL/46 MISSING/
  18 N-A/2 UNCLEAR
- 158 test cases across 11 areas + 5 harness lanes + CI fixes
- Machine-readable C->TC traceability matrix (100 rows)

Refs #3084

* docs(plans): design + issue split for cass-parity session test suite (wave 1)

Design doc maps the research artefact (#148) onto verified repo reality:
- terraphim_agent/terraphim_sessions are workspace-local here (path dep
  1.21.2) — research decision R1 (registry-vs-local canary) dissolves
- Cursor connector exists with 15 tests (research assumed partial)
- native-ci.yml already --all-targets; remaining CI gap = --all-features lane
- search_nfr bench not carried into this repo -> PR-5 ports it (refs #3014)

5 sequenced PRs -> issues terraphim-clients #150..#154.
Refs #3084

* test(sessions): parity harness — shared fixture builders + --all-features CI lane

Closes terraphim-clients#150 (wave 1 PR-1 of the cass-parity suite).
Design: docs/plans/design-session-test-suite-2026-09.md.
Research: docs/plans/research-session-test-parity-2026-09.md.

- search_tests_support.rs (cfg(test)): deterministic Session/Thesaurus
  builders + claude-jsonl/aider-history corpus writers so parity tests
  never read real user session stores
- CI: terraphim_sessions --all-features lane (cursor/codex/extras +
  search-index score module are invisible to default/enrichment lanes)

Verified: 74 default / 88 enrichment / 122 all-features tests green;
clippy -D warnings clean.

* test(sessions): hybrid KG-boost ordering suite (P0 parity rows)

Closes terraphim-clients#151 (wave 1 PR-2). Covers parity rows
C07(p)/C62/T15/T18/E17a from the cass-parity research artefact:
search_with_thesaurus/search_sessions_hybrid had zero tests.

8 tests (feature-gated enrichment):
- boost promotes enriched session above raw-BM25 leader (ordering only,
  never score equality — fusion math is implementation detail)
- boost monotone in thesaurus match count
- None/empty-thesaurus degrade to plain BM25 (identical ordering)
- empty query/corpus edge cases, deterministic ordering
- unenriched sessions keep BM25 relative order; no-boost query still
  returns BM25 results

Verified: 96 tests green (--features enrichment); 74 default;
clippy -D warnings clean.

* test(sessions): import contracts — global limit, auto-import single-attempt

Closes the import-contract portion of terraphim-clients#152 (wave 1 PR-3a):
- import_all global-limit truncation via native connector (3-session
  tempdir corpus, limit=2 -> 2 sessions)
- auto-import single-attempt contract: consecutive cache-touching calls
  observe the same session set (no re-import); count is env-dependent by
  design so only the stability contract is asserted

Both hermetic (tempdir corpora via ImportOptions.path). Full connector
hermetic suites follow in the next commit on this branch.

* test(sessions): per-connector hermetic import suites (cline, aider)

Completes terraphim-clients#152 (wave 1 PR-3):
- cline: taskHistory.json + api_conversation_history.json tempdir corpus
  -> 1 session, role order asserted; empty dir imports nothing (the
  taskHistory/api path had zero import tests before)
- aider: .aider.chat.history.md tempdir corpus via ImportOptions.path
  (detection walk bounding already covered by #123 fix; import itself
  was untested)

Verified: 135 tests green with --all-features; clippy -D warnings clean.

* test(agent): REPL/CLI session contract tests (robot JSON, exit-4, flag order)

Closes terraphim-clients#153 (wave 1 PR-4). First session integration
tests in the agent crate — covers parity rows C40(p)/C80(p)/C87(p)/
C89(p)/C99(p) from the cass-parity research artefact.

7 tests in crates/terraphim_agent/tests/sessions_cli_contract.rs:
- sessions sources membership JSON (claude-code-native always present)
- machine-mode empty search: exit 4 AND zero-payload (pairing rule)
- --robot after subcommand rejected with ERROR_USAGE(2) (root-flag rule)
- search/list/stats JSON shape contracts (session_output structs)
- human-mode no-match text path

All hermetic via tests/support/cli_test_env.rs (temp HOME).

* test(agent): DOCS-DRIFT probes + NFR bench wiring (port #3014 into clients)

Closes terraphim-clients#154 (wave 1 PR-5, final).

DOCS-DRIFT probes (tests/sessions_docs_drift.rs):
- CLAUDE_SESSIONS_DIR documented-but-unimplemented: env var has NO effect
  on discovery (asserted); doc decision: fix the skill docs
- robot capabilities advertise supported_formats [json,jsonl,minimal,table]
  but CLI OutputFormat enum is human|json|json-compact — drift pinned,
  fix deliberate
- /sessions import removal: CLI rejects as unrecognized subcommand;
  REPL parser returns the explanatory message

NFR bench: port benches/search_nfr.rs from terraphim-ai@8fb947863
(10K deterministic sessions, criterion) + criterion dev-dep +
[[bench]] required-features search-index. Measured locally:
search_sessions_10k ~89ms [88.2-91.4] — within the <100ms G1 NFR.

Verified: 135 sessions tests green (--all-features); agent clippy clean.

* docs(verification): wave-1 verification report for session-test-suite-2026-09

Records the 5/5 merge status, local verification evidence, G1 NFR
measurement (89ms/10K), docs-drift findings, and the wave-2 backlog.
Refs #3084.

* fix(tests): address PR-review P2 findings — panic-safe tempdirs, drop dead-code shim

Follow-up to the structural PR reviews (wave 1, comments on #156/#158/#160):
- service.rs import-limit test: tempfile::tempdir() (panic-safe) instead of
  manual std::env::temp_dir + remove_dir_all
- sessions_docs_drift.rs: remove unused apply_hermetic_env import + the
  _unused() shim

Remaining review findings tracked as wave-2 backlog (nightly bench job,
p50 regression test, seeded-corpus entry-shape assertions, auto-import
hermeticity rewrite).

* docs(terraphim_agent): document LearningStore hybrid scoring (Refs #850)

Add an Unreleased entry for the LearningStore::query_relevant graph-rank
hybrid path landed in 13c5a36. The trait impl now ranks candidates by
RoleGraph::query_graph rank when a role graph is configured, with the
min_trust and applicable_agents gates preserved, and falls back to
substring text matching otherwise.

* style: cargo fmt on sessions_cli_contract.rs (fmt gate caught drift)

* docs(verification): addendum — structured PR review round results for wave 1

All five wave-1 PRs reviewed (structural-pr-review), scores recorded,
mechanical fixes merged (#161), wave-2 backlog enumerated. Pre-existing
PRs #147 (merged after rebase+review) and #146 (scope mismatch flagged)
also covered. Refs #3084.

* feat(sessions): redact API keys and secrets in session import connectors Refs terraphim/terraphim-ai#1986

Add crates/terraphim_sessions/src/redaction.rs with redact_session_content() which
applies regex patterns for AWS keys, OpenAI/GitHub tokens, Bearer tokens, and
connection strings. Call redact_sessions() in import() of all five connectors
(native, aider, cline, opencode, codex) before returning.

Make regex a mandatory dep (was optional/aider-connector-gated) since redaction
applies to all connectors regardless of feature flags. Fix pre-existing clippy
collapsible_if warnings in service.rs (two if-let nesting blocks).

63 tests pass, 8 new redaction tests, 1 doctest.

* fix(redaction): compile redaction patterns once via OnceLock

PR-review P1 fix on the revived #34 branch: redact_session_content
recompiled all 10 regexes per message, making auto-import over a real
corpus (151MB ~/.claude) pathologically slow (test hang >5min diagnosed
via spindump: Regex::new dominating redact_sessions). Patterns now build
once per process via OnceLock. Auto-import contract test completes in
seconds; 85 lib tests green; clippy -D warnings clean.

* test(procedure): add CLI integration tests for learn procedure from-session Refs #2350

Adds three integration tests to procedure_cli_tests.rs (all feature-gated
behind repl-sessions):

- procedure_from_session_extracts_non_trivial_commands: verifies that from-session
  filters trivial commands (cd) and failed commands (exit_code != 0), auto-generates
  a title, and reports correct step and command counts.

- procedure_from_session_deduplicates_on_repeat: verifies that running from-session
  twice with the same session results in one procedure via save_with_dedup().

- procedure_from_session_missing_id_fails: verifies that a missing session ID causes
  a non-zero exit code.

The core implementation (extract_bash_commands_from_session, from_session_commands,
FromSession CLI variant, and the existing unit test) was already in place. This
commit provides the CLI-level evidence required by AC 6 of terraphim-ai#2350.

Co-Authored-By: Terraphim AI <team@terraphim.ai>

* fix(tests): make from-session CLI tests platform-hermetic + align dedup assertions

Revival of #21 (terraphim-ai#2350 coverage), two fixes required before merge:
- Fixture write: dirs-5.0.1 macOS ignores XDG_CACHE_HOME (cache_dir =
  HOME/Library/Caches), Linux honours it — write sessions.json to all
  platform-mirrored cache variants (rule from the parity design doc).
  Original test false-failed on macOS / false-passed on Linux CI.
- Dedup assertion aligned with current save_with_dedup semantics: merge
  only happens for high-confidence existing procedures; a fresh 0%-
  confidence procedure does not merge, so two identical runs yield two
  procedures (was: older always-merge behaviour).

15/15 procedure_cli_tests green; fmt clean.

* docs(verification): addendum 2 — PR queue triage, revived PRs, >=4/5 gate results

* ci: add manual publish-registry workflow for terraphim registry Refs terraphim/terraphim-ai#2515

Adds a workflow_dispatch-triggered workflow to publish any crate to the
terraphim Gitea cargo registry. Requires the secret
CARGO_REGISTRIES_TERRAPHIM_TOKEN (package:write scope) to be set in
the repo CI secrets.

To publish terraphim_sessions 1.20.4 and unblock issue #2515:
1. Admin: set CARGO_REGISTRIES_TERRAPHIM_TOKEN secret
2. Trigger this workflow with crate=terraphim_sessions
3. Merge terraphim-agents PR #60

* feat(sessions): implement expand subcommand (Task 2.6.4) Refs terraphim/terraphim-ai#2134

Add SessionsSub::Expand variant to the terraphim-agent CLI:
- SessionsSub::Expand { id, context_lines } added to enum
- SessionExpandOutput + ExpandedMessage types added to session_output mod
- Handler in offline/async path: loads from disk cache then get_session()
- Handler in server/rt.block_on path: auto-imports via list_sessions() then get_session()
- Human-readable output prints all messages with role headers
- Machine-readable output (--format json) wraps via robot envelope
- context_lines field reserved for future --query-aware context expansion
- 2 new unit tests verify JSON serialisation of output types
- cargo clippy -D warnings: clean
- cargo fmt --check: clean
- 465 unit tests pass (cross_mode_consistency_test pre-existing polyrepo infra gap)

* style: cargo fmt + drop unused cache_dir binding in dedup test

* feat(learn): auto-suggest corrections from KG in capture_failed_command Refs terraphim/terraphim-ai#1648

When ToolPreference corrections exist in the storage directory, search the
compiled corrections thesaurus for patterns matching the failing command or
error text. If a match is found, set learning.correction to the suggested
replacement so that `terraphim-agent learn list` surfaces it immediately.

The auto-suggest is non-blocking: if the corrections directory is empty or
the thesaurus search fails, capture continues without setting the field.

Adds test_capture_sets_correction_when_kg_match_found regression test.

* feat(learn): implement KG auto-suggest corrections in capture_failed_command

When capture_failed_command annotates entities from a failed command, it
now calls compile_corrections_to_thesaurus to check whether any matched
entity has a known ToolPreference correction. The first match is set as
learning.correction so that terraphim-agent learn list surfaces the
suggestion without any manual intervention.

Adds suggest_correction_from_entities as a public helper for unit testing
and adds test_capture_sets_correction_when_kg_match_found covering match,
no-match, and empty-entity cases.

Refs terraphim/terraphim-ai#1648

Co-Authored-By: Terraphim AI <noreply@terraphim.ai>

* fix(learn): resolve duplicate test name breaking PR #29 build

The terraphim_agent bin test target failed to compile because
`test_capture_sets_correction_when_kg_match_found` was defined twice
(a merge/copy-paste artefact). The two tests cover different behaviour:
one drives capture_failed_command end-to-end via the compiled thesaurus,
the other unit-tests suggest_correction_from_entities directly. Rename
the latter to test_suggest_correction_from_entities_matches_tool_preference
and apply rustfmt. Both tests are retained.

Unblocks native-ci / build on task/1648 (PR #29).

Refs #30
Refs terraphim/terraphim-ai#1648

* fix(learn): borrow corrections thesaurus in find_matches (1.21.x borrowed-&Thesaurus family)

* feat(learnings): implement CorrectionEvent CLI and security hardening Refs #2083

- Restructure `learn correction` into sub-subcommands:
  - `learn correction add --original X --corrected Y --correction-type T` (records a CorrectionEvent)
  - `learn correction list [--filter-type T] [--global]` (lists stored corrections, optionally filtered by type)
- Add YAML frontmatter injection protection: sanitise_yaml_value() strips newlines from hostname, session_id, working_dir, id, and tags written to .md files
- Add 64 KiB per-field size limit in capture_correction() to prevent disk exhaustion
- Escape backticks in markdown body so inline code blocks are never broken
- Add #[serde(default)] to CorrectionEvent.tags for backward-compatible deserialisation
- Add four new security-focused tests: YAML injection via hostname/session_id, oversized input rejection, backtick escaping

* feat(terraphim_lsp): add KG analysis engine with Aho-Corasick term matching Refs #2669

Adapted to current main: terraphim_automata at registry 1.21.0 with
borrowed-&Thesaurus find_matches API; terraphim_negative_contribution
and terraphim_types at current path/registry versions.

* test(agent): REPL /sessions parser contract tests (adapted from #39, owner decision A)

Land the still-valid parser tests from stale PR #39 per owner decision A:
- /sessions import stays REMOVED (auto-import replaced it); the removal
  message itself is pinned as a contract test
- REPL has no expand alias (expand is CLI-only via #165); show/get
  aliases pinned instead
- 13 parser-contract tests: list (+ --source/--limit), ls alias, search
  (+missing-query error), show/get, sources/detect, /session singular,
  missing/unknown subcommand errors

Closes terraphim/terraphim-ai#2435 (test half). Supersedes #39.

* feat(agent): add shared-learning to default features (Refs terraphim-ai#2516, owner decision A)

learn shared list/promote/import/stats subcommands are now available in
production builds without --features shared-learning. The collapsible_if
clippy fixes from the original stale commit are already on main (let-chains).

Verified with default features only: 498 lib tests, 9 shared_learning_cli
tests, clippy -D warnings clean, fmt clean, CLI sanity (learn shared list)
exits 0.

* refactor(agent): bin reuses lib modules — eliminates twin module tree

The bin re-declared lib modules (mod client; mod robot; ...) so every
lib item was compiled twice: reachable (no warnings) in the lib target,
privately-duplicated (dead_code warnings -> silenced by 152
#[allow(dead_code)] across the workspace) in the bin target.

main.rs now imports the lib modules (terraphim_agent::{...}) instead of
re-declaring them; main-only modules (listener, shell_dispatch,
kg_validation, session_output, test mods) stay local. learnings:capture
list_learnings is now re-exported from the learnings root.

This removes the need for most dead_code suppressions in the agent
crate. Follow-ups: TSA + remaining crates.

* refactor(robot): drop dead_code allows on submodules (twin-tree fix made them unnecessary)

* refactor(agent): purge remaining dead_code suppressions — wire, read, or delete

Agent crate now has ZERO #[allow(dead_code)] (was 87):
- capture.rs: delete old auto-extract/suggest pipeline (score_entry_relevance,
  TranscriptEntry, ScoredEntry, suggest_learnings, auto_extract_corrections,
  contains_correction_phrase, extract_command_from_input) — superseded by
  SharedLearningStore::suggest hybrid scoring (13c5a36)
- install.rs: delete unwired uninstall_hook/is_hook_installed/
  get_installation_status CLI plumbing
- hook.rs: read the serde(flatten) extra map in from_json (debug log) —
  it exists for unknown-field forward-compat
- executor.rs: drop never-read api_client field + unused with_api_client
- validator.rs: drop dead determine_execution_mode wrapper
- hybrid.rs: drop never-read vm_for_unknown setting
- listener.rs: drop dead run_once/handoff_issue wrappers
- shell_dispatch.rs: move MAX_OUTPUT_BYTES into test scope, delete unused
  DISPATCH_TIMEOUT_SECS

Workspace total: 152 -> 67 allows (TSA crate + test files remain).

* refactor(tsa): bin reuses lib modules + purge dead correlation code

- main.rs now uses terraphim_session_analyzer::{...} instead of
  re-declaring the module tree (same twin-tree fix as the agent crate)
- delete calculate_agent_tool_correlations + CorrelationData from main.rs
  (TODO on the fn said superseded by Analyzer::calculate_agent_tool_correlations)
- 52 dead_code allows removed; TSA: 128+2+20 tests green, fmt clean

* refactor: purge all production dead_code suppressions (closes #5, #1)

Workspace #[allow(dead_code)]: 152 -> 11 (all 11 in test binaries,
documented shared-support pattern; production source: ZERO).

Key changes:
- agent: twin module tree fixed — bin imports lib modules instead of
  re-declaring them (client.rs 30 allows + robot/mod 5 were silencing
  the bin duplicate); deleted genuinely-dead auto-extract pipeline in
  capture.rs (superseded by SharedLearningStore::suggest hybrid scoring),
  unwired uninstall/status hook fns, unused executor api_client field,
  dead determine_execution_mode wrapper, never-read vm_for_unknown,
  dead listener wrappers; hook.rs ToolInput.extra now read (debug log)
- tsa: bin reuses lib modules; deleted calculate_agent_tool_correlations
  + CorrelationData (TODO said superseded by Analyzer impl)
- sessions: serde DTO fields dropped where never read (msg_type,
  ComposerTab.model) — serde ignores unknown fields, behavior-safe
- update: deleted hand-rolled is_newer_version (superseded by semver
  static fn); test migrated to is_newer_version_static
- mcp_server: frecency field now read via getter + startup debug log

All gates green: workspace clippy --all-targets --all-features -D
warnings clean; fmt clean; 143+493+130+15+128 sessions/agent/update/tsa
test suites pass.

* fix(agent): derive thesaurus_matched from the boundary-aware matcher

thesaurus_matched was built with a naive substring scan
(`query.to_lowercase().contains(key)`) independently of the automaton, so
any thesaurus term appearing inside a longer query word was reported: the
two-letter term `ce` matched `con(ce)pt` and surfaced in the robot-mode
search envelope for a query that matched nothing.

Derive it from the same matcher that produces concepts_matched, at both the
offline and server (run_server_command) sites, so the two fields cannot
disagree.

Note this only closes the agent-local half. concepts_matched comes from
terraphim_automata, pinned here to 1.21.0 from the Gitea registry, and stays
wrong until terraphim-core#65 is published and the pin bumped.

Refs #197

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U7LFsSpVkhACQS6HYM42JH

* fix(agent): honour --fail-on-empty in server mode

`run_server_command` destructured `fail_on_empty: _`, so the flag was
silently ignored whenever the agent talked to a server: an empty result set
exited 0 instead of 4, and only the offline path behaved as documented.

Capture `results_count` before `res.results` is consumed by the robot
formatter, and apply the same exit check the offline handler uses.

Refs #197

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U7LFsSpVkhACQS6HYM42JH

* fix(redaction): apply central redaction before every persistence boundary (Refs #178)

Establishes a single canonical redaction policy at every persistence
boundary in terraphim_agent::shared_learning and terraphim_sessions::cla.

Pinned 7-call-site coverage:
- markdown_store::save / save_to_shared (redaction in to_markdown)
- wiki_sync::sync_learning (covers sync_all_learnings + sync_batch
  via the single funnel point)
- from_normalized_session / from_normalized_message in terraphim_sessions::cla
  (covers both ClaClaudeConnector and ClaCursorConnector via the shared
  helper)

Implementation:
- New terraphim_sessions::redaction module with inline SECRET_PATTERNS
  list + drift-guard test asserting parity with the canonical
  terraphim_agent::learnings::redaction.
- New terraphim_agent::shared_learning::redaction module with the same
  drift guard, exposed via 'pub use redaction::redact_secrets'.
- regex promoted from optional to direct dep on terraphim_sessions
  (central redaction is now a baseline requirement).
- Logging-hygiene tests asserting no tracing::warn! / tracing::debug! /
  dbg! carries the pre-redaction body field.

Tests:
- terraphim_sessions: 76 lib tests pass (incl. drift guard + 1 new
  real-fixture connector test, no mocks).
- terraphim_agent --features shared-learning: 308 lib tests pass (incl.
  new redaction module + 7 new boundary tests).

Closes #178 (canonical redaction policy).
Refs PR #22 (will be closed as superseded once this lands; PR #22's
4 surviving non-redaction P1s become a smaller follow-on).

* fix(security): validate wiki_page_name + XSS strip + component validation (Refs #22)

Ports the 3 surviving P1 security findings from PR #22 (closed as
superseded by PR #200). Each P1 lands with real-fixture tests
(no mocks per AGENTS.md) and bounded scope per Phase 2 design.

P1-1 XSS strip (Refs #22): strip <script>, <iframe>, <object>,
<embed> blocks from wiki markdown body. Case-insensitive, handles
self-closing forms, unclosed tags, nested tags, attributes, and
multibyte content safely. 13 unit tests cover adversarial variants.

P1-2 Page-name validation (Refs #22): wiki_page_name must match
^[a-zA-Z0-9_][a-zA-Z0-9_-]{0,199}$ — leading '-' rejected …
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant