Repository navigation
feat(acceptance): validate the public entry points against live infrastructure - #34
Merged
Merged
Conversation
AlexMikhalev
force-pushed
the
feat/public-release-acceptance
branch
2 times, most recently
from
September 27, 2026 11:36
741d0fd to
caa7791
Compare
…structure
Adds the repeatable acceptance run and records what the family publish proved.
scripts/acceptance-public-release.py is read-only and reports pass, fail or
not-executed per entry point, so an environment limitation is never counted as
success. It runs the documented installer against a scratch directory and
checks the installed binary reports the channel version, verifies every
advertised archive against the manifest size and SHA-256, resolves the
Homebrew formula, and reports the crates.io versions a cargo install would
serve.
The plans under docs/plans/ gain the measured result of the family publish.
A dry run of the full client crate family on main (run 36314289075) reaches
terraphim_config 1.20.4 and fails to compile:
unresolved import terraphim_automata::parse_markdown_directives_dir
the item is gated behind the `fs-traversal` feature
crates.io's terraphim_automata is 1.21.1, which gates that symbol, while
crates.io's terraphim_config is 1.20.4, which does not enable the feature, and
the client crates require config 1.20.2. No client crate is publishable until
that is repaired, and the repair lives in terraphim-core and
terraphim-config-persistence, not here.
Run 36314289075 also confirmed that publish-crates.yml refuses terraphim_agent
before mutating any manifest (#95), and that terraphim_update 1.20.2 and
terraphim_command_runtime 0.1.0 already exist on crates.io.
Verified: acceptance script compiles; record manifest check passes (3 binaries
at 1.21.16, 20 archives byte-verified); the failing dry run is cited by id.
AlexMikhalev
force-pushed
the
feat/public-release-acceptance
branch
from
September 27, 2026 13:12
caa7791 to
ea6d880
Compare
AlexMikhalev
added a commit
that referenced
this pull request
Oct 4, 2026
* fix(ci): add --features enrichment to Gitea native CI Refs #2171
Mirror the GitHub CI enrichment steps in the Gitea native-ci workflow:
- cargo clippy -p terraphim_sessions --features enrichment -- -D warnings
- cargo test -p terraphim_sessions --features enrichment --lib --no-fail-fast
All 67 tests pass locally including the 3 enrichment tokio tests and
2 concept unit tests that were silently skipped before.
* test(grep): add default-feature zero-chunk smoke guard
Adds an explicit, named CI guard for the silent zero-chunk regression
documented in terraphim/terraphim-ai#3025 / #4325. A default-feature
build of terraphim_grep must return non-zero chunks for a matching
query; the test fails loudly if `code-search` is ever removed from the
`default` feature set (which would compile search_code() to a no-op
stub returning Ok(vec![]) -- success-with-zero-items).
- crates/terraphim_grep/tests/default_feature_smoke.rs: new integration
test. Distinct from no_thesaurus_cli.rs (KG-absent fallback) -- this
test's single purpose is the default-feature contract.
- .github/workflows/ci.yml: named, visible smoke-test step after the
workspace test run (functional gating already existed via
`cargo test --workspace`; this makes it explicit and documented).
Verified bi-directionally this session (real binaries, no mocks):
- default features (code-search on): chunks_returned=1 -> PASS
- code-search removed from default: chunks_returned=0 -> FAIL with the
named regression message -> revert -> PASS again.
Gates: fmt clean; clippy -p terraphim_grep --all-targets -D warnings 0;
full crate test suite green.
Refs terraphim/terraphim-ai#4325 (cross-repo: terraphim_grep lives in
terraphim-clients, extracted via #1910).
Co-Authored-By: Claude <noreply@anthropic.com>
* docs(quality): capture verification + validation reports for PRs #44, #45, #49, #51, #52, #59, #60
- PR #44: Insufficient KG propagation (#2721)
- PR #45: rust-engineer shortname fix (#2723)
- PR #49: CI enrichment feature (#2171)
- PR #51: thesaurus NotFound ERROR suppress (#48)
- PR #52: Gitea CI enrichment feature (#2171)
- PR #59: grep default-feature smoke test (#4325)
- PR #60: terraphim_grep crates.io publishable metadata (#58)
Also refreshes Cargo.lock to register env_logger as a terraphim_agent
runtime dependency introduced by PR #51 (Cargo.lock had drifted from
Cargo.toml since the merge).
Refs #108
* feat(memory): scaffold terraphim-agent memory CLI namespace with 10 subcommands Refs #1899
Add top-level command wrapping the eight-stage agentic memory
lifecycle behind a single discoverable CLI surface.
Subcommands:
- capture (routes to learn hook) - validate (rubric scorer stub)
- distill (routes to learn compile) - retire (learned-rules stub)
- scope (role/project boundaries) - rubric (6-dimension diagnostic)
- provenance (routes to sessions) - second-run (token delta)
- retrieve (routes to search)
- apply (routes to terraphim_hooks)
Seven subcommands are routing stubs delegating to existing handlers.
Three (validate, rubric, second-run) are placeholder stubs for net-new
code to be wired in subsequent steps.
Implemented in crates/terraphim_agent/src/main.rs following existing
Command/Subcommand enum pattern (Memory variant in Command enum,
MemorySub enum with clap derive, run_memory_command async handler).
Tests: cargo check passes, 246 lib tests pass, all 10 subcommands
respond to --help and execute their handler stubs.
* feat(memory): wire terraphim_agent_evolution into CLI with capture/list/show/export/scope Refs #1899
Add terraphim_agent_evolution v1.20.2 from terraphim registry as dependency.
Implement real memory lifecycle operations backed by AgentEvolutionSystem:
- capture: write MemoryItem to evolution store with provenance tags
- list: enumerate memory items with optional item_type filter
- show: display full details of a memory item or lesson by ID
- export: dump memory items + lessons as JSON or markdown
- scope: show role/project KG boundaries from ~/.config/terraphim/kg
capture/scope are now real implementations instead of routing stubs.
Three new subcommands (list, show, export) added for evolution inspection.
Tests: 246 lib tests pass. cargo check clean with 0 warnings.
* feat(memory): implement rubric scorer, second-run signal, policy doc, README Refs #1899
Implement all remaining memory lifecycle components:
Rubric subcommand:
- 6-dimension rule-based scorer: faithfulness, scope, provenance,
actionability, decay, risk
- Composite score with weighted average (faithfulness 0.30,
actionability 0.25, scope 0.15, provenance 0.10, decay 0.10,
risk 0.10 inverted)
- Markdown readout with dimension table, top-3 offenders,
recommended retirements
- RubricScore struct + score_memory_item/scoring helpers
Validate subcommand:
- Scores all, specific, or most recent memory items
- Per-item scores with composite average
Retire subcommand:
- Writes retirement proposals to ~/.config/terraphim/
(learned-rules-retirements.md) with CTO approval flag
Second-run subcommand:
- Reads ADF artefact directory (~/.cache/terraphim/adf-artefacts/)
- Finds JSON run metrics files per Gitea issue
- Computes token delta, retry delta, wall-time delta
- Emits structured JSON with interpretation
MEMORY_POLICY.md:
- Public commons vs permissioned memory boundary
- Storage location table
- Enforcement via memory scope --check
README:
- Full memory command reference (13 subcommands documented)
- Link to MEMORY_POLICY.md
Tests: 246 lib tests pass. cargo check clean.
* feat(memory): add cross-invocation persistence via JSON file store Refs #1899
Add load_evolution()/save_evolution() helpers that persist
AgentEvolutionSystem state as JSON to:
~/.config/terraphim/evolution/cli-agent.json
- MemoryState and LessonsState are serialized/deserialized
- Captured items survive CLI invocations
- List/show/export/rubric consume persisted state
- save_evolution() called after every capture mutation
Verified: capture → list shows item, capture → rubric scores,
capture → export produces markdown. All 246 tests pass.
* fix(memory): address PR review findings -- P1 capture semantics, P2 dedup scoring Refs #1899
P1: Capture now stores provenance_tag as a tag, not as content.
Content gets a descriptive string. Tags field receives
'provenance:<tag>' entries. Help text updated accordingly.
P2: Extract compute_decay() and compute_risk() shared functions
to eliminate duplicate scoring logic between score_memory_item
and the rubric retirement filter. Add TERRAPHIM_ADF_ARTEFACTS_DIR
env var override for second-run artefact path.
* fix(grep): resolve project thesaurus by role shortname
* fix(grep): rank KG matches above substring metadata
* feat(grep): add update commands via shared updater
* fix(update): rotate zipsign release verifier key
* fix(update): preserve release asset name for verification
* fix(agent): use clients release assets for updates
* fix(rebase): repair post-rebase fallout in PR #61
The rebase of task/1899-memory-lifecycle-cli onto current main
d54f28f left several files in a non-compiling state due to upstream
renames (terraphim_update, terraphim_automata 1.21.0 borrows &Thesaurus,
thesaurus registry pins from Refs #112).
- crates/terraphim_grep/Cargo.toml: drop the path-dep duplicate
left by the rebase; keep the Refs #112 registry-bearing entry.
- crates/terraphim_grep/src/hybrid_searcher.rs:
- bind the kg_concepts result of search_kg (was dropped by an
orphan semicolon);
- pass &thesaurus to find_matches (1.21.0 borrows it);
- iterate over all search_paths, apply boost_chunks_with_kg
before truncating, so KG-ranked chunks survive the candidate
cut.
- crates/terraphim_grep/src/main.rs: restore the function
signatures that the conflict resolution collapsed (push_unique_candidate,
discover_project_thesaurus) and add the missing #[test] attributes
on the two post-resolution tests.
- crates/terraphim_update/src/signature.rs: remove an orphan closing
brace left after the conflict on get_embedded_public_keys.
Verified locally: cargo check --workspace --all-features, cargo clippy
--workspace --all-features --all-targets -- -D warnings, and cargo test
on the affected crates (terraphim_grep, terraphim_update, terraphim_sessions)
all pass. Pre-existing Refs #113 failures in cross_mode_consistency_test,
mcp_server integration tests, and terraphim_update::test_resolve_asset_url_against_local_manifest
are unchanged.
Refs #1899
* fix(update): revert redundant with_repo to unbreak packaged install gate
PR #61 added UpdaterConfig::with_repo and called it from agent + grep
with the constructor's own default values (terraphim/terraphim-clients).
The packaged_install_graph_regression test (cargo package + cargo install
--path) resolves terraphim_update from the terraphim registry, where the
published 1.20.2 lacks with_repo, so the install failed to compile.
The design's own rollback plan says to revert with_repo if no consumer
uses it functionally. Here every caller passed the same value the
constructor already sets, so the calls were no-ops; the method had no
real consumer and would have required a registry publish to keep CI
green. Reverting restores install-graph correctness and aligns the
branch with main's updater API.
Refs #1899
* docs(quality): add PR #61 verification report
Captures the 8 native-ci steps, gitea native-ci status check, traceability
matrix for Refs #1899 / #95 / #62 / #112, and the defect register (the
with_repo rollback plus two rebase-fallout defects resolved before
merge). Refs #1899
* docs(quality): add PR #61 validation report
Acceptance criteria trace from Refs #1899 / #95 / #62 / #112, performance
and security review notes, defect register (the three closed defects from
the verification report plus follow-ups). Refs #1899
* docs(quality): PR #84 verification + validation -- documents 57 pre-existing failures exposed by --all-targets
Refs #84, closes part of #108 (campaign summary)
Findings (full report in .quality/pr-84-{verification,validation}.md):
- 57 unique test failures surface on Linux CI when --all-targets replaces --lib
- 37 are pre-existing (also fail on macOS main baseline); 20 are Linux-only
- Pre-existing failures are tracked by #113 (PR #135 in progress) and
separately by the missing docs/src/kg fixture path bug
- PR's empirical claim 'does not destabilise the pipeline' does not hold on
the Gitea runner because is_ci_environment() does not recognise the runner
(no CI=true / GITHUB_ACTIONS / dockerenv / root)
- Recommendation: do not merge PR #84 as-is; merge PR #135 first or extend
PR #84 with env: CI: true plus minimal #[ignore] for the MCP server-binary
tests, plus fix the wrong <workspace>/docs/src/kg path in replace_feature_tests
* test(fixtures): sync terraphim_server test fixtures from terraphim-ai v1.21.3 (Refs #113)
The integration tests under crates/terraphim_agent/tests/ that depend
on a real terraphim_server binary reference fixtures under
terraphim_server/{default,fixtures}/ that lived only in the
terraphim-ai repo. Sync the minimal set required by the 7 #113 tests
(2 in cross_mode_consistency_test, 2 in integration_tests, 3 in
kg_ranking_integration_test) so the tests can run from terraphim-clients
without fetching from terraphim-ai at test time.
Files synced (verbatim copy from terraphim-ai v1.21.3):
- terraphim_server/default/terraphim_engineer_config.json
- terraphim_server/fixtures/cross_mode_test_config.json
- terraphim_server/fixtures/term_to_id.json
- terraphim_server/fixtures/thesaurus_Default.json
- terraphim_server/fixtures/haystack/Engineer_thesaurus.json
- terraphim_server/fixtures/haystack/System Operator_thesaurus.json
- terraphim_server/fixtures/haystack/*.md (16 files)
Source: https://git.terraphim.cloud/terraphim/terraphim-ai @ v1.21.3
Refs #113
* fix(agent): make PreToolUse hook rewriting opt-in, default-on guard
The PreToolUse hook silently rewrote destructive commands when their
substring matched a thesaurus entry (e.g. `rm -rf /tmp/foo` could
become `rm -Readiness Feedback /tmp/foo`). This is dangerous because:
* Substitution was always-on, never warned the user.
* A typo'd KG entry or stray synonym could mutate an unrelated
destructive command, blowing away files instead of running the
user-intended action.
This commit flips the default to safe-by-construction:
* New `--rewrite` flag (default false) opts in to KG substitution.
When false, the hook still probes for replacements but emits a
`warnings` array entry describing the suppressed match.
* Guard check defaults to ON for `pre-tool-use` (still off for
`post-tool-use`, `pre-commit`, `prepare-commit-msg` because they
fire after execution or on text inputs).
* `--no-with-guard` flag explicitly disables the guard (clap does
not auto-derive `--no-with-guard`; explicit override required).
* `with_guard` and `no_with_guard` are mutually exclusive via
`conflicts_with`.
Adds `tests/hook_safety.rs` with five regression tests covering:
* `rm -rf /tmp/foo` passes through with a warning (the original bug).
* `rm -rf /` is denied by the default guard.
* `--rewrite` enables substitution as before.
* `--no-with-guard` suppresses the guard envelope.
* `/tmp/` allowlist still allows `rm -rf /tmp/foo`.
Refs #126
* feat(agent): --explain for guard, README robot-mode fix, priority docs
Refs #127, #129
#127 (README robot-mode example)
The example `terraphim-agent search "retry policy" --robot --format json`
was wrong. `--robot` and `--format` are global flags on the top-level
Cli struct and must precede the subcommand. Placed them after the
subcommand clap fails with `error: unexpected argument '--robot' found`.
Updated the README with both the corrected example and a parenthetical
explaining the failure mode.
#129 (guard --explain + priority docs)
* New `--explain` flag on `terraphim-agent guard` that prints the
per-stage evaluation trace (allowlist > destructive > suspicious >
default). With `--json` it emits a structured `GuardTrace`; without
`--json` the trace goes to stderr in a human-readable form.
* New `CommandGuard::check_with_trace` returns a `GuardTrace`
containing the final `GuardResult` plus per-stage outcomes
(allow/block/sandbox/continue/no_match) and matched terms.
* README now documents the priority order with a worked example
showing the allowlist short-circuiting destructive.
* Added `tests/guard_priority.rs` with 5 regression tests:
- allowlist_short_circuits_before_destructive
- destructive_short_circuits_before_suspicious
- default_allow_path_emits_default_stage
- explain_exits_one_on_block
- explain_text_output_is_human_readable
Both tests (`hook_safety` and `guard_priority`) and the existing
`learn_no_service_tests` pass. Clippy clean with -D warnings.
* test(grep): pin --search-only flag and LLM-client skip
Refs #128
The installed `~/.cargo/bin/terraphim-grep` was 1.21.11 while the
workspace source is 1.21.13. The 1.21.11 binary is missing the
`--search-only` flag (Refs terraphim-clients#81) and several other
recent fixes (chunks_returned counter, etc.). Rebuilt and installed
the 1.21.13 release binary.
Adds `tests/search_only_flag.rs` with three regression tests so CI
catches a future drift between source and installed binary:
* search_only_flag_is_accepted -- the flag is parsed.
* search_only_skips_llm_client_with_openrouter_key_present --
even with a stray API key, --search-only must skip the LLM client
build (asserted by checking stderr for the
"skipping LLM client setup" debug log).
* help_documents_search_only_flag -- defensive: --help must mention
--search-only so users can discover it.
Clippy clean with -D warnings.
* feat(robot): mark REPL-only commands in schema output
Refs #131
`terraphim-agent robot schemas` previously listed every REPL command
(name, description, args, flags, examples, response_schema) without
flagging which ones are actually available as top-level CLI
subcommands. Consumers parsing the schema expected parity with
`terraphim-agent --help`, which doesn't list REPL-only entries like
`vm` or `chat`. The mismatch was confusing.
Adds a `repl_only: bool` field to `CommandDoc` (defaults to `false`
via serde). Marks the two REPL-only commands called out in #131:
* `vm` -- always REPL-only (firecracker-gated, no top-level CLI).
* `chat` -- REPL-only, feature-gated behind `repl-chat`.
All other 13 commands keep `repl_only: false`. Comments next to the
literals document why each is REPL-only.
Adds `tests/robot_schemas.rs` with four regression tests:
* every_command_has_repl_only_field
* top_level_cli_commands_are_not_repl_only
* vm_is_marked_repl_only
* chat_is_marked_repl_only (tolerates missing chat when the
`repl-chat` feature is off in the test binary)
All 4 tests pass with `--features repl-chat`. Clippy clean with
-D warnings.
* docs: ADRs for guard priority and pretool rewrite, reference, blog posts
Refs #130, #132, #133
#130 (reference doc)
* docs/agent-reference.md — 271-line reference enumerating all 22
top-level `terraphim-agent` subcommands. The README's "Key
Commands" table only listed 8; the remaining 14
(roles, config, kg, extract, replace, validate, suggest,
interactive, repl, setup, check-update, update, listen, cache)
now have one worked example each.
* Quick-reference family grouping at the bottom.
* Cross-links to the two ADRs and the design doc.
* Cross-links to the five new blog posts.
#132 (five blog posts)
* docs/blog/terraphim-agent-sessions.md — Claude Code / Cursor /
Aider import flow, bounded walker, JSON output schema.
* docs/blog/terraphim-agent-setup.md — onboarding wizard and the
10 role templates.
* docs/blog/terraphim-agent-robot-mode.md — JSON output, exit
codes (0..7), self-describing schemas, global flag caveat.
* docs/blog/terraphim-agent-shared-learning.md — markdown-backed
BM25-deduped learning store with trust levels.
* docs/blog/terraphim-update-r2-backend.md — R2 manifest format,
verification path, fallback chain, backend overrides.
* README "Further reading" section links to all five plus the
reference doc.
#133 (two ADRs)
* adr/ADR-002-guard-priority-order.md — rationale for the
allowlist > destructive > suspicious > default priority with
fail-open per stage. Includes worked examples and the four
alternatives that were rejected (destructive-first, voting,
most-specific-wins, etc.).
* adr/ADR-003-pretool-hook-rewrite.md — rationale for making
thesaurus substitution opt-in via `--rewrite`. Documents the
two modes (default warn-only vs opt-in substitution), the
residual risk of warnings being dropped by future agent
runtimes, and the four alternatives that were rejected
(always-off, always-on, sentinel-in-command, default-on).
Both ADRs follow the same Context / Decision / Consequences /
Alternatives / References format as ADR-001.
* feat(robot): add CLI chat schema, mark summarize REPL-only
The structural-pr-review for #134 flagged two cross-file
inconsistencies in [
{
"aliases": [
"q",
"query",
"find"
],
"arguments": [
{
"description": "Search query text",
"name": "query",
"required": true,
"type": "string"
}
],
"description": "Search documents using semantic and keyword matching",
"examples": [
{
"command": "/search async error handling",
"description": "Basic search"
},
{
"command": "/search database migration --role DevOps --limit 5",
"description": "Search with role and limit"
}
],
"flags": [
{
"default": "current",
"description": "Role context for search",
"name": "--role",
"short": "-r",
"type": "string"
},
{
"default": "10",
"description": "Maximum results to return",
"name": "--limit",
"short": "-l",
"type": "integer"
},
{
"default": "false",
"description": "Enable semantic search",
"name": "--semantic",
"type": "boolean"
},
{
"default": "false",
"description": "Include concept matches",
"name": "--concepts",
"type": "boolean"
}
],
"name": "search",
"response_schema": {
"properties": {
"concepts_matched": {
"items": {
"type": "string"
},
"type": "array"
},
"results": {
"items": {
"properties": {
"id": {
"type": "string"
},
"preview": {
"type": "string"
},
"rank": {
"type": "integer"
},
"score": {
"type": "number"
},
"title": {
"type": "string"
},
"url": {
"type": "string"
}
},
"type": "object"
},
"type": "array"
},
"total_matches": {
"type": "integer"
}
},
"type": "object"
}
},
{
"aliases": [
"c",
"cfg"
],
"arguments": [
{
"description": "Subcommand: show, set",
"name": "subcommand",
"required": true,
"type": "string"
}
],
"description": "View and modify configuration",
"examples": [
{
"command": "/config show",
"description": "Show current configuration"
},
{
"command": "/config set selected_role Engineer",
"description": "Set configuration value"
}
],
"flags": [],
"name": "config",
"response_schema": {
"properties": {
"config": {
"type": "object"
}
},
"type": "object"
}
},
{
"aliases": [
"r"
],
"arguments": [
{
"description": "Subcommand: list, select",
"name": "subcommand",
"required": true,
"type": "string"
}
],
"description": "Manage roles",
"examples": [
{
"command": "/role list",
"description": "List available roles"
},
{
"command": "/role select Engineer",
"description": "Select a role"
}
],
"flags": [],
"name": "role",
"response_schema": {
"properties": {
"current_role": {
"type": "string"
},
"roles": {
"items": {
"type": "string"
},
"type": "array"
}
},
"type": "object"
}
},
{
"aliases": [
"g",
"kg"
],
"arguments": [],
"description": "Display knowledge graph concepts",
"examples": [
{
"command": "/graph",
"description": "Show top concepts"
},
{
"command": "/graph --top-k 20",
"description": "Show top 20 concepts"
}
],
"flags": [
{
"default": "10",
"description": "Number of top concepts to show",
"name": "--top-k",
"short": "-k",
"type": "integer"
}
],
"name": "graph",
"response_schema": {
"properties": {
"concepts": {
"items": {
"properties": {
"count": {
"type": "integer"
},
"term": {
"type": "string"
}
},
"type": "object"
},
"type": "array"
}
},
"type": "object"
}
},
{
"aliases": [],
"arguments": [
{
"description": "Subcommand: list, pool, status, metrics, execute, agent, tasks, allocate, release, monitor",
"name": "subcommand",
"required": true,
"type": "string"
}
],
"description": "Manage Firecracker VMs",
"examples": [
{
"command": "/vm list",
"description": "List VMs"
},
{
"command": "/vm execute python print('hello')",
"description": "Execute code in VM"
}
],
"flags": [
{
"description": "VM identifier",
"name": "--vm-id",
"type": "string"
}
],
"name": "vm",
"response_schema": {
"properties": {
"status": {
"type": "string"
},
"vms": {
"type": "array"
}
},
"type": "object"
}
},
{
"aliases": [
"h",
"?"
],
"arguments": [
{
"description": "Command to get help for",
"name": "command",
"required": false,
"type": "string"
}
],
"description": "Show help information",
"examples": [
{
"command": "/help",
"description": "Show all commands"
},
{
"command": "/help search",
"description": "Get help for search"
}
],
"flags": [],
"name": "help",
"response_schema": {
"properties": {
"commands": {
"type": "array"
},
"help_text": {
"type": "string"
}
},
"type": "object"
}
},
{
"aliases": [],
"arguments": [
{
"description": "Subcommand: capabilities, schemas, examples",
"name": "subcommand",
"required": true,
"type": "string"
}
],
"description": "Robot mode commands for AI agents",
"examples": [
{
"command": "/robot capabilities",
"description": "Get capabilities"
},
{
"command": "/robot schemas search",
"description": "Get schema for search"
}
],
"flags": [
{
"default": "json",
"description": "Output format: json, jsonl, minimal, table",
"name": "--format",
"short": "-f",
"type": "string"
}
],
"name": "robot",
"response_schema": {
"type": "object"
}
}
]:
P1: `Command::Chat` in main.rs is a top-level CLI subcommand
gated by `--features llm` (default-on), but the only
`chat` schema entry was the REPL `chat` (gated by
`--features repl-chat`) marked `repl_only: true`.
A downstream agent introspecting via `robot schemas` would
conclude there is no top-level `chat` subcommand when there
actually is one. Added a separate `CommandDoc` for the CLI
`chat` (gated by `#[cfg(feature = "llm")]`) with
`repl_only: false`, the required `prompt` argument, and
`--role` / `--model` flags. The REPL entry is unchanged.
P2: `summarize` was marked `repl_only: false` but it is
REPL-only (registered in `repl::commands`, no top-level
CLI subcommand). Flipped to `repl_only: true`.
Both fixes are now visible in `robot schemas` output. The
`chat` descriptions are also disambiguated (CLI one-shot vs
REPL interactive) so a downstream consumer can pick the right
entry without reading the `repl_only` field first.
Refs terraphim-clients#134 P1, P2 (summarize)
Refs docs/plans/design-terraphim-grep-agent-fixes-2026-08-30.md
* test(agent): pin chat and summarize repl_only correctness
Pins the cross-file consistency for #134:
- `top_level_cli_commands_are_not_repl_only`: now includes
`chat` and filters by `repl_only == false` so the assertion
picks the CLI entry (and is not double-counted by the REPL
one in `repl-chat` builds).
- `repl_chat_is_marked_repl_only` (renamed from
`chat_is_marked_repl_only`): filters by `repl_only == true`
so it survives a `repl-chat` build that also contains the
new CLI `chat` entry.
- `cli_chat_is_marked_not_repl_only` (new): asserts the CLI
`chat` is present with `repl_only: false` and that the
`prompt` argument is required. Catches a future edit that
silently drops the required argument or flips the flag.
- `summarize_is_marked_repl_only` (new): asserts `summarize`
is `repl_only: true` when present (gated by `repl-chat`).
All four pass in default builds and in `--features repl-chat`
builds (6/6 tests in each).
Refs terraphim-clients#134 P1, P2 (summarize)
* refactor(agent): extract GuardTrace::print, make check delegate to check_with_trace
Two related simplifications in `terraphim-agent guard`:
P2.2: The `--explain` trace formatting was duplicated verbatim
across `run_offline_command` (lines ~2018-2048) and
`run_server_command` (lines ~4928-4952) — about 30 lines
each, differing only in the `*fail_open` deref. Drift risk.
Extracted `GuardTrace::print(json: bool)` in
`guard_patterns.rs`; the two call sites collapse to
`trace.print(*json)?`.
P2.3: `CommandGuard::check` and `CommandGuard::check_with_trace`
had ~95% identical bodies (`check` returns `GuardResult`,
`check_with_trace` returns `GuardTrace`). Made `check`
a one-line wrapper:
`pub fn check(&self, command: &str) -> GuardResult {
self.check_with_trace(command).result
}`
The trace is cheap to build (`Vec<4>` plus three
Aho-Corasick matches), so centralising the pipeline
eliminates ~70 lines of duplication.
Verified: 5/5 `guard_priority` tests, 5/5 `hook_safety` tests,
6/6 `robot_schemas` tests, 3/3 `search_only_flag` tests pass.
Clippy clean with `-D warnings` on
`--all-targets --features repl-chat`.
Refs terraphim-clients#134 P2.2, P2.3
* docs: clarify REPL vs CLI chat in robot-mode blog and reference
The robot-mode blog post and the 271-line agent reference both
documented `chat` as a single REPL-only command behind
`--features repl-chat`, omitting the top-level CLI `chat`
subcommand (gated by `--features llm`, default-on).
Update both to:
- Mention both flavours and the `repl_only` flag that
distinguishes them.
- Show both `chat` entries in the `robot schemas` example
output (with `false` and `true` rows).
- Note that `--features repl-chat` transitively enables `llm`,
so a `repl-chat` build carries two `chat` schemas.
- Add a `chat` (CLI) section to `docs/agent-reference.md`
before the existing `chat` (REPL-only) section, with the
one-shot `prompt` argument and `--role`/`--model` flags.
Refs terraphim-clients#134 P1
* ci(native-ci): build terraphim_server from terraphim-ai and run #113 integration tests (Refs #113)
terraphim_server is not a workspace member of terraphim-clients; it
lives in the private terraphim-ai repo. Build it from the v1.21.3 git
tag with cargo install --git so the runner has a real binary on disk,
then export TERRAPHIM_SERVER_BIN and run the three integration test
targets whose ensure_server_binary() / server_binary_path() helpers
look up the binary by that env var first.
Three landmines in 'cargo install --git', in order:
1. Multiple-binary repo. terraphim-ai has firecracker, dsm,
gitea_runner, merge_coordinator, server, eval_check -- refuses --bin
unless a positional <PACKAGE> selects one. Fix: trailing 'terraphim_server'.
2. Isolated context. cargo install does NOT inherit the workspace's
.cargo/config.toml, but terraphim-ai v1.21.3 has [patch.crates-io]
entries with registry = 'terraphim'. Without the registry declared,
parsing the cloned manifest fails: 'registry index was not found
in any configuration: terraphim'. Fix: --config 'registries
.terraphim.index=...' and --config 'registry.global-credential
-providers=[cargo:token]'.
3. Patch resolution. The v1.21.3 patches use caret ranges ('version
= "1.20.2"'), and the registry now publishes both 1.20.2 and
1.21.0. cargo refuses: 'patch for terraphim_automata resolved to
more than one candidate: 1.20.2, 1.21.0'. Fix: --locked honours
the v1.21.3 Cargo.lock, which pins terraphim_automata and
terraphim_types to exactly 1.20.2.
Also patch kg_ranking_integration_test.rs::ensure_server_binary() to
honour TERRAPHIM_SERVER_BIN (mirrors the cross_mode_consistency_test
helper; the kg_ranking helper had the right error message but never
actually read the env var).
This makes the server-binary-dependent integration tests run on every
push instead of silently failing fast because the binary is missing:
- crates/terraphim_agent/tests/cross_mode_consistency_test.rs (2)
- crates/terraphim_agent/tests/integration_tests.rs (5)
- crates/terraphim_agent/tests/kg_ranking_integration_test.rs (3)
Verified locally: install succeeds (4m08s cold), 2/2 + 5/5 + 1/3 pass
on the install side. The remaining 2 kg_ranking failures are pre-existing
test-data setup issues (docs/src haystack does not exist; tracked in
#84's 57 pre-existing failures).
Refs #113
* style(robot,grep,agent): apply rustfmt to cherry-picked commits (Refs #135b)
Run 278 / job 61817 failed cargo fmt --all -- --check on the
scope-creep commits cherry-picked from PR #135. The diffs are
all line-wrapping cleanups that rustfmt wants:
- crates/terraphim_agent/src/main.rs (GuardTrace::print refactor)
- crates/terraphim_agent/tests/guard_priority.rs
- crates/terraphim_agent/tests/hook_safety.rs
- crates/terraphim_agent/tests/robot_schemas.rs
- crates/terraphim_grep/tests/search_only_flag.rs
cargo fmt --all brings the tree back into compliance; no
semantic changes.
Refs #135b
* test(kg-ranking): point haystack at synced fixtures, replace missing 'python' term
Refs #113
The test_config.json haystacks referenced docs/src/, which exists in
terraphim-ai but not in this repo (terraphim-clients). After syncing
fixtures from terraphim-ai v1.21.3, the engineer's markdown content
lives at terraphim_server/fixtures/haystack/. Pointing both the
"Test Engineer" and "Default" Ripgrep haystacks at that path lets
the Ripgrep scorer pick up actual .md content instead of returning
zero results.
test_term_specific_boosting also searched for the term 'python', which
is not present in the terraphim-ai v1.21.3 fixture corpus (only
'rust', 'machine_learning', and 'neural_networks' markdown files are
synced). Swapping 'python' for 'neural networks' keeps the same test
intent (three varied terms) while using terms that genuinely exist
in the haystack.
Verified locally with TERRAPHIM_SERVER_BIN pointing at the cached
cargo install binary:
- test_role_switching: ok
- test_term_specific_boosting: ok (3/3 terms returned results)
- test_knowledge_graph_ranking_impact: ok
* test(cross-mode): un-ignore test_role_consistency_across_modes via per-role pre-warm (Refs #113b)
Removes the last `#[ignore]` in the integration suite. The cold-cache
slowness is paid exactly once per role, so warming up each role before
the timing-critical loop makes the 30s default ApiClient timeout
sufficient on CI. Results from the warm-up queries are discarded; only
the count-consistency assertions in the real loop are checked.
The warm-up uses `limit: Some(1)` so each call only pays the cold-cache
cost, not full pagination. If even a warm-up query times out, the test
fails loudly with the underlying transport error (no silent reliance on
a longer timeout).
Verified locally:
- test_mode_specific_verification ... ok (unchanged)
- test_cross_mode_consistency ... ok (unchanged)
- test_role_consistency_across_modes ... ok (un-ignored)
- 3 passed; 0 failed; 0 ignored; total 25.7s
* test(cross-mode): drop unnecessary explicit deref flagged by clippy
Removes the `*` deref of `warm_role` in the warm-up SearchQuery
construction. `warm_role: &&str` already auto-derefs to `&str` at
the `RoleName::new` call site; clippy::explicit_auto_deref (implied
by -D warnings) refused the explicit deref.
* fix(terraphim_agent): filter sub-word matches in extract (Refs #46)
Re-applies 31baedf8 onto current main. extract reported phantom term
labels (e.g. 'learning path' for the two-letter abbreviation 'lp'
matching inside 'aLPha') and started paragraphs at mid-word byte
offsets, because short thesaurus terms match as substrings via
Aho-Corasick with no word-boundary check.
extract_paragraphs now drops matches not flanked by word boundaries,
and labels each surviving paragraph with the surface form actually
present (Matched.term) rather than the concept expansion. Also
corrects the REPL/MCP extract path which shares this method.
Adds unit tests in service.rs covering: standalone word matches,
sub-word rejection, text-edge cases, and the full phantom-label /
mid-word-offset regression which is the headline defect in #46.
Note: the regression test borrows the thesaurus (`&thesaurus`) because
terraphim_automata 1.21.0 takes `&Thesaurus`. The original commit
passed an owned Thesaurus, which fails to compile against 1.21.0.
* test: provision docs/src/kg fixture + extend manifest fixture (Refs #84) (#141)
* test(terraphim_update): derive manifest fixture + document KG format (Refs #141) (#145)
* ci(terraphim-clients): run all workspace targets in test gate
Second per-repo remediation for terraphim/terraphim-agents#91.
Replace the library-only test gate with all-targets testing so
binaries, examples, and integration tests in tests/*.rs are
exercised. Regression context: 2026-07-31 (cargo test --lib
silently excluded the integration suite). ADF was not restarted;
no orchestrator configuration changed.
Scope: workflow yml only. Mirrors terraphim/terraphim-ai#3159.
* Fix #142: repoint find_files KG-scorer fixture at mcp_server (#12)
Repoint KG-scorer fixture at mcp_server.
find_files_with_kg_scorer_boosts_matching_paths built a thesaurus
containing the term 'automata' and asserted that a path under
crates/terraphim_automata/ would be boosted to the top of the
results. The terraphim_automata crate does not exist in this workspace,
so the assertion always failed.
Repoint the fixture at 'mcp_server' (which has a real
crates/terraphim_mcp_server/ directory) and update the assertion to
look for the matching path segment. The thesaurus still exercises the
KG-scorer boosting path; only the keyword and assertion substring
change.
Refs #142.
(cherry picked from commit b1c8247e0f59d339e6aa2d95307aa47647e568b7)
* Fix #143: hermetic MCP stdio tests (#13)
Make test_tools_list and test_all_mcp_tools deterministic and hermetic:
- Spawn the server with cwd set to a unique temp dir so
terraphim_config::project::discover() cannot walk up to a host
.terraphim/.
- Drain stderr on a background thread so the OS pipe buffer never
fills and SIGPIPEs the server mid-test.
- Drop the leading '--' separator before --verbose (clap Args::parse
rejects '--').
- Send the notifications/initialized frame between initialize and
tools/list.
- Switch test_all_mcp_tools to lightweight tools (json_decode,
find_files, grep_files) instead of the KG-backed tools whose
ensure_thesaurus_loaded walk hangs in CI when the path is empty.
- Move shared helpers into tests/support/mod.rs and silence the
per-binary dead-code warnings.
Refs #143.
(cherry picked from commit 5f54243b7f629a2806af99d369e7778dff721716)
* Fix #144: hermetic user_prompt_submit tests via TERRAPHIM_DEFAULT_DATA_PATH (#14)
The user-prompt-submit hook path uses LearningCaptureConfig::default() which
resolves global_dir via dirs::data_dir(). On macOS/Windows that ignores
XDG_DATA_HOME and returns $HOME/Library/Application Support, so the test
that set HOME and XDG_DATA_HOME never found the file it expected.
Production:
- Honour TERRAPHIM_DEFAULT_DATA_PATH in Default::default().
- storage_location() short-circuits to global_dir when the env var is set.
Test:
- Rewrite user_prompt_submit_tests to use support::cli_test_env helpers.
- No mocks, no #[ignore], no timeout increases.
All 4 tests pass on macOS.
Refs #144.
(cherry picked from commit 2ceda189769ae95f31a1a5f00e45c4aca3eaeb2a)
* test: resolve workspace binaries via CARGO_BIN_EXE, honour TERRAPHIM_SERVER_BIN
The terraphim_mcp_server and terraphim_agent integration tests located
their binaries by guessing <workspace>/target/debug/<name>. That guess is
wrong whenever CARGO_TARGET_DIR is set, which is how the native runner
builds (~/.cargo/build/by-runner/...), so under the all-targets gate every
stdio-driven MCP test and the extract-validation suite failed with
"binary not found" even though cargo had just built the binary.
Use env!("CARGO_BIN_EXE_<name>") instead: cargo sets it for a package's
own integration tests and guarantees the binary is built first, under any
target dir. TERRAPHIM_MCP_SERVER_BIN / TERRAPHIM_AGENT_BIN remain as
explicit overrides.
server_binary() in terraphim_agent's test support now honours
TERRAPHIM_SERVER_BIN, matching the other server-dependent tests, so the
CI-installed terraphim_server (not a workspace member, Refs #113) is used
by extract_functionality_validation too.
Refs #91
* ci: provision terraphim_server before the all-targets test gate
Run the #113 cargo install step before the workspace test step and export
TERRAPHIM_SERVER_BIN on it. With --all-targets the workspace step now
executes server_mode_tests, cross_mode_consistency_test, integration_tests
and kg_ranking_integration_test, all of which need the prebuilt server;
with --lib they were never compiled, so the ordering did not matter.
The focused #113 re-runs are kept for fast failure attribution.
Refs #91
* test: make hook_safety and search_only_flag hermetic
Both suites passed on developer machines and failed on the Gitea runner
(run 29586) because they read the host environment:
- hook_safety spawned terraphim-agent with the host's HOME, so the
PreToolUse thesaurus came from ~/.config/terraphim. On one machine a
personal KG term happened to match "-rf" (rewriting it to "Readiness
Feedback"); on the runner nothing matched, so the "KG-replaceable"
warning and the --rewrite substitution never fired. The tests now run
under support::cli_test_env::apply_hermetic_env (fixture role config,
KG at tests/test_kg) and a fixture concept tests/test_kg/trash.md maps
"rm -rf" to "trash" so the replacement path is genuinely exercised.
- search_only_flag asserts on a debug!-level line. terraphim-grep's
tracing filter defers to RUST_LOG when set, and the runner exports
RUST_LOG=info, which hid the line. The tests now pin
RUST_LOG=info,terraphim_grep=debug on the spawned process.
Verified with the full workflow step list under RUST_LOG=info and an
empty HOME (CARGO_HOME/RUSTUP_HOME kept): 91 binaries, 2231 passed, 0 failed, 3 ignored.
Refs #91
* docs(plans): add cass-parity session-search test plan + traceability matrix
Research artefact from the disciplined research pass (2026-09-03):
- 100-capability cass v0.6.11 catalog (C01-C100) benchmarked against
terraphim session search; verdicts 0 FULL/34 PARTIAL/46 MISSING/
18 N-A/2 UNCLEAR
- 158 test cases across 11 areas + 5 harness lanes + CI fixes
- Machine-readable C->TC traceability matrix (100 rows)
Refs #3084
* docs(plans): design + issue split for cass-parity session test suite (wave 1)
Design doc maps the research artefact (#148) onto verified repo reality:
- terraphim_agent/terraphim_sessions are workspace-local here (path dep
1.21.2) — research decision R1 (registry-vs-local canary) dissolves
- Cursor connector exists with 15 tests (research assumed partial)
- native-ci.yml already --all-targets; remaining CI gap = --all-features lane
- search_nfr bench not carried into this repo -> PR-5 ports it (refs #3014)
5 sequenced PRs -> issues terraphim-clients #150..#154.
Refs #3084
* test(sessions): parity harness — shared fixture builders + --all-features CI lane
Closes terraphim-clients#150 (wave 1 PR-1 of the cass-parity suite).
Design: docs/plans/design-session-test-suite-2026-09.md.
Research: docs/plans/research-session-test-parity-2026-09.md.
- search_tests_support.rs (cfg(test)): deterministic Session/Thesaurus
builders + claude-jsonl/aider-history corpus writers so parity tests
never read real user session stores
- CI: terraphim_sessions --all-features lane (cursor/codex/extras +
search-index score module are invisible to default/enrichment lanes)
Verified: 74 default / 88 enrichment / 122 all-features tests green;
clippy -D warnings clean.
* test(sessions): hybrid KG-boost ordering suite (P0 parity rows)
Closes terraphim-clients#151 (wave 1 PR-2). Covers parity rows
C07(p)/C62/T15/T18/E17a from the cass-parity research artefact:
search_with_thesaurus/search_sessions_hybrid had zero tests.
8 tests (feature-gated enrichment):
- boost promotes enriched session above raw-BM25 leader (ordering only,
never score equality — fusion math is implementation detail)
- boost monotone in thesaurus match count
- None/empty-thesaurus degrade to plain BM25 (identical ordering)
- empty query/corpus edge cases, deterministic ordering
- unenriched sessions keep BM25 relative order; no-boost query still
returns BM25 results
Verified: 96 tests green (--features enrichment); 74 default;
clippy -D warnings clean.
* test(sessions): import contracts — global limit, auto-import single-attempt
Closes the import-contract portion of terraphim-clients#152 (wave 1 PR-3a):
- import_all global-limit truncation via native connector (3-session
tempdir corpus, limit=2 -> 2 sessions)
- auto-import single-attempt contract: consecutive cache-touching calls
observe the same session set (no re-import); count is env-dependent by
design so only the stability contract is asserted
Both hermetic (tempdir corpora via ImportOptions.path). Full connector
hermetic suites follow in the next commit on this branch.
* test(sessions): per-connector hermetic import suites (cline, aider)
Completes terraphim-clients#152 (wave 1 PR-3):
- cline: taskHistory.json + api_conversation_history.json tempdir corpus
-> 1 session, role order asserted; empty dir imports nothing (the
taskHistory/api path had zero import tests before)
- aider: .aider.chat.history.md tempdir corpus via ImportOptions.path
(detection walk bounding already covered by #123 fix; import itself
was untested)
Verified: 135 tests green with --all-features; clippy -D warnings clean.
* test(agent): REPL/CLI session contract tests (robot JSON, exit-4, flag order)
Closes terraphim-clients#153 (wave 1 PR-4). First session integration
tests in the agent crate — covers parity rows C40(p)/C80(p)/C87(p)/
C89(p)/C99(p) from the cass-parity research artefact.
7 tests in crates/terraphim_agent/tests/sessions_cli_contract.rs:
- sessions sources membership JSON (claude-code-native always present)
- machine-mode empty search: exit 4 AND zero-payload (pairing rule)
- --robot after subcommand rejected with ERROR_USAGE(2) (root-flag rule)
- search/list/stats JSON shape contracts (session_output structs)
- human-mode no-match text path
All hermetic via tests/support/cli_test_env.rs (temp HOME).
* test(agent): DOCS-DRIFT probes + NFR bench wiring (port #3014 into clients)
Closes terraphim-clients#154 (wave 1 PR-5, final).
DOCS-DRIFT probes (tests/sessions_docs_drift.rs):
- CLAUDE_SESSIONS_DIR documented-but-unimplemented: env var has NO effect
on discovery (asserted); doc decision: fix the skill docs
- robot capabilities advertise supported_formats [json,jsonl,minimal,table]
but CLI OutputFormat enum is human|json|json-compact — drift pinned,
fix deliberate
- /sessions import removal: CLI rejects as unrecognized subcommand;
REPL parser returns the explanatory message
NFR bench: port benches/search_nfr.rs from terraphim-ai@8fb947863
(10K deterministic sessions, criterion) + criterion dev-dep +
[[bench]] required-features search-index. Measured locally:
search_sessions_10k ~89ms [88.2-91.4] — within the <100ms G1 NFR.
Verified: 135 sessions tests green (--all-features); agent clippy clean.
* docs(verification): wave-1 verification report for session-test-suite-2026-09
Records the 5/5 merge status, local verification evidence, G1 NFR
measurement (89ms/10K), docs-drift findings, and the wave-2 backlog.
Refs #3084.
* fix(tests): address PR-review P2 findings — panic-safe tempdirs, drop dead-code shim
Follow-up to the structural PR reviews (wave 1, comments on #156/#158/#160):
- service.rs import-limit test: tempfile::tempdir() (panic-safe) instead of
manual std::env::temp_dir + remove_dir_all
- sessions_docs_drift.rs: remove unused apply_hermetic_env import + the
_unused() shim
Remaining review findings tracked as wave-2 backlog (nightly bench job,
p50 regression test, seeded-corpus entry-shape assertions, auto-import
hermeticity rewrite).
* docs(terraphim_agent): document LearningStore hybrid scoring (Refs #850)
Add an Unreleased entry for the LearningStore::query_relevant graph-rank
hybrid path landed in 13c5a36. The trait impl now ranks candidates by
RoleGraph::query_graph rank when a role graph is configured, with the
min_trust and applicable_agents gates preserved, and falls back to
substring text matching otherwise.
* style: cargo fmt on sessions_cli_contract.rs (fmt gate caught drift)
* docs(verification): addendum — structured PR review round results for wave 1
All five wave-1 PRs reviewed (structural-pr-review), scores recorded,
mechanical fixes merged (#161), wave-2 backlog enumerated. Pre-existing
PRs #147 (merged after rebase+review) and #146 (scope mismatch flagged)
also covered. Refs #3084.
* feat(sessions): redact API keys and secrets in session import connectors Refs terraphim/terraphim-ai#1986
Add crates/terraphim_sessions/src/redaction.rs with redact_session_content() which
applies regex patterns for AWS keys, OpenAI/GitHub tokens, Bearer tokens, and
connection strings. Call redact_sessions() in import() of all five connectors
(native, aider, cline, opencode, codex) before returning.
Make regex a mandatory dep (was optional/aider-connector-gated) since redaction
applies to all connectors regardless of feature flags. Fix pre-existing clippy
collapsible_if warnings in service.rs (two if-let nesting blocks).
63 tests pass, 8 new redaction tests, 1 doctest.
* fix(redaction): compile redaction patterns once via OnceLock
PR-review P1 fix on the revived #34 branch: redact_session_content
recompiled all 10 regexes per message, making auto-import over a real
corpus (151MB ~/.claude) pathologically slow (test hang >5min diagnosed
via spindump: Regex::new dominating redact_sessions). Patterns now build
once per process via OnceLock. Auto-import contract test completes in
seconds; 85 lib tests green; clippy -D warnings clean.
* test(procedure): add CLI integration tests for learn procedure from-session Refs #2350
Adds three integration tests to procedure_cli_tests.rs (all feature-gated
behind repl-sessions):
- procedure_from_session_extracts_non_trivial_commands: verifies that from-session
filters trivial commands (cd) and failed commands (exit_code != 0), auto-generates
a title, and reports correct step and command counts.
- procedure_from_session_deduplicates_on_repeat: verifies that running from-session
twice with the same session results in one procedure via save_with_dedup().
- procedure_from_session_missing_id_fails: verifies that a missing session ID causes
a non-zero exit code.
The core implementation (extract_bash_commands_from_session, from_session_commands,
FromSession CLI variant, and the existing unit test) was already in place. This
commit provides the CLI-level evidence required by AC 6 of terraphim-ai#2350.
Co-Authored-By: Terraphim AI <team@terraphim.ai>
* fix(tests): make from-session CLI tests platform-hermetic + align dedup assertions
Revival of #21 (terraphim-ai#2350 coverage), two fixes required before merge:
- Fixture write: dirs-5.0.1 macOS ignores XDG_CACHE_HOME (cache_dir =
HOME/Library/Caches), Linux honours it — write sessions.json to all
platform-mirrored cache variants (rule from the parity design doc).
Original test false-failed on macOS / false-passed on Linux CI.
- Dedup assertion aligned with current save_with_dedup semantics: merge
only happens for high-confidence existing procedures; a fresh 0%-
confidence procedure does not merge, so two identical runs yield two
procedures (was: older always-merge behaviour).
15/15 procedure_cli_tests green; fmt clean.
* docs(verification): addendum 2 — PR queue triage, revived PRs, >=4/5 gate results
* ci: add manual publish-registry workflow for terraphim registry Refs terraphim/terraphim-ai#2515
Adds a workflow_dispatch-triggered workflow to publish any crate to the
terraphim Gitea cargo registry. Requires the secret
CARGO_REGISTRIES_TERRAPHIM_TOKEN (package:write scope) to be set in
the repo CI secrets.
To publish terraphim_sessions 1.20.4 and unblock issue #2515:
1. Admin: set CARGO_REGISTRIES_TERRAPHIM_TOKEN secret
2. Trigger this workflow with crate=terraphim_sessions
3. Merge terraphim-agents PR #60
* feat(sessions): implement expand subcommand (Task 2.6.4) Refs terraphim/terraphim-ai#2134
Add SessionsSub::Expand variant to the terraphim-agent CLI:
- SessionsSub::Expand { id, context_lines } added to enum
- SessionExpandOutput + ExpandedMessage types added to session_output mod
- Handler in offline/async path: loads from disk cache then get_session()
- Handler in server/rt.block_on path: auto-imports via list_sessions() then get_session()
- Human-readable output prints all messages with role headers
- Machine-readable output (--format json) wraps via robot envelope
- context_lines field reserved for future --query-aware context expansion
- 2 new unit tests verify JSON serialisation of output types
- cargo clippy -D warnings: clean
- cargo fmt --check: clean
- 465 unit tests pass (cross_mode_consistency_test pre-existing polyrepo infra gap)
* style: cargo fmt + drop unused cache_dir binding in dedup test
* feat(learn): auto-suggest corrections from KG in capture_failed_command Refs terraphim/terraphim-ai#1648
When ToolPreference corrections exist in the storage directory, search the
compiled corrections thesaurus for patterns matching the failing command or
error text. If a match is found, set learning.correction to the suggested
replacement so that `terraphim-agent learn list` surfaces it immediately.
The auto-suggest is non-blocking: if the corrections directory is empty or
the thesaurus search fails, capture continues without setting the field.
Adds test_capture_sets_correction_when_kg_match_found regression test.
* feat(learn): implement KG auto-suggest corrections in capture_failed_command
When capture_failed_command annotates entities from a failed command, it
now calls compile_corrections_to_thesaurus to check whether any matched
entity has a known ToolPreference correction. The first match is set as
learning.correction so that terraphim-agent learn list surfaces the
suggestion without any manual intervention.
Adds suggest_correction_from_entities as a public helper for unit testing
and adds test_capture_sets_correction_when_kg_match_found covering match,
no-match, and empty-entity cases.
Refs terraphim/terraphim-ai#1648
Co-Authored-By: Terraphim AI <noreply@terraphim.ai>
* fix(learn): resolve duplicate test name breaking PR #29 build
The terraphim_agent bin test target failed to compile because
`test_capture_sets_correction_when_kg_match_found` was defined twice
(a merge/copy-paste artefact). The two tests cover different behaviour:
one drives capture_failed_command end-to-end via the compiled thesaurus,
the other unit-tests suggest_correction_from_entities directly. Rename
the latter to test_suggest_correction_from_entities_matches_tool_preference
and apply rustfmt. Both tests are retained.
Unblocks native-ci / build on task/1648 (PR #29).
Refs #30
Refs terraphim/terraphim-ai#1648
* fix(learn): borrow corrections thesaurus in find_matches (1.21.x borrowed-&Thesaurus family)
* feat(learnings): implement CorrectionEvent CLI and security hardening Refs #2083
- Restructure `learn correction` into sub-subcommands:
- `learn correction add --original X --corrected Y --correction-type T` (records a CorrectionEvent)
- `learn correction list [--filter-type T] [--global]` (lists stored corrections, optionally filtered by type)
- Add YAML frontmatter injection protection: sanitise_yaml_value() strips newlines from hostname, session_id, working_dir, id, and tags written to .md files
- Add 64 KiB per-field size limit in capture_correction() to prevent disk exhaustion
- Escape backticks in markdown body so inline code blocks are never broken
- Add #[serde(default)] to CorrectionEvent.tags for backward-compatible deserialisation
- Add four new security-focused tests: YAML injection via hostname/session_id, oversized input rejection, backtick escaping
* feat(terraphim_lsp): add KG analysis engine with Aho-Corasick term matching Refs #2669
Adapted to current main: terraphim_automata at registry 1.21.0 with
borrowed-&Thesaurus find_matches API; terraphim_negative_contribution
and terraphim_types at current path/registry versions.
* test(agent): REPL /sessions parser contract tests (adapted from #39, owner decision A)
Land the still-valid parser tests from stale PR #39 per owner decision A:
- /sessions import stays REMOVED (auto-import replaced it); the removal
message itself is pinned as a contract test
- REPL has no expand alias (expand is CLI-only via #165); show/get
aliases pinned instead
- 13 parser-contract tests: list (+ --source/--limit), ls alias, search
(+missing-query error), show/get, sources/detect, /session singular,
missing/unknown subcommand errors
Closes terraphim/terraphim-ai#2435 (test half). Supersedes #39.
* feat(agent): add shared-learning to default features (Refs terraphim-ai#2516, owner decision A)
learn shared list/promote/import/stats subcommands are now available in
production builds without --features shared-learning. The collapsible_if
clippy fixes from the original stale commit are already on main (let-chains).
Verified with default features only: 498 lib tests, 9 shared_learning_cli
tests, clippy -D warnings clean, fmt clean, CLI sanity (learn shared list)
exits 0.
* refactor(agent): bin reuses lib modules — eliminates twin module tree
The bin re-declared lib modules (mod client; mod robot; ...) so every
lib item was compiled twice: reachable (no warnings) in the lib target,
privately-duplicated (dead_code warnings -> silenced by 152
#[allow(dead_code)] across the workspace) in the bin target.
main.rs now imports the lib modules (terraphim_agent::{...}) instead of
re-declaring them; main-only modules (listener, shell_dispatch,
kg_validation, session_output, test mods) stay local. learnings:capture
list_learnings is now re-exported from the learnings root.
This removes the need for most dead_code suppressions in the agent
crate. Follow-ups: TSA + remaining crates.
* refactor(robot): drop dead_code allows on submodules (twin-tree fix made them unnecessary)
* refactor(agent): purge remaining dead_code suppressions — wire, read, or delete
Agent crate now has ZERO #[allow(dead_code)] (was 87):
- capture.rs: delete old auto-extract/suggest pipeline (score_entry_relevance,
TranscriptEntry, ScoredEntry, suggest_learnings, auto_extract_corrections,
contains_correction_phrase, extract_command_from_input) — superseded by
SharedLearningStore::suggest hybrid scoring (13c5a36)
- install.rs: delete unwired uninstall_hook/is_hook_installed/
get_installation_status CLI plumbing
- hook.rs: read the serde(flatten) extra map in from_json (debug log) —
it exists for unknown-field forward-compat
- executor.rs: drop never-read api_client field + unused with_api_client
- validator.rs: drop dead determine_execution_mode wrapper
- hybrid.rs: drop never-read vm_for_unknown setting
- listener.rs: drop dead run_once/handoff_issue wrappers
- shell_dispatch.rs: move MAX_OUTPUT_BYTES into test scope, delete unused
DISPATCH_TIMEOUT_SECS
Workspace total: 152 -> 67 allows (TSA crate + test files remain).
* refactor(tsa): bin reuses lib modules + purge dead correlation code
- main.rs now uses terraphim_session_analyzer::{...} instead of
re-declaring the module tree (same twin-tree fix as the agent crate)
- delete calculate_agent_tool_correlations + CorrelationData from main.rs
(TODO on the fn said superseded by Analyzer::calculate_agent_tool_correlations)
- 52 dead_code allows removed; TSA: 128+2+20 tests green, fmt clean
* refactor: purge all production dead_code suppressions (closes #5, #1)
Workspace #[allow(dead_code)]: 152 -> 11 (all 11 in test binaries,
documented shared-support pattern; production source: ZERO).
Key changes:
- agent: twin module tree fixed — bin imports lib modules instead of
re-declaring them (client.rs 30 allows + robot/mod 5 were silencing
the bin duplicate); deleted genuinely-dead auto-extract pipeline in
capture.rs (superseded by SharedLearningStore::suggest hybrid scoring),
unwired uninstall/status hook fns, unused executor api_client field,
dead determine_execution_mode wrapper, never-read vm_for_unknown,
dead listener wrappers; hook.rs ToolInput.extra now read (debug log)
- tsa: bin reuses lib modules; deleted calculate_agent_tool_correlations
+ CorrelationData (TODO said superseded by Analyzer impl)
- sessions: serde DTO fields dropped where never read (msg_type,
ComposerTab.model) — serde ignores unknown fields, behavior-safe
- update: deleted hand-rolled is_newer_version (superseded by semver
static fn); test migrated to is_newer_version_static
- mcp_server: frecency field now read via getter + startup debug log
All gates green: workspace clippy --all-targets --all-features -D
warnings clean; fmt clean; 143+493+130+15+128 sessions/agent/update/tsa
test suites pass.
* fix(agent): derive thesaurus_matched from the boundary-aware matcher
thesaurus_matched was built with a naive substring scan
(`query.to_lowercase().contains(key)`) independently of the automaton, so
any thesaurus term appearing inside a longer query word was reported: the
two-letter term `ce` matched `con(ce)pt` and surfaced in the robot-mode
search envelope for a query that matched nothing.
Derive it from the same matcher that produces concepts_matched, at both the
offline and server (run_server_command) sites, so the two fields cannot
disagree.
Note this only closes the agent-local half. concepts_matched comes from
terraphim_automata, pinned here to 1.21.0 from the Gitea registry, and stays
wrong until terraphim-core#65 is published and the pin bumped.
Refs #197
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U7LFsSpVkhACQS6HYM42JH
* fix(agent): honour --fail-on-empty in server mode
`run_server_command` destructured `fail_on_empty: _`, so the flag was
silently ignored whenever the agent talked to a server: an empty result set
exited 0 instead of 4, and only the offline path behaved as documented.
Capture `results_count` before `res.results` is consumed by the robot
formatter, and apply the same exit check the offline handler uses.
Refs #197
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U7LFsSpVkhACQS6HYM42JH
* fix(redaction): apply central redaction before every persistence boundary (Refs #178)
Establishes a single canonical redaction policy at every persistence
boundary in terraphim_agent::shared_learning and terraphim_sessions::cla.
Pinned 7-call-site coverage:
- markdown_store::save / save_to_shared (redaction in to_markdown)
- wiki_sync::sync_learning (covers sync_all_learnings + sync_batch
via the single funnel point)
- from_normalized_session / from_normalized_message in terraphim_sessions::cla
(covers both ClaClaudeConnector and ClaCursorConnector via the shared
helper)
Implementation:
- New terraphim_sessions::redaction module with inline SECRET_PATTERNS
list + drift-guard test asserting parity with the canonical
terraphim_agent::learnings::redaction.
- New terraphim_agent::shared_learning::redaction module with the same
drift guard, exposed via 'pub use redaction::redact_secrets'.
- regex promoted from optional to direct dep on terraphim_sessions
(central redaction is now a baseline requirement).
- Logging-hygiene tests asserting no tracing::warn! / tracing::debug! /
dbg! carries the pre-redaction body field.
Tests:
- terraphim_sessions: 76 lib tests pass (incl. drift guard + 1 new
real-fixture connector test, no mocks).
- terraphim_agent --features shared-learning: 308 lib tests pass (incl.
new redaction module + 7 new boundary tests).
Closes #178 (canonical redaction policy).
Refs PR #22 (will be closed as superseded once this lands; PR #22's
4 surviving non-redaction P1s become a smaller follow-on).
* fix(security): validate wiki_page_name + XSS strip + component validation (Refs #22)
Ports the 3 surviving P1 security findings from PR #22 (closed as
superseded by PR #200). Each P1 lands with real-fixture tests
(no mocks per AGENTS.md) and bounded scope per Phase 2 design.
P1-1 XSS strip (Refs #22): strip <script>, <iframe>, <object>,
<embed> blocks from wiki markdown body. Case-insensitive, handles
self-closing forms, unclosed tags, nested tags, attributes, and
multibyte content safely. 13 unit tests cover adversarial variants.
P1-2 Page-name validation (Refs #22): wiki_page_name must match
^[a-zA-Z0-9_][a-zA-Z0-9_-]{0,199}$ — leading '-' rejected …
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Adds a repeatable acceptance run for the public release entry points, and
records what an attempt to publish the dependency family actually proved.
scripts/acceptance-public-release.pyRead-only, and reports
pass,failornot-executedper entry point, so anenvironment limitation is never counted as success. Checks:
binary must report the channel version and the run must report checksum
verification
cargo installwould serveExit codes:
0all executed checks passed,1a check failed,2theinstaller check failed,
3nothing was reachable.Current live result:
The installer check reports
not-executedrather than a false pass untilterraphim/terraphim-ai#964 is merged, because the published script is still the
pre-fix one. The Homebrew check reports the same on a host without
brew.What the family publish proved
Publishing the dependency family was attempted against live infrastructure. A
dry run of the full client family on
main(run 36314289075)
reaches
terraphim_config 1.20.4and fails to compile:The chain, verified end to end:
cargo installon a client crate resolves against crates.ioterraphim_configandterraphim_persistenceat 1.20.4,while the client crates require 1.20.2 (pinned exactly, because 1.20.4 is
yanked on the Gitea registry)
terraphim_automatais 1.21.1, which gatesparse_markdown_directives_dirbehindfs-traversal, a featureterraphim_config1.20.4 does not enableterraphim_config1.20.4 fails, so no client crate ispublishable
This is the failure the
[patch.crates-io]block already documents, reachedthrough the publish pipeline itself. The repair lives in
terraphim-coreandterraphim-config-persistence, not here.The same run confirmed two smaller facts:
publish-crates.ymlrefusesterraphim_agentbefore mutating any manifest (#95), as designed; andterraphim_update1.20.2 andterraphim_command_runtime0.1.0 already exist oncrates.io, so they are no-ops in a future
crate_list.Contents
scripts/acceptance-public-release.py(new)docs/plans/— the research and design documents for the public-releaseitems, including the addendum recording the above
Testing
python3 scripts/acceptance-public-release.pyruns against live infrastructure;the run above is its actual output. The script is also exercised with
--skip-installerand--skip-crates.