Skip to content

sec: sec-af's security audit, built into codeaf as its second program - #1781

Merged
ZeroPoint95 merged 26 commits into
devfrom
zeropoint95/sec
Oct 7, 2026
Merged

ZeroPoint95 merged 26 commits into
devfrom
zeropoint95/sec

Conversation

@ZeroPoint95

@ZeroPoint95 ZeroPoint95 commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator
1781

Followed by #1784, which builds pr-af's code review into codeaf as /review on this branch's shared agent loop (internal/agentsession). It is stacked on this branch and retargets to dev when this merges.

What changed

sec-af, AgentField's security audit, is built into codeaf as its second carried program: /sec in the chat, codeaf sec at a shell, via: "sec" from a proposal. It is the senior-dev treatment applied to sec-af: the Go port is copied in once at sec-af's tag codeaf-absorb (47d57d7) as internal/secaf and frozen there, and every model call it makes goes through the run's model API, priced and held to the run's ceiling. It reads the repository with four read-only tools (read_file, list_dir, glob, grep), changes no files, and lands text: an account in the conversation and sec-report.md / .json / .sarif in the task's record folder. The chat then offers a senior-dev fix, and starts nothing.

  • Scope: /sec audits the whole repository; /sec changes (or changes since <ref>) audits a branch's changes, read in the folder under audit, with the hunters told what changed; quick / thorough set depth.
  • Its phases keep sec-af's names everywhere (recon, hunt, prove, remediate, report); its page says which hunter found what; machinery words are kept out of what a person reads.
  • Ceilings are per program now (Delegate.Unattended), not senior-dev's chosen by name in nine places. sec's are $5 and four hours. A program with none is held to the conversation's limits, and its ending names them.
  • A program whose brief is words (Delegate.Words) is handed exactly those words from a proposal, as from a typed command. The composed brief opens on WHAT THE PERSON ASKED FOR, and sec read that as its scope.
  • A report program's ending: the account's first line says what it found (the line the project index keeps), its index row points at its record folder, and a run cut by its time ceiling says where it was cut.
  • Model API: a call held only by calls still in flight is answered 429 + Retry-After + X-Codeaf-Held, and the client waits it out. 402 is kept for a ceiling truly reached. This used to refuse concurrent calls at $0 spent.
  • A program's ending reaches the person even when the turn it wakes cannot answer: the account is written into the conversation once.
  • The markup detector no longer treats prose marks (/, backticks, punctuation) as fences around a tool's name. A summary naming …/tasks/1/… in a table was cut as leaked markup four times in a row.
  • Fixes to sec-af's own bugs found on the way: compliance mapping failed every audit that named frameworks, and PR mode read the diff from the wrong folder.
  • Manual: a new sec page, plus updates to delegates, commands, tasks, models-and-cost and running-from-the-terminal, with probes for each.
  • Budgets: SIZE-BUDGET rises by exactly what linking internal/secaf measured (+2,512,992 bytes), and the prefix waivers by sec's guide bytes; both are recorded in PERF.md.

Design notes: docs/design/security-audit/ABSORB.md, docs/design/delegate/PROTOCOL.md.

Two mid-branch commits (6338d7e66, f5568bc49) do not compile on their own; d35442a1d restores them, and the tip builds clean. The branch should be squash-merged.

How it was checked

  • make pr-ready (on 04dd61f90): the light gate, the law tests and the whole internal/session suite (8 shards) passed. Two failures were this branch's and are fixed in the last two commits: TestEveryCompiledModuleHasItsSection (the notices now carry sec's new modules, all MIT, BSD-3 or Apache-2.0) and TestRegistryCoversEveryUserFacingEnvironmentPin (CODEAF_RECORDS is registered as plumbing). Five more fail identically on dev 7e3033564 on this laptop and are not this branch's: TestGrepWithoutRipgrepFindsWhatRipgrepWouldFind (exec/bare), TestTheReleaseSurfaceTestSurvivesEveryLedgerState (release), and in tui3 TestANewAtTokenWaitsForTheWalkInFlightAndIsSettledByItsFollowUp, TestAtBoxesShareOnePendingRecentCatalog and TestRecentWalksNeverOverlapAndCoalesceOpeningBursts. The classifier did not probe the base for these itself because they exceeded its five-failure cap, so I ran each one at dev by hand.
  • internal/secaf: a stub-model end-to-end audit (whole pipeline, changes nothing, writes all three report files); scope reading; changes mode; account and report wording, including the first line and the cut-by-ceiling sentence.
  • internal/session: TestAProposedProgramWhoseBriefIsWordsIsHandedOnlyItsWords reproduces the live failure (fails without the fix, handing over the composed brief) and checks that the index row points at the record folder; TestAProgramsEndingIsSaidWhenTheTurnItWokeCannotAnswer; and the bare-start, ceilings and report-ending tests.
  • internal/provider: TestMachineryLeakLeavesAToolsNameInProseAlone, using the exact reply that was cut live; modelapi's held-call test.
  • internal/run: TestEstimatedRefusalLandsAsConversationCostLimit, which this branch had broken and now fixes.
  • Live: a whole-repository standard audit of a mid-sized Rust repository took 1,952 calls, cost $2.40 and ran 1h57m. That run is why the time ceiling went to four hours, and it found the brief and markup bugs above.

Checklist

  • A change entry — docs/changes/unreleased/1781-sec.md
  • The manual knows about it — sec.md is new, and the delegates/commands/tasks/models-and-cost/terminal pages are updated, with probes
  • No new line in .github/known-red.txt
  • Only my own paths are staged — no git add -A

Not in this PR: sec-af's codeaf-absorb tag exists only in the local sec-af checkout and has not been pushed.

🤖 Generated with Claude Code

ZeroPoint95 and others added 15 commits October 5, 2026 18:33
… senior-dev's

Every unattended ceiling was senior-dev's, chosen by its name in nine
places, so a second program started from the chat ran on whatever the
conversation had left. A program now names its own (Delegate.Unattended);
senior-dev's are unchanged.

A program that lands text had no ending of its own: the wake turn sent the
model looking for a branch and a worktree that were never cut. It now reads
a report as a finished answer, asks for a short summary and the program's
own offer (Delegate.FollowUp), and starts nothing. A program is told its
record folder (CODEAF_RECORDS) for the files its report points at.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sec-af was an AgentField node: a control plane carried its calls, a router
key it held paid for its models, and a coding-agent binary ran each of its
agent sessions. It is now copied into codeaf once, at sec-af's tag
codeaf-absorb (47d57d7), as internal/secaf, and runs only through codeaf:
/security-audit in the chat, codeaf security-audit at a shell, and
propose_task with via "security-audit".

Its algorithm is kept: the phases, the hunters, the four-agent proof
chain and the prompts. What it runs on is codeaf's (internal/secaf/backing):
model calls go to the run's model API, each agent session is a read-only
loop of four tools with a schema-checked answer, and calls between its
reasoners stay in the process. It changes nothing in the folder; its report
goes to the task's record folder and its account to the conversation, which
offers to hand confirmed findings to senior-dev.

An audit of the changes (the branch since its base, uncommitted work
included) tells the hunters what changed instead of filtering a whole-
repository scan afterwards. Naming compliance frameworks no longer fails
every audit at its end.

A program may now run bare on a default brief, show its own arguments on
its row, and say no ceiling it does not have. The prefix waivers and the
size budget rise by exactly what the program measured (PERF.md).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A call the ceiling holds only because calls in flight have reserved
  what is left is answered 429 with Retry-After and X-Codeaf-Held, and
  opens no turn; 402 stays for a ceiling truly reached. A program that
  makes many calls at once was told its ceiling was reached at $0 spent.
  sec-af's client waits a held call out within its own bounds.
- A program names the flag that carries its model (Delegate.ModelFlag);
  the shell resolved only senior-dev's --high.
- A shell run of a program that answers prints its answer, not what a
  tree program's model claimed; an unset ceiling is not said as $0.00.
- A tool with no required argument sent required: null, which a strict
  server refuses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A typed /security-audit quick was titled "quick"; a program may now title
its typed runs (Delegate.Title), and the audit's say what it audits. Each
agent's "starting" note duplicated its session's line and is left off the
page; the hunters' are kept. The protocol spec and the programs page say
the ceiling's held answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
/security-audit is /sec, codeaf security-audit is codeaf sec, and
via: "security-audit" is via: "sec". Its manual page is sec, its narrow
badge [s], and a shell run's records go under ~/.codeaf/v3/carried/sec/. The
report's files keep their descriptive names (security-audit.md, .json,
.sarif). The shorter name takes eleven bytes off the fixed prefix and the
waivers come down with it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The page headed its steps MAP, TEST and FIX while sec-af's own notes under
them said RECON, PROVE and REMEDIATION. The row, the headings and the notes
now all use sec-af's phase names (recon, hunt, prove, remediate, report),
and its agents keep their own names except where they are banned words,
which are reworded in the record itself so a shell run says them the same
way. The CWE expansion note, which claimed a widening the hunters never
see, is left off, and the dedup note no longer speaks of fingerprints. The
verdict agent, a single call rather than a session, now has its line, so
the proof chain shows all four agents.

security-audit.md was sec-af's report: Verdict: inconclusive, not
exploitable, Cost: $0.00, Commit: HEAD, Provider: harness. It is written by
sec in the account's words (confirmed, likely, unclear, ruled out) with
each finding's trace, attack, fix and patch, and only figures that were
measured. The JSON and SARIF keep sec-af's field names.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
f609ea5ec renamed security-audit to sec but staged these two files before
their last edits, so it and b403e776e name programguide.Sec while the file
still declared SecurityAudit, and the entry still said security-audit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d what

The report files were security-audit.* and the SARIF named its tool SEC-AF,
with sec-af's link, sec-af/ rule ids and properties. They are
sec-report.md, .json and .sarif (and sec-compliance.md); the SARIF names
sec, codeaf's build and home, and sec/ ids. The JSON and the compliance
report carry the run's own cost and agent count where sec-af left zeros.
sec-af's writers keep their own identity by default, so its goldens hold.

Every hunter's sessions were "hunt location scanner" and "hunt finding
enricher", so the hunt read as two lines repeated eleven times. Each says
its hunter and what it found: "injection hunter · scan" with how many
places, and "injection hunter · app/views.py:4 · <finding>" with its
severity. The hunters' start notes, which those lines replace, are off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A two-hour sec run on furrow ended, the turn it woke read the report and
began a good summary, and every model it was offered was cut as "the
model's own internal markup". The partial summary was kept as an
interrupted message the chat does not draw, so the person saw a done card
and nothing else.

- The markup detector read a tool's name between two prose marks as tool
  grammar: the report's path, …/tasks/1/sec-report.md, spells the tool
  tasks between two slashes, and a findings table put the reply at a tenth
  symbols. Path separators, backticks, emphasis and punctuation are no
  longer fence material; the leak shapes (bars, brackets, quotes) still are.
- A turn woken by a program's ending that cannot finish an answer now
  writes the program's own account into the conversation as the session's
  line, once per run.
- A sec run its time ceiling cuts says so, in which phase, and what was
  left undone, in its account, ending and report; the demotion notes that
  cut produced read as findings staying unclear, not verifier_error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A standard audit of furrow made 1,952 calls, spent $2.40 of its $5 and was
cut by two hours in prove with no fixes written: time is the ceiling it
meets first, so the hours double and the dollars stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ints at its report

A run the chat proposed was handed the composed brief, whose first line is
WHAT THE PERSON ASKED FOR, and sec reads its scope off the first line: a
proposed `whole repository thorough` ran at standard depth, and a proposed
`changes` would have audited the whole repository. A program whose brief is
words (Delegate.Words) is now handed those words alone.

The project index kept only "Security audit of the whole repository." of a
run and pointed at the repository, so another conversation searched the disk
for the report and opened a different run's first. The account's first line
now says what it found, and a report program's row points at its record
folder.

And a program with no ceilings of its own is held to the conversation's,
whose ending names them: it had read "the run's $0.00 limit".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
internal/secaf compiles in santhosh-tekuri/jsonschema (its answers' schema
checks) and invopop/jsonschema with what they pull in; each has its section,
regenerated by codeaf-notices.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…'s stop told as theirs

The change entry's `pr:` still said 1757: the rename was committed without
the edit to the field. Two pointers named docs/design/sec/ABSORB.md, which is
docs/design/security-audit/ABSORB.md. Delegate.Unattended's comment gave an
audit a quarter of an hour; it now points at the two programs' own figures.

A program's ending the wake turn did not answer was always told as a failed
reply, and a person's stop ends that turn the same way. A stop now keeps the
account and says the person stopped the answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ZeroPoint95
ZeroPoint95 marked this pull request as ready for review October 6, 2026 01:05
ZeroPoint95 and others added 6 commits October 6, 2026 09:44
…gram's work

internal/secaf/backing and internal/secaf/appx are codeaf's own, not sec-af's,
and /pr runs on them too, so they live at internal/agentsession and
internal/agentsession/appx. The loop told every agent it was one of a
security audit; the program now names its work (Config.Work), and sec's is
"a security audit", so sec's prompts are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eply fails one agent

codeaf puts `sec did not finish:` (or which ending it was) in front of a
program's message, and sec's messages opened on its name too, so they read
`sec did not finish: sec did not start: …`. They now give the reason alone.

A session whose model kept answering with no choices came back from
App.Harness as an error, and a program that reads an agent's error as fatal
ended a whole review on it (found by /pr's live run). Only what ends every
call — the ceiling, a refused key, the caller's context — is an error now;
anything else is that session's failed result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A dollar ceiling is reached by a refused call, and codeaf says
`sec reached the run's dollar ceiling of $5.00: sec said …`; only the time
ceiling reads `sec stopped on its own ceiling: …`. And ABSORB.md no longer
says /pr is in this tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One cap for every agent, sec-af's fifty turns and thirty minutes, cut sec's
location scanners at fifty while its context profiler needed eleven, and /pr's
reviewers spent fifty turns where a dozen would have done. The shared loop
takes a program's per-agent bounds (Config.Limits), and sec's come from what
the owner's furrow audit measured, keyed by each agent's scratch folder:
75 turns and 20 minutes for a location scanner down to 30 and 5 for the
context profiler. --max-turns and --session-wall at a shell are one figure
for every agent; unset, each agent has its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a match far into one

read_file cut a line at 2,000 bytes, which could split a character, and said
only `[line cut]`; grep showed a matching line's first 240 characters, so a
match past them came back without its text. Nothing past the cut was
reachable, and /pr's agents, whose context file was one line of JSON, read on
to their turn cap. read_file now cuts on a character, says how long the line
is and to grep in it, and grep shows the text around the match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…unset one is refused

agentsession began as sec's, and its numbers were sec-af's: New filled fifty
turns, thirty minutes and eight sessions and calls, RunSession fell back to
fifty turns again, and two follow-ups, the context room, the answer-now
message, every tool's caps and the client's retries were constants. A second
program on it would have run on figures tuned for another's agents without
saying so. Now New and RunSession refuse any figure left unset, naming it
(Config.Sessions, Calls, MaxTurns, SessionWall, and Policy: FollowUps,
ContextChars, AnswerNow and Tools), NewClient takes the retries, and sec
states today's figures as its own in internal/secaf/limits.go. The owner
chose this so /sec's tuning never silently becomes /pr's, or the reverse.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ZeroPoint95 and others added 5 commits October 6, 2026 15:47
The pr program was renamed /review on the owner's call (#1784); three
comment and doc lines here still said /pr.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dev baa2d7f's page carries a program's guide on the lean arm too, so sec
costs the lean prefix 396 bytes where it cost 229; both waivers are dev's
plus exactly that (57,614 and 49,986), recorded in PERF.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ZeroPoint95
ZeroPoint95 merged commit 741f1b9 into dev Oct 7, 2026
9 checks passed
@ZeroPoint95
ZeroPoint95 deleted the zeropoint95/sec branch October 7, 2026 18:08
ZeroPoint95 added a commit that referenced this pull request Oct 7, 2026
#1781 landed on dev as one commit whose tree is exactly zeropoint95/sec's
tip, which this branch already carries, so every conflict is that same
code on both sides; each resolves to this branch's version. The merged
tree is identical to the one validated on dev (pr-ready, builds, size).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant