Skip to content

review: pr-af's code review, built into codeaf as its third program - #1784

Merged
ZeroPoint95 merged 48 commits into
devfrom
zeropoint95/pr
Oct 7, 2026
Merged

ZeroPoint95 merged 48 commits into
devfrom
zeropoint95/pr

Conversation

@ZeroPoint95

@ZeroPoint95 ZeroPoint95 commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Renamed from /pr to /review on the owner's call: the chat command, the shell verb, via, the manual page, the guide, the report files and the endings. internal/praf keeps its name, after the project it was copied from, as internal/secaf does.

Retargeted to dev after #1781 merged. dev (with #1781 squash-merged as 741f1b9ee) is merged into this branch with a plain merge, so the diff against dev is only this program. Squash-merge it.

1784

What changed

pr-af, AgentField's pull-request reviewer, is built into codeaf as its third carried program: /review in the chat, codeaf review at a shell, via: "review" from a proposal. It follows the same approach as sec in #1781. pr-af's Go port is copied in once, at pr-af's tag codeaf-absorb (b70667e), as internal/praf and frozen there. It runs on the shared agent loop (internal/agentsession), and every model call goes through the run's model API, priced and held to the run's ceiling. Its reviewers read the pull request's checkout with the four read-only tools. It lands text: an account in the conversation, and review-report.md / .json in the task's record folder. The chat then offers to post the review or to hand the blocking findings to senior-dev, and starts neither until the person says yes.

  • Scope: /review <pull request> [focus] takes a URL, owner/repo#N or #N. A bare /review reviews the current branch's open pull request, found with gh pr view or, failing that, the upstream remote and the GitHub API. Focus words become pr-af's review hints.
  • The chat hears it by its command and its work, not the word "review" (Delegate.Asked, new and optional). "review" is a word people use for much else, so the bare word names nothing ("review this function" stays in the conversation). /review always asks for it, and so does a message that points at a pull request (PR, pull request, its link, owner/repo#N) and asks for it to be looked over ("review PR 123", "take a look at my PR"), which keeps the feel PR had when the program was called pr. The turn-back for a proposal without via then says which work was heard. senior-dev and sec set no Asked and are heard exactly as before.
  • Posting is its own run, on a per-run yes. A review posts nothing. /review post <review-report.json> posts the saved review on the commit that was reviewed, and the chat proposes it only after the person agrees. GitHub refuses a request for changes on the author's own pull request, so that case is posted as a comment and the ending says so.
  • GitHub auth: GH_TOKEN, then GITHUB_TOKEN, then gh auth token; with none, public repositories are read anonymously. git gets the token through an extraheader, never in a URL. The checkout lives in a temporary folder of its own that is removed when the run ends, and checkouts left by killed runs are swept after a day.
  • Every limit is review's own (internal/praf/limits.go): each agent's turns and time (10–30 turns, 5–15 minutes), plus concurrency, retries and the session policy. Since sec: sec-af's security audit, built into codeaf as its second program #1781 the shared loop has no defaults, and a praf test proves the config review hands it is complete. --max-turns, --session-wall, --sessions, --model and --light override for one run. The sessions are told their work is "a code review".
  • Ceilings: $5 and two hours (Delegate.Unattended). A review cut by its time ceiling still reports and says where it was cut, and it keeps three minutes back to write the review.
  • One failed session fails one reviewer, not the review. The account says how many agent sessions and single calls failed, and a review where every session failed ends as a failure.
  • Changes to pr-af's own code on the way:
    • The JSON context files the reviewers read are written one value to a line. codeaf cuts a line at 2,000 characters, so pr-af's single-line JSON was unreadable past that, and reviewers spent their turns re-reading it.
    • The review is a single pass.
    • The HITL loop, GitHub App auth, the SDK harness and every PR_AF_* variable are gone.
  • Manual: a new review page (internal/manual/chat/review.md, its probes keyed review), plus updates to commands, delegates, tasks, models-and-cost and running-from-the-terminal, with nine probes.
  • Budgets: SIZE-BUDGET rises by exactly what linking internal/praf measured (+1,032,784 bytes on darwin/amd64 with furrow staged), and the prefix waivers by review's guide (+203 bytes in both arms, over sec's as re-measured on dev baa2d7f0f); both are in PERF.md.

Design notes: docs/design/pr-review/ABSORB.md (what was copied, what changed, and why), and docs/design/pr-review/PR-AF-GO-README.md (pr-af's own README, kept for reference).

How it was checked

  • BASE=origin/dev make pr-ready passed on this branch's tree as merged with dev 741f1b9ee (16e937b82): the light gate, the law tests, the touched packages and the whole internal/session suite (8 shards), every selected test on its first run. On this Mac, scripts/one-suite_test.sh (needs flock) and cmd/codeaf's TestRunSurfaceWiresTheDeferredLaunchCheckAndInstallerThroughRealInit fail the same on a clean dev; neither is this branch's.
  • GitHub CI runs on this PR now that its base is dev.
  • An independent review from another session (build on darwin and windows, vet, laws, and the touched packages) found four issues, all fixed on this branch: a bare /review post started a review; the manual promised a progress count the page never drew; --sessions 0 ended with agentsession's sentence instead of review's; and one manual quote did not match its message.
  • A final-pass review from the same session, on the tip before dev was merged in, found nothing blocking: builds on darwin and windows, vet, laws, size, and every touched package.
  • make test-laws passed.
  • GOOS=windows go build ./... passes; GOARCH=amd64 make size on 16e937b82: 72,234,096 bytes, under the 72,395,784 budget.
  • internal/praf:
    • a stub-model end-to-end review (checkout, sessions, pipeline, account, both report files, nothing posted);
    • brief reading (URL, owner/repo#N, #N, bare, focus, post <path>);
    • posting, including the own-pull-request fallback;
    • the account's wording, the time-ceiling cut and the all-sessions-failed ending;
    • the per-agent limits and overrides;
    • that the sessions' config passes agentsession.New;
    • what asks for a review (asksForReview): a pull request pointed at and asked to be looked over, never either alone.
  • internal/session: TestAProgramNamedByAnEverydayWordIsHeardByItsCommandAndItsWork, covering the bare word, the command, the work, the name after an ask word, the work's turn-back, and senior-dev unchanged beside it.
  • Live, at a shell (as codeaf pr, before the rename) on Make GH_TOKEN optional in the Go node manifest pr-af#73, on deepseek/deepseek-v4-flash-0731. The whole review ran in 55m14s: 575 calls, $0.34, 6 findings (0 blocking, 3 important, 3 suggestions), and nothing was posted. Each finding had a location and a fix. Earlier live runs found the bugs fixed in this branch:
    • sessions were told they were part of a security audit;
    • one reply with no choices ended the whole review;
    • endings repeated their prefix;
    • the single-line JSON context files above, which drove every session to its turn cap.

Not yet checked: a review started from the chat, watching the turn it wakes; and /review post against a real pull request, which is covered only by tests. The per-agent limits are a choice of depth, not a measurement, and a small pull request still takes close to an hour, so tuning them is the next step.

Checklist

  • A change entry: docs/changes/unreleased/1784-review.md
  • The manual knows about it: review.md is new, the commands/delegates/tasks/models-and-cost/terminal pages are updated, and there are probes
  • No new line in .github/known-red.txt
  • Only my own paths are staged, no git add -A

Not in this PR: pr-af's codeaf-absorb tag and the commit it points at exist only in a local pr-af checkout and are not pushed.

🤖 Generated with Claude Code

ZeroPoint95 and others added 30 commits October 5, 2026 18:33
… senior-dev's

Every unattended ceiling was senior-dev's, chosen by its name in nine
places, so a second program started from the chat ran on whatever the
conversation had left. A program now names its own (Delegate.Unattended);
senior-dev's are unchanged.

A program that lands text had no ending of its own: the wake turn sent the
model looking for a branch and a worktree that were never cut. It now reads
a report as a finished answer, asks for a short summary and the program's
own offer (Delegate.FollowUp), and starts nothing. A program is told its
record folder (CODEAF_RECORDS) for the files its report points at.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sec-af was an AgentField node: a control plane carried its calls, a router
key it held paid for its models, and a coding-agent binary ran each of its
agent sessions. It is now copied into codeaf once, at sec-af's tag
codeaf-absorb (47d57d7), as internal/secaf, and runs only through codeaf:
/security-audit in the chat, codeaf security-audit at a shell, and
propose_task with via "security-audit".

Its algorithm is kept: the phases, the hunters, the four-agent proof
chain and the prompts. What it runs on is codeaf's (internal/secaf/backing):
model calls go to the run's model API, each agent session is a read-only
loop of four tools with a schema-checked answer, and calls between its
reasoners stay in the process. It changes nothing in the folder; its report
goes to the task's record folder and its account to the conversation, which
offers to hand confirmed findings to senior-dev.

An audit of the changes (the branch since its base, uncommitted work
included) tells the hunters what changed instead of filtering a whole-
repository scan afterwards. Naming compliance frameworks no longer fails
every audit at its end.

A program may now run bare on a default brief, show its own arguments on
its row, and say no ceiling it does not have. The prefix waivers and the
size budget rise by exactly what the program measured (PERF.md).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A call the ceiling holds only because calls in flight have reserved
  what is left is answered 429 with Retry-After and X-Codeaf-Held, and
  opens no turn; 402 stays for a ceiling truly reached. A program that
  makes many calls at once was told its ceiling was reached at $0 spent.
  sec-af's client waits a held call out within its own bounds.
- A program names the flag that carries its model (Delegate.ModelFlag);
  the shell resolved only senior-dev's --high.
- A shell run of a program that answers prints its answer, not what a
  tree program's model claimed; an unset ceiling is not said as $0.00.
- A tool with no required argument sent required: null, which a strict
  server refuses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A typed /security-audit quick was titled "quick"; a program may now title
its typed runs (Delegate.Title), and the audit's say what it audits. Each
agent's "starting" note duplicated its session's line and is left off the
page; the hunters' are kept. The protocol spec and the programs page say
the ceiling's held answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
/security-audit is /sec, codeaf security-audit is codeaf sec, and
via: "security-audit" is via: "sec". Its manual page is sec, its narrow
badge [s], and a shell run's records go under ~/.codeaf/v3/carried/sec/. The
report's files keep their descriptive names (security-audit.md, .json,
.sarif). The shorter name takes eleven bytes off the fixed prefix and the
waivers come down with it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The page headed its steps MAP, TEST and FIX while sec-af's own notes under
them said RECON, PROVE and REMEDIATION. The row, the headings and the notes
now all use sec-af's phase names (recon, hunt, prove, remediate, report),
and its agents keep their own names except where they are banned words,
which are reworded in the record itself so a shell run says them the same
way. The CWE expansion note, which claimed a widening the hunters never
see, is left off, and the dedup note no longer speaks of fingerprints. The
verdict agent, a single call rather than a session, now has its line, so
the proof chain shows all four agents.

security-audit.md was sec-af's report: Verdict: inconclusive, not
exploitable, Cost: $0.00, Commit: HEAD, Provider: harness. It is written by
sec in the account's words (confirmed, likely, unclear, ruled out) with
each finding's trace, attack, fix and patch, and only figures that were
measured. The JSON and SARIF keep sec-af's field names.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
f609ea5ec renamed security-audit to sec but staged these two files before
their last edits, so it and b403e776e name programguide.Sec while the file
still declared SecurityAudit, and the entry still said security-audit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d what

The report files were security-audit.* and the SARIF named its tool SEC-AF,
with sec-af's link, sec-af/ rule ids and properties. They are
sec-report.md, .json and .sarif (and sec-compliance.md); the SARIF names
sec, codeaf's build and home, and sec/ ids. The JSON and the compliance
report carry the run's own cost and agent count where sec-af left zeros.
sec-af's writers keep their own identity by default, so its goldens hold.

Every hunter's sessions were "hunt location scanner" and "hunt finding
enricher", so the hunt read as two lines repeated eleven times. Each says
its hunter and what it found: "injection hunter · scan" with how many
places, and "injection hunter · app/views.py:4 · <finding>" with its
severity. The hunters' start notes, which those lines replace, are off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A two-hour sec run on furrow ended, the turn it woke read the report and
began a good summary, and every model it was offered was cut as "the
model's own internal markup". The partial summary was kept as an
interrupted message the chat does not draw, so the person saw a done card
and nothing else.

- The markup detector read a tool's name between two prose marks as tool
  grammar: the report's path, …/tasks/1/sec-report.md, spells the tool
  tasks between two slashes, and a findings table put the reply at a tenth
  symbols. Path separators, backticks, emphasis and punctuation are no
  longer fence material; the leak shapes (bars, brackets, quotes) still are.
- A turn woken by a program's ending that cannot finish an answer now
  writes the program's own account into the conversation as the session's
  line, once per run.
- A sec run its time ceiling cuts says so, in which phase, and what was
  left undone, in its account, ending and report; the demotion notes that
  cut produced read as findings staying unclear, not verifier_error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A standard audit of furrow made 1,952 calls, spent $2.40 of its $5 and was
cut by two hours in prove with no fixes written: time is the ceiling it
meets first, so the hours double and the dollars stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ints at its report

A run the chat proposed was handed the composed brief, whose first line is
WHAT THE PERSON ASKED FOR, and sec reads its scope off the first line: a
proposed `whole repository thorough` ran at standard depth, and a proposed
`changes` would have audited the whole repository. A program whose brief is
words (Delegate.Words) is now handed those words alone.

The project index kept only "Security audit of the whole repository." of a
run and pointed at the repository, so another conversation searched the disk
for the report and opened a different run's first. The account's first line
now says what it found, and a report program's row points at its record
folder.

And a program with no ceilings of its own is held to the conversation's,
whose ending names them: it had read "the run's $0.00 limit".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
internal/secaf compiles in santhosh-tekuri/jsonschema (its answers' schema
checks) and invopop/jsonschema with what they pull in; each has its section,
regenerated by codeaf-notices.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…'s stop told as theirs

The change entry's `pr:` still said 1757: the rename was committed without
the edit to the field. Two pointers named docs/design/sec/ABSORB.md, which is
docs/design/security-audit/ABSORB.md. Delegate.Unattended's comment gave an
audit a quarter of an hour; it now points at the two programs' own figures.

A program's ending the wake turn did not answer was always told as a failed
reply, and a person's stop ends that turn the same way. A stop now keeps the
account and says the person stopped the answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gram's work

internal/secaf/backing and internal/secaf/appx are codeaf's own, not sec-af's,
and /pr runs on them too, so they live at internal/agentsession and
internal/agentsession/appx. The loop told every agent it was one of a
security audit; the program now names its work (Config.Work), and sec's is
"a security audit", so sec's prompts are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eply fails one agent

codeaf puts `sec did not finish:` (or which ending it was) in front of a
program's message, and sec's messages opened on its name too, so they read
`sec did not finish: sec did not start: …`. They now give the reason alone.

A session whose model kept answering with no choices came back from
App.Harness as an error, and a program that reads an agent's error as fatal
ended a whole review on it (found by /pr's live run). Only what ends every
call — the ceiling, a refused key, the caller's context — is an error now;
anything else is that session's failed result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A dollar ceiling is reached by a refused call, and codeaf says
`sec reached the run's dollar ceiling of $5.00: sec said …`; only the time
ceiling reads `sec stopped on its own ceiling: …`. And ABSORB.md no longer
says /pr is in this tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One cap for every agent, sec-af's fifty turns and thirty minutes, cut sec's
location scanners at fifty while its context profiler needed eleven, and /pr's
reviewers spent fifty turns where a dozen would have done. The shared loop
takes a program's per-agent bounds (Config.Limits), and sec's come from what
the owner's furrow audit measured, keyed by each agent's scratch folder:
75 turns and 20 minutes for a location scanner down to 30 and 5 for the
context profiler. --max-turns and --session-wall at a shell are one figure
for every agent; unset, each agent has its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a match far into one

read_file cut a line at 2,000 bytes, which could split a character, and said
only `[line cut]`; grep showed a matching line's first 240 characters, so a
match past them came back without its text. Nothing past the cut was
reachable, and /pr's agents, whose context file was one line of JSON, read on
to their turn cap. read_file now cuts on a character, says how long the line
is and to grep in it, and grep shows the text around the match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…unset one is refused

agentsession began as sec's, and its numbers were sec-af's: New filled fifty
turns, thirty minutes and eight sessions and calls, RunSession fell back to
fifty turns again, and two follow-ups, the context room, the answer-now
message, every tool's caps and the client's retries were constants. A second
program on it would have run on figures tuned for another's agents without
saying so. Now New and RunSession refuse any figure left unset, naming it
(Config.Sessions, Calls, MaxTurns, SessionWall, and Policy: FollowUps,
ContextChars, AnswerNow and Tools), NewClient takes the retries, and sec
states today's figures as its own in internal/secaf/limits.go. The owner
chose this so /sec's tuning never silently becomes /pr's, or the reverse.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pr-af (github.com/Agent-Field/pr-af, go/ at b70667e, tag codeaf-absorb) is
copied once as internal/praf and frozen there, the way sec-af became
internal/secaf. Its algorithm and prompts are unchanged; what it ran on is
taken out:

- The SDK's agent and harness packages are gone. internal/praf/appx declares
  the review's own seam (Harness, AI, Note) with HarnessOptions{Cwd, Label}
  and HarnessResult; only sdk/go/ai remains.
- The orchestrator runs once. pr-af's HITL loop waited on a control plane for
  a person's approval before posting; inside codeaf every review is a dry run
  and posting is its own step on the person's yes. hitl, the HITL config and
  the Pause verb are not copied.
- No PR_AF_* environment is read. The provider and harness-binary machinery
  is deleted; the evidence pack and post-worthiness gate keep their shipped
  defaults; budget caps come from the caller; the clone folder and GitHub
  token are handed in (orch.Access), and a review with nothing to clone is
  refused instead of falling back to the working directory.
- The GitHub token never reaches the disk: it rides git's environment as an
  Authorization header instead of the clone's remote URL. GitHub App sign-in
  (and golang-jwt) is not carried.
- Progress prints go to stderr: stdout is the protocol stream.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pr is the third program codeaf carries, beside senior-dev and sec. It reviews
one GitHub pull request — its link, owner/repo#N, #N against the folder's
origin, or with no brief the current branch's open pull request (gh pr view,
else GitHub asked for the branch's upstream) — in a checkout of its own that
it removes when it ends, on sec's agent sessions over the run's model API.

- Words: the program is handed the pull request and focus alone, never a
  composed brief whose quoted conversation could name another pull request.
- Its account's first line says what it found; the findings follow one per
  line, blocking first, with where and a fix. pr-report.md and pr-report.json
  go to the record folder, the JSON holding the GitHub review pr-af built and
  the commit it reviewed.
- It never posts during a review. FollowUp offers posting and a senior-dev fix
  and waits; on the person's yes the chat proposes `post <report>`, a run that
  posts the saved review as it stands on the reviewed commit.
- Unattended ceilings $5 and two hours; the pipeline's budget gate sits at 80%
  of the dollar ceiling, three minutes are kept back to write the review, and a
  review its time ceiling cut says where.
- The page shows each stage once and every finished step with what it found;
  code review's machinery words are reworded wherever a person reads them.
- GitHub access: GH_TOKEN, GITHUB_TOKEN, then `gh auth token`.
- Manual: chat/pr.md, the program lists, nine probes. Prefix caps rise by the
  206 bytes pr's guide costs; SIZE-BUDGET by the 1,032,768 bytes it links
  (PERF.md). docs/design/pr-review/ABSORB.md records the copy.

The change entry's number is a placeholder (9999) until the draft pull
request exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sec's agent loop moved to internal/agentsession (11d0b3b), and a program now
names the work its sessions are part of. pr's sessions were told they were
"one agent of a security audit"; they are told "a code review". The size
budget is re-measured on the rebased tree: pr costs 1,032,784 bytes on
darwin/amd64, so SIZE-BUDGET is 72,395,784. The prefix caps do not move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…reason once

A live review of Agent-Field/pr-af#73 ended with nothing after 389 calls:
one reviewer's session got a reply with no choices, the session loop
answered it as an error, and pr-af's pipeline reads an error from a reviewer
as the end of the review, so every reviewer still running was cancelled. A
session error is now that reviewer's failed result, which the pipeline
degrades and the run counts; only the dollar ceiling, a refused key or the
run's stop end every call.

codeaf already says a program's name and how it ended before its sentence
(`pr did not finish: …`), so pr's sentences no longer open with "pr did not
finish", "pr did not start" or "pr could not"; the run read
"pr did not finish: pr did not finish: …".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t is swept

A live review of a two-file pull request (Agent-Field/pr-af#73) reached its
challenge pass after an hour: nine of its twelve agent sessions read until the
loop made them answer at turn fifty, at about twenty seconds a turn. A
review's prompts already carry the diff and the touched files, so pr's
sessions now take at most twenty turns (`--max-turns` overrides); one at its
limit is still made to answer.

A run killed outright left its checkout in the temporary folder twice. Each
run now removes codeaf-pr-* checkouts last changed a day or more ago.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pr-af writes a large context (the lenses' analysis, the findings the evidence,
challenge and cross-reference passes weigh) to a file and points the agent at
it, as json.dumps' single line. codeaf's sessions read a file a line at a
time and cut a line at 2,000 characters, so a context of tens of thousands of
characters showed its first 2,000 — which fits the live review's lens and
evidence sessions reading the repository to their turn limit. The same JSON
is written indented; the prompts are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The owner's call: each agent runs under limits suited to its own work, as
sec's do (agentsession.Config.Limits, 1453508). A review's sessions are
keyed on the label their reasoner gives them, now constants in reasoners
(LabelReviewer, LabelLens, …), and internal/praf/limits.go names each one's
turns and wall; a test holds every row to a label a reasoner carries.
--max-turns and --session-wall default to 0, each agent's own, and override
every agent when set. The blanket 20-turn default is gone.

The figures are a first cut from the live review of Agent-Field/pr-af#73
whose sessions could not read their context files; they are reset from the
next measured run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The live review read "3 of 41 model calls failed" beside "575 model calls":
the 41 are its agent sessions and single calls, each many model turns, and
the note now says so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The manual and the absorb note still said every reviewer took at most
twenty turns; each agent now has its own turns and time, and the shared
agent loop holds none of pr's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ZeroPoint95 and others added 3 commits October 6, 2026 12:34
agentsession now holds no defaults (bedb052). pr states its own
concurrency, retries and session policy beside its per-agent limits,
with today's figures, and a test proves the config it hands the loop is
complete.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ZeroPoint95 and others added 2 commits October 6, 2026 13:37
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t draws

From the REVIEWER session's review of #1784:
- `/pr post` with nothing after it ran a review of the current branch;
  posting is now its own field, so it ends asking for the report.
- The manual and phaseTracker promised a 'reviewing, 3 of 8 done' count
  nothing drew; the count is gone, and the page's own lines are described.
- --sessions below one is refused in pr's words before anything runs,
  not with agentsession's.
- The tokenless post message drops the backticks a chat line prints raw,
  so it matches the manual's quote.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ZeroPoint95
ZeroPoint95 marked this pull request as ready for review October 6, 2026 19:05
ZeroPoint95 and others added 2 commits October 6, 2026 15:47
The owner's call: the program is review everywhere a person meets it.
The chat command, the shell verb, the via name, its manual page
(review.md), its guide (programguide.Review), its report files
(review-report.md/.json), its record folder (carried/review), its
checkout prefix and its endings ('review did not finish: …') all move.
internal/praf keeps its name: it is pr-af, the project it was copied
from, as internal/secaf is sec-af.

The longer name costs four bytes in every prefix; the guide gives back
seven ('then any focus'), so both caps are three bytes lower than pr's
were: 57,787 and 50,022.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ZeroPoint95 ZeroPoint95 changed the title pr: pr-af's code review, built into codeaf as its third program review: pr-af's code review, built into codeaf as its third program Oct 6, 2026
ZeroPoint95 and others added 10 commits October 6, 2026 16:51
The pr program was renamed /review on the owner's call (#1784); three
comment and doc lines here still said /pr.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
'review' is a word people use for much else, so hearing it as the
program's name claimed 'review this function'. A program may now say
what in a message asks for its work (Delegate.Asked); when it does, the
chat hears its command (/review) and that work, never the bare word.
review's work is a message that points at a pull request (PR, pull
request, its link, owner/repo#N) and asks for it to be looked over, so
'review PR 123' and 'take a look at my PR' ask for it as 'PR' named pr
before, and 'open a PR' or 'review my essay' do not. The turn-back for a
message heard by its work says which work was heard. senior-dev and sec
set no Asked and are heard exactly as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dev baa2d7f's page carries a program's guide on the lean arm too, so sec
costs the lean prefix 396 bytes where it cost 229; both waivers are dev's
plus exactly that (57,614 and 49,986), recorded in PERF.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ropoint95/pr

The prefix waivers conflicted: sec re-measured its cost on today's dev
(396 bytes in both arms). review's guide still costs 203 in both, so the
waivers are sec's plus 203: 9_817 and 18_689 (caps 57,817 and 50,189).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…int95/pr

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ZeroPoint95 added a commit that referenced this pull request Oct 7, 2026
…#1781)

* delegate: a program's ceilings and its report ending are its own, not senior-dev's

Every unattended ceiling was senior-dev's, chosen by its name in nine
places, so a second program started from the chat ran on whatever the
conversation had left. A program now names its own (Delegate.Unattended);
senior-dev's are unchanged.

A program that lands text had no ending of its own: the wake turn sent the
model looking for a branch and a worktree that were never cut. It now reads
a report as a finished answer, asks for a short summary and the program's
own offer (Delegate.FollowUp), and starts nothing. A program is told its
record folder (CODEAF_RECORDS) for the files its report points at.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* security-audit: sec-af, the security auditor, built into codeaf

sec-af was an AgentField node: a control plane carried its calls, a router
key it held paid for its models, and a coding-agent binary ran each of its
agent sessions. It is now copied into codeaf once, at sec-af's tag
codeaf-absorb (47d57d7), as internal/secaf, and runs only through codeaf:
/security-audit in the chat, codeaf security-audit at a shell, and
propose_task with via "security-audit".

Its algorithm is kept: the phases, the hunters, the four-agent proof
chain and the prompts. What it runs on is codeaf's (internal/secaf/backing):
model calls go to the run's model API, each agent session is a read-only
loop of four tools with a schema-checked answer, and calls between its
reasoners stay in the process. It changes nothing in the folder; its report
goes to the task's record folder and its account to the conversation, which
offers to hand confirmed findings to senior-dev.

An audit of the changes (the branch since its base, uncommitted work
included) tells the hunters what changed instead of filtering a whole-
repository scan afterwards. Naming compliance frameworks no longer fails
every audit at its end.

A program may now run bare on a default brief, show its own arguments on
its row, and say no ceiling it does not have. The prefix waivers and the
size budget rise by exactly what the program measured (PERF.md).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* security-audit: the fixes its first real runs asked for

- A call the ceiling holds only because calls in flight have reserved
  what is left is answered 429 with Retry-After and X-Codeaf-Held, and
  opens no turn; 402 stays for a ceiling truly reached. A program that
  makes many calls at once was told its ceiling was reached at $0 spent.
  sec-af's client waits a held call out within its own bounds.
- A program names the flag that carries its model (Delegate.ModelFlag);
  the shell resolved only senior-dev's --high.
- A shell run of a program that answers prints its answer, not what a
  tree program's model claimed; an unset ceiling is not said as $0.00.
- A tool with no required argument sent required: null, which a strict
  server refuses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* security-audit: titled by what it audits, quieter page, change entry

A typed /security-audit quick was titled "quick"; a program may now title
its typed runs (Delegate.Title), and the audit's say what it audits. Each
agent's "starting" note duplicated its session's line and is left off the
page; the hunters' are kept. The protocol spec and the programs page say
the ceiling's held answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* sec: security-audit is called sec

/security-audit is /sec, codeaf security-audit is codeaf sec, and
via: "security-audit" is via: "sec". Its manual page is sec, its narrow
badge [s], and a shell run's records go under ~/.codeaf/v3/carried/sec/. The
report's files keep their descriptive names (security-audit.md, .json,
.sarif). The shorter name takes eleven bytes off the fixed prefix and the
waivers come down with it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* sec: one vocabulary on its page, and a report in the account's words

The page headed its steps MAP, TEST and FIX while sec-af's own notes under
them said RECON, PROVE and REMEDIATION. The row, the headings and the notes
now all use sec-af's phase names (recon, hunt, prove, remediate, report),
and its agents keep their own names except where they are banned words,
which are reworded in the record itself so a shell run says them the same
way. The CWE expansion note, which claimed a widening the hunters never
see, is left off, and the dedup note no longer speaks of fingerprints. The
verdict agent, a single call rather than a session, now has its line, so
the proof chain shows all four agents.

security-audit.md was sec-af's report: Verdict: inconclusive, not
exploitable, Cost: $0.00, Commit: HEAD, Provider: harness. It is written by
sec in the account's words (confirmed, likely, unclear, ruled out) with
each finding's trace, attack, fix and patch, and only figures that were
measured. The JSON and SARIF keep sec-af's field names.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* sec: the guide constant and change entry the rename left unstaged

f609ea5ec renamed security-audit to sec but staged these two files before
their last edits, so it and b403e776e name programguide.Sec while the file
still declared SecurityAudit, and the entry still said security-audit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* sec: its files and SARIF say sec, and its hunt says which hunter found what

The report files were security-audit.* and the SARIF named its tool SEC-AF,
with sec-af's link, sec-af/ rule ids and properties. They are
sec-report.md, .json and .sarif (and sec-compliance.md); the SARIF names
sec, codeaf's build and home, and sec/ ids. The JSON and the compliance
report carry the run's own cost and agent count where sec-af left zeros.
sec-af's writers keep their own identity by default, so its goldens hold.

Every hunter's sessions were "hunt location scanner" and "hunt finding
enricher", so the hunt read as two lines repeated eleven times. Each says
its hunter and what it found: "injection hunter · scan" with how many
places, and "injection hunter · app/views.py:4 · <finding>" with its
severity. The hunters' start notes, which those lines replace, are off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* a program's ending reaches the person even when the reply to it fails

A two-hour sec run on furrow ended, the turn it woke read the report and
began a good summary, and every model it was offered was cut as "the
model's own internal markup". The partial summary was kept as an
interrupted message the chat does not draw, so the person saw a done card
and nothing else.

- The markup detector read a tool's name between two prose marks as tool
  grammar: the report's path, …/tasks/1/sec-report.md, spells the tool
  tasks between two slashes, and a findings table put the reply at a tenth
  symbols. Path separators, backticks, emphasis and punctuation are no
  longer fence material; the leak shapes (bars, brackets, quotes) still are.
- A turn woken by a program's ending that cannot finish an answer now
  writes the program's own account into the conversation as the session's
  line, once per run.
- A sec run its time ceiling cuts says so, in which phase, and what was
  left undone, in its account, ending and report; the demotion notes that
  cut produced read as findings staying unclear, not verifier_error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* sec: a run nobody limited gets four hours, not two

A standard audit of furrow made 1,952 calls, spent $2.40 of its $5 and was
cut by two hours in prove with no fixes written: time is the ceiling it
meets first, so the hours double and the dollars stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* sec: a proposed run reads the words it was asked with, and its row points at its report

A run the chat proposed was handed the composed brief, whose first line is
WHAT THE PERSON ASKED FOR, and sec reads its scope off the first line: a
proposed `whole repository thorough` ran at standard depth, and a proposed
`changes` would have audited the whole repository. A program whose brief is
words (Delegate.Words) is now handed those words alone.

The project index kept only "Security audit of the whole repository." of a
run and pointed at the repository, so another conversation searched the disk
for the report and opened a different run's first. The account's first line
now says what it found, and a report program's row points at its record
folder.

And a program with no ceilings of its own is held to the conversation's,
whose ending names them: it had read "the run's $0.00 limit".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* notices: the modules sec brought in

internal/secaf compiles in santhosh-tekuri/jsonschema (its answers' schema
checks) and invopop/jsonschema with what they pull in; each has its section,
regenerated by codeaf-notices.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* config: CODEAF_RECORDS is launch plumbing

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* changelog: sec's entry carries its pull request's number

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* sec: the review's fixes — the entry's number, ABSORB's path, a person's stop told as theirs

The change entry's `pr:` still said 1757: the rename was committed without
the edit to the field. Two pointers named docs/design/sec/ABSORB.md, which is
docs/design/security-audit/ABSORB.md. Delegate.Unattended's comment gave an
audit a quarter of an hour; it now points at the two programs' own figures.

A program's ending the wake turn did not answer was always told as a failed
reply, and a person's stop ends that turn the same way. A stop now keeps the
account and says the person stopped the answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* agentsession: sec's agent loop moves out of secaf, and names each program's work

internal/secaf/backing and internal/secaf/appx are codeaf's own, not sec-af's,
and /pr runs on them too, so they live at internal/agentsession and
internal/agentsession/appx. The loop told every agent it was one of a
security audit; the program now names its work (Config.Work), and sec's is
"a security audit", so sec's prompts are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* sec and agentsession: an ending says sec's name once, and one empty reply fails one agent

codeaf puts `sec did not finish:` (or which ending it was) in front of a
program's message, and sec's messages opened on its name too, so they read
`sec did not finish: sec did not start: …`. They now give the reason alone.

A session whose model kept answering with no choices came back from
App.Harness as an error, and a program that reads an agent's error as fatal
ended a whole review on it (found by /pr's live run). Only what ends every
call — the ceiling, a refused key, the caller's context — is an error now;
anything else is that session's failed result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* sec: the manual's ceiling endings are the two codeaf prints

A dollar ceiling is reached by a refused call, and codeaf says
`sec reached the run's dollar ceiling of $5.00: sec said …`; only the time
ceiling reads `sec stopped on its own ceiling: …`. And ABSORB.md no longer
says /pr is in this tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* agentsession and sec: each agent's turns and time are its own

One cap for every agent, sec-af's fifty turns and thirty minutes, cut sec's
location scanners at fifty while its context profiler needed eleven, and /pr's
reviewers spent fifty turns where a dozen would have done. The shared loop
takes a program's per-agent bounds (Config.Limits), and sec's come from what
the owner's furrow audit measured, keyed by each agent's scratch folder:
75 turns and 20 minutes for a location scanner down to 30 and 5 for the
context profiler. --max-turns and --session-wall at a shell are one figure
for every agent; unset, each agent has its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* agentsession: a long line says how to reach its rest, and grep shows a match far into one

read_file cut a line at 2,000 bytes, which could split a character, and said
only `[line cut]`; grep showed a matching line's first 240 characters, so a
match past them came back without its text. Nothing past the cut was
reachable, and /pr's agents, whose context file was one line of JSON, read on
to their turn cap. read_file now cuts on a character, says how long the line
is and to grep in it, and grep shows the text around the match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* agentsession: machinery only — every figure is the program's, and an unset one is refused

agentsession began as sec's, and its numbers were sec-af's: New filled fifty
turns, thirty minutes and eight sessions and calls, RunSession fell back to
fifty turns again, and two follow-ups, the context room, the answer-now
message, every tool's caps and the client's retries were constants. A second
program on it would have run on figures tuned for another's agents without
saying so. Now New and RunSession refuse any figure left unset, naming it
(Config.Sessions, Calls, MaxTurns, SessionWall, and Policy: FollowUps,
ContextChars, AnswerNow and Tools), NewClient takes the retries, and sec
states today's figures as its own in internal/secaf/limits.go. The owner
chose this so /sec's tuning never silently becomes /pr's, or the reverse.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* agentsession, ABSORB: /pr is called /review

The pr program was renamed /review on the owner's call (#1784); three
comment and doc lines here still said /pr.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* prefix: sec's guide costs both prefixes 396 bytes on today's dev

dev baa2d7f's page carries a program's guide on the lean arm too, so sec
costs the lean prefix 396 bytes where it cost 229; both waivers are dev's
plus exactly that (57,614 and 49,986), recorded in PERF.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Base automatically changed from zeropoint95/sec to dev October 7, 2026 18:08
#1781 landed on dev as one commit whose tree is exactly zeropoint95/sec's
tip, which this branch already carries, so every conflict is that same
code on both sides; each resolves to this branch's version. The merged
tree is identical to the one validated on dev (pr-ready, builds, size).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ZeroPoint95
ZeroPoint95 merged commit 6246d71 into dev Oct 7, 2026
7 checks passed
@ZeroPoint95
ZeroPoint95 deleted the zeropoint95/pr branch October 7, 2026 18:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant