I build small, verifiable systems around AI agents.
权限感知的检索、工具边界、人工复核、可恢复的执行与评测回归。公开仓库包括个人复现、自研工具和开源贡献,每一项都标明证据边界。
I prefer explicit evidence over broad claims: implementation, local verification, upstream acceptance, and real production use are different milestones.
Stack LangGraph · FastAPI · Pydantic · Chroma · MCP / A2A · OpenTelemetry
- Reliable agent workflows — tool boundaries, human review, failure handling, and observable execution.
- Grounded retrieval systems — evidence-carrying answers, access-aware retrieval, and explicit evaluation boundaries.
- Developer tooling — focused changes with regression coverage and reviewable behavior.
enterprise-service-desk-agent-history — an internal service-desk agent evolved in four stages: RAG → MCP read-only tool → A2A local validation → multi-agent collaboration.
Permission-aware retrieval · human-approval resume (LangGraph) · idempotent ticket creation · tests and version tags.
Personal reproduction with fictional data. Scope and limits are documented in the repo.
reuse-before-build — a dependency-free Agent Skill that helps coding agents inspect existing implementations, tests, and prior decisions before choosing what to Take, Borrow, or Build.
Dated evaluation records, diffs, test replays, compatibility notes, and documented failures — rather than a success-rate claim.
Merged upstream pull requests across 11 projects. Highlights:
| Project | Contribution | PR |
|---|---|---|
| pandas | Preserved the intended AssertionError for nested sequence length mismatches, so unhashable nested values no longer mask it. Regression tests across low-level assertions, Series, and DataFrame. |
#69015 |
| pnpm | Fixed configured .js pnpmfile loading so the module format follows the nearest package.json. CommonJS and ES module regression tests. |
#15152 |
| NVIDIA/SkillSpector | Exposed MCP rug-pull and least-privilege match details in JSON reports; redacted URL credentials in pattern across report formats. Regression tests. |
#641 |
| PR-Agent | Preserved request-specific extra_instructions during code-suggestion reflection, so scoring sees the context used in generation. Passed as untrusted user-prompt data; tests for context transport, empty-context compatibility, and token-budget enforcement. |
#3987 |
| Docling | Fixed the AsciiDoc backend emitting an extra empty table for incomplete tables; unified end-of-document table flushing to preserve captions. Regression tests. | #4300 |
Other merged PRs (6)
- Headroom #3778 — keep unmarked HTML tool output intact when no retrieval marker exists.
- Kornia #5300 — keep inverse mask sampling overrides local, so state does not leak across images.
- OpenMAIC #1601 — request-ID correlation and server-side 5xx logging for persistence routes.
- ECC #3184 — fix project-scoped Claude hooks in ESM projects via managed CommonJS boundaries.
- Archify #437 — report the resolved quality profile in exported SVG metadata.
- Laya #240 — keep an empty
Router.preload([])selection from loading all checkpoints.
- Evidence before claims.
- Reuse before building.
- Small changes with explicit verification.
- Human review at consequential boundaries.
- A passing demo is not production evidence.



