Development Rules
Overview
This repo is Orphus: an agent harness whose agents deliberate in rooms that live outside their context windows. It began as a fork of Atomic, itself a fork of pi, so most of the tree is vendored upstream code. Workspace packages carry the @orphus/* scope: the rename is deliberate and means git merge upstream/main now conflicts on import lines, so upstream changes are cherry-picked rather than merged wholesale.
The packages this project exists for:
@orphus/roundtableinpackages/roundtable— the Orphus contribution. Rooms and the context-window contract: the budgeted digest algorithm (digest.ts), the local-socket broker and its client (broker/), theroundtableandmemorytools, the declarative role manifest and launcher (roles/,bin/orphus-roles.ts), the discussion-etiquette skill, and the no-model demos. When a change here is not obviously about rooms, digests, roles, or memory, it probably belongs in the vendored tree instead.@orphus/fleetinpackages/fleet— Orphus-authored orchestration on top of rooms and subagents. Shareable fleet blueprints (*.fleet.yaml: teams of agent definitions with pre-assigned skills and a delegation mode each), the/fleetand/fleetsetupcommands, thefleetintrospection tool, and thefleet-orchestrationandkie-ai-mediaskills. It executes nothing itself — members run via thesubagenttool and deliberate in roundtable rooms. When a change is about how members run rather than how a fleet is described and briefed, it belongs inpackages/subagentsor the vendored tree.@orphus/transcribeinpackages/transcribe— Orphus-authored, derived from pi-transcribe (MIT, with attribution and a pinned upstream-sync record inUPSTREAM.md). Local dictation: the versioned six-request JSON-Lines worker/helper protocol, ABI and build-hash verification, and the consent-then-checksum model catalog. It is not bundled and not registered as a builtin — the native miniaudio/transcribe.cpp artifacts for the eight release targets are not built in this repository, so both channels fail closed. That is the intended state, not a bug to fix: the ABI is pinned innative/ABI.md, and wiring it up means building those artifacts, not removing the guard.
Inherited from Atomic, and mostly left alone:
@orphus/coding-agentinpackages/coding-agent— the coding-agent CLI, which builds theorphusbinary. The only independently published package.@orphus/workflowsinpackages/workflows— a first-party extension for Atomic/pi that brings multi-stage, DAG-driven workflow execution to agent sessions.@orphus/subagentsinpackages/subagents— builtin subagent orchestration, reusable agent definitions, skills, prompts, chains, and foreground/background execution.@orphus/mcpinpackages/mcp— builtin MCP adapter extension that exposes MCP servers as agent tools.@orphus/web-accessinpackages/web-access— builtin web search, URL fetching, GitHub repository, PDF, and video extraction tools.@orphus/intercominpackages/intercom— builtin coordination channel for parent/child and cross-session agent communication.
Default completion route: for substantial verifiable build/change/fix/refactor work in this tree, use the builtin Goal workflow as Orphus’s native core loop, not as an optional skill. Goal freezes a validated depth-tree plan before implementation, assigns disjoint ownership and dependency edges, rolls all ready leaves concurrently up to its configured bound, requires exact independent per-check evidence for every leaf, and keeps final completion behind three decorrelated reviewers and the reducer. Stay inline only for tiny deterministic read-only or handoff work where the team loop would add no proof.
Companion packages under packages/* ship as raw TypeScript (no compile step) and are bundled into @orphus/coding-agent at build time rather than published independently. The coding-agent package follows upstream pi’s compiled-package layout.
Minimal-change principle (KISS) — read this first
Default discipline: ponytail — the
bundled skill that enforces everything in this section as a ladder (YAGNI → reuse →
stdlib → native platform → installed dependency → one line → minimum code). It ships
with every Orphus session via @orphus/subagents; agents working on this repository
load it at full intensity for any coding task, and /ponytail lite|full|ultra
switches level. Vendored from
dietrichgebert/ponytail (MIT).
Fix the actual problem with the smallest correct change. Do not rewrite files, and do not add speculative hardening for issues that cannot occur in this codebase. Don’t reinvent the wheel or burn tokens rewriting file after file when the fix is usually a few lines.
- Verify before fixing. Reproduce a reported issue, or trace that an existing guard already prevents it — the manifest
NAME_PATTERNvalidation, the librarian writer convention (a coordination check against accidental concurrent writes, not a security boundary — seedocs/memory.md), the local same-machine socket trust boundary. Refute non-issues instead of patching them. A review tool flagging something is a hypothesis, not a fact. - Prefer a one-line fix to a subsystem. Weigh diff size against real risk reduced. The
client.tsunhandled-rejection fix isvoid registered.catch(() => {}), not a rewrite. - Reuse what exists — sibling modules (e.g. intercom’s reconnect guard), existing helpers, validation already in place — rather than building new machinery.
- Don’t over-test. Skip a disproportionate harness for a low-severity edge; a clear comment can be the right call. But every real fix gets a regression test proven to fail without it.
- Precedent: a broad review once flagged 15 issues here; verification confirmed 2. The minimal fix was ~50 lines; the kitchen-sink alternative was ~1,150. Ship the 50.
Reuse check — query the graph before you write
“Reuse what exists” is only enforceable if you can find what exists. This tree is large (a vendored Atomic/Pi fork plus the Orphus packages), so grepping for a plausible function name is not evidence that one does not already exist.
GitNexus indexes the repository
into a knowledge graph and serves it over MCP; it is declared in the committed
.mcp.json, so any agent session that reads that file gets the tools. Index once
per checkout:
npx gitnexus@1.6.9 analyze . --skip-agents-md # writes .gitnexus/ (gitignored); ~3 min for this tree
npx gitnexus@1.6.9 status # re-run analyze when this reports "stale"
--skip-agents-md is not optional politeness: a bare analyze appends a
45-line block to AGENTS.md and CLAUDE.md whose MUST/NEVER framing
contradicts this file (the graph is a lookup tool, not an authority), leaving a
fresh clone with a dirty tree in the two files that define how to work here.
Before adding a function, a helper, or a module:
-
context <symbol>— who calls it, what it calls, which flows it sits in. -
impact <symbol>— blast radius, before changing anything shared. -
cypher "<query>"— when you need to search rather than start from a known symbol. Node tables areFunction,Method,Class,Interface,File, and the location property isfilePath(notpathorfile_path):npx gitnexus@1.6.9 cypher --repo orphus "MATCH (m:Method) WHERE m.filePath CONTAINS 'roundtable/broker/' \ RETURN m.name AS name, m.filePath AS file, m.startLine AS line ORDER BY file, line"--repo orphusmatters once the machine-wide GitNexus store holds more than one repository: without it the CLI errors with “Multiple repositories indexed” instead of defaulting to the current checkout.
The graph is the anti-duplication gate: nothing gets recreated that the tree already has. It is a lookup tool, not an authority — it can be stale, and a confident-looking answer still needs the file read before you act on it. When the graph and the source disagree, the source wins and the index needs a re-run.
Two limits worth knowing before you trust an empty result:
queryneeds the LadybugDB FTS extension, whichanalyzedownloads on first use. Behind a proxy that blocks it, indexing still succeeds and the graph is complete, butquerysilently returns zero matches with awarningfield rather than an error — which reads exactly like “no such code exists”. Usecypherfor search in that case, and checkdoctorif you are unsure.- Files over 512 KB are skipped (currently one: the
0002rebrand patch). RaiseGITNEXUS_MAX_FILE_SIZEif you need them indexed.
Definition of done — housekeeping is part of the change
A change is not finished when the code works. It is finished when the repository is in a state the next person can pick up without archaeology. Every piece of work closes with these, in order:
- Clean up after yourself. Delete the scratch file, the debug logging, the commented-out alternative, the branch you no longer need, the dependency you stopped using. If the change made something obsolete — a script nothing calls, a doc describing a path that moved — remove it in the same change rather than leaving it for someone to discover and be unsure about.
- Update the documentation the change invalidates. Not “add docs” as a
ritual: find what is now wrong. A renamed flag, a changed default, a new
action on a tool, a step that is no longer needed. Stale documentation is
worse than none, because it is trusted. The relevant files are usually
docs/, the tool reference, andpackages/coding-agent/docsfor anything user-facing. - Update the README if the change alters what Orphus is or how it is run. Not every change touches it. A new tool action, a changed command, a moved directory, or a different install step does.
- Record user-visible changes in the changelog —
packages/*/CHANGELOG.md, under[Unreleased]. CI configuration, repository automation, agent instructions, and documentation-only work are infrastructure and stay out of it. See the Changelog section below. - Verify, and say what you ran.
npm run checkplus the suites your change touches, with the actual output — not the commands you intended to run.
The test for whether documentation needed updating is not “did I add a feature”. It is: would someone following the current docs now be misled? If yes, that is part of this change, not a follow-up.
Bundled skills follow the writing-for-agents method (context pointers, the
two loads, the information hierarchy — see “Creating skills” in
packages/coding-agent/docs/skills.md). A skill’s description is a context
pointer paid for on every turn: front-load the leading word, one trigger per
branch, and prefer deletion to explanation in the body.
At every phase or milestone boundary — a release cut, a feature arc landing
across several PRs, a “what’s next” pause — run a docs sync pass as its own
step, not per-PR: reread README.md, docs/getting-started.md, and the docs of
whatever the arc touched as a NEW USER would, against what main actually does
now. Per-PR housekeeping catches what one change invalidates; it reliably
misses what an arc of changes adds up to (an installer landing in one PR and
self-update in another means “how do I upgrade?” belongs in three places no
single PR owned). Also refresh GitNexus (--skip-agents-md) at these
boundaries.
Tech Stack
This repo runs a hybrid toolchain, matching upstream earendil-works/pi task for task.
Each tool is used where it is actually better, rather than one runtime being mandated
everywhere. Where the split differs from pi, the reason is written down.
| Task | Tool | Why |
|---|---|---|
| Dependency install | npm ci --ignore-scripts |
package-lock.json is the single verified lockfile. npm ci refuses to install when it and package.json disagree; nothing enforced that while two lockfiles coexisted |
| Supply-chain gate | committed .npmrc |
save-exact=true, plus a 2-day release-age cooldown declared under both min-release-age (pi’s spelling, which npm 10 ignores) and minimumReleaseAge (npm’s own key, npm 11.6+). .github/dependabot.yml carries the matching cooldown, scoped to the github-actions ecosystem |
| Build | npm run build |
tsgo, not Bun; no behaviour change |
| Lint / format | biome check (npm run check, npm run format) |
pi’s rule set exactly: recommended preset plus the same six overrides. Tab indent width 3, line width 120 |
| Typecheck / check | npm run check (biome + tsc --noEmit + shrinkwrap check) |
pi runs biome + tsgo here |
| Root test suites | vitest --run --project {unit,integration,ci} |
pi uses vitest for its workspace tests, with a shared vitest.base.ts setting only resolve.alias |
packages/coding-agent suite |
vitest --run |
already parity; it now runs under Node rather than bun --bun, SQLite selectors included |
| Script tests | node --test scripts/*.test.mjs |
pi parity. Scripts Node can run are tested with Node’s own runner |
| Repository scripts | bun run scripts/*.ts |
Bun executes .ts directly and resolves .js specifiers to .ts source with no loader hook. Bare node cannot; scripts meant for node --test are .mjs |
| Binary compilation | bun build --compile |
Cross-compiles the single-file executables; upstream pi uses Bun for exactly this step too. Bun pinned to 1.3.14 |
| npm-package smoke tests | Node (node-version: 22 in CI, matching pi) |
test/integration/installed-package-node-extensions.test.ts verifies the shipped atomic bin under #!/usr/bin/env node, which is how npm installs run it |
| Registry publish | npm publish --provenance |
npm’s OIDC-signed provenance lives in the npm CLI, and npm trusted publishing requires a GitHub-hosted runner |
What actually gates a pull request here. .github/workflows/ci.yml — and only that
file. It runs three ubuntu-latest jobs: verify (biome, tsc, the shrinkwrap check, the
coding-agent build, the roundtable tests, the demo’s digest bound, the manifest plan, and
the long-context baseline check), suites (the inherited unit suite and the CI contract
tests), and review-gate (fails a PR whose automated review reported passing but was
actually skipped — see CI Docs).
The inherited Atomic workflows — test.yml, publish.yml, warm-toolchain-cache.yml —
describe a nine-context matrix with full Windows coverage on Blacksmith runners, and
test/ci/test-workflow-topology.test.ts still asserts that shape. None of them run.
Blacksmith runners are registered to the upstream organization and never pick up jobs on
this repository, so those workflows are disabled at the repository level rather than
rewritten. publish.yml and warm-toolchain-cache.yml are kept byte-identical to upstream,
except that Dependabot moves their action SHA pins (the repo opted into that in dependabot.yml).
test.yml cannot be: it carries the rebrand’s ORPHUS_REQUIRE_* env-var names, and it has
also fallen behind upstream’s own later edits to it. Read them as a record of upstream’s
topology, not as this repository’s gate, and do not optimize for check contexts that will
never report.
The practical consequence: this fork has no Windows CI. prek.toml records a
Windows-only line-ending bug that reached main because of it. Treat a
platform-sensitive change as unverified on Windows until someone runs it there.
- TypeScript ≥ 5.x (strict,
noUnusedLocals,noUnusedParameters) @sinclair/typeboxfor schema definitionsjitifor runtime TS loading where needed
Quick Reference
Commands
npm ci --ignore-scripts— install dependencies frompackage-lock.jsonnpm install <pkg>— add a dependency;.npmrcapplies the 2-day release-age gate andsave-exactnpm run check— biome (--error-on-warnings),tsc --noEmit, and the published-shrinkwrap check.npm run typecheckis the typecheck alonenpm run demo— the scripted three-agent discussion over the real broker socket. No model. Asserts the late-joiner digest stays under 40% of the raw transcript, so it fails rather than merely reportingnpm run demo:loop— the same, extended through export → memory ingest → recall by a fresh sessionnpm run evals:baseline— the long-context baseline: what an oversized tool result costs the parent’s context window, measured through the real spill path. Fails on regression against the committedevals/longcontext/scorecard.json. Deterministic and model-free; the model-backed task families are deliberately kept out so CI can gate on this half. Seeevals/longcontext/README.mdnpm run roles— turnorphus.roles.yamlinto launch commands (--format plan|json|sh|tmux|orca)npx vitest --run --project unit test/unit/roundtable-— the Orphus tests alone, in secondsnpm run test:unit,npm run test:integration,npm run test:ci-contracts,npm run test:allnpm run test --workspace=@orphus/coding-agent— the coding-agent vitest suite, under Nodenpm run test:scripts—node --test scripts/*.test.mjsnpm run hooks:install,npm run hooks:runbun run scripts/<name>.ts— repository scripts stay on Bun; see the Tech Stack table- Git hooks are configured in
prek.toml;npm installruns the rootpreparescript to install hooks withprek install --prepare-hooksusingdefault_install_hook_types.
Do not run yarn install or pnpm install, and do not reintroduce bun install: each
writes a competing lockfile that npm ci neither reads nor verifies, and bypasses the
.npmrc release-age gate. bun.lock and packageManager: bun@… were removed for this
reason. Bun remains a declared engine and is still the right tool for the rows above that
name it.
Best Practices
- Avoid ambiguous types like
anyandunknown. Use specific types instead. - Source files use
.jsimport extensions (TypeScript ESM convention). The repo ships as.tsfiles; Bun resolves.jsspecifiers to the underlying.tssource directly — no loader hook required. atomic’s loader follows the same convention as pi. - Do not add a build step (
dist/,tsconfig.build.json, etc.) topackages/workflows; it distributes raw TypeScript and the host loads it directly.packages/coding-agentis copied from upstream pi and keeps its existing build setup. - When using skills, if you see a frontmatter of
metadata: internalset totrue(if missing assumefalse), that means the skill is for internal developers of this package. If this flag is omitted, the skill is meant for consumers/everyday users.
Design Context
For the Orphus contribution — why the digest is extractive and model-free, why cursors are
keyed by role name, why the broker is separate from intercom’s, and the trust boundary —
read packages/roundtable/DESIGN.md.
The root DESIGN.md is a different document: it is Atomic’s inherited TUI design-token
spec (palette, typography, spacing), and says nothing about rooms. PRODUCT.md is likewise
Atomic’s product brief.
Issues and pull requests
Follow CONTRIBUTING.md for external-contributor coordination, issue assignment, and pull request guidance.
Testing
Use npm run test:unit (or test:integration, test:all) and make use of your tdd skill to write high quality tests. The suites run under vitest; the assertion style stays node:assert/strict:
import { test } from "vitest";
import assert from "node:assert/strict";
test("hello world", () => {
assert.equal(1, 1);
});
Replacing Bun globals in tests
Root suites run under Node, so Bun.* and import.meta.dir are unavailable. Use
test/helpers/runtime.ts rather than reaching for node:fs/node:child_process directly —
several of the replacements differ in ways that fail silently:
| Bun | Helper | Trap it closes |
|---|---|---|
Bun.sleep |
sleep |
— |
import.meta.dir |
moduleDir(import.meta.url) |
— |
Bun.file(p).text()/.json()/.exists() |
readText / readJson / fileExists |
readJson<T>() returns unknown by default; Bun.file().json() returned any |
Bun.write |
writeFileEnsuringDir |
Bun.write creates parent directories; fs.writeFile throws ENOENT |
Bun.spawnSync |
spawnSyncCollect |
Node returns status, not exitCode; a raw port makes every assert.equal(r.exitCode, 0) compare against undefined |
Bun.spawn |
spawnProcess |
Node has no .exited promise and no web-stream stdio, and reports a missing binary asynchronously rather than throwing |
process.execPath when spawning Bun |
bunExecutable() |
Under vitest process.execPath is Node, so a .ts child, bun run, or a Bun-specific -e script silently runs under the wrong runtime |
Shipped packages/ code that uses Bun.* behind isBunBinary/isBundledBuild is out of
scope and must not be edited to suit the test runner. One shipped module —
packages/web-access/subprocess.ts — calls Bun.spawn/Bun.sleep unguarded, because it
only ever runs inside the Bun-compiled binary. test/unit/web-access-subprocess.test.ts
therefore calls installBunGlobal(), which substitutes the helpers above for those two
globals. All five test names and their assertions survive, and what they cover — bounded
draining, the byte cap, the timeout, the kill path — is unchanged. What they no longer
cover is Bun’s own spawn implementation, which is what the shipped binary runs on. The
alternative, re-execing the file under Bun, would collapse five names into one wrapper
assertion; that is a worse trade, but the gap is real and is stated here rather than only in
the helper.
SQLite selectors run on either runtime
src/core/tools/resource-selectors.ts loads node:sqlite first and falls back to
bun:sqlite. node:sqlite is unflagged from Node v22.13.0 and is what upstream pi uses;
Bun 1.3.14 does not ship it (oven-sh/bun#32498 is merged but unreleased), and the shipped
binary is Bun-compiled, so the fallback is what keeps that binary working. When Bun releases
node:sqlite, both runtimes take the first branch and the fallback can be deleted.
better-sqlite3 was evaluated and rejected: it segfaults Bun 1.3.14 on construction, which is
worse than a catchable missing-module error.
Test fixtures go through packages/coding-agent/test/helpers/sqlite.ts, which mirrors that
preference order behind the bun:sqlite-shaped API the suites were written against. Never
reintroduce a soft guard (if (!sqlite) return, or a ? it : it.skip) — that is how one test
skipped and eleven kept passing with every assertion dead. test/ci/ci-workflow-contracts.test.ts
enforces the loader order and rejects those guards.
Per-test timeout policy
- The suite-wide per-test budget is 30000 ms, declared once as
TEST_TIMEOUT_MSintest/helpers/test-timeout.tsand applied by the rootvitest.config.tsto all three projects.test/ci/ci-workflow-contracts.test.tsenforces that the threetest:*scripts each select a project and that all three resolve to that one value. - Do not restate the budget in a package script, in
.github/workflows/test.yml, or inbunfig.toml. The contract test rejects a--timeoutflag in any script, and Bun ignores[test] timeoutin bunfig anyway — it looks correct and does nothing. - One platform-neutral value, never a Windows-only branch. A Windows-only bump would leave Linux as the only place the budget is enforced and hide Windows regressions until they were far worse. (
packages/coding-agent/vitest.config.tskeeps its own pre-existing 90 s Windows branch, local to that project.) - Add an explicit third-argument timeout only for a test whose cost is structural (a full builtin-package loader reload, a real CLI child process, a real
vitestchild, atscinvocation, a built-package install). Name the constant and keep it at the call site —REAL_VITEST_SUITE_TIMEOUT_MSintest/unit/flaky-test-suite-runner.test.tsis the pattern; a bare120_000says nothing about why the cost is structural rather than a slow test nobody fixed. Never restate the default value — an explicit timeout that merely repeats it silently lowers that test’s budget when the default rises. scripts/run-flaky-test-suite.tsscores every duration against that test’s effective timeout: warn at 40 % of budget, fail the step at 70 %. Every attempt is scored, so a fast bounded retry cannot hide a first attempt that burned a test’s headroom. It always writes the per-test duration table to.ci-diagnostics/<suite>-durations.md, on green runs too. If it fails your test, make the test faster or justify a structural explicit timeout — do not raise the shared default.- The gate reads vitest’s JSON reporter, which the wrapper requests alongside the default one so the step log stays readable. The reporter emits a record per test, so the gate now scores the whole suite rather than the 97 % that printed a duration under Bun’s stdout, and
blind— tests ran, no durations — finally means the harness broke. A report that is missing or unreadable counts as blind too: an unparsable report measures exactly as much as one that was never written. - The gate reads a budget only from a vitest invocation, following one
npm run <script>indirection intopackage.jsonand then into the config that script selects. A leadingbun/bunxis the runtime rather than the command and is stepped over. Any other wrapped command leaves the gate disabled rather than scoring output against a budget nothing enforced. Explicit per-test budgets are matched by the fully qualifiedscope > name, so a declaration insidedescribenever lends its budget to a same-named test in another scope. - Do not raise
WARN_RATIO. The move to vitest made the heaviest tests materially slower (vite transform cost:coding-agent builtin resources > loads builtin pi package resourceswent 622 ms → ~10 s), and the slowest unit test now sits just under the 40 % warn line. A loaded or Windows runner may start warning. That is the gate working as designed — make the test faster or justify a structural explicit timeout.
Load sensitivity
vitest runs test files in parallel by default, and this repository deliberately sets no
pool, maxWorkers, poolOptions, or fileParallelism — pi sets none either. A test that
only passes on an idle machine is a bug in that test. Fix it where it lives: give the real
work headroom and derive the assertion from a named constant (see STALLED_ATTEMPT_CAP_MS in
test/unit/subagents-attempt-watchdog-helpers.ts). Do not skip it, do not serialize the
suite, and do not shard — test/ci/test-workflow-topology.test.ts forbids
--parallel|--shard|--concurrent|--max-concurrency for exactly this reason.
Ambient provider credentials
A test that only passes on a machine without provider credentials is a bug in
that test, exactly like load sensitivity. Model-world fixtures build a real
ModelRuntime over pi-ai’s full builtin provider list, and some providers
authenticate from the environment alone (amazon-bedrock via the AWS default
chain, google-vertex via ADC) — so a developer’s configured AWS CLI, or a
sandbox proxy injecting dummy AWS keys, silently makes “no models available”
fixtures see a 114-model catalog and spawned CLI children dispatch real
provider requests. packages/coding-agent/test/provider-env-scrub.ts (wired as
that project’s vitest setupFiles) deletes the ambient credential variables
before any test module loads; fixtures that need a credential set their own
afterwards. When adding a suite outside that project that touches model
availability, scrub the same list rather than assuming a bare environment.
Hook name compatibility
Use beforeAll/afterAll for once-per-suite setup/teardown and beforeEach/afterEach for
per-test hooks. before/after are not exported.
Code Quality
- Frequently run
npm run check(typecheck plus the shrinkwrap check).npm run typecheckis the typecheck alone. - Avoid
anyandunknowntypes. - Modularize code and avoid re-inventing the wheel. Use functionality of libraries and SDKs whenever possible.
Debugging
You are bound to run into errors when testing. As you test and run into issues/edge cases, address issues in a file you create called issues.md to track progress and support future iterations. Delegate to the debugging sub-agent for support. Delete the file when all issues are resolved to keep the repository clean.
Docs
Relevant resources (use your playwright-cli skill if the information is not available in the local docs):
- Bun (runtime + test runner):
oven-sh/bun - Pi:
earendil-works/pi - TypeScript:
microsoft/TypeScript - Schema tooling:
@sinclair/typeboxfor runtime-validated schemasjitifor on-demand TS loading
Coding Agent Configuration Location
atomic:
- global:
- Linux/MacOS:
~/.atomic/agent/ - Windows:
%HOMEPATH%\.atomic\agent\\
- Linux/MacOS:
- extensions:
~/.atomic/agent/extensions/<name>/ - local:
.atomic/in the project directory
Releasing
Atomic uses a versionless release-base flow: supported bases keep packages/*/package.json at 0.0.0; scripts/cut-release.ts materializes the real version only on a tagged detached Release <version> commit with harmless immutable Release-base-ref/Release-base-sha trailers. Pushing the version tag directly starts publish.yml. Its lightweight integrity job checks that the source resolves to the tag commit, packages/coding-agent/package.json equals the tag, and the subject is Release <version>. Build jobs produce and smoke-test native modules and archives; a draft GitHub Release is staged before OIDC-only npm publication and undrafted only after npm succeeds. publish-npm alone receives id-token: write under npm-publish; release staging, undrafting, and failed-draft cleanup alone receive contents: write. Configure npm trusted publishers with filename publish.yml and environment npm-publish.
Cut and publish a release with:
bun run scripts/cut-release.ts 0.8.31 --base main --push
The selected base is never advanced by the version stamp. The script resolves its exact refs/heads/... ref on origin, creates the release commit in a detached git worktree, records the base trailers, tags it, and abandons the worktree. The tag push is the publication signal. The publisher deliberately does not validate or allowlist those trailers; its integrity boundary is the tag/package-version/commit-subject match.
Agent publishing requests
If a user asks to publish a release or prerelease, route the request through the repository-local publish-release Atomic workflow:
- Ask for the version only when it was not supplied. Stable releases use
MAJOR.MINOR.PATCH; prereleases useMAJOR.MINOR.PATCH-alpha.REVISIONwith revision starting at 1. - Infer release versus prerelease from a valid supplied version; ask only when it is ambiguous or invalid. Use the requested
base_ref, defaulting to the short branch namemainwhen omitted. - For non-main bases, require the branch to be protected with the repository’s required CI checks before using it as the selected release base.
- Launch one
publish-releaseworkflow run withtarget_version,release_kind, andbase_ref. Do not duplicate its Git, PR, tag, or publishing actions inline. - The workflow creates
[release|prerelease]/<version>from the selected base, updates relevant changelogs without bumping package versions, validates and commits the changes, pushes the branch, and opens the PR. - It watches required CI until every required check reaches a terminal state, treating an admin merge of the PR as approval to proceed; check failures or an expired watch window stop the run with evidence.
- After checks pass, it merges the exact verified PR head, switches to the selected base, and fast-forwards from
origin/<base_ref>. - It runs
bun run scripts/cut-release.ts <version> --base <base_ref> --push --yes, which stamps only the detached release commit and pushes the tag. That tag push automatically startspublish.yml; the workflow does not manually dispatch normal publication. - It watches the matching
Publish <version>action until it completes. Failure or an expired watch window stops the run with evidence; success returns a concise release summary.
Docs
- ALWAYS keep the user-facing docs in
packages/coding-agent/docsup-to-date with the latest changes after you make changes. Prefer to keep other docs up-to-date as well, but the coding-agent docs are the most important since they are user-facing and often consulted by users and other agents. - To update docs, prefer using your
release-docsworkflow to thoroughly update all relevant docs with the latest changes. If you need to make a quick fix or update, you can also edit the markdown files directly, but make sure to keep them comprehensive and up-to-date.
Changelog
Location: packages/*/CHANGELOG.md (each package has its own)
Format
Use these sections under ## [Unreleased]:
### Breaking Changes- API changes requiring migration### Added- New features### Changed- Changes to existing functionality### Fixed- Bug fixes### Removed- Removed features
Rules
- Package changelogs are user-facing release notes. Add entries only for changes to shipped package behavior, APIs, features, or user-visible fixes.
- CI configuration, release/publish pipelines, repository automation, maintainer scripts, and agent-instruction changes are infrastructure-level changes. Do not add them to
packages/*/CHANGELOG.mdunless they also change the behavior of a shipped package for users. - In particular, changing how a release is tagged, dispatched, built, verified, or published does not itself warrant a package changelog entry.
- Before adding entries, read the full
[Unreleased]section to see which subsections already exist - New entries ALWAYS go under
## [Unreleased]section - Append to existing subsections (e.g.,
### Fixed), do not create duplicates - NEVER modify already-released version sections (e.g.,
## [0.12.2]) - Each version section is immutable once released
- When updating the changelog entry you should:
- Carefully note key features that were added for a particular
prereleaserevision and for eachreleaseversion changelog you should note every key feature that was introduced in the cumulativeprerelease(s) that led up to therelease. - Do NOT be lazy and avoid saying something like: “Bumped package version for the Atomic prerelease.” That is not helpful to users and does not provide any information on what was actually changed.
- The changelog should be a comprehensive and detailed summary of all the key features, bug fixes, breaking changes, and other relevant information about the
release/prereleasethat would be helpful for users.
- Carefully note key features that were added for a particular
Attribution
- Internal changes (from issues):
Fixed foo bar ([#123](https://github.com/earendil-works/pi-mono/issues/123)) - External contributions:
Added feature X ([#456](https://github.com/earendil-works/pi-mono/pull/456) by [@username](https://github.com/username))
Versionless release bases & bumping
main and supported workstream bases are versionless: every packages/*/package.json (plus package-lock.json workspace entries, the @orphus/natives dependency pin, packages/natives/native/index.js checks, and the Cargo manifests/lock) stays at the 0.0.0 placeholder. Do not bump the version on a release base.
scripts/bump-version.ts is the low-level stamper that rewrites every versioned manifest. It is invoked by scripts/cut-release.ts inside a throwaway worktree at the exact remote base SHA to materialize the real version on the tagged release commit. You normally never run it directly against a release base; the only direct use is resetting the placeholder if it ever drifts:
# stamp a real version onto the off-base tag commit (preferred; explicit base shown)
bun run scripts/cut-release.ts 0.1.0 --base main
bun run scripts/cut-release.ts 0.1.0-alpha.1 --base main
# low-level: reset main back to the versionless placeholder
bun run scripts/bump-version.ts 0.0.0 && npm install --package-lock-only --ignore-scripts
CI
An overview of CI is described here: CI Docs.
Note: npm provenance publishing uses GitHub OIDC trusted publishing and must not configure a static npm credential.
Tips
- The workflows extension is bundled into
@orphus/coding-agent. For local development against upstream pi, symlinkpackages/workflowsinto~/.pi/agent/extensions/workflowsif you want host-level discovery outside Atomic. - Rely on agent skills to provide information on best practices during implementation. Here is a short list of Agent Skills that are incredibly relevant to this project that you should try to use when applicable:
- bun
- gh-commit
- gh-create-pr
- prek
- typescript-advanced-types
- typescript-expert
- Ask for clarity if you are unsure about a change. The developer is your best friend and oftentimes can clarify intent.
- When modifying this extension, follow pi’s extension and SDK conventions.
<EXTREMELY_IMPORTANT>
@orphus/workflows ships raw .ts files with no build step — do NOT introduce dist/, tsconfig.build.json, outDir, or any bundling.
Install with npm ci --ignore-scripts, and add dependencies with npm install. Never run
yarn install or pnpm install, and do not bring back bun install: each writes a competing
lockfile that npm ci neither reads nor verifies, and bypasses the min-release-age gate in
the committed .npmrc. package-lock.json is the only lockfile, and it is also the input to
the shrinkwrap published inside @orphus/coding-agent.
Bun is still required, and still correct, for three things: compiling release binaries with
bun build --compile, running scripts/*.ts, and running the Bun-hosted test fixtures. See
the Tech Stack table for the full split.
</EXTREMELY_IMPORTANT>