Recorded engineering

A Todo CLI session, checked against its audit

A real AiOrch Python demo: three implementation agents, review revisions, a resolved merge conflict and an independent audit that corrected the reported test count.

The recorded task

On 24 April 2026, an owner-run AiOrch session built a Python Todo CLI with SQLite persistence, data validation, an argparse interface and packaging. It is a software demonstration, with a public demo repository, rather than a customer deployment.

The session assigned three implementation roles: core logic, CLI presentation, and packaging/documentation. Core logic and packaging began together; the CLI agent started after its core dependency was approved. The work used separate branches, not three workers editing one checkout.

Review changed the implementation

All three implementation agents received changes-requested feedback and subsequently recorded approved final verdicts. The responses covered scope cleanup, CLI error handling, and packaging/documentation corrections. Approval is a review result, not a claim that every possible defect was eliminated.

StepRecorded outcome
Implementation3 agents assigned; all 3 eventually approved
Branch integrationA CLI branch conflict was recorded and later resolved
Integrated result3 of 3 branches marked merged inside the session
Independent auditA separate session used 4 audit/follow-up agents
GitHub deliveryA public pull request exists; it was still open when checked

The record contains session-resume events and conflict resolution. It does not establish a run without operator intervention, nor identify the actor for every resume.

The audit corrected the test count

The parent final report claimed 36 tests. The independent audit reran python -m pytest tests/ -v and recorded 35 passed, with exit status zero. It also recorded a successful CLI import check. This case uses the audited count rather than repeating the initial report.

The audit identified report/specification discrepancies and prepared follow-up fixes, including date-model consistency and empty-update behavior. Its session completed with four approved agents. The inspected audit record did not show a merged audit result, so completion of that audit must not be read as proof that every follow-up was integrated into the public PR.

Evidence has layers. The parent timeline recorded an integration-test pass, while the audit noted incomplete detail in the saved integration log. The separate audit test run is the clearer evidence for the 35-test count.

Inspect the public handoff

Demo-projects pull request #5 was created on 24 April 2026. GitHub still reported it as open, with no merge timestamp, when checked on 5 October 2026. The integrated session branch and a merged GitHub PR are different delivery states.

The historical run reported planning with Opus, implementation roles using Kimi, Sonnet and Haiku labels, and review using a Codex default label. Exact model revisions and the historical AiOrch build were not established from the inspected record.

What this example establishes

This is evidence of task decomposition, dependency-aware work, review revisions, conflict resolution, an independent audit and a public handoff. It is not a benchmark against another tool or proof of production reliability.

The original task file was removed during the development review; the audit reconstructed the specification from saved session configuration. A reusable run should preserve its approved specification as an audit artifact.

Usage data was missing for several agents. The dashboard’s displayed zero cost is not evidence that the work was free; no total cost or time-saving claim is made here. Historical test results were inspected, not rerun during this publication.

Learn how to configure these controls in review and validation, or start with the installation guide.

Run your first session →