<aside> 🗓️
Diary log entries 25/09/2026 to 28/09/2026, in original chronological order. 11 entries.
</aside>
Layer 1 covers the core Pydantic data models (Message, Ticket, Channel, RoleConfig) and the declarative ticket finite-state-machine transition table. RoleConfig gained a personality field and an is_preset flag to support user-created custom roles; the FSM gained a requires_approval flag on transitions to support human intervention.
Config layer: config/roles.yaml, config/settings.yaml, a loader module connecting both to the Layer 1 models, and a pyproject.toml declaring dependencies without installing anything, since the dev environment remains intentionally parked.
An independent review of the newly-rebuilt diagram found three real correctness errors, all traced to the same root cause: the diagram had drawn "In Progress" as several separate boxes spread across a branching column layout, when the actual finite state machine has exactly one IN_PROGRESS state with four incoming transitions and two outgoing ones. Specifically, the changes_requested transition was visually stacked beneath the failed-test retry column, making it look like it originated from a failed test rather than from Review; neither retry-loop re-entry into In Progress showed the code_submission transition looping back to Awaiting Test; and the escalation section opened with its own visually disconnected In Progress box.
Lesson: a branching box-grid layout is the wrong shape for any state that has multiple incoming edges feeding back into itself, since drawing it more than once inevitably implies it's a different state each time. Fixed with a single-column main path plus an explicit list of every loop-back/exception transition as "A leads to B via trigger" rows, the same representation now used as the table format on this page and the Brief page.
Built the control loop that actually drives a ticket through the FSM. Three pieces: a SQLite-backed event log (EventLog, using the stdlib sqlite3 module, no new dependency) whose messages table is literally the append-only stream the comm protocol describes, with a ticket-state projection alongside it; an Agent/AgentContext/AgentResponse interface that concrete LLM-backed agents will implement once model selection is resolved, deliberately decoupled from the loop mechanics so those mechanics could be built and tested now; and the step/run/resume functions that invoke the active role's agent, log its message, resolve the trigger it declares into an FSM transition, own retry-count bookkeeping (the agent reports tests_failed, but whether that stays within budget or forces retry_cap_exceeded is orchestrator state, not an agent decision), and pause at requires_approval transitions in intervention mode for a human decision to resume with.
Building this surfaced a real bug in the Layer 1 FSM transition table: retry_cap_exceeded was defined as a transition from IN_PROGRESS, but the only place it's meant to fire, per RETRY_LOOP_EDGE and the module's own comments, is in place of tests_failed on the AWAITING_TEST -> IN_PROGRESS edge. No transition existed from AWAITING_TEST to ESCALATED at all, so the orchestrator could never actually reach escalation via the retry cap. Fixed by moving that transition's from_state to AWAITING_TEST. Same root cause as the earlier FSM diagram bugs: IN_PROGRESS is a single state with several incoming edges, and it's easy to attach a transition to the wrong one of them.
This is also the first layer with a passing end-to-end test suite (8 tests, stdlib unittest, covering the happy path to Done, the retry-cap escalation path through to a CTO-issued Halted, and intervention-mode pause/resume) and the first time the declared dependencies (pydantic, pyyaml) were actually installed into a venv and run, rather than only syntax-checked with py_compile.
Built the fixed tool interface agents use instead of raw shell access: SandboxController (run_tests/run_command/read_file/write_file/git_diff/git_commit). DockerAttemptSandbox is the Docker-backed implementation: one container per ticket-attempt, non-root, no network egress, CPU/memory/pids limits, with a fresh checkout cloned from a persistent host-side GitWorkspace rather than a bind mount. FakeSandboxController covers fast tests that don't need a real container. docker/sandbox.Dockerfile defines the base image, with git and pytest baked in at build time rather than installed at container run time, since attempt containers have no network egress by design.
Before building this, resolved a real design decision: should awaiting_test re-invoke the engineering LLM agent to call run_tests() and self-report the result, or should the orchestrator call it directly as a deterministic step? Chose the latter. Letting the agent whose code is being tested also be the one certifying whether it passed is a genuine task-verification failure mode (MAST FC3), and it matches how SWE-bench/TheAgentCompany's own harnesses work: the harness verifies, the agent doesn't self-certify. awaiting_test no longer invokes an Agent at all; AgentContext now carries the current attempt's sandbox so the engineering agent can call its tools during in_progress, and the loop creates a fresh sandbox on entering in_progress and closes it once awaiting_test resolves, win or lose.
Docker is installed on the dev machine (29.7.2), but the daemon socket wasn't accessible yet (permission denied) as of this session. The group-membership fix was applied but hadn't taken effect in-session by the time this layer shipped. tests/test_docker_sandbox_integration.py exercises the real Docker path end to end and self-skips gracefully when the daemon or the built image isn't available, so the rest of the suite (16 tests) runs green against FakeSandboxController regardless.
Deferred, not built here: real "attach to existing repo" wiring (GitWorkspace takes an already-local repo path, not a user-supplied remote), and hidden-vs-visible test isolation (depends on the evaluation harness's benchmark adaptation, not yet built).
Noticed the Diary and the Jira board didn't reference each other at all: two independent logs of the same work rather than a linked pair. Fixed one direction cleanly: this page now links the matching ticket key inline wherever a log entry or status-table row maps clearly to one ticket. Some entries don't map to a single ticket (the brief-writing and FSM-bug-fix entries span work that was never separately ticketed) and were left as-is rather than forcing a link.
For the other direction, the first attempt overshot: a Web Link back to this page was added to thirteen tickets individually, when what was actually wanted was a single reference ticket. Corrected after the clarification: FYP-21, a new ticket under the Documentation and report epic, is the intended single place collecting Brief/Diary/GitHub/board links, alongside one deliberate exception: the "Keep project diary updated" ticket also carries a Diary Web Link, since that is literally its job. The stray links on every other ticket couldn't be removed via the available Atlassian API (no delete operation exists for a Jira issue's Web Links) and were removed by hand in the Jira UI instead. Confirmed clean by re-checking all thirteen.