<aside> 🗓️
Diary log entries 23/09/2026 to 25/09/2026, in original chronological order. 8 entries.
</aside>
Started from the supervisor's original pitch, a multi-agent LLM system mimicking a software company, and the project brief document brought to the first working session. Confirmed the pre-existing decisions: five agents across three tiers (CTO, Product, and three Engineering), a single code path for attaching to or creating a repository, Slack-style channel-based communication rather than a free-for-all chat, and a presentation-agnostic event log underneath whatever visualisation gets built. Explicitly rejected a pixel-art "walk around the office" visualisation as premature.
Designed the communication protocol from four open questions (message format, routing, turn-taking, context passing). Landed on a unified typed Message record that doubles as the event log, orchestrator-mediated publish/subscribe routing, a per-ticket finite state machine for turn-taking, and structured handoff notes instead of full chat-history replay for context passing.
Designed the execution sandbox: one Docker container per ticket attempt, torn down after each attempt, with a fixed tool interface rather than raw shell access.
A follow-up pass surfaced two gaps and closed them: a persistent per-run git workspace beneath the ephemeral sandbox containers, and isolating hidden acceptance tests from the agent-visible test runner. Also confirmed: retry cap of 3 before escalation, strictly sequential one-ticket-at-a-time concurrency, a CTO-then-Product requirement intake pipeline, git branch-per-ticket with the Review to Done transition acting as the actual merge, and SQLite over Postgres for the event log and task board.
Decided observation needed to be live, not just after-the-fact log inspection. Settled on a minimal default view (chat feed) with additional detail gated behind a settings panel, split into two independent toggle groups: showcasing the multi-agent "society" dynamics for the viva, and surfacing technical metrics for debugging/evaluation.
Compared the working design against the CS3IP Project Definition Form as actually submitted to Ziyang (dated 24/09/2026) and found four real divergences, plus one softer emphasis mismatch. In every one of the four hard clashes, the form turned out to be correct and the working design had drifted away from it:
The unresolved fifth point: the form states that formal repeated experiments are not the core deliverable, while the evaluation design reads as fairly rigorous, flagged for Ziyang's input rather than a unilateral call.
Drafted and sent an email to Ziyang asking directly whether the evaluation methodology should stay a lightweight validation layer or whether the more rigorous approach already designed is worth keeping despite the form's framing. Awaiting his reply.
Compiled the full technical design into a single detailed brief document, intended both as standalone documentation and as source material for a future project poster. Added a Related Work section engaging with three papers Ziyang assigned as background reading, plus verified citations for MAST, TheAgentCompany, and SWE-bench.