Log

Week of Built Broke Next
21/09/2026 Confirmed the foundational architecture decisions and finalized the Project Brief. Built Layers 1 through 5 of the orchestrator: core Pydantic models and the ticket FSM, role and settings config, the orchestrator loop with its SQLite event log, the Docker sandbox controller, and the FastAPI and websocket dashboard with a chat feed and kanban board. Completed the evaluation harness (FYP-19) across all five of its own layers, including two real tasks ported from TheAgentCompany. Shipped the custom roles editor and kanban intervention controls, fixed Docker daemon access, and brought the README up to date. The ticket FSM had a transition wired to the wrong state. A newer pytest version wrapped hidden-test output in colour codes the parser could not read. The repeat-with-notes action never actually reached the next agent. Fixing Docker access surfaced a far more serious bug underneath it: the sandbox image never set a git identity, so every ticket-attempt commit had been failing silently since Layer 4. Wire up a real Claude-backed agent and get supervisor confirmation on the evaluation scope.
28/09/2026 Wired up the real Claude-backed agent for every stage of the ticket lifecycle and ran the first real end-to-end ticket through the live pipeline. Found and fixed the Review to Done approval never actually merging code, despite being documented as if it did. Built the hierarchical-versus-flat comparison runner and LLM rubric judge for Option B and ran it live against real tasks. Built the human-eval study delivery mechanism. Redesigned the dashboard's visual system and made the kanban board interactive. Started the top-down 3D office visualization, getting its static shell built, then the five agents seated at their home stations, then navigation and the walk cycle so they can be sent anywhere in the office, and then a behaviour model driven by the orchestrator's own ticket states, including what the office looks like when a ticket retries, times out, escalates or halts, then a library of six scenarios that can be played, paused, restarted and sped up on demand, and finally click-to-inspect with an event log, which completes the visualization as planned bar the deliberately deferred live-data layer. The engineering agent could declare a submission complete without ever committing its work. Running Option B for real surfaced five further bugs: an unbounded retry loop on review rejection, no cap on escalation-override cycles, no spend cap before a real-money run, a passing test run wrongly read as a failure, and a sandbox container leaking when a turn raised mid-way through. Start writing the Term 1 report; the 3D office view is built as far as planned, with only the deferred live-stream layer left, start writing the Term 1 report, and wait on the first supervisor meeting.
05/10/2026 Built the project poster, from studying eleven example posters and choosing what goes on it, through a bare first layout and two rebuilds to a symmetric page in A1 proportions with a numbered reading order. Drew the three diagrams, took the dashboard screenshots in the light theme, cleaned the 3D office image and later replaced it with a frame showing the agents walking to the drawing board, recoloured the Aston logo for a dark background, and exported a PDF. Measured the layout against the centre line and fixed the small misalignments it found. Drafted the questions to ask at the supervisor meeting. The second layout looked right but had no structure for the eye to follow, so it had to be rebuilt. Adding spacing pushed the footer off the page more than once. Several header attempts for the logo and QR code had to be undone. Two claims on the poster were looser than my diary supports and were reworded. Supervisor meeting on 08/10/2026, 3:30pm to 4pm, for ethics approval on the human evaluation study and the poster's requirements. Check the cost claim for each team, the reference style and that the QR code scans, before the poster is printed. Start writing the Term 1 report.