<aside> 📌
Foundational, frozen reference. This page is the project's architecture reference and is not expected to change further. Ongoing decisions, research, and scope changes are tracked in the companion CS3IP Project Diary page instead.
</aside>
| Student | Waleed Ahmed, 240048282 | Course | BSc Computer Science with International Foundation Year |
|---|---|---|---|
| Supervisor | Ziyang Wang | Module | CS3IP Individual Project, Aston University |
| Document date | 25/09/2026 | Status | Foundational. Architecture finalised, no open design questions remaining. |
Large language models are now highly capable at individual software engineering sub-tasks: writing code, following instructions, producing designs, analysing data. But almost all existing tools treat this as a single conversation with a single assistant. That is not how software is actually built. A working system is produced by a team: product managers scoping work, engineers implementing it, testers verifying it, all communicating and iterating together, often with disagreement, rework, and escalation along the way. There is no accessible tool that tests whether a coordinated team of specialised LLM agents, each modelled on a distinct role in a real software company, can take a high-level requirement and autonomously turn it into a working, tested system.
This project investigates that gap directly. Rather than building a single, more capable coding agent, it builds an organisation of agents, a virtual software company, and studies how well that organisation performs the earliest and most expensive stage of software delivery: turning an ambiguous requirement into a working, verified product.
For a solo developer, small team, or research group, on-demand access to a full engineering department without hiring or coordinating one has direct commercial relevance for solo founders, agencies, and product teams. The project also sits inside active multi-agent LLM research: it engages with the growing literature on how role structure, communication design, and personality affect the output quality of orchestrated LLM teams, and it deliberately builds on an existing failure taxonomy (Cemri et al.'s MAST framework) rather than inventing failure categories from scratch.
The project does not aim to compete with industry-scale tools such as ChatDev or MetaGPT, nor does it treat formal repeated experimentation as its primary deliverable. The priority is a strong, functioning system whose design decisions can be explained and defended individually. Function comes before visual polish, and every architectural choice is made for a stated, defensible reason.
Three papers assigned as supervisor-directed background reading each map onto a specific design decision in this project, in some places confirming the chosen approach and in others motivating a deliberate departure from it.
Park et al. (2023) demonstrate that believable, emergent social behaviour in simulated agents arises from three components: a natural-language memory stream recording every experience, a reflection layer that periodically synthesises those memories into higher-level insights, and a planning process that retrieves both when deciding what to do next. Starting from a single seed constraint, their agents autonomously organise a party, spreading invitations and coordinating attendance without being scripted to do so.
This project deliberately does not adopt the full memory-and-reflection architecture. Where Park et al.'s agents pursue open-ended social goals with no fixed completion criterion, this project's agents work against concrete, checkable units of work, tickets with explicit acceptance criteria and, where available, hidden tests. That difference in task structure is what justifies replacing a persistent per-agent autobiographical memory with the lighter structured handoff note passed at each transition (Section 4.3.5): sufficient context to complete a bounded task, at a fraction of the token cost a full memory-and-reflection cycle would add across the many tickets processed during evaluation.
Li et al. (2024) simulate a hospital in which every patient, nurse, and doctor is an LLM-powered agent, with doctor agents improving over tens of thousands of simulated patient encounters and, once evolved, outperforming prior state-of-the-art medical agents on the external MedQA benchmark.
The role-specialisation pattern and the strategy of validating the system against an external, published benchmark rather than a bespoke one both carry over directly to this project's design (Section 4 and Section 6.5, respectively, the latter through TheAgentCompany). Where this project departs is on agent evolution: Agent Hospital's doctor agents accumulate experience and change behaviour over repeated encounters, whereas this project's agents use fixed, tiered models per role for the full duration of any evaluation run, since the evaluation goal is to characterise a fixed architecture's failure modes and compare it against a single-agent baseline, which requires the architecture under test to stay constant across runs.
Schmidgall et al. (2025) structure autonomous research assistance as a staged pipeline (literature review, experimentation, then report writing) with a human able to give feedback at each stage boundary, and report that this human involvement measurably improves output quality, alongside an 84% cost reduction achieved partly through deliberate model selection at each stage.
Both findings map closely onto decisions already made here. The staged pipeline with feedback at defined boundaries is structurally the same pattern as this project's ticket finite-state-machine with intervention gated at specific transitions (Sections 4.3.4 and 4.6). The cost-efficiency result reinforces this project's own tiered model allocation strategy (Section 7): cheaper models for high-volume Engineering-tier work, a stronger model reserved for CTO-tier architectural decisions.