While using the Crustdata AI Assistant on their live website for a personal project, I ran into a bug: after a few back-and-forth messages, the conversation broke with a 413 (Content Too Large) error.
Digging into it, the root cause became clear: every message sends the entire conversation history to the backend, which then passes it fully to the LLM on each request.
This creates two compounding problems:
| Issue | Impact |
|---|---|
| Token bloat | LLMs charge per input/output token. Resending the full chat on every turn means paying to reprocess old messages again and again — costs scale badly as conversations grow. |
| Accuracy risk | Longer contexts increase the chance of hallucination — the model has more irrelevant noise to sift through to find what's relevant now. |
| Hard failure | Eventually, the payload exceeds the request size limit entirely — the 413 error — breaking the experience for any user with a longer conversation. |
This isn't a one-off glitch — it's a structural issue in how the assistant handles "memory": by just replaying the entire conversation, every time.
To show what a fix could look like, I built a custom Crustdata AI Assistant MVP — a rebuild of the same assistant, using a two-tier memory + LLM routing system, orchestrated with LangGraph.
User Message
│
▼
┌─────────────────────┐
│ Cheap/Fast LLM │ → Extracts user intent & preferences
└─────────┬───────────┘
│
▼
Stored in Postgres (structured memory)
│
▼
┌─────────────────────┐
│ Expensive/Main LLM │ → Receives a short, distilled context
└─────────┬───────────┘ (not the full conversation)
▼
Response to User
Core mechanics: