Fixing a real bug on crustdata.com's AI Assistant, and cutting LLM costs by ~50% in the process


Problem Statement

While using the Crustdata AI Assistant on their live website for a personal project, I ran into a bug: after a few back-and-forth messages, the conversation broke with a 413 (Content Too Large) error.

Digging into it, the root cause became clear: every message sends the entire conversation history to the backend, which then passes it fully to the LLM on each request.

This creates two compounding problems:

Issue Impact
Token bloat LLMs charge per input/output token. Resending the full chat on every turn means paying to reprocess old messages again and again — costs scale badly as conversations grow.
Accuracy risk Longer contexts increase the chance of hallucination — the model has more irrelevant noise to sift through to find what's relevant now.
Hard failure Eventually, the payload exceeds the request size limit entirely — the 413 error — breaking the experience for any user with a longer conversation.

This isn't a one-off glitch — it's a structural issue in how the assistant handles "memory": by just replaying the entire conversation, every time.


Solution

To show what a fix could look like, I built a custom Crustdata AI Assistant MVP — a rebuild of the same assistant, using a two-tier memory + LLM routing system, orchestrated with LangGraph.

How it works:

User Message
     │
     ▼
┌─────────────────────┐
│   Cheap/Fast LLM      │  → Extracts user intent & preferences
└─────────┬───────────┘
          │
          ▼
   Stored in Postgres (structured memory)
          │
          ▼
┌─────────────────────┐
│  Expensive/Main LLM   │  → Receives a short, distilled context
└─────────┬───────────┘     (not the full conversation)
          ▼
     Response to User

Core mechanics:

  1. Intent extraction layer — A cheap, fast LLM silently reads every incoming message in the background and extracts the user's current intent/preferences (not for generating the response, just for understanding).
  2. Structured memory (Postgres) — Extracted intent is stored as structured data, not raw chat logs — so retrieval is fast and cheap.
  3. Distilled context injection — Instead of resending the full conversation, the main (expensive) LLM only receives a short summary of "what the user wants right now" — dramatically shrinking the input payload.
  4. Data hygiene / cron cleanup — A scheduled cron job wipes anonymous user data older than 24 hours, keeping the database lean and privacy-friendly by default.
  5. Optional insights pipeline — Instead of deleting that data, it can be rerouted to a data/analytics layer (e.g. Databricks) to surface patterns in what users are actually asking for — turning "chat data" into a product insight engine.