Skip to content

Reliable answers

The assistant is at its sharpest in a short, focused conversation about one topic. The longer and fuller a chat becomes — many questions, many documents, several topics mixed together — the more the assistant has to hold in view at once. At some point it can start to mix up details or “fill in” a passage that isn’t literally there. This isn’t a PrudAI quirk: it applies to every current AI model and it’s measurable — the longer the conversation, the higher the chance of such mistakes.

The good news: this is largely in your own hands.

This is by far the most effective habit. Begin a new chat for each new topic or part of the case file. A short, focused chat keeps the assistant anchored to exactly the documents that matter for that question — and that is precisely what stops it from mixing up sources.

Rules of thumb:

  • New topic? New chat. Don’t carry one chat from the liability question to legal costs to the writ of summons.
  • Keep the number of documents per chat focused — only what’s relevant to that question, not your entire case file in one running conversation.
  • Notice a chat getting “tired” (answers become vaguer, or the assistant refers to something you don’t recognise)? Start again in a fresh chat. You lose nothing: your documents and projects stay put.

Treat every answer as a draft. Click the source reference under an answer and read the source text yourself before you use anything in your advice or court document. See a reference to a document or passage you don’t recognise? Don’t trust it — start a new chat and ask again with only the relevant document attached. More on this: Citations & source availability and Chat.

If the assistant hasn’t actually retrieved a document, it should say so — not guess at its contents. If you notice it quoting a passage that’s wrong anyway, give the message a 👎 with a short note. That signal reaches us and helps us sharpen the assistant further.

This isn’t a PrudAI quirk but a well-documented property of every current AI model:

  • Long inputs become less reliable. Independent research tested 18 leading models from different providers and found performance degrades steadily as the input grows — even on simple tasks.
  • Information in the “middle” is used less well. Models use information at the beginning and end of a long context better than information buried in the middle.
  • Long conversations amplify this. Across multiple turns, accuracy drops by roughly 39% on average versus a single focused question. A key cause: when a model is missing something, it sometimes “fills in” an assumption and then anchors on it — exactly the pattern you want to avoid.

This is precisely why “a new chat per topic” works: you keep the context short, so the model never falls into that pattern.