Answers carry sources
Every response names the file and location it came from, so you can open the original. This is what determines whether an answer is citable.
Search and answer within your own papers and materials, with context and sources attached to the response.
Once a collection passes a few hundred documents, the problem is no longer having the material but locating it. You remember a figure from some paper without remembering which one; comparing the same metric across several reports means opening them one at a time. Filename search does not help, because what you need is inside the content.
Knowledge base Q&A turns your uploaded material into something you can ask questions of. You ask; it searches within your own documents, reads the relevant passages, answers, and tells you which file and section the answer came from. When the material does not contain the answer, it says so instead of filling the gap with general knowledge that merely sounds right.
Every response names the file and location it came from, so you can open the original. This is what determines whether an answer is citable.
When nothing in the base supports an answer, it says the material is not there rather than substituting the model general knowledge. For research use this distinction is the whole point.
Ask questions that span several documents, such as the value of one metric across multiple reports, and get them located individually and laid out side by side.
The base is persistent. Add new material whenever, without disturbing what is already indexed, and coverage improves as you use it.
Which studies in this base report side effects for this intervention, and what were their sample sizes?
The matching papers listed with the side effects and sample size each reported, and the section of each paper it came from.
What timeout did last year technical spec set for this interface? Give me the source.
The value with the exact document passage, and where several documents disagree, the inconsistency is pointed out too.
It grows continuously and material can be added at any time. Retrieval still runs across the whole base as it grows rather than favouring recent uploads.
The two are kept separate by design: content found in your base is answered with sources, and when nothing is found it states that the base does not contain it. That separation is what makes it usable for research.
Primarily PDF, plus common document and spreadsheet formats. Scanned files go through text recognition first.
It belongs to your own knowledge base, is used only to answer your questions, and is not used to train models.