PDF Knowledge Base Q&A

Search and answer within your own papers and materials, with context and sources attached to the response.

Once a collection passes a few hundred documents, the problem is no longer having the material but locating it. You remember a figure from some paper without remembering which one; comparing the same metric across several reports means opening them one at a time. Filename search does not help, because what you need is inside the content.

Knowledge base Q&A turns your uploaded material into something you can ask questions of. You ask; it searches within your own documents, reads the relevant passages, answers, and tells you which file and section the answer came from. When the material does not contain the answer, it says so instead of filling the gap with general knowledge that merely sounds right.

What it can do

Answers carry sources

Every response names the file and location it came from, so you can open the original. This is what determines whether an answer is citable.

Absence is reported as absence

When nothing in the base supports an answer, it says the material is not there rather than substituting the model general knowledge. For research use this distinction is the whole point.

Cross-document comparison

Ask questions that span several documents, such as the value of one metric across multiple reports, and get them located individually and laid out side by side.

Built up over time

The base is persistent. Add new material whenever, without disturbing what is already indexed, and coverage improves as you use it.

How it works

  1. Create and uploadSet up a base by topic and upload PDFs and other material in bulk; each file is parsed and indexed.
  2. Ask and retrieveQuestions are asked in natural language and matched against passage meaning, not just keywords.
  3. Read, then answerRetrieved passages are read before the answer is composed, and the answer arrives with its sources.
  4. Follow up and go deeperFollow-up questions build on the answer, and the agent can consolidate everything on a topic into one document.

Worked examples

Finding evidence in your own library

Which studies in this base report side effects for this intervention, and what were their sample sizes?

The matching papers listed with the side effects and sample size each reported, and the section of each paper it came from.

Checking internal documentation

What timeout did last year technical spec set for this interface? Give me the source.

The value with the exact document passage, and where several documents disagree, the inconsistency is pointed out too.

Frequently asked questions

How much material can a base hold?

It grows continuously and material can be added at any time. Retrieval still runs across the whole base as it grows rather than favouring recent uploads.

Will it answer from the model general knowledge instead?

The two are kept separate by design: content found in your base is answered with sources, and when nothing is found it states that the base does not contain it. That separation is what makes it usable for research.

Which file formats are supported?

Primarily PDF, plus common document and spreadsheet formats. Scanned files go through text recognition first.

Is my material private?

It belongs to your own knowledge base, is used only to answer your questions, and is not used to train models.