A Retrieval-Augmented Generation (RAG) chatbot that can answer questions from research papers and PDFs.
Runs fully offline using Ollama with open-source models (Mistral, LLaMA, Phi).
Large Language Models (LLMs) like ChatGPT are powerful, but they suffer from:
- ❌ Hallucination – making up facts when they don’t know the answer
- ❌ Outdated knowledge – frozen at training time
- ❌ No access to private data – can’t directly read your PDFs, manuals, or research papers
Retrieval-Augmented Generation (RAG) solves this problem:
- Retrieve – Search for relevant passages from your own documents using embeddings & vector databases.
- Augment – Provide those passages as context to the LLM.
- Generate – The LLM answers the question grounded in the retrieved context.
In short: RAG makes LLMs more accurate, up-to-date, and customizable to your data.
This project is a PDF-based AI research assistant.
You upload PDFs (e.g., research papers, policies, technical docs), and the chatbot:
- Splits PDFs into text chunks.
- Embeds them using SentenceTransformers (BAAI/bge-small-en-v1.5).
- Stores vectors in FAISS (fast similarity search).
- When you ask a question:
- The system retrieves the most relevant chunks.
- Sends them to a local LLM (Mistral, LLaMA, or Phi via Ollama).
- The LLM generates an answer grounded in the PDFs.
I have also built an evaluation pipeline that measures:
- 🔍 Recall improvement compared to simple keyword search
- ⏱️ Query latency (average ~1.2s per query with Mistral 7B)
- 📊 Visual results with recall and latency comparison charts
This project is a measurable, benchmarked RAG system.
flowchart TD
subgraph Ingestion["Ingestion Pipeline"]
A[PDFs in /data] --> B[Text Extraction - PyPDF]
B --> C[Chunking - RecursiveCharacterTextSplitter]
C --> D[Embeddings - BGE-small SentenceTransformers]
D --> E[FAISS Index - vector_store]
end
subgraph QueryFlow["Query Flow"]
U[User Question] --> Q[Embed Query]
Q --> R[FAISS Retrieve Top-K]
R --> CXT[Build Context with Chunks]
CXT --> P[Prompt Builder]
P --> LLM[Ollama LLM - Mistral or LLaMA or Phi]
LLM --> ANS[Grounded Answer with Citations]
end
E -.-> R
ANS --> UI[Streamlit UI]
sequenceDiagram
participant User
participant UI as Streamlit UI
participant FAISS as FAISS Index
participant Ollama as Ollama LLM
User->>UI: Ask Question
UI->>FAISS: Retrieve Top-K Chunks
FAISS-->>UI: Return Relevant Chunks
UI->>Ollama: Send Prompt + Context
Ollama-->>UI: Return Answer
UI-->>User: Display Grounded Answer
- 📄 Upload one or more PDFs and query them in natural language
- 🔍 Semantic search with SentenceTransformers embeddings
- 🧠 Local LLM inference with Ollama (Mistral / LLaMA / Phi) or Cloud LLM inference with Gemini
- 📊 Evaluation framework with QA datasets to measure recall & latency
- ⚡ Runs fully offline (after models are pulled)
- Vector DB: FAISS
- Embeddings: BAAI/bge-small-en-v1.5 (SentenceTransformers)
- LLM Inference: Ollama (Mistral 7B / LLaMA 3B / Phi-3) / Gemini (gemini-2.5-flash)
- Frontend: Streamlit
- Backend: Python
- Deployment: Docker-ready
rag-pdf-assistant/
│── app.py # Streamlit frontend
│── rag.py # Core RAG logic
│── ingest.py # PDF ingestion + FAISS index builder
│── eval_rag.py # Evaluation script for recall/latency
│── qa.json # Sample QA pairs for benchmarking
│── data/ # Your PDFs go here
│── vector_store/ # FAISS index storage
│── benchmarks/ # Evaluation results
│── requirements.txt # Python dependencies
│── README.md # Project documentation
pip install -r requirements.txtDownload from ollama.ai and pull a model:
ollama pull mistral:7b-instruct-q4_K_M
# or lighter models based on your system:
ollama pull llama3.2:3b-instruct-q4_K_M
ollama pull phi3:3.8-mini-128k-instructCreate a .env file:
GENERATOR=local_ollama
EMBED_MODEL=BAAI/bge-small-en-v1.5
OLLAMA_MODEL=mistral:7b-instruct-q4_K_MPut your PDFs in the data/ folder.
python ingest.py --pdfs "data/*.pdf" --out vector_storestreamlit run app.pyAccess at: [(https://rag-based-pdf-assistant.streamlit.app/)]
Our Retrieval-Augmented Generation (RAG) pipeline significantly improves accuracy compared to traditional keyword-based retrieval:
- 35% higher recall compared to keyword search
- Average latency ~1.2s/query (Mistral 7B on CPU/GPU hybrid)
- Add rerankers for improved context retrieval
- Integrate Online LLMs like OpenAI, Claude etc.,
- Support multimodal PDFs (figures + text)
- Make it availble for document formats other than PDFs like Word Doc etc...
- Docker Compose for one-line deployment
PurnaChander Konda