PvdAI
AI-powered document browser for the articles of association of the Dutch Labour Party (PvdA).
Party statutes are long, dense, and written in legal Dutch. PvdAI lets members ask questions in plain language and get answers they can actually understand.
PvdAI makes the 188-page Statuten en reglementen PvdA 2023 accessible through a split-screen interface: a document browser on the left and an AI chat on the right. Ask questions in plain language and get answers at B1 reading level with references to specific articles.
- RAG-powered Q&A — answers grounded in the actual document via semantic search over embeddings
- Conversational context — follow-up questions use conversation history with automatic query rewriting
- Interactive document browser — collapsible table of contents with on-demand section loading
- Document search — full-text search with text snippet previews and highlighted matches
- Clickable article references — AI responses link directly to referenced articles in the browser (including plural "Artikelen X, Y, Z" references)
- Copy & share — copy or share AI responses directly from the chat
- Shareable questions — link to a question via URL query parameter for auto-submission on load
- Dark/light mode — system-aware with manual toggle
- Mobile-first — responsive panel switching with accessible touch targets
- Accessible — WCAG AA contrast compliance, 44px touch targets,
prefers-reduced-motionsupport,aria-liveannouncements, focus-visible styles, andinertattribute on hidden panels - SEO-friendly — server-rendered summary and TOC, sitemap, robots.txt, IndexNow, FAQPage schema, SearchAction, canonical URL; dedicated
/veelgestelde-vragenand/privacybeleidpages - Privacy-first — questions not stored or used for training (
store: false); IP addresses HMAC-SHA256 hashed with daily rotation before storage; GA4 advertising features disabled - Rate limiting — 20 questions/day per IP via Redis (Vercel/Upstash), with remaining count shown to users; IPs stored only as HMAC hashes
- Error handling — custom 404 page, error boundaries, incomplete stream detection
| Layer | Technology |
|---|---|
| Framework | Next.js 16 (App Router) |
| UI | React 19, SCSS Modules |
| AI | OpenAI SDK — text-embedding-3-small + Responses API |
| Language | TypeScript 5 |
| Hosting | Vercel |
| Rate Limiting | Redis via Upstash (in-memory fallback for local dev) |
| Testing | Vitest |
| CI | GitHub Actions — lint, test, and npm audit (CVE check) on push/PR; Dependabot weekly dependency updates |
| Database | None — static JSON files |
git clone https://github.com/julianaijal/PvdAI.git
cd PvdAI
cp .env.example .env.local # then add your OpenAI API key
npm install
npm run devOpen http://localhost:3000.
| Variable | Required | Description |
|---|---|---|
OPENAI_API_KEY |
Yes | OpenAI API key |
REDIS_URL |
No | Upstash Redis URL. Falls back to in-memory rate limiting |
IP_HASH_SECRET |
Yes (prod) | Secret key for HMAC-SHA256 IP hashing. Generate with openssl rand -hex 32 |
The project uses static JSON instead of a database. Run these once:
# 1. Parse the PDF into structured data
npx tsx scripts/parse.ts
# Output: data/structure.json, data/toc.json, data/chunks.json
# 2. Generate vector embeddings for semantic search
npx tsx scripts/embed.ts
# Output: data/embeddings.bin + data/embeddings.meta.json (gitignored)npm testnpm run build # runs embed.ts as prebuild step, then next buildWhen the source PDF changes:
npm run content-update # parse + embed + notify search engines via IndexNowgraph TB
subgraph Client["Browser"]
ML["MainLayout<br/><small>split-screen · mobile toggle</small>"]
DB["DocumentBrowser<br/><small>TOC tree · article viewer · search</small>"]
CH["Chat<br/><small>messages · starter questions · article links</small>"]
TT["ThemeToggle<br/><small>dark / light</small>"]
ML --> DB
ML --> CH
ML --> TT
end
subgraph Server["Next.js Server"]
PAGE["page.tsx<br/><small>SSR · reads toc.json</small>"]
ASK["/api/ask<br/><small>POST · RAG endpoint</small>"]
SEC["/api/section<br/><small>GET · 24h cache</small>"]
STR["/api/structure<br/><small>GET · full doc</small>"]
end
subgraph Pipeline["Data Pipeline <small>(build time)</small>"]
PDF["PvdA PDF<br/><small>188 pages</small>"]
PARSE["parse.ts"]
EMBED["embed.ts"]
PDF --> PARSE
PARSE --> EMBED
end
subgraph Data["Static JSON"]
TOC["toc.json"]
STRUCT["structure.json"]
CHUNKS["chunks.json"]
EMB["embeddings.bin"]
end
subgraph Lib["lib/"]
OAI["openai.ts<br/><small>client singleton</small>"]
EMBLIB["embeddings.ts<br/><small>dot-product search</small>"]
RL["ratelimit.ts<br/><small>Redis / in-memory</small>"]
end
subgraph External["External Services"]
OPENAI["OpenAI API<br/><small>embeddings + responses</small>"]
REDIS["Upstash Redis<br/><small>rate limiting</small>"]
end
PAGE --> TOC
PARSE --> TOC
PARSE --> STRUCT
PARSE --> CHUNKS
DB -->|"fetch section"| SEC
DB -->|"search"| STR
CH -->|"question"| ASK
SEC --> STRUCT
STR --> STRUCT
ASK --> EMBLIB
ASK --> OAI
EMBLIB --> EMB
EMBLIB --> CHUNKS
OAI --> OPENAI
RL --> REDIS
ASK --> RL
EMBED --> EMB
style Client fill:#f8f9fa,stroke:#dee2e6
style Server fill:#e8f4f8,stroke:#bee5eb
style Pipeline fill:#fff3cd,stroke:#ffc107
style Data fill:#d4edda,stroke:#28a745
style Lib fill:#e2e3f1,stroke:#6c757d
style External fill:#f8d7da,stroke:#dc3545
app/
├── page.tsx Server component — reads toc.json at build time
├── layout.tsx Root layout with metadata and theme setup
├── globals.scss CSS variables, PvdA branding, dark/light themes
├── robots.ts Dynamic robots.txt
├── sitemap.ts Dynamic sitemap
├── api/
│ ├── ask/route.ts POST — RAG endpoint (embed → search → respond)
│ ├── section/route.ts GET — on-demand section content (24h cache)
│ └── structure/route.ts GET — full structure.json for client-side search
├── components/
│ ├── MainLayout/ Split-screen layout, mobile panel toggle
│ ├── DocumentBrowser/ Collapsible TOC tree, article viewer, document search
│ ├── Chat/ Message history, starter questions, copy/share, article links
│ └── ThemeToggle/ Dark/light mode switch
lib/
├── openai.ts OpenAI client singleton
├── embeddings.ts Dot-product search over L2-normalized vectors
├── ratelimit.ts Redis rate limiter with in-memory fallback (20/day)
├── logger.ts Structured JSON logger (Vercel-friendly)
└── schemas.ts Zod schemas for API request validation
scripts/
├── parse.ts PDF → structured JSON (pure JS, no system deps)
├── embed.ts Chunks → binary vector embeddings (runs at build time)
└── indexnow.ts Notify search engines of content changes
tests/ Unit tests (Vitest)
.github/workflows/ci.yml CI pipeline (lint + test on push/PR)
data/ Generated data files (embeddings gitignored)
- User question is embedded with
text-embedding-3-small - For follow-up questions, the query is automatically rewritten using conversation history for better context
- Top 5 chunks found via dot-product search against L2-normalized embeddings; chunks below cosine similarity 0.35 are discarded
- Chunks + question sent to OpenAI Responses API with a Dutch B1-level system prompt
- Response streamed back as SSE; article references become clickable links in the UI
Ask a question about the articles of association.
Request:
{
"question": "Hoe word ik lid van de PvdA?",
"history": []
}Response (SSE stream):
data: {"delta": "Je kunt lid worden als je 16 jaar of ouder bent..."}
data: [DONE]
Headers: X-RateLimit-Remaining — questions left today.
Errors:
| Status | Body | Cause |
|---|---|---|
400 |
{"error": "..."} |
Invalid request (missing/malformed question) |
429 |
{"error": "Je hebt het maximum aantal vragen voor vandaag bereikt..."} |
Rate limit exceeded (20/day per IP) |
500 |
{"error": "Er is een fout opgetreden bij het verwerken van je vraag."} |
Internal server error |
Load section content on demand.
Response: Full section object with nested children. Cached for 24 hours via Cache-Control.
Returns the full document structure for client-side search.
- Questions are not stored and not used for AI training — OpenAI API called with
store: false - IP addresses are never stored in plaintext — hashed with HMAC-SHA256 (server secret + daily date) before Redis storage; hashes expire after 24h and rotate daily
- GA4 advertising features disabled (
allow_google_signals: false,allow_ad_personalization_signals: false) - Full details at pvdai.tech/privacybeleid
Contributions are welcome. Please open an issue first to discuss what you'd like to change.
This project is licensed under the MIT License.
PvdAI is an independent open-source project. It is not affiliated with, endorsed by, or an official product of the Partij van de Arbeid (PvdA). AI-generated answers may contain inaccuracies. Always consult the official document for authoritative information.
