Automatically shrink LLM prompts by up to 70% while preserving semantic meaning. Measure token savings, similarity scores, and real-time API cost reduction.
| Feature | Description |
|---|---|
| Extractive Compression | TextRank (sumy) — picks the most informative sentences |
| Abstractive Compression | BART-large-CNN — rewrites and condenses content |
| Semantic Similarity | all-MiniLM-L6-v2 cosine similarity scoring |
| Cost Analytics | Token delta + USD saved across 5 LLM pricing tiers |
| Async FastAPI | Non-blocking inference via thread pool |
client (Astro)
│ /api/* → proxy
↓
server (FastAPI)
├── /compress POST — main endpoint
└── /health GET — liveness check
↓
CompressionService
├── TextCleaner — normalise + remove filler
├── Compressor — extractive | abstractive
├── SimilarityModel — sentence-transformers
├── Tokenizer — tiktoken cl100k_base
└── CostAnalyzer — per-model USD savings
prompt-compression/
├── client/ # Astro frontend
│ └── src/
│ ├── components/
│ │ ├── PromptInput.astro
│ │ ├── CompressionControls.astro
│ │ ├── ResultsPanel.astro
│ │ └── CostDashboard.astro
│ ├── pages/index.astro
│ ├── lib/api.ts
│ └── styles/global.css
│
├── server/ # FastAPI backend
│ └── app/
│ ├── main.py
│ ├── core/
│ │ ├── compressor.py # TextRank + BART
│ │ ├── similarity.py # SentenceTransformers
│ │ ├── tokenizer.py # tiktoken
│ │ └── cost.py # USD analytics
│ ├── models/schemas.py
│ ├── services/compression_service.py
│ └── utils/text_cleaner.py
│
└── benchmarks/
├── long_prompts.json
├── run_benchmark.py
└── results.csv # generated
cd server
python -m venv venv
# Windows
venv\Scripts\activate
# macOS/Linux
source venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload --port 8000Server runs at http://localhost:8000 · Open /docs for interactive Swagger UI.
cd client
npm install
npm run devUI runs at http://localhost:4321
{
"prompt": "Your long prompt text here…",
"method": "extractive", // "extractive" | "abstractive"
"target_ratio": 0.5 // 0.1 (very aggressive) → 0.9 (light)
}Response:
{
"original_tokens": 412,
"compressed_tokens": 187,
"compression_ratio": 0.454,
"similarity_score": 0.921,
"cost_saved_percent": 54.6,
"compressed_prompt": "…compressed text…",
"method_used": "extractive",
"latency_ms": 243.1
}python benchmarks/run_benchmark.pyRuns all 5 sample prompts × 2 methods × 3 ratios = 30 runs.
Results saved to benchmarks/results.csv.
cd server
docker build -t promptzip-server .
docker run -p 8000:8000 promptzip-serverModels are pre-downloaded at build time — container starts instantly.
| Model | Purpose | Size | Device |
|---|---|---|---|
facebook/bart-large-cnn |
Abstractive summarisation | ~1.6 GB | CPU |
all-MiniLM-L6-v2 |
Semantic similarity | ~90 MB | CPU |
cl100k_base (tiktoken) |
Token counting | tiny | CPU |
RAM requirement: ≥ 8 GB. 16 GB recommended for BART + MiniLM simultaneously.