Skip to content

About

A web app that compresses LLM prompts by up to 70% using extractive or abstractive methods, with semantic similarity scoring and API cost savings analytics.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Prompt Compression Engine

Automatically shrink LLM prompts by up to 70% while preserving semantic meaning. Measure token savings, similarity scores, and real-time API cost reduction.

Features

Feature Description
Extractive Compression TextRank (sumy) — picks the most informative sentences
Abstractive Compression BART-large-CNN — rewrites and condenses content
Semantic Similarity all-MiniLM-L6-v2 cosine similarity scoring
Cost Analytics Token delta + USD saved across 5 LLM pricing tiers
Async FastAPI Non-blocking inference via thread pool

Architecture

client (Astro)
   │   /api/* → proxy
   ↓
server (FastAPI)
   ├── /compress     POST  — main endpoint
   └── /health       GET   — liveness check
        ↓
   CompressionService
   ├── TextCleaner     — normalise + remove filler
   ├── Compressor      — extractive | abstractive
   ├── SimilarityModel — sentence-transformers
   ├── Tokenizer       — tiktoken cl100k_base
   └── CostAnalyzer    — per-model USD savings

Project Structure

prompt-compression/
├── client/                  # Astro frontend
│   └── src/
│       ├── components/
│       │   ├── PromptInput.astro
│       │   ├── CompressionControls.astro
│       │   ├── ResultsPanel.astro
│       │   └── CostDashboard.astro
│       ├── pages/index.astro
│       ├── lib/api.ts
│       └── styles/global.css
│
├── server/                  # FastAPI backend
│   └── app/
│       ├── main.py
│       ├── core/
│       │   ├── compressor.py    # TextRank + BART
│       │   ├── similarity.py    # SentenceTransformers
│       │   ├── tokenizer.py     # tiktoken
│       │   └── cost.py          # USD analytics
│       ├── models/schemas.py
│       ├── services/compression_service.py
│       └── utils/text_cleaner.py
│
└── benchmarks/
    ├── long_prompts.json
    ├── run_benchmark.py
    └── results.csv            # generated

Quickstart

1 — Server (Python)

cd server
python -m venv venv
# Windows
venv\Scripts\activate
# macOS/Linux
source venv/bin/activate

pip install -r requirements.txt
uvicorn app.main:app --reload --port 8000

Server runs at http://localhost:8000 · Open /docs for interactive Swagger UI.

2 — Client (Astro)

cd client
npm install
npm run dev

UI runs at http://localhost:4321

API Reference

POST /compress

{
  "prompt": "Your long prompt text here…",
  "method": "extractive",      // "extractive" | "abstractive"
  "target_ratio": 0.5          // 0.1 (very aggressive) → 0.9 (light)
}

Response:

{
  "original_tokens":    412,
  "compressed_tokens":  187,
  "compression_ratio":  0.454,
  "similarity_score":   0.921,
  "cost_saved_percent": 54.6,
  "compressed_prompt":  "…compressed text…",
  "method_used":        "extractive",
  "latency_ms":         243.1
}

Benchmarking

python benchmarks/run_benchmark.py

Runs all 5 sample prompts × 2 methods × 3 ratios = 30 runs.
Results saved to benchmarks/results.csv.

Docker (Server)

cd server
docker build -t promptzip-server .
docker run -p 8000:8000 promptzip-server

Models are pre-downloaded at build time — container starts instantly.

Models Used

Model Purpose Size Device
facebook/bart-large-cnn Abstractive summarisation ~1.6 GB CPU
all-MiniLM-L6-v2 Semantic similarity ~90 MB CPU
cl100k_base (tiktoken) Token counting tiny CPU

RAM requirement: ≥ 8 GB. 16 GB recommended for BART + MiniLM simultaneously.

About

A web app that compresses LLM prompts by up to 70% using extractive or abstractive methods, with semantic similarity scoring and API cost savings analytics.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages