Kubernetes-native token cost control plane for LLM APIs.
Pario sits between your applications and LLM providers (OpenAI, Anthropic, etc.) to give you full visibility and control over token spend — budgets, caching, routing, and real-time observability.
- Transparent Proxy — drop-in replacement for OpenAI and Anthropic API endpoints with SSE streaming support
- Token Tracking — per-key, per-model usage tracking with session detection
- Token Budgets — per-team, per-app, per-model spend limits
- Semantic Caching — deduplicate similar prompts (SQLite local, Redis distributed)
- Smart Routing — route requests across models with fallback chains
- Cost Attribution — team/project cost breakdowns with per-model pricing
- Audit Log — opt-in full request/response logging for compliance and debugging
- MCP Server — expose stats, budgets, costs, and audit data to AI agents via Model Context Protocol
- Live Observability —
pario topfor real-time token usage, Prometheus metrics
┌─────────────┐ ┌───────────────────────────────────────┐ ┌──────────────┐
│ Your App │────▶│ Pario │────▶│ LLM Provider│
│ │◀────│ proxy · budget · cache · router │◀────│ (OpenAI etc)│
└─────────────┘ └───────────────────────────────────────┘ └──────────────┘
│ │ │
┌────▼───┐ ┌────▼───┐ ┌───▼────┐
│Metrics │ │ SQLite │ │ CRDs │
│(Prom) │ │ Cache │ │(Budget)│
└────────┘ └────────┘ └────────┘
# Install (once releases are available)
# brew install pario-ai/tap/pario
# Or build from source
make build
./bin/pario --helpmake build # compile to bin/pario
make test # run tests
make lint # run golangci-lint
make clean # remove build artifactsTBD