Skip to content
#

prompt-compression

Here are 137 public repositories matching this topic...

14-stage Fusion Pipeline for LLM token compression — reversible compression, AST-aware code analysis, intelligent content routing. Zero LLM inference cost. MIT licensed.

  • Updated Apr 1, 2026
  • Python

Local proxy that compresses your LLM API requests so you pay less, with no change to the answers. Trims wasted tokens from prompts, history, tool output, and code before they're sent: -31% input / -74% output, measured live. Any provider, no extra model calls. Also an MCP server and embeddable library (Rust, Python, Ruby, Kotlin, Swift, JS/TS).

  • Updated Sep 1, 2026
  • Rust

A curated list of strategies, tools, papers, and resources for reducing LLM token costs and improving efficiency in production.

  • Updated Sep 6, 2026
terseai

Claude Code & Cursor token usage monitor and cost tracker for macOS & Windows. Live token/cost/burn monitoring across 8 AI coding agents, a budget circuit breaker that stops runaway agents before the next API call, an MCP manager, and 40-70% on-device prompt compression.

  • Updated Sep 6, 2026
  • JavaScript

Add this topic to your repo

To associate your repository with the prompt-compression topic, visit your repo's landing page and select "manage topics."

Learn more