MMT Token Optimizer reduces unnecessary prompt and context usage across AI agents, RAG and API workloads while protecting required facts, instructions and outcomes. Published lab tests across 7 provider routes show measured input-token reductions from 6.54% to 22.17%, with before/after evidence and semantic outcome guards.
Hi Product Hunt 👋
I’m Bakker, founder of MMTNEXUS.
We built MMT Token Optimizer around one question:
**Can we reduce LLM prompt and context usage without removing the facts, instructions, and outcomes the model actually needs?**
Instead of treating a shorter prompt as success, the optimizer follows a guarded flow:
**Analyze → Reduce → Validate → Run → Measure**
Our published lab cases across 7 provider routes currently show measured input-token reductions ranging from **6.54% to 22.17%**, including:
• Anthropic: 22.17%
• Google Gemini: 18.21%
• OpenAI: 17.91%
• OpenRouter: 17.91%
• DeepSeek: 17.14%
• Mistral: 17.04%
• xAI: 6.54%
These are measured test cases, not guaranteed savings for every workload.
The area I’m most interested in is **AI agents, RAG, long conversation history, and high-volume API workflows**, where context can grow quickly.
For perspective, an agent workload processing 1M input tokens/day would avoid about **5.37M input tokens/month** at a 17.91% reduction.
We’ve opened the product as a limited free public beta so developers can test it with their own workloads.
I’d especially appreciate feedback on:
• workloads where optimization works well
• cases where the system should refuse to reduce context
• agent/RAG edge cases
• what evidence you would want before trusting optimization in production
Thanks for testing it — and please challenge the methodology. That feedback is more valuable to us than simply getting another signup.