MorrowCache - Skip the chat call when the question is the same

by•
MorrowCache is a drop-in OpenAI-compatible proxy. Change baseURL—keep your model and API keys. Agents rephrase until done; each try is a full chat bill. Reuse cached completions when a same-intent judge matches (cloud Jev or local Kev/Laya)—not cosine alone. Demo: 2817ms → 423ms on a paraphrase HIT. Fail-open on judge errors. Freshness refuses stale "latest" hits. Streaming & tools bypass. npx /jevcache@latest start · morrowcache.vercel.app/agent-setup

Add a comment

Replies

Best
Maker
📌
Hey PH 👋 I'm Kushal, maker of MorrowCache. The pain: same user question, different wording—or an agent that rephrases until the answer "looks done"—and you pay for every chat completion again. Vector-ish caches miss paraphrases or merge the wrong prompts. MorrowCache is a local OpenAI-compatible proxy. Point your client at , keep your upstream model. Before calling the model, a same-intent adjudicator (TypeSafe Jev by default, or local Kev/Laya) decides whether to reuse the cached answer. Not "similar text" as the only signal. What I'd call out upfront: • Fail-open — if the judge errors, the real model still runs • Freshness — time-sensitive asks can refuse a stale HIT • Streaming + tool calls bypass the cache (by design in v0) Quick try (~5 min, Node 22.5+): npx /jevcache@latest start Demo with no real API key: MOCK_UPSTREAM=1 npx /jevcache@latest start --demo Then send a prompt + a paraphrase → expect MISS then HIT (X-Jevcache headers or /stats). Easiest in Cursor / Claude Code: copy the setup prompt from Recorded pair on the site: 2817ms MISS → 423ms HIT at intent 0.94 — one example, not a guarantee for every workload. I'd love feedback from anyone running agents or support bots: where would same-intent caching help, and where would you never trust it? Happy to dig into architecture in the comments.