Congrats on the launch, Otsar. Private inference is quietly becoming a purchase requirement rather than a nice-to-have. We build AI for healthcare and the first question in every security review is where the tokens go, so zero retention plus optional on-prem is exactly the right combination. One question the page does not answer: would you sign a BAA for HIPAA-regulated customers, and do you have SOC 2 or similar in place or on the roadmap? If yes, there is a whole segment of health tech builders currently stuck between closed APIs and self-hosting that you could own.
Report
Congrats on launch @emirsoyturk I’ve been using it for the past week, and I’ve genuinely been impressed with the results. The speed improvements on GLM 5.2 are noticeable right away, and it’s already become part of my workflow. Great product.
Report
Latency felt noticeably snappier than what I usually get from other inference providers, and I love that nothing gets stored. Refreshing to see a team actually prioritize both speed and privacy at the same time.
@selim321016 Amazing! We are working hard to optimize models specifically for long-context agentic workflows.
Report
the zero-retention, no-training-on-your-code angle is the whole pitch for me. most devs just assume their prompts end up as training data somewhere and live with it. if you can actually prove retention is zero, that's a real reason to switch, not just a nice-to-have
@alex_watson2110 Thanks for the comment! Privacy is part of our core beliefs and has been at the heart of what we have built over the past years at MoonMath and Ingonyama.
Private inference is one of the most underrated problems in the agent space right now. I build in a compliance-heavy vertical, so "your code never leaves your control" is a real buying trigger, not a nice-to-have.
How are you handling the latency tradeoff vs. hosted frontier models?
@clemente_lopez1 100% private inference is a need. In terms of latency, there is no need to trade it for privacy! We don't trade any privacy for latency and our latency is top tier.
Report
Love that there's no request retention here, that's a real differentiator for anyone dealing with sensitive data. One thing that would help me evaluate it faster though: a small public latency dashboard comparing Zro to other inference providers on common open models. Even just a simple weekly update with p50 and p95 numbers across regions would make it way easier to decide if it's worth migrating workloads over.
Hi @otsar_shalmoni , Congrats on the launch. Just wanted to ask - what sort of performance metrics are you getting right now and what are you expecting under sustained load? TTFT/TPS etc? Will you be including performance metrics on the website - and can subscribers be a part of the new model optimization experiments?
ClinicFrame
Congrats on the launch, Otsar. Private inference is quietly becoming a purchase requirement rather than a nice-to-have. We build AI for healthcare and the first question in every security review is where the tokens go, so zero retention plus optional on-prem is exactly the right combination. One question the page does not answer: would you sign a BAA for HIPAA-regulated customers, and do you have SOC 2 or similar in place or on the roadmap? If yes, there is a whole segment of health tech builders currently stuck between closed APIs and self-hosting that you could own.
Latency felt noticeably snappier than what I usually get from other inference providers, and I love that nothing gets stored. Refreshing to see a team actually prioritize both speed and privacy at the same time.
Zro
@selim321016 Amazing! We are working hard to optimize models specifically for long-context agentic workflows.
the zero-retention, no-training-on-your-code angle is the whole pitch for me. most devs just assume their prompts end up as training data somewhere and live with it. if you can actually prove retention is zero, that's a real reason to switch, not just a nice-to-have
Zro
@alex_watson2110 Thanks for the comment! Privacy is part of our core beliefs and has been at the heart of what we have built over the past years at MoonMath and Ingonyama.
ClinicFrame
Private inference is one of the most underrated problems in the agent space right now. I build in a compliance-heavy vertical, so "your code never leaves your control" is a real buying trigger, not a nice-to-have.
How are you handling the latency tradeoff vs. hosted frontier models?
Zro
@clemente_lopez1 100% private inference is a need. In terms of latency, there is no need to trade it for privacy! We don't trade any privacy for latency and our latency is top tier.
Love that there's no request retention here, that's a real differentiator for anyone dealing with sensitive data. One thing that would help me evaluate it faster though: a small public latency dashboard comparing Zro to other inference providers on common open models. Even just a simple weekly update with p50 and p95 numbers across regions would make it way easier to decide if it's worth migrating workloads over.
Zro
@selinkutbag9gj Good idea, thanks for the feedback!
Hi @otsar_shalmoni ,
Congrats on the launch. Just wanted to ask - what sort of performance metrics are you getting right now and what are you expecting under sustained load? TTFT/TPS etc? Will you be including performance metrics on the website - and can subscribers be a part of the new model optimization experiments?