Releval is an API-first platform for evaluating, tracking, and improving search relevance. Connect popular search engines like Elasticsearch, OpenSearch, Solr, Vespa, or your own custom search APIs or search pages; measure quality with standard Information Retrieval metrics; automate judgments with LLM-as-a-judge using popular providers; compare ranking changes; run evaluations in CI; and turn real user behaviour into a continuous relevance feedback loop.
I built Releval because improving search relevance is still far more fragmented and manual than it should be.
I've spent years building and operating search systems, and the same problem appears repeatedly: teams know search quality matters, but they often lack a reliable way to measure whether a change actually made it better. Evaluation data ends up spread across scripts, notebooks, spreadsheets, dashboards, and one-off internal tools. Experiments are difficult to reproduce, judgments are expensive to collect, and relevance regressions are often discovered only after users notice them.
Large technology companies often solve this by building dedicated evaluation platforms internally. They have the engineering capacity, specialist teams, and budgets to create tooling tailored to their own search systems. For smaller companies, building and maintaining that same infrastructure can be difficult to justify. Search may be critical to the customer experience, but relevance tooling is rarely part of the company’s core product, so it competes with customer-facing features for limited engineering time and investment.
Existing commercial and open-source tools also tend to solve only one part of the workflow. A team might have one tool for offline metrics, another for analytics, a custom script for comparing rankings, and a separate process for gathering judgments. Connecting everything requires significant engineering effort before the team can begin improving search.
Releval aims to make that workflow accessible without every company having to build its own internal platform.
The goal is to give search teams a practical place to define representative queries, run them against real search systems, judge the results, compare ranking changes, track quality over time, and connect offline measurements with production behaviour. It works with the search stack a team already has rather than requiring a particular engine, model, or AI provider.
I also want evaluation to become part of normal engineering practice. Relevance tests should be repeatable, automatable, and capable of running in CI like other quality checks. Search changes should be supported by data rather than intuition alone.
This early launch is about learning whether the broader search community experiences the same problems, and whether a focused platform can make relevance evaluation substantially easier.
I built Releval because improving search relevance is still far more fragmented and manual than it should be.
I've spent years building and operating search systems, and the same problem appears repeatedly: teams know search quality matters, but they often lack a reliable way to measure whether a change actually made it better. Evaluation data ends up spread across scripts, notebooks, spreadsheets, dashboards, and one-off internal tools. Experiments are difficult to reproduce, judgments are expensive to collect, and relevance regressions are often discovered only after users notice them.
Large technology companies often solve this by building dedicated evaluation platforms internally. They have the engineering capacity, specialist teams, and budgets to create tooling tailored to their own search systems. For smaller companies, building and maintaining that same infrastructure can be difficult to justify. Search may be critical to the customer experience, but relevance tooling is rarely part of the company’s core product, so it competes with customer-facing features for limited engineering time and investment.
Existing commercial and open-source tools also tend to solve only one part of the workflow. A team might have one tool for offline metrics, another for analytics, a custom script for comparing rankings, and a separate process for gathering judgments. Connecting everything requires significant engineering effort before the team can begin improving search.
Releval aims to make that workflow accessible without every company having to build its own internal platform.
The goal is to give search teams a practical place to define representative queries, run them against real search systems, judge the results, compare ranking changes, track quality over time, and connect offline measurements with production behaviour. It works with the search stack a team already has rather than requiring a particular engine, model, or AI provider.
I also want evaluation to become part of normal engineering practice. Relevance tests should be repeatable, automatable, and capable of running in CI like other quality checks. Search changes should be supported by data rather than intuition alone.
This early launch is about learning whether the broader search community experiences the same problems, and whether a focused platform can make relevance evaluation substantially easier.