AnySearch - Real-time structured search trusted by agents and developers

A search tool for agents, not a search box. AI agents are only as good as the information they receive. When connected to AnySearch, your agent gets filtered, de-duplicated, and structured information from trusted sources searched in parallel, helping it produce more reliable results. Free to start.

Add a comment

Replies

Best

Hey Product Hunt 👋Grant here from the AnySearch team.

We’re a team of AI developers and engineers building search infrastructure specifically for AI agents.

Traditional search was built for people. We built search for AI agents.

People skim links, compare sources, and decide what to trust. Agents don't.

AnySearch delivers real-time structured search that agents and developers can trust.

When that information is stale, incomplete, poorly routed, or buried in messy HTML, the final output becomes less reliable. Sometimes agents search again and again. Sometimes they confidently build on weak context. Either way, the workflow breaks.

That's the problem we built AnySearch to solve.

🧠 Understands what a query is asking for

🔍 Searches trusted sources in parallel

🚫 Filters SEO spam, ads, and duplicate results

📄 Returns clean, structured information for agents

Why developers use it:

· Fewer repeated search calls

· Less HTML cleanup

· Cleaner context for models

· More reliable agent outputs

Works with your existing workflows

AnySearch is available through:

· Skill

· MCP

· API

Install AnySearch in your agent👇

1. Go to

2. Click Add to Agent in the top right corner.

3. Select Skill, copy the prompt, and paste it into your agent.

Your agent will handle the installation automatically.

Try AnySearch today, add it to your agent, and get started for free.

We'd love your feedback.

how do you handle sites that usually break standard crawlers? like cloudflare challenges.

 Good question! We don't try to solve challenges — we focus on not triggering them. Real browser engine, proper fingerprinting, human-like patterns, etc.

It works on most sites that trip up standard crawlers. For the really aggressive ones, honestly no silver bullet — but those are rare in practice.

 Congratulations

 Congrats. For a developer building multi-step agents that must cite or justify actions, how does AnySearch handle provenance and source confidence over a chain of reasoning?

   Great question. We think of AnySearch less as “the agent says trust me” and more as an evidence layer for the agent.

For each step, the agent can pull structured results with source attribution and confidence/freshness signals, then carry those references forward into its own justification. So instead of only getting a final answer, you can keep a trail of which sources supported which intermediate claim or action.

We’re also careful not to pretend confidence is magic. It’s a source/result-level signal that developers can combine with their own policies, especially when sources conflict or when a human review step is needed.

Congrats on the launch! What is the source of your search results, and what's the frequency of your updates?

 AnySearch features self data sources, covering both general and vertical search.

 Thanks! Short answer: AnySearch isn’t just one static index with a single “updated every X hours” schedule.

We route each query to the best source mix for the intent: live web, vertical sources like academic/legal/security/finance/code, and direct URL extraction when the agent needs to inspect a page. For freshness, it depends on the source: web and page extraction are queried live, while vertical sources follow the update cadence of the underlying data provider.

We also return source attribution and metadata where available, so agents can cite the source instead of treating the result as a black box.

The MCP + skill install path is the right call — "paste the prompt and the agent installs it" is exactly how agent-facing infra should distribute.

Two things I'd want to know before wiring it into an agent loop: what does structured actually mean here — a fixed result schema, or can the calling agent shape it per query (product/price/availability fields for a commerce lookup vs citation-style for research)?

And what latency should we budget? Parallel trusted-source search + dedup + structuring sounds like real work per call, and for a multi-step agent that searches five or six times per task, p95 per call matters more than anything.

Congrats on the launch

 Thanks! To answer both:

Structured output: we have a fixed schema per result (title, url, context) in markdown, plus a JSON API for full structured data. The schema is consistent, so your agent can parse reliably without per-query shaping.

Latency: p50 is ~1s, p95 under 5s. For a multi-step agent loop, I'd budget 5-8s per call to be safe. We're actively working on bringing p95 down.

 thanks for the details around percentiles, will check the product out.

The "real-time structured search trusted by agents" framing raises the question of what structured actually means here. Is the output schema defined by the caller, something like a typed JSON spec the agent passes in, or is AnySearch inferring structure from the query and returning whatever shape seems right? That distinction matters a lot when an agent is downstream and needs to reliably parse the response without a validation step. Also curious how you handle sources that block crawlers aggressively, since real-time web search quality tends to fall apart fast on exactly the sites that have the freshest information.

 The structured content fed back to the agent by AnySearch is in Markdown format, with a single core objective: the search results should be directly used by Agents. This means that the returned context is clean, comes with precise citations, and the agent can infer upon receiving it, rather than spending hundreds of tokens guessing first.

 Spot on—these are arguably the two biggest bottlenecks for agent builders right now!

1. On the "Structured" output: Currently, AnySearch infers the structure and returns Entity-Enriched Markdown (cleanly separating core facts, citations, and metadata). We standardize this to drastically reduce LLM cognitive load, giving your agent a reliable format without needing a heavy validation step. That said, letting developers pass strict JSON schemas is a fantastic idea, and we’re actively exploring it for our roadmap.

2. On aggressive anti-bot blocking: You're absolutely right—generic scraping fails when you need fresh data the most. We solve this structurally: rather than fighting Cloudflare as a headless browser, our Smart Intent Routing bypasses this entirely by tapping into directly integrated, authoritative vertical feeds (finance, health, cyber, etc.).

Really appreciate you digging into the architecture!"

 We don’t try to solve challenges, we just try not to trigger them in the first place, which gets us a lot further than standard crawlers in practice. By structured, I just mean a consistent markdown format (title, url, context) that works well for agents, plus a JSON API if you want actual structured data on the developer side.

The deduplication across parallel sources is the part I'm most curious about. In practice, the same story or data point gets syndicated everywhere, and agents end up with 4 near-identical chunks in context that all look authoritative. How are you handling that, string similarity or something semantic? And how do you manage source trust when a "trusted source" is just wrong about something recent?

 

Thank you for the thoughtful comment — really sharp questions.

Improving the information density of AI Search results is genuinely one of our core objective functions. On the deduplication front, our approach falls into two categories:

1. URL-level dedup. Straightforward — if the same URL is recalled multiple times, we deduplicate at that layer.

2. Information-density ranking (this is the important one). The real challenge isn't URL duplication — it's exactly what you described: different sources reporting the same story with different phrasing but near-zero marginal information gain. To address this, we've built our own information-density ranking algorithm. For multi-intent queries, rather than optimizing purely for relevance, we optimize for maximizing information density when organizing the final result set. This involves jointly accounting for relevance, cross-passage redundancy, information entropy, and a corresponding length penalty — all to ensure that every token the user consumes delivers the largest possible information set, rather than having the same story rehashed three different ways.

As for the dynamic trust problem you raised — when an "authoritative source" gets breaking news wrong — you've called out a genuinely hard one. We're actively researching dynamic trust mechanisms for this, but honestly, it's very difficult to strike the right balance across quality, latency, and cost. We're committed to solving it and plan to roll out improvements in future releases.


Thanks again for the engagement!

 The info-density ranking approach is interesting, optimizing across relevance + redundancy + entropy is basically MMR with extra steps, curious if you're doing that at retrieval time or as a reranking pass after. On the trust problem, the honest answer might just be recency weighting with a decay curve and accepting that no static trust score survives a fast-moving news cycle. Hard to get right without adding latency.

 

We mainly do this at the ranking stage — or what I'd call the "selection" stage.

My take is that in AI/agent scenarios, the traditional notion of "ranking order" doesn't really hold anymore. Users don't need a sorted list to scan top-to-bottom; what they need is subset selection from a large recall pool. That essentially becomes the second stage of search in the agent era: it's no longer ranking, it's subset selection.

On the latency front, we try to keep it lean on two fronts: use word-level features wherever possible instead of document-level recomputation, and avoid re-running models on signals that were already computed in the first stage. The core idea is to keep the subset selection step as lightweight as possible by front-loading computation into the reusable stage.

Finally gave this a spin with a small research task and was honestly surprised how clean the results came back, already sorted and grouped so my agent didn't have to do the busywork.

 Thank you for support, AnySearch is so necessary for agents.

This is not a gimmick. It is a prerequisite for more reliable agents.

 AnySearch isn't a flashy skill, it's foundational. Without clean, verified context, every downstream reasoning step is built on sand. Appreciate you seeing the deeper point here!

 Couldn’t agree more. Better agents need better inputs first. That’s a big part of why we built AnySearch — to help agents work with more reliable information before they start reasoning or acting. Thanks for putting it so clearly.

the dedup and structuring part is the interesting bit to me, most "search for agents" pitches i've seen just wrap a regular search API. what happens when two trusted sources genuinely disagree on a fact, does anysearch pick one and present it as ground truth or does it surface both and let the agent reason about it

 
That's exactly the part that makes this interesting to me. Most "search for agents" pitches I've seen are just thin wrappers around a standard search API. The hard problem isn't "find relevant pages" — it's packing the most useful information into a limited context window.

So we didn't spend much effort on URL dedup — that's a relatively shallow problem. The real challenge is exactly what you described: different sources telling the same story with different wording, where reading three of them yields near-zero incremental information but burns through thousands of tokens.

The solution is a custom information-density ranking algorithm we built in-house. For multi-intent queries, the objective function isn't pure relevance — it's maximizing the information density of the final result set. Specifically, ranking jointly accounts for:

- Relevance — obviously can't go off-topic

- Cross-passage redundancy — how much a new passage overlaps with what's already in the selected set; high overlap gets down-weighted

- Information entropy — the passage's intrinsic information richness and distinctiveness

- Length penalty — given equal information value, shorter wins

This naturally answers your question — when two trusted sources genuinely disagree, it means each carries independent semantic information that the other lacks. They aren't redundant with each other. The algorithm won't pick one — both get included and surfaced. The agent sees a multifaceted, tension-rich information landscape and gets to reason, compare, and weigh things itself, rather than being fed a sanitized "canonical answer" with the contradictions already smoothed over.

At its core, what we're optimizing for is that every token the agent reads delivers maximum marginal information gain — not the same story rehashed three different ways. Disagreement between sources is actually one of the highest-information-density scenarios you can encounter. Suppressing it would actively degrade the value of the result set.

Makes sense, and I like that you're not trying to force a single "canonical" answer out of the model. My only worry with keeping disagreement in the set is cost - if two sources contradict each other you're now paying for both plus whatever reasoning tokens the agent burns reconciling them. Do you cap how many "genuinely conflicting" passages get surfaced per query, or is that left uncapped and trusted to the length penalty to sort itself out?

 There's no specific cap — we keep as much diversity or disagreement as possible without exceeding the num_results the user requested.

got it, so it's whichever comes first - hitting the diversity target or hitting num_results. makes sense for a general default, just wondering if there's ever a case where a user wants strictly ranked results with zero diversity weighting (like a legal or medical search where you want the single best match, not variety)

Given that a significant amount of low-quality content can rank highly on Google, having an agent that can accurately distinguish and prioritize high-quality information is essential. How do you ensure that poor-quality content is effectively filtered out, while only reliable and valuable content is selected and integrated? And congratulations on your launch!

 Thanks for comment. We have resolved these issues through algorithms, I think the content in the Gallery can answer your question.

 Thank you! And yes, that’s one of the core problems we’re trying to solve.

We don’t rely on ranking alone, because high-ranking does not always mean high-quality. AnySearch looks at a mix of signals: source type, originality, freshness, relevance, duplication patterns, and whether other credible sources support or contradict the same claim.

A big part of the work is reducing obvious noise: ads, SEO farms, repeated copies, and pages that look optimized for ranking rather than usefulness. But we also try not to over-filter, because sometimes a niche source can be valuable even if it is not a big authority.

So the goal is not “hide everything except one perfect answer.” It’s to give agents cleaner, structured evidence with attribution and confidence signals, so they can reason from better inputs and still know where the information came from.

The part that stands out is filtering duplicate results and SEO content first.

 definitely

 Yep! AnySearch uses its retrieval and ranking pipeline to filter out low-quality, duplicate, and SEO-heavy content.

123
•••
Next