I kept watching my research agent burn through 6+ search iterations on queries that should take 2.
Been running a research agent (LangGraph + GPT-4o + Playwright) for competitive analysis. The pattern that kept hitting me:
Hop 1: works fine. Agent searches, gets results, moves on.
Hop 2: usually fine. Agent finds a reference, chases it.
Hop 3: this is where it falls apart.
What I'm seeing at hop 3+:
The agent re-queries with slightly different phrasing, gets near-identical results, doesn't recognize the overlap
Context window is now 40k+ tokens of raw search results and the model starts ignoring its own instructions
It either loops (same query, rephrased) or gives up and synthesizes from whatever's in the window, including irrelevant stuff
When it does find a new source, it loses track of which claim came from which URL
I tried:
Hard-capping iterations at 5 (just makes it give up with a worse answer)
Adding "stop when you have enough" to the system prompt (the model doesn't know when it has "enough")
Compressing context every 3 turns (helps cost, doesn't help the agent realize it's going in circles)
What's actually working for people? Is the problem the search tool, the stop condition, the context management, or all three? I'm starting to think the issue is that the agent is doing two jobs at once — researching AND evaluating whether to stop — and it's bad at the second one because it's already committed to the first.
Replies
Be the first to reply
Have a question or a thought to share? Add a comment above to start the conversation.