How much of your Cursor rework is the agent guessing the wrong part of a screenshot?

by•

Half the threads here are about accuracy and fewer iterations. One source of rework I rarely see named: when I paste a screenshot, Cursor has to guess which element on a busy screen I actually meant, and it edits the wrong one. Then I'm re-prompting, which is its own iteration tax.

Curious how much of your rework is this specific thing versus raw model quality. And what's your fix, crop hard, or describe the element in text?

220 views

Add a comment

Replies

Best

I think the screenshot is often just the symptom.

The bigger issue is that the agent lacks the project context behind the screenshot.

Even if it correctly identifies the button or component, it still has to guess:

  • why you're changing it,

  • what decision led to it,

  • what other areas it impacts,

  • what constraints already exist,

  • and what "correct" actually means for this project.

In my experience, a lot of rework comes from intent reconstruction rather than image understanding.

The screenshot tells the agent where.

The project context tells it why.

 Good distinction, and I agree the "why" is underrated. I'd split it in two though. The project-wide why, constraints, what "correct" means here, lives in repo context or a rules file, not the screenshot. The immediate why for one change, "make this element do X," should travel attached to the element, not get rebuilt from a paragraph. A raw screenshot gives neither, it doesn't even pin down "where" on a busy screen. Mark the element plus one line of intent and you've collapsed the local why and the where into one object. The project-level why is a real, separate layer on top. Where do you keep your project context now, a rules file or in each prompt?

 Alexander do we know each other? Have you worked in Switzerland before?

 :), no, never worked in Switzerland, must be a doppelganger out there. Where are you based? Always happy to compare notes with someone thinking about this at the right level.

 Geneva but relocating now to Spain... we shall compare where we are with both apps... to answer your question : I don't think it belongs in either. Rules files are useful for relatively static guidance, and prompts are ephemeral. Neither is a good place for information that changes every day.

I've been experimenting with keeping project context as a structured operational layer alongside the repository: active decisions, implementation history, ownership, impact relationships, verification evidence, and current constraints. The agent retrieves only what's relevant to the task instead of carrying a giant prompt or rules file around.

So my answer would be: neither a rules file nor each prompt. It lives as project intelligence that's continuously updated as the project evolves.

When prompting Cursor and providing an image for a feature change, I tend to explain the feature, where it's located, the function it performs (or should), the change, and the expected outcome. I also ask the agent to confirm it understands the given task and to ask probing questions if the instructions aren't clear.

 Solid routine, especially asking it to confirm and ask back. The part that jumps out is "where it's located." That's the piece you're spelling out in words because the screenshot doesn't carry it. On a busy screen that's also where it slips, the agent reads your description but still has to map it onto the right element in the image. Does the confirm-and-ask-back step actually catch the wrong-element cases for you, or does it still pick the neighbor sometimes?

 .... I have a high rate of the agent picking the right item(s). The thing that gets me... how the agent will loop (thinking in a loop) or some of the sub-agents spun from the main agent will spiral in different directions. It's inside the little swarm of sub-agent workers that I really have to buckle down and keep them on task. Personally, I'm not crazy about the swarm effect (not sure what others call it).

 "swarm effect" is a good name for it. Agreed that's a different beast from the wrong-element thing, it's orchestration drift, sub-agents wandering off. One small overlap: when each sub-agent re-reads the screenshot, they can each interpret it a little differently, so a shared structured reference at least removes that one axis of divergence. The looping itself is on the harness, not the input though. How do you keep them on task, tighter scoping per sub-agent?

 .... I've had to hit the stop button and tell the main agent to stop spinning up agents, and that it needs permission to add subagent helpers. But even then, it will do that from time-to-time. I really dislike when it has 4-6 spinning off at once.

What I think helps me, I always ask the main agent to not act on anything (make changes) until I've approved the plan. In my Cursor Rules folder, I have a ReadMe document (MD) that tells the agent it can't make changes until we've created a plan, and that I must approve said plan.


How do you handle the swarm effect?

crop hard plus a red arrow or rectangle on top. crop alone is not enough on a busy screen because the agent still has to guess which element matters.

other small thing that has cut my rework. paste two screenshots in the same message: the current state and a quick sketch of the target. the diff between them does more work than a paragraph of description.

side note: a structured 'what i actually selected' signal solves this category of problem. same shape as how recruiters mis-read resumes. the fix in both is to make the implicit explicit.

 "Make the implicit explicit" is the whole thing in four words. Crop-plus-arrow matches what I keep landing on too, crop narrows it, the arrow says which one, because crop alone still leaves the guess. And the two-screenshots trick is smart, current vs target sketch as a diff instead of a paragraph. That's intent shown instead of described. Does the target sketch need to be close, or does a rough box-and-label do the job?

  rough box-and-label works about 70 percent of the time. label has to be specific. not 'header' but 'profile dropdown at the top right of the header.' that level of specificity is where the box wins over a polished sketch.

when box-and-label is not enough: when layout is part of the answer (spacing, hierarchy, alignment between unrelated elements). then a closer-to-final sketch beats a box because the model uses pixel relationships as a proxy for design intent.

shortcut: if you find yourself describing what should NOT change, that is when to use the sketch. if you only need to point at the one thing to change, box and label.

 Sharpest breakdown in the thread. The "if you're describing what should NOT change, use the sketch" rule is the part I'm stealing. It maps to a clean line: pointing at one element to change is a box-and-label job, structured and exact. But when the intent is the relationships, spacing, hierarchy, alignment across unrelated elements, the pixels carry meaning a single reference can't, so the sketch wins. Different kind of intent, different tool. Appreciate you spelling it out.

 

yes. and the rephrase you just did is what the documentation should say. the diff between 'point at one element' and 'show intent across relationships' is also why box-and-label maps badly to figma redesign work but cleanly to copy or color or icon swaps.

one more layer worth naming: the box-and-label rule breaks when the change requires a counterfactual. 'this button should not move when the modal opens.' there is nothing to point at. the absence is the spec. that one wants two screenshots or a recording.

your tool sits at a useful seam.

I’ve noticed the same thing when working on larger SwiftUI screens. I usually crop screenshots aggressively or describe the exact UI element I want modified. That significantly reduces incorrect edits and saves a lot of back-and-forth iterations.

 SwiftUI is where I feel it worst, no DOM, no Figma, so the agent has only the pixels to go on. Cropping helps but it's a tax you pay every single time, and aggressive crops sometimes cut the surrounding context the agent actually needs to place the change. Describing the element is better, but that's the part I always get lazy about mid-flow.

What stuck for me was capturing the element and the intent once as structured text, instead of re-cropping or re-describing each round. Same thing you're doing by hand, minus the by-hand. I ended up building a small tool for it (SlimSnap) since I was doing it constantly on Mac UIs.

 That's an interesting approach.

There is no cropping.

It just zooms into the map perfectly.

I downladed the entire sold property db, and migrated it into a Progess db, then addded some indexes etc to speed the queries up

 Ha, sounds like a different build. When you're working on that UI in Cursor though, do you hit the wrong-element guessing, or does the map case dodge it?

I haven't encountered these issues with Claude Code. I recommend using Claude Code instead of Cursor.