First, write your tests, then allow the agent to emerge: has this approach worked for you?
Previously, I'd present a requirement to the agent called Cursor, and then accept anything it returned to me. Recently, I have been reversing that process a bit. First, I write up three test cases using natural language such as "User whose card expired should see retry message" and request that those pass.
It makes a difference in two ways: The generated code is evaluated against something I determined in advance, rather than simply judged by appearance. When the agent starts drifting into refactoring other files for no good reason, failing tests tell us very quickly.
On the minus side, there is a cost of about ten minutes' thinking time before each feature, which I think is reasonable.
I don't know if this is applicable to large-scale projects. If your code base has a lot more than a couple dozen files in it, is it really sustainable, or just another thing that the agent will keep breaking?
Replies
Be the first to reply
Have a question or a thought to share? Add a comment above to start the conversation.