How I use Claude Code and Cursor on real projects: guardrails, and what breaks
By Lal Chand, founder of Codic Systems
Published Last updated 5 min read
I use Claude Code and Cursor most days. I also don't trust either of them.
Those two things sit together fine. A tool can save me hours and still need supervision. This is how I use them on client work, the guardrails I rely on, and where they fall over.
What I use them for
They are good at the work that is tedious and low in risk. Writing a first draft of a form. Explaining a function in an old codebase I've just inherited. Renaming something across forty files. Turning a repeated pattern into a helper. Drafting tests for edge cases I've already thought of.
They are weaker at decisions. Where data lives, how permissions work, what happens when a payment fails halfway. Those decisions come from understanding the business, and the tools don't have that unless I give it to them. So I decide, then ask the tool to help me build what I decided.
The guardrails
I read every change. Not skim. Read. The tool can touch many files quickly, which makes it tempting to accept a large diff on trust. I don't. If a change is too big to review properly, I ask for smaller ones.
I give it something to check against. Anthropic's own guidance for Claude Code puts this first: give it a check it can run, such as tests, a build, a linter or a screenshot. Without one, "looks done" is the only signal it has, and I become the only checker. With one, it can run the test, read the failure and try again.
I plan before I build. For anything touching several files or code I don't know, I ask for a plan first and edit the plan before any code is written. Claude Code has a plan mode for exactly this. For a one-line fix, I skip it. The overhead isn't worth it.
I keep standing instructions short. Claude Code reads a file called CLAUDE.md at the start of every session, and Cursor reads project rules from .mdc files in .cursor/rules. Both are useful for things the tool cannot guess: how to run the tests, naming conventions, folders it must not touch. Both fail the same way, which is by getting too long. Anthropic's guidance says a bloated file causes the tool to ignore your actual instructions. I prune mine like code.
I use hooks for rules that must hold every time. An instruction in a file is advice. A hook is a script that runs deterministically. If something must always happen, such as running the linter after an edit, a hook is the right place for it.
I keep secrets out. No keys, no passwords, no customer records in prompts or in files the tool reads. I use environment variables and made-up sample data. This isn't clever. It's the cheapest protection available.
I keep it on a short lead on production. The tool can run commands. I limit which ones it may run without asking, and I don't hand it credentials for live systems.
What breaks
Plausible code that misses the edge case. This is the most common failure. The code runs, the happy path works and the demo looks fine. Then someone submits an empty field, a very long name or a date in another format. Anthropic's guide describes the same thing as a gap between trusting the output and verifying it. The fix is to name the awkward cases in the prompt and test them.
Tests that agree with the bug. If the tool writes the code and then writes the tests, the tests may simply describe whatever the code does, bugs included. I write the important test cases myself, or at least decide the expected results before the tool sees the code.
Invented details. Function names that don't exist, options a library doesn't have, a config key from a different version. They look right, and they fail at runtime. I check anything unfamiliar against the real documentation.
Context that fills up. Long sessions pile up files, output and half-abandoned ideas. The documentation says performance drops as the context window fills, and I see it: the tool forgets a rule it followed an hour ago. If I have corrected the same mistake twice, I stop, clear the session and write a better starting prompt.
Large, confident refactors. Ask for a small change and you may get a tidy rewrite of three files you didn't mention. I tell it the scope, and I check the diff for files I didn't expect to see.
Fixing the symptom. Faced with a failing build, a tool can silence the error instead of fixing the cause. I ask for the root cause, and I look at what changed.
Two tools, one rule
I use Claude Code when I want to work in the terminal across a whole project, with commands and git in the loop. I use Cursor when I'm editing inside the editor and want suggestions at the cursor. They overlap. The choice is mostly about where I'm already working.
The rule is the same for both: the tool proposes, I decide.
A note on cost of mistakes
How tightly I supervise depends on what is at stake. A throwaway script for my own use gets a quick look. Code that handles money, personal data or access gets slow, line by line review, and I am happy to say no to a suggestion that I cannot fully explain. The tools make it cheap to produce code, which makes the reviewing, not the typing, the scarce part of the job.
Why I wrote this down
Clients ask whether I use AI tools. The honest answer is yes, with rules. I'd rather explain those rules up front than have anyone discover later that code was written without a human reading it.
If the answer to "who reviewed this?" is nobody, the work isn't finished.
Questions this article answers
Do you let AI write code for client projects?
AI tools draft and edit code, but I review every change before it ships and I run the tests myself. I treat the output as a colleague's first draft, not a finished piece of work.
What are CLAUDE.md and Cursor rules?
Both are files of standing instructions. Claude Code reads CLAUDE.md at the start of every session. Cursor reads project rules from .mdc files in .cursor/rules, or from an AGENTS.md file.
What breaks most often when using AI coding tools?
Code that looks right and misses an edge case, tests that confirm the bug instead of catching it, invented function names, and long sessions where the tool forgets earlier instructions.
Should you put client data or secrets into an AI tool?
No. Keep secrets, credentials and personal data out of prompts and out of the files the tool can read. Use environment variables and fake sample data.
How do you check AI-generated code is correct?
Give the tool something to verify against, like tests, a build or a linter, then read the diff yourself. If you cannot verify a change, do not ship it.