Vybes is now valued at $10M.Read more

How to build

Should AI Coding Agents Touch Production? Rules for Founders

AI coding agents can now run shell commands, apply migrations and open pull requests on their own. Here are the rules that keep them away from your customers' data until they have earned the access.

Nate Laquis10 min read

The Agent Has Your Credentials Now

A year ago, an AI coding tool suggested code and you pasted it in. Today the same tools run commands. They read your repo, install packages, execute test suites, apply database migrations and open pull requests while you are in a meeting. That is a real productivity jump, and it changes the risk in a way most founders have not priced in.

When an agent runs in your terminal, it inherits whatever your terminal can reach. If your shell has a production database URL in an environment variable, the agent can use it. If your cloud CLI is logged in as an admin, the agent is an admin. It does not need to be malicious for this to hurt you. It only needs to misread a task, decide a failing migration is easier to fix by dropping a table, and have the permission to do so.

Founders who have shipped with vibe coding tools tend to learn this the hard way, usually once. A deleted staging table is an annoying afternoon. A deleted production table with no recent backup is a company-threatening week. The rules below exist so you only ever have the first kind of story.

None of this is an argument against agents. We use them every day and our delivery speed depends on them. The argument is that access should be earned in steps, the same way you would treat a new contractor on day one. If you want the broader picture of getting AI-built code ready for real users, start with our guide to taking vibe-coded software to production quality, then come back here for the access rules.

Developer terminal with code open on a laptop screen

Rule One: Production Credentials Never Live Where the Agent Runs

This is the rule that prevents most disasters, and it is the one most small teams break. Look at your laptop or your dev container right now. Search your shell profile, your .env files and your cloud CLI config for anything pointing at production.

If you find a production connection string, the agent can find it too. Agents search the repo and the environment to understand context, and they will happily read a variable called DATABASE_URL without knowing which database it points to. The fix is boring and it works:

  • Keep development and staging credentials in the local environment. Nothing else.
  • Store production secrets in your host's secret manager (Vercel, Fly, AWS Secrets Manager, Doppler, 1Password) and inject them only at deploy time, from the CI system, never from a developer machine.
  • Log your cloud CLI into a role that can read staging and nothing else. Production access goes through a separate login that you use deliberately, in a different terminal, with no agent running.
  • If a human ever needs to run a production query, use a read-only replica credential first.

The test is simple. Open a fresh agent session on your machine and ask it to list every credential it can find. If any of them touch production, fix that before you do anything else this week.

Teams that skip this step usually say the same thing afterward: "We knew we should, but it was faster to keep the prod URL handy." It was faster right up until it was not.

Rule Two: Give the Agent Its Own Database, Not a Copy of Yours

Most agent horror stories involve a database. The agent was asked to fix a migration, the migration failed in a confusing way, and the agent chose a destructive path to get the tests green. Your job is to make sure the worst thing it can destroy is something you can recreate in two minutes.

The cheapest way to do that is a branchable database. Neon, PlanetScale and Supabase all let you create a throwaway copy of your schema, and with some seed data, in seconds. Point the agent at that branch for every task and delete the branch when the pull request merges. On a typical startup stack, this costs a few dollars a month and removes an entire category of risk.

A few details matter more than they look:

  • Seed data should be fake. Do not clone production into a dev database for convenience. Real customer emails in a dev database leak into logs, screenshots and agent transcripts that get sent to a model provider. Generate realistic fake records with a seed script and keep that script in the repo.
  • Migrations run through your normal tool. Prisma, Drizzle, Rails or Alembic should own schema changes. If the agent edits tables with raw SQL in a console, there is no record and no way to replay it on another environment.
  • Reviewing a migration is a human job. Read every generated migration before it merges. Look for DROP, TRUNCATE, column type changes and anything that rewrites a large table. A migration that locks a ten-million-row table for several minutes will take your app down even if it is technically correct.

If you are running something like prisma migrate reset or db push --force in the same environment as live data, stop. Those commands exist for disposable databases, and an agent will reach for them the moment a migration history drifts. Our deeper notes on safe schema changes are in the database migration strategies guide.

Rule Three: Least Privilege for Every Tool the Agent Can Call

Modern agents do not just run shell commands. They connect to tools through protocols like MCP, so they can query a database, read your issue tracker, post to Slack or call your payments API. Every connected tool is a permission, and most founders accept the default scope without reading it.

Go through each connection and ask what the narrowest useful version looks like:

  • A database tool should connect with a role that can SELECT on the tables it needs and nothing more. If the agent only has to read data to debug, it should not hold write access.
  • A GitHub token should be limited to the one repo, with permission to push branches but not to merge into main or change repo settings.
  • A Stripe key for agent work should be a test-mode key. Live-mode keys do not belong in a development session under any circumstances.
  • An email or messaging integration should point at a sandbox inbox. An agent that "tests the welcome email" against your real customer list is a story that has happened.

There is also a security angle beyond honest mistakes. Agents read untrusted text: issue descriptions, web pages, package READMEs, error messages from third-party APIs. Any of that text can contain instructions, and a model with powerful tools may follow them. This is prompt injection, and it is much worse when the agent holds broad credentials. We cover the attack patterns in detail in our piece on AI agent security, prompt injection and data leaks, and the short version is that limiting what the agent can do is a better defense than hoping it cannot be tricked.

Code on a monitor representing permissions and access control in a software project

Rule Four: Agents Open Pull Requests, Humans Merge Them

The simplest control in the whole system is also the most effective. The agent never pushes to main. It works on a branch, opens a pull request, and a person reads the diff before anything ships. Turn on branch protection in GitHub so this is enforced by the platform and not by good intentions.

Founders sometimes resist this because it feels slow, especially solo founders who are the only reviewer. It is not slow when the diff is small. The trick is to keep agent tasks small enough to review in five minutes. A prompt like "refactor the billing module" produces a 2,000-line diff that nobody reads carefully. A prompt like "add a retry to the webhook handler and a test for it" produces 40 lines you can actually understand.

Some practical review habits we use:

  1. Read the tests first. If the agent changed or deleted a test to make it pass, that is a red flag, and it happens more than you would expect.
  2. Scan for new dependencies. Agents add packages freely, and every package is code you now trust. Check that the name is spelled correctly, since look-alike package names are a known attack.
  3. Search the diff for hardcoded secrets, disabled auth checks and commented-out validation.
  4. Run the branch yourself on a preview deploy and click through the one flow that changed.

Run continuous integration on every pull request, with type checking, linting and your test suite. A green pipeline does not prove the change is right, but a red one saves you from reading code that does not even compile. If you do not have CI yet, setting it up is a one-day job and the most valuable thing you can do before leaning harder on agents.

Rule Five: Back Up Like the Agent Will Eventually Get It Wrong

Plan as though a destructive mistake will happen once, not as though it never will. That reframes backups from an ops chore into the thing that decides whether a mistake costs you an hour or a quarter.

Check these four items on your production database today:

  • Automated backups are on. Most managed providers include daily snapshots on paid plans. Some free tiers do not. Confirm which you have.
  • Point-in-time recovery is enabled. Daily snapshots can still cost you up to a day of orders. Point-in-time recovery lets you rewind to a minute before the mistake. On Postgres hosts this is typically an add-on costing tens of dollars a month, which is cheap against a lost day of revenue.
  • You have restored from a backup at least once. A backup you have never restored is a hope. Do a restore into a scratch database and confirm the data is there and the app boots against it. Put it on the calendar twice a year.
  • Backups live outside the account that could be compromised. If an attacker or a bad script can delete your database and your backups with the same credential, you have one copy, not two.

File storage deserves the same thinking. If customers upload documents to a bucket, turn on versioning so an overwritten or deleted object can be recovered. It costs very little and removes a painful class of mistake.

A Staged Path for Giving Agents More Access

Blanket bans waste the speed you are paying for, and blanket trust is how the stories above happen. A staged approach works better, and it mirrors how you would onboard a new engineer.

Stage one: sandbox only

The agent works in a local or containerized environment with a throwaway database and fake data. It can run anything it wants, including destructive commands, because nothing real is at stake. Most teams should spend their first month here, building a feel for how the tool fails.

Stage two: staging with review

The agent opens pull requests that deploy to a staging environment wired to fake data and test-mode integrations. Humans review and merge. Production remains untouched by anything that runs on an agent's machine.

Stage three: read-only production

For debugging, give the agent a read-only credential on a replica, with sensitive columns masked or excluded. This is valuable because many production bugs only show up with real data shapes. It stays read-only, and queries are logged.

Stage four: narrow, audited write access

Only after months of clean history, consider letting an agent perform a specific, bounded write action, such as running a documented data-fix script that a human has approved. The agent runs one named script. It does not get a general write credential. Every action is logged and attributable.

Plenty of teams never need stage four. We rarely do. The first three stages deliver most of the speed gain with a fraction of the exposure, and if you are weighing whether an agent-heavy workflow suits your product at all, our comparison of vibe coding and traditional development is a good companion read.

What to Do If You Already Gave Out Too Much

Most founders reading this have already broken at least one rule. That is normal, and the cleanup is a single afternoon, so here is the order we would work in.

  1. Rotate every production credential that has ever sat on a developer machine. Database passwords, API keys, cloud access keys, signing secrets. Assume they have been read by tools, written to transcripts or synced to cloud backups. Rotation costs an hour and closes the question for good.
  2. Separate the environments. Create a staging database and staging deploy if you do not have them. Move the agent's working credentials to those.
  3. Turn on branch protection. Require a pull request and one approval on main. Block force pushes.
  4. Verify backups and do a test restore. Do this before the next agent session, not after.
  5. Write a one-page access policy. Who and what can touch production, how, and with which credential. A page is enough, and it forces the decisions into the open.

If your codebase was mostly generated by AI tools and you are not sure it can be trusted with real customers, that is a different question than access control. Sometimes the right call is a cleanup pass, and sometimes it is a rebuild, and our decision framework for rebuilding a vibe-coded app walks through how to tell which.

Analytics dashboard on a laptop used to monitor a production application

What This Costs and Where to Start

The good news for a startup budget is that the controls above are cheap. A branchable database runs from free to roughly $20 to $70 a month depending on provider and size. A secret manager is free to $10 per user. Point-in-time recovery on managed Postgres commonly adds $20 to $100 a month. Continuous integration on GitHub Actions is free for small teams and a few dollars beyond that. Call it a few hundred dollars a month at the high end for a serious product, and often under $100.

The main cost is attention. Expect a solo technical founder to spend one to two days getting environments separated, credentials rotated and CI running. A non-technical founder working with a developer should expect that developer to need about the same, and should ask for it by name. If a freelancer or agency tells you they work directly against production because it is quicker, treat that as a warning about how they will handle everything else.

If you only do three things this week, do these: remove production credentials from every machine where an agent runs, require pull request review on main, and perform a test restore of your database. Those three remove most of the risk, and none of them slow your team down in a way you will notice after the first few days.

If you want a second set of eyes on your setup, or you are not sure whether your AI-built product is safe to put in front of paying customers, we can look at the repo, the environments and the access model with you. Book a free strategy call.

AI coding agents production accessAI agent guardrailsvibe coding safetystaging environmentsdatabase permissions

Software people love.

Thirty minutes. You leave with a plan and a number.

Book a call