Should you ban AI-generated code?

Researchers read 281 AI policies from the most popular open source projects. 83% allow AI in code, and 188 anti-slop measures came out of them. I went through the findings as a fractional CTO and asked which of those rules a startup should actually copy.

Should You Ban AI-Generated Code? A Fractional CTO Take on 281 Policies

Image: METAHEURISTIC

Alexander Myasoedov

+Alexander Myasoedov Alexander writes about the operational side of shipping production AI - agents, retrieval, evals, and the guardrails that keep them from going sideways.

Should you ban AI-generated code? For most teams, no. My bar is simpler: if you can’t tell whether a change was generated by an agent in one minute or written by hand over six hours, it passes, and how it was made stops mattering. A ban does make sense in a few places, like core systems and niche domains where models have little to learn from. I came to this reading the landscape of AI policies in popular open source projects by Hora, Robbes and Zacchiroli, then answering an interviewer who kept asking me, as a fractional CTO, whether a 20 or 30 person startup should adopt the same rules.

Should you ban AI-generated code? The short answer

Bar
Indistinguishable
If you can't tell a one-minute agent change from six hours of careful human work, it passes the quality bar.
Ban
Core and niche only
Compilers, Bitcoin-class core tech and domains the model has few examples of. Elsewhere, guide the agent.
Policy
Ownership and authorship
You can generate anything. You still own what you ship and stand behind every claim in it.

What did the research find?

The paper covers 281 AI contribution policies from the 2,000 most popular GitHub repositories plus 36 well-known projects and organizations. 83.3% of the policies permit or encourage AI in code contributions and 14.9% forbid it. 48.8% require contributors to disclose AI usage. Across the policies the authors counted 188 anti-slop countermeasures: closing pull requests is 48.4% of them, banning or blocking users 20.8%, and disallowing autonomous agents 13.8%. Rust Analyzer doesn’t allow autonomous agents to open pull requests or issues, Bitcoin says pull requests should not be opened or driven by autonomous agents, Apache Flink closes unrefined AI-looking PRs without review, and DuckDB forbids AI-generated comments when talking to maintainers.

Open source maintainers wrote these rules for strangers sending code over the internet. A startup has employees, a shared codebase and people who know each other, so I went through them one by one.

Figure - What open source projects enforce, and what I'd keep in a startup
View:
Press Play. The bars are the paper's numbers, and each one has a different base: disclosure is a share of all 281 policies, no AI in communication a share of the policies that address communication, and the last three a share of the 188 anti-slop countermeasures. Switch to Startup take and watch which rules survive the move into a private team.

Should you ban AI-generated contributions?

Whether we should or shouldn’t ban using AI is a controversial topic, and it’s buried in shame and blame of developers who feel like they’re cheating the system they use to generate the code. And we’ve seen a lot of slop grenades: easy, low-effort pieces of added complexity and abstraction put to the codebase and to your system, which clearly take a lot of effort to review. There’s an asymmetry between the effort to create and the effort to review, process and integrate. There’s also maintenance overhead. Even if the review passed, maintaining a low-quality generated solution or a bad decision can escalate and snowball into processes, procedures, instructions and on-call activities that would be a burden for the team. I wrote about how that builds up in what is a slop grenade.

So would banning generated code solve it? Potentially, yes, if you have high-performance engineers who have a lot of context and can do as much work as AI. In certain projects the complexity just doesn’t fit in the AI context window, so no agent can produce good code there. There are also domains like niche languages. Imagine someone working on the OCaml compiler. LLMs just don’t have enough examples of how you’d write a compiler like that to actually contribute to it. Yes, it can potentially provide some solution to the problem, but generally that would be a sloppish solution and not good enough.

But I also see examples where a pretty experienced engineer can use AI and produce 2x more outcomes for the team and the company, which is also great. That’s the same number I landed on in the 100x AI engineer is a myth.

How do you tell a good AI pull request from slop?

The most distinguishing factor is whether you can tell generated code from code written by a human being. Take two versions of a change. One is an artifact generated by AI in one minute. The other was crafted manually, by thoughtful work, reviewing and writing it over six hours. If you can’t tell the difference, that means the change set is high quality and good enough. That should be the de facto standard for any pull request.

Sometimes the signal is obvious. You see a bloated pull request with 10K lines of changes and you can say there’s no way a human being produced that many lines within the last 24 hours. That’s an AI indicator, and that’s not a good pull request to review. On the other side, a pull request with 20 lines of code where you can’t tell whether a human or an agent did it, with sufficient tests and a description that all makes sense, is a good solution.

The indistinguishability bar If a reviewer can’t tell an agent’s one-minute change from six hours of careful manual work, the change passes. A 10K-line PR from the last 24 hours fails before anyone reads it.

Should a startup require AI disclosure?

Picture a company that requires every developer to add a tag to the commit message saying it was co-authored by a certain model, from a certain provider, or that it was written by a human. Over time you’d collect that across multiple projects in your main branches, so the data stays in git history. What would it give us? On a long horizon, some level of meta-analytics: what we were using to write our code, and whether commits from a certain model were low or high quality. Is that useful? Sometimes it can be.

Would it be my first priority? I don’t think so. It’s a burden for the team and carries zero information. Going back to my bar: if you can’t distinguish an artifact made by a human from one made by an agent, it passes the quality bar, and you don’t care how it was generated or how many hours or tokens the developer burned to get there. Maybe it was a slow model that spent 10 hours of inference, maybe it was a person typing the solution with one finger. Nobody cares, as long as it works and passes the quality bar.

And if you have no disclosure policy at all, why not? I think the information is a bit excessive. It can give you some level of control, but it reminds me of an organization I used to collaborate with that ran over 200 developers and required every developer to fill in fields on the Jira ticket. How many hours did you spend? How many hours did you estimate? What was the root cause, across 16 different options? Do you think the scope was right at the beginning? It was a list of eight fields you had to fill manually, and it was a brutal overhead to the process. That burden existed so a manager could feel in control of the metrics and say at the end of the quarter that we improved one of them by 20%, so we were successful. I’d try to avoid that scenario.

Where’s the line on autonomous agents?

The interviewer’s scenario: an engineer delegates a task to Claude Code, the agent implements it, writes tests, runs them and opens a PR. The engineer reviews the diff, sees the tests passed and approves it, without fully understanding every detail.

I’d say full delegation should not be used in some projects at all. It should be banned for core projects, rare and core technology like Bitcoin or compilers. There, developers should guide the agent to write the code, not delegate the feature end to end, especially in domains where the model has very little knowledge. It should stay totally within the harness and guidance of the developer.

I can see use cases where you can use it, for things that have visual feedback, like frontend. My hypothesis is that the huge decline in demand for frontend engineers in the recent year happened because the harness for frontend work is a very low bar. You can generate the visual changes, the API calls, a typical frontend application. I’m not talking about complex applications with lots of animation or WebGL. For the typical one, reviewing the diff and testing the outcome is enough, and that’s not the case for a core project. Generally I’d be inclined to ban full delegation and keep human-guided generation.

Full delegation is fine

  • Visual feedback you can check
  • Typical frontend, API calls, UI changes
  • Review the diff, test the outcome

Guide the agent instead

  • Core and rare technology
  • Compilers, Bitcoin-class systems
  • Domains the model barely knows

Should reviewers close AI pull requests without review?

48.4% of the countermeasures in the paper are about closing pull requests, and Apache Flink lets maintainers close AI-looking contributions without review: excessive scaffolding, ineffective tests, bloated descriptions. The interviewer asked what I’d tell a CTO whose team spends half the week reviewing AI-assisted PRs.

In a private organization you can’t just close a pull request. But you can say we need to rewrite it for quality purposes. I agree that reviewers shouldn’t nitpick every single line an AI assistant generated. You can say this feels like a very sloppy solution, hard to review, and we need to redesign it so it’s easier to review and easier to implement. I think a CTO-level organization should derive metrics from that. Define a label for pull requests that need to be reworked due to slop.

What I’d definitely avoid is a private leaderboard of developers, where you report people for generating slop and they become candidates for a performance improvement plan because they often use AI irresponsibly, without awareness of other people’s time, of what it takes them to review and understand the problem, or of the project. I’d avoid that kind of surveillance and reporting. Keeping metrics on review iterations and rework is important. I covered the reviewer’s side of this in code review is the interview now.

Splitting a large change into small pull requests and stacking them with conditional merge also helps, and GitHub recently released that as stacked pull requests. But there’s a counterexample. A purposely complex feature with a sloppish top-level design can be split into chunks that each make sense and get merged, and altogether, at the integration layer, the solution is very low quality and a bad decision. Keep these scenarios in mind and try to avoid them too.

Rework label, not a leaderboard Label PRs that need a rewrite because of slop and track review iterations and rework. Don’t turn the label into a list of people to report.

Are repo skills another slop grenade?

Adding too many skills to the repo is another way to install a slop grenade, because those skills can be very controversial inside their instructions. Say you have a large corpus of guidance and suggestions on how you’re supposed to do things. There’s no way to guarantee that all of them will be applied, that all of them make sense, that they don’t contradict each other, and that they all fit a given feature.

I’d generally keep skills either local to developers or very limited and concise. You don’t want the whole 20 pages of description on how to optimize performance. State the facts instead: we don’t keep unnecessary copies of objects in memory, we don’t do unnecessary memory allocations, and so on.

How do you stop architectural drift?

The interviewer’s scenario: six months later, decisions nobody explicitly made have calcified into architecture. Each PR was small and the tests passed, but the big picture eroded.

That’s the example I already mentioned, where a complex feature gets split into tiny pull requests that calcify the architecture and ruin the overall design. That’s why I’d keep review quality as a PR metric. Record whether it was easy to review, hard to review, or difficult. It’s also fine to put a label that says I reviewed this and I can approve it, it seems okay, but I have zero understanding, and it’s hard for me to follow due to my lack of domain knowledge or expertise in this repo. You might be examining it from a security perspective, or something else, without being a subject matter expert.

For the systemic, structural issues that drifted through code review, I’d track them down from the history of pull requests. Link the pull requests and see what the chain of reviewers was and how exactly it happened. Was it a huge feature that was split into chunks, or a large chunk of changes committed at once and not properly reviewed? What was the root cause? This will happen from time to time, and that’s exactly why you sometimes need to run due diligence on your entire architecture, checking whether it’s aligned with the implementation. Sometimes it’s fine to rewrite the service from scratch, to simplify it and make it much cleaner in scope and closer to the initial intent.

Easy
Easy to review
The change was clear, and the approval means what it says.
Hard
Hard to review
It passed, but it was hard or difficult to follow.
Blind
Approved, zero understanding
Fine to say. You checked it from a security angle, without domain expertise in this repo.

Should AI write PR descriptions and review comments?

The paper found that 27.8% of the policies that address communication forbid AI in it. DuckDB, for example, requires human-written comments and discussions. The worry is an agent writing a beautiful long description of a change, and a reviewer spending time on it only to realize it doesn’t match the implementation.

I think review comments should definitely be prohibited from being written by AI, since that introduces an incorrect feedback loop. AI assistance for things like grammar, yes. For the pull request itself, the title is fine, but the description should not be written by AI. It’s fine to use AI to fix the structure and format of the words, but initially it should be written by a human. So AI is the editor, not the author of the description, and the human needs to stand behind every claim.

Then the interviewer pushed back on purpose. I said that if you can’t tell the difference, I don’t care how the code was produced, but I drew a hard line on descriptions. Imagine two engineers. One writes a description by hand that’s superficial and misses a key trade-off. The other has Claude draft it and checks every line. The second one is better in practice, so does my policy reward the worse behavior? Should the rule be about authorship, or about verified responsibility for the content?

I think it’s a combination of both. We should have verified authorship and responsibility for what you actually ship. Ownership plus accountability.

Editor, not author AI can fix grammar, structure and format. A human writes the description and the review comments, and stands behind every claim in them.

Do you need a written AI coding policy?

The last scenario: you join a 30-engineer startup as a fractional CTO. There’s no AI policy, most engineers already use Claude Code or Cursor, and the CEO asks you to write one.

I definitely wouldn’t introduce it as a Markdown file like AGENTS.md. This is an engineering culture policy, not a policy you hand to an agent. It’s nice to introduce it at an all-hands or as a concise one-page guidance doc that any human can understand. Then a reviewer can say this change doesn’t pass the guideline, and here are the criteria, out of the ten we have, for why it doesn’t pass the quality bar. That’s how I’d take it.

So, should you ban AI-generated code?

Mostly no. Ban full delegation on core systems and in domains the model doesn’t know, and let everything else meet one bar: the reviewer can’t tell how it was made, and the person who shipped it owns it. Ownership was there before the AI era. Now it’s more concentrated, because you can produce anything, but you still need to own the solution. That’s the bottom line of any AI policy: ownership and authorship. If you’re the CTO writing that one-pager for a team that’s already deep in Claude Code, that’s the kind of call a fractional CTO helps make, and momentum-driven code reviews covers keeping the review queue moving once the PRs pile up.

Follow in Google

Make Metaheuristic a preferred source.

One tap and posts like this one surface higher in your Top Stories.

Work with us

Production AI, with guardrails.

Start with a fixed-scope AI Workflow Audit. We map the opportunity and quote a build.

Start a Discovery Sprint →