Run several coding agents on one repo, and their pull requests collide
Running several coding agents in parallel feels like free throughput. Each one takes a task, opens a pull request, and moves on. The cost shows up later, at merge time.
A study of agent pull requests on GitHub measured it directly (arXiv). The authors took pairs of pull requests that were open on the same repository at the same time, 747 pairs, one per repository, and replayed the actual three-way git merges. When the two pull requests came from two different agents, 41.7% hit a textual merge conflict. When both came from the same agent, it was 19.8%.
Two things stand out. Conflicts are common even when one agent opens both pull requests, so this is not only a coordination problem between tools. And mixing agents roughly doubles the rate. The paper does not say why; our guess is that each agent has its own idea of where code belongs.
Why this matters more than it looks
A merge conflict is not just a few minutes of fixing. Someone has to understand both changes to resolve it correctly, and with agent-written code that someone often did not write either side. A conflict resolved in a hurry is a common way for a correct change to become an incorrect one.
What reduces it
- Give each agent its own area. Split work by directory, module or service, not by ticket. In the study, 84% of conflicted files were source code, not dependency files.
- Keep each pull request small and short-lived. The longer two branches stay open together, the more they drift.
- Merge often. Rebase or merge the main branch into each agent's branch before it opens a pull request, and again before it merges.
- Use one agent per task type where you can. The study's same-agent rate was about half the cross-agent rate.
- Isolate the working copies. Separate git worktrees keep agents from overwriting each other's files on disk, though they do not stop conflicts at merge time.
What the study does not tell you
It measures textual conflicts on replayed merges in GitHub repositories (the AIDev-pop dataset), and cross-agent pairs were rare: 122 of 2,807 repositories had one. It does not tell you how often a clean merge still produced wrong behaviour, and your own rate will depend on how your code is laid out. Measure it on your repo before you scale up the number of agents.
The takeaway
More agents in parallel only adds throughput if their work does not overlap. Plan the split before you add the second agent.
Found this useful?
Related posts
A 16.9 MB model can do the job. Here is how to tell when.
Last week (October 2) Cactus released Whistle, an open speech recognition model in a single 16.9 MB file ([post](https://cactuscompute.com/blog/whistle)). It runs on a CPU with no dependencies, transcribes seven language…
Your agent graded its own homework. Use a second model as the reviewer.
A coding agent finishes a change and reports that it is done and the tests pass. Often both are true, and the change is still wrong: the tests check what the agent thought the code should do, not what you needed.
A small model checking our docs was a coin flip. Giving it the docs fixed most of that.
A Hacker News thread this week argued that agents don't need memory, they need documentation ([discussion](https://news.ycombinator.com/item?id=49945933)). We had just measured a small version of that question, so here i…
We scanned our own agents' Claude Code transcripts for secrets. What we found, and what we changed.
We scanned our coding agents' Claude Code transcripts for secrets, then changed how prompts reach the model. What we found, measured, and what still fails.