Every merge queue product sells you batching — group ten pull requests, run CI once, merge all ten — and under agent load that pitch quietly inverts, because the same agents that doubled your pull request volume also raised the per-PR failure rate that batching raises to the tenth power. At a 20% queue-failure rate, a batch of ten passes 10.7% of the time and costs you almost a full CI run per merged change, which is what you were paying before you bought anything. Two capabilities actually change the arithmetic: bisecting a failed batch instead of discarding it, and running independent lanes for changes that cannot affect each other. Shortlist on those two and the rest of the comparison collapses into configuration.
At a glance
Four products, ordered by how much machinery they put between "approved" and "on main". The first is free and in the box; the other three exist because the first one stops at a specific place.
| Product | Failed-batch handling | Lane derivation | Shape |
|---|---|---|---|
| GitHub merge queue | Remove the failing PR from the queue; no bisection | None — one queue per protected branch | Native, free, zero deployment |
| Trunk Merge Queue | Move the batch to a bisection queue; split, retest, reuse prior passing results | Dynamic lanes from impacted build-graph targets (Bazel, Nx) | SaaS, CI-agnostic |
| Mergify | Parallel/speculative checks on temporary draft PRs; batch size 1–128 or dynamic {min, max} | Scope-aware batching — group changes touching the same declared areas | SaaS, queue plus merge governance |
| Aviator MergeQueue | Optimistic validation — wait out a suspected flake by checking later batches rather than resetting | Dynamic queues from affected targets | SaaS, with a self-hosted option |
A note on sourcing before we go further: much of the public comparison material on this category is written by the vendors about each other, and it is partisan in both directions. Everything stated here as fact comes from each product's own documentation of its own behaviour. Where a claim is about a competitor, we have left it out.
The arithmetic that decides the shortlist
A merge queue serializes: take the next pull requests, test them against the current tip plus everything ahead of them, merge what passes. One CI run per merged change is the baseline cost, and batching is the universal answer to that cost.
If each pull request independently fails in the queue with probability p, a batch of b passes with probability (1−p)b. Without bisection, one failure discards the whole batch:
| Per-PR queue-failure rate | Batch of 4 passes | Batch of 10 passes | CI runs per merge, b=10 |
|---|---|---|---|
| 5% — well-tested human PRs | 81.5% | 59.9% | 0.17 |
| 10% | 65.6% | 34.9% | 0.29 |
| 20% — typical agent-authored, unrebased | 41.0% | 10.7% | 0.93 |
The bottom-right cell is the whole argument. At an agent failure rate, a batch of ten costs nearly a full CI run per merged pull request — you have bought the batching machinery, added a queue-reset failure mode, and landed back at the unbatched baseline.
And p is not inherited. Agent-authored changes fail in the queue more often for structural reasons: the agent verified against the base commit it checked out rather than the tip that exists when the PR reaches the front, it edits files it has no local context for, and it opens PRs faster than the tree settles, so cross-agent conflicts are ordinary. Volume and p rise together — which is why the intuitive response to volume, bigger batches, is backwards.
Bisection is what rescues high-p batching. Split the failed batch, test the halves, recurse: isolating one culprit in a batch of b costs roughly 2·log₂(b) extra runs instead of re-running everything, and the innocent pull requests do not return to the end of the line. With result caching on top, halves that already passed inside a larger group need not run again.
GitHub merge queue — the right answer below a computable threshold
What it does
It forms merge groups: your pull request is grouped with the tip of the target branch and everything ahead of it in the queue, and required status checks run against that group. A build-concurrency setting caps how many merge_group webhooks dispatch at once, between 1 and 100, which is how you throttle concurrent CI. A separate setting governs whether pull requests with failing required checks may be grouped behind a passing one, or whether every PR in a group must be green.
Where it stops
On failure — a required check red against the merge group, a conflict with the base, a timeout against the configured wait, or a branch-protection failure it cannot resolve — the pull request is removed from the queue. There is no bisection of a failed group and no derivation of independent lanes from a build graph. One protected branch, one line.
When that is fine
When p is low and CI is fast, the machinery the others sell has nothing to do. Compute your own threshold rather than guessing: if your expected CI runs per merged PR at your measured p and your tolerable batch size is already under about 0.5, and your queue wait is dominated by CI duration rather than by depth, the native queue is not the constraint and replacing it buys you a bill. This is the common case for a single repository with a fifteen-minute pipeline, even at meaningful agent volume.
Trunk Merge Queue — bisection as the headline
Failure isolation
Trunk batches compatible pull requests and, when a batch fails, moves it into a separate queue for bisection: the batch is split in various ways and retested in isolation until the culprits are identified, with prior passing results reused so the isolation pass does not re-run what already went green. The healthy pull requests in the batch keep moving.
Lanes
Trunk infers parallel lanes dynamically from the targets a change actually impacts, through build-system integration with Bazel and Nx. Changes with no overlap test concurrently and merge independently; overlapping ones are grouped and tested together. The prerequisite is a build system that can answer "what does this diff affect" — without one, lane derivation degrades to a hand-maintained path map.
Who it fits
A monorepo where a large fraction of pull requests touch disjoint areas, and where p is high enough that discarding whole batches is the dominant cost. That is the shape agent-heavy repositories converge on.
Mergify — the knobs are the product
Batching and speculation
Mergify exposes batch size directly: a fixed integer from 1 to 128, or an object {min, max} for dynamic batching that uses min when parallel-check slots are plentiful and grows toward max to drain a backlog. Parallel (speculative) checks work by creating temporary draft pull requests that combine queued changes with the base branch and running CI on each concurrently, with max_parallel_checks capping how much concurrency you ask your CI to absorb.
Ordering and scoping
Each batch is seeded with the pull request next to merge — highest priority, oldest among equals — and batching candidates are ranked by shared scopes, so changes touching the same declared areas group together. Priority rules support interrupting running checks for a higher-priority change. Queue freezing is handled through a Pause API with a flag governing whether checks continue to run while paused.
Who it fits
Teams who want to tune this rather than accept a policy, and teams who want the queue and the merge-governance rules in one configuration file. The dynamic {min, max} batch size is the setting the arithmetic above is about — but note it tunes on backlog depth, not on p, so set the ceiling from your measured failure rate yourself.
Aviator — flakes and monorepo scale
Affected targets
Aviator builds around what a change touches: dynamic queues derived from affected targets, so non-overlapping pull requests run independently, while overlapping ones are optimistically stacked and tested together as in parallel mode.
Optimistic validation
This is the distinctive piece. When a test fails inside a batch in parallel mode, rather than resetting the queue immediately, Aviator waits to see whether later batches pass — use_optimistic_validation, with optimistic_validation_failure_depth controlling how far it looks. The effect is to stop a flaky required check from repeatedly discarding healthy work.
Deployment
CI-agnostic via status checks, with first-class integration for GitHub Actions, Buildkite, CircleCI, Jenkins, GitLab CI and Argo, and a self-hosted option — the one in this group, and the deciding factor if your source of truth cannot talk to a third-party SaaS.
The caveat that comes with it
Tolerating flakes and tolerating real failures are the same behaviour seen from two angles. Optimistic validation buys throughput against a genuinely flaky check and, configured aggressively, delays your discovery of a genuinely broken one. Treat it as a stopgap attached to a tracked quarantine list with an owner, not as a substitute for fixing the check.
When to pick which
| Situation | Pick | Because |
|---|---|---|
| One repo, CI under ~15 min, measured p under 10% | GitHub merge queue | Free, native, and the extra machinery has nothing to do at your failure rate. |
| Monorepo with a real build graph, high agent share | Trunk or Aviator | Lane derivation from impacted targets is the only thing that removes false serialization. |
| High p, no build graph, you want to tune | Mergify | Explicit batch sizing and speculative checks without requiring Bazel or Nx. |
| A required check you cannot de-flake this quarter | Aviator | Optimistic validation is built for exactly this, with an explicit failure-depth setting. |
| Source of truth cannot reach third-party SaaS | Aviator | The self-hosted option in this group. |
Whichever you pick, the highest-leverage change is not in the product. It is re-verifying each pull request against the queue head at enqueue time, so a correct patch tested against stale state fails cheaply on a branch rather than expensively in a serialized queue. That is one pipeline change, it usually produces the largest single drop in p, and it makes every product on this list cheaper to run.
FAQ
Is review really not the bottleneck?
Review is slower — median review time is up roughly 5× in Faros AI's 2026 study of about 22,000 developers, and agentic pull requests wait around 5.3× longer for pickup per LinearB's 2026 benchmarks. But 31% more pull requests now merge with no review at all, and the incidents-to-PR ratio more than tripled. A bottleneck holds work back; this system routed around its own gate. The automated gate is the one still applying uniformly.
How do I measure my p without deploying a queue?
Take fifty pull requests merged last month, check each out at the tip of main as it was at their merge moment, and run the affected tests. The failure fraction is your p. Split it by author type — human versus agent — because the two numbers are usually far apart and only the blend matters for batch sizing.
Do I need Bazel or Nx to get lanes?
For derived lanes, effectively yes — the products infer targets from a build graph. Without one, a hand-maintained path-to-lane map is a worse approximation and still far better than one global line. Keep a small set of integration checks that run in every lane, because a wrong lane boundary merges two changes that do interact.
Can agents merge through the queue themselves?
They should, and they should have no other route. Remove the admin-merge permission from agent tokens rather than writing a policy about it — one direct push to main invalidates every speculative run in flight.
What stops an agent from deleting a test to get through?
Nothing in the queue. A diff-level rule: an agent fixing a queue ejection may not modify test files unless the ejection was itself a test failure it is legitimately updating, and any deleted assertion triggers human review regardless.
Further reading
On this wiki:
- Merge Queues for Agent-Authored Changes — the build order, the enqueue-time rebase, and the four metrics to track.
- CI Repair Agents — where a queue ejection should be routed.
- Patch Generation & Test-Driven Loops — lowering p at the source.
- Background Coding Agents — where the volume comes from.