AI Blog

GitHub Merge Queue vs Trunk vs Mergify vs Aviator

Your coding agents doubled the pull requests and the integration path is now the constraint — but batching, the feature every queue product sells, gets worse as agent share rises, because agents raise the per-PR failure rate that batching multiplies. Only two capabilities change the arithmetic: bisecting a failed batch, and deriving independent lanes from what a change actually touches. Shortlist on those; everything else is configuration.

By Agentic AI Wiki 13 min read

Every merge queue product sells you batching — group ten pull requests, run CI once, merge all ten — and under agent load that pitch quietly inverts, because the same agents that doubled your pull request volume also raised the per-PR failure rate that batching raises to the tenth power. At a 20% queue-failure rate, a batch of ten passes 10.7% of the time and costs you almost a full CI run per merged change, which is what you were paying before you bought anything. Two capabilities actually change the arithmetic: bisecting a failed batch instead of discarding it, and running independent lanes for changes that cannot affect each other. Shortlist on those two and the rest of the comparison collapses into configuration.

At a glance

Four products, ordered by how much machinery they put between "approved" and "on main". The first is free and in the box; the other three exist because the first one stops at a specific place.

ProductFailed-batch handlingLane derivationShape
GitHub merge queueRemove the failing PR from the queue; no bisectionNone — one queue per protected branchNative, free, zero deployment
Trunk Merge QueueMove the batch to a bisection queue; split, retest, reuse prior passing resultsDynamic lanes from impacted build-graph targets (Bazel, Nx)SaaS, CI-agnostic
MergifyParallel/speculative checks on temporary draft PRs; batch size 1–128 or dynamic {min, max}Scope-aware batching — group changes touching the same declared areasSaaS, queue plus merge governance
Aviator MergeQueueOptimistic validation — wait out a suspected flake by checking later batches rather than resettingDynamic queues from affected targetsSaaS, with a self-hosted option
Merge queue feature matrix Four products — GitHub merge queue, Trunk, Mergify and Aviator — scored weak, medium or strong on batch control, failed-batch isolation, lane derivation, flake tolerance and self-hosting. Merge queue capability matrix BATCH CONTROL FAILED-BATCH ISOLATION LANE DERIVATION FLAKE TOLERANCE SELF-HOSTED GitHub Merge groups only Eject the PR None None N/A — native Trunk Compatible batching Bisection + caching Impacted targets Via isolation SaaS Mergify 1–128 or {min,max} Parallel checks Declared scopes Pause / priority SaaS Aviator Parallel mode Optimistic retry Affected targets Failure depth knob Yes Strong Medium Weak or absent Columns two and three are the ones that change the asymptotics; the rest is configuration.
Where each one leans hardest. The second column is the one that changes the asymptotics.

A note on sourcing before we go further: much of the public comparison material on this category is written by the vendors about each other, and it is partisan in both directions. Everything stated here as fact comes from each product's own documentation of its own behaviour. Where a claim is about a competitor, we have left it out.

The arithmetic that decides the shortlist

A merge queue serializes: take the next pull requests, test them against the current tip plus everything ahead of them, merge what passes. One CI run per merged change is the baseline cost, and batching is the universal answer to that cost.

If each pull request independently fails in the queue with probability p, a batch of b passes with probability (1−p)b. Without bisection, one failure discards the whole batch:

Per-PR queue-failure rateBatch of 4 passesBatch of 10 passesCI runs per merge, b=10
5% — well-tested human PRs81.5%59.9%0.17
10%65.6%34.9%0.29
20% — typical agent-authored, unrebased41.0%10.7%0.93

The bottom-right cell is the whole argument. At an agent failure rate, a batch of ten costs nearly a full CI run per merged pull request — you have bought the batching machinery, added a queue-reset failure mode, and landed back at the unbatched baseline.

And p is not inherited. Agent-authored changes fail in the queue more often for structural reasons: the agent verified against the base commit it checked out rather than the tip that exists when the PR reaches the front, it edits files it has no local context for, and it opens PRs faster than the tree settles, so cross-agent conflicts are ordinary. Volume and p rise together — which is why the intuitive response to volume, bigger batches, is backwards.

Recovering from a failed batch of eight, with and without bisection A batch of eight pull requests fails because one is bad. Without bisection the whole batch is discarded and requeued, costing one wasted run and eight PRs of lost position. With bisection the batch is split in halves repeatedly until the culprit is isolated, costing about three extra runs while the seven healthy pull requests continue. One bad PR in a batch of eight 1 2 3 4 5 6 ✕ 7 8 one CI run, red Without bisection whole batch discarded — all eight return to the queue requeue, retry Cost: 1 wasted run, 7 innocent PRs lose their position, and the next batch inherits the same bad PR With bisection 1 – 4 retested: green 5 – 8 retested: red run 2 5 – 6: red 7 – 8: green run 3 5 ✓ 6 ✕ run 4 — culprit isolated, 1–5 and 7–8 merge Discard is linear in batch size; bisection is about 2·log₂(b) extra runs and with result caching, the halves that already went green inside a larger group need not run again
Same failure, two recoveries. Bisection turns a linear penalty into a logarithmic one and keeps the healthy PRs moving.

Bisection is what rescues high-p batching. Split the failed batch, test the halves, recurse: isolating one culprit in a batch of b costs roughly 2·log₂(b) extra runs instead of re-running everything, and the innocent pull requests do not return to the end of the line. With result caching on top, halves that already passed inside a larger group need not run again.

GitHub merge queue — the right answer below a computable threshold

What it does

It forms merge groups: your pull request is grouped with the tip of the target branch and everything ahead of it in the queue, and required status checks run against that group. A build-concurrency setting caps how many merge_group webhooks dispatch at once, between 1 and 100, which is how you throttle concurrent CI. A separate setting governs whether pull requests with failing required checks may be grouped behind a passing one, or whether every PR in a group must be green.

Where it stops

On failure — a required check red against the merge group, a conflict with the base, a timeout against the configured wait, or a branch-protection failure it cannot resolve — the pull request is removed from the queue. There is no bisection of a failed group and no derivation of independent lanes from a build graph. One protected branch, one line.

When that is fine

When p is low and CI is fast, the machinery the others sell has nothing to do. Compute your own threshold rather than guessing: if your expected CI runs per merged PR at your measured p and your tolerable batch size is already under about 0.5, and your queue wait is dominated by CI duration rather than by depth, the native queue is not the constraint and replacing it buys you a bill. This is the common case for a single repository with a fifteen-minute pipeline, even at meaningful agent volume.

Trunk Merge Queue — bisection as the headline

Failure isolation

Trunk batches compatible pull requests and, when a batch fails, moves it into a separate queue for bisection: the batch is split in various ways and retested in isolation until the culprits are identified, with prior passing results reused so the isolation pass does not re-run what already went green. The healthy pull requests in the batch keep moving.

Lanes

Trunk infers parallel lanes dynamically from the targets a change actually impacts, through build-system integration with Bazel and Nx. Changes with no overlap test concurrently and merge independently; overlapping ones are grouped and tested together. The prerequisite is a build system that can answer "what does this diff affect" — without one, lane derivation degrades to a hand-maintained path map.

Who it fits

A monorepo where a large fraction of pull requests touch disjoint areas, and where p is high enough that discarding whole batches is the dominant cost. That is the shape agent-heavy repositories converge on.

Mergify — the knobs are the product

Batching and speculation

Mergify exposes batch size directly: a fixed integer from 1 to 128, or an object {min, max} for dynamic batching that uses min when parallel-check slots are plentiful and grows toward max to drain a backlog. Parallel (speculative) checks work by creating temporary draft pull requests that combine queued changes with the base branch and running CI on each concurrently, with max_parallel_checks capping how much concurrency you ask your CI to absorb.

Ordering and scoping

Each batch is seeded with the pull request next to merge — highest priority, oldest among equals — and batching candidates are ranked by shared scopes, so changes touching the same declared areas group together. Priority rules support interrupting running checks for a higher-priority change. Queue freezing is handled through a Pause API with a flag governing whether checks continue to run while paused.

Who it fits

Teams who want to tune this rather than accept a policy, and teams who want the queue and the merge-governance rules in one configuration file. The dynamic {min, max} batch size is the setting the arithmetic above is about — but note it tunes on backlog depth, not on p, so set the ceiling from your measured failure rate yourself.

Aviator — flakes and monorepo scale

Affected targets

Aviator builds around what a change touches: dynamic queues derived from affected targets, so non-overlapping pull requests run independently, while overlapping ones are optimistically stacked and tested together as in parallel mode.

Optimistic validation

This is the distinctive piece. When a test fails inside a batch in parallel mode, rather than resetting the queue immediately, Aviator waits to see whether later batches pass — use_optimistic_validation, with optimistic_validation_failure_depth controlling how far it looks. The effect is to stop a flaky required check from repeatedly discarding healthy work.

Deployment

CI-agnostic via status checks, with first-class integration for GitHub Actions, Buildkite, CircleCI, Jenkins, GitLab CI and Argo, and a self-hosted option — the one in this group, and the deciding factor if your source of truth cannot talk to a third-party SaaS.

The caveat that comes with it

Tolerating flakes and tolerating real failures are the same behaviour seen from two angles. Optimistic validation buys throughput against a genuinely flaky check and, configured aggressively, delays your discovery of a genuinely broken one. Treat it as a stopgap attached to a tracked quarantine list with an owner, not as a substitute for fixing the check.

When to pick which

SituationPickBecause
One repo, CI under ~15 min, measured p under 10%GitHub merge queueFree, native, and the extra machinery has nothing to do at your failure rate.
Monorepo with a real build graph, high agent shareTrunk or AviatorLane derivation from impacted targets is the only thing that removes false serialization.
High p, no build graph, you want to tuneMergifyExplicit batch sizing and speculative checks without requiring Bazel or Nx.
A required check you cannot de-flake this quarterAviatorOptimistic validation is built for exactly this, with an explicit failure-depth setting.
Source of truth cannot reach third-party SaaSAviatorThe self-hosted option in this group.
The three measurements that decide the choice Three columns — per-PR queue-failure rate, CI wall-clock, and conflict domain — each with what it determines about the merge queue configuration and how to measure it from data a team already has. Measure three things; the product choice follows Queue-failure rate p CI wall-clock Conflict domain Decides: batch size, and whether bisection is mandatory Decides: whether queue depth or run time dominates the wait Decides: whether lanes buy you anything Measure: retest 50 merged PRs at the tip they merged onto Measure: p50 required-check duration on the default branch Measure: share of PR pairs whose touched paths do not overlap Low p and fast CI: the native queue is not your constraint High p: buy bisection. Wide conflict domain and a build graph: buy lanes. Both: you are in the monorepo case.
Three measurements decide this, and you can take all three from data you already have.

Whichever you pick, the highest-leverage change is not in the product. It is re-verifying each pull request against the queue head at enqueue time, so a correct patch tested against stale state fails cheaply on a branch rather than expensively in a serialized queue. That is one pipeline change, it usually produces the largest single drop in p, and it makes every product on this list cheaper to run.

FAQ

Is review really not the bottleneck?

Review is slower — median review time is up roughly 5× in Faros AI's 2026 study of about 22,000 developers, and agentic pull requests wait around 5.3× longer for pickup per LinearB's 2026 benchmarks. But 31% more pull requests now merge with no review at all, and the incidents-to-PR ratio more than tripled. A bottleneck holds work back; this system routed around its own gate. The automated gate is the one still applying uniformly.

How do I measure my p without deploying a queue?

Take fifty pull requests merged last month, check each out at the tip of main as it was at their merge moment, and run the affected tests. The failure fraction is your p. Split it by author type — human versus agent — because the two numbers are usually far apart and only the blend matters for batch sizing.

Do I need Bazel or Nx to get lanes?

For derived lanes, effectively yes — the products infer targets from a build graph. Without one, a hand-maintained path-to-lane map is a worse approximation and still far better than one global line. Keep a small set of integration checks that run in every lane, because a wrong lane boundary merges two changes that do interact.

Can agents merge through the queue themselves?

They should, and they should have no other route. Remove the admin-merge permission from agent tokens rather than writing a policy about it — one direct push to main invalidates every speculative run in flight.

What stops an agent from deleting a test to get through?

Nothing in the queue. A diff-level rule: an agent fixing a queue ejection may not modify test files unless the ejection was itself a test failure it is legitimately updating, and any deleted assertion triggers human review regardless.

Further reading

On this wiki:

Product documentation: