All four of these libraries are Apache-2.0, and all four implement PPO and GRPO, so the two things teams actually compare decide nothing. The decision that matters is structural: is your environment a separately scheduled participant in the rollout, or a callback buried inside the generator? That one choice sets your step time the moment a tool call takes ninety seconds — and every fix for it buys throughput by training on data the current policy did not produce. So pick on which staleness control you are handed, not on whose speedup headline is larger, because ROLL Flash's reported 2.72× on an agentic benchmark and SkyRL-Agent's 1.55× are measurements of different environments rather than a ranking.
At a glance
These are post-training libraries for reinforcement learning on language models, all four now carrying explicit support for multi-turn, tool-using rollouts. Read the table for shape, not for a winner.
| Project | Origin | Rollout shape | Leans hardest on |
|---|---|---|---|
| verl | ByteDance Seed, now community-maintained | Trainer — Generator, tools inside the loop | Breadth: FSDP and Megatron, vLLM and SGLang, the largest ecosystem of forks. |
| AReaL | RL Lab, Ant Research | Trainer — Generator — Environment | Asynchrony as the primitive: interruptible rollout, in-flight weight updates, explicit staleness control. |
| ROLL | Alibaba | Trainer — Generator — Environment | Scale-out scheduling: queue policies and environment-level asynchronous execution. |
| SkyRL | NovaSky / Sky Computing Lab, UC Berkeley | Modular, environment layer optional | Real long-horizon agent tasks on an existing harness, with a separate gym and serving layer. |
One more framing before the detail. Three of the four describe their rollout data flow as trainer, generator and environment; verl's canonical description is trainer and generator, with custom environments reached through tools. That is not a deficiency — it is the right shape for verifiable-reward work, where the "environment" is a grader that returns in milliseconds. It becomes a deficiency precisely when the environment is a repository, a browser or a sandbox.
The axis that decides it: where the environment sits
Single-turn RLVR has a comfortable loop: sample completions, score them with a verifier, update. The verifier is fast and deterministic, and synchronous execution costs you almost nothing. Agentic RL breaks exactly one assumption in that loop, and it breaks it badly. A rollout is now twenty to fifty turns, each turn may execute a test suite, drive a browser or wait on a container, and the per-turn latency distribution has a tail measured in minutes.
If the environment is a callback inside the generator, the generation slot is held for the whole turn. Your step time becomes the slowest trajectory in the batch, and the batch is not done until the straggler is. This is the dominant cost in agentic RL and it is not a tuning problem — a single ninety-second test run can idle a GPU through what should have been thousands of decode steps.
If the environment is separately scheduled, three things become possible at once, and they are the real content of every "async" claim in this space:
- Overlap. Generation for trajectory B proceeds while trajectory A waits on its tool. SkyRL-Agent reports 1.55× throughput from a fine-grained asynchronous pipeline that does precisely this.
- Independent concurrency and timeouts. The environment pool scales on its own axis, which matters because tool concurrency and GPU concurrency have nothing to do with each other.
- Partial progress is addressable. A trajectory can be paused, resumed, or dropped mid-flight without stalling a generation slot — the precondition for everything in the next section.
The architectural question to ask a candidate library is therefore not "does it support tools" — all four do. It is: can the environment block a GPU? If yes, your effective throughput is set by your worst tool, and no amount of batch tuning recovers it.
verl — the default, and the one you will most likely extend
What it is
verl (also published as HybridFlow) was initiated by ByteDance's Seed team and is now maintained as a community project. It is the breadth option: FSDP and Megatron for training, vLLM and SGLang for generation, PPO, GRPO, DPO and supervised fine-tuning in one codebase, Ray for orchestration. If you are reading a paper that post-trained an open-weights model this year, there is a good chance the code is verl or a fork of it.
Where it sits on the axis
verl's rollout data flow is canonically trainer and generator, with custom environments reached through its tool interface, and a fully asynchronous mode available as an opt-in rather than the organising principle. For verifiable-reward work that is the correct trade: you get on-policy data, a simpler failure surface, and no off-policy loss to reason about. For a twenty-turn repository task it means you are either running the opt-in async path or paying the straggler tax.
The real reason to pick it
Gravity, and it is a legitimate reason. The surrounding ecosystem is the largest of the four — VerlTool exists specifically to make agentic tool-use RL first-class on top of verl, SkyRL's early work was built on it, and the fork you need for your exotic requirement probably already exists. If your team's constraint is "we need to reproduce a published result and then change one thing", start here and do not overthink it.
The honest caveat
Breadth has a cost in surface area. Four training backends times two inference engines times several algorithm families is a lot of configuration that can be individually correct and jointly wrong, and the fully-async path is the least travelled of those combinations. Budget time for a shakeout run that measures nothing but throughput and policy lag.
AReaL — asynchrony as the primitive, not a mode
What it is
AReaL comes from the RL Lab at Ant Research and was published as a large-scale asynchronous RL system for language reasoning. Its distinguishing property is that asynchrony is the design centre rather than an option: generation streams, rewards are computed as trajectories complete, and the trainer never waits for a synchronised batch.
The three mechanisms worth knowing
- Interruptible rollout. When a new policy version lands, rollout workers pause or interrupt generation in flight, load the new parameters, and resume the unfinished trajectories. The alternative — discarding in-flight work on every update — is what makes naive async expensive at long horizons.
- Explicit staleness control. A rollout controller bounds how far behind the current policy the data in flight may be. This is the single most important configuration value in asynchronous RL and AReaL treats it as a first-class dial rather than an emergent property of queue sizes.
- A decoupled PPO objective. Off-policy data needs an objective that tolerates it. Pairing a staleness bound with a loss designed for that bound is what keeps the throughput win from quietly becoming a quality loss.
Who should pick it
Teams whose environments have long, heavy-tailed latency and who intend to run at a scale where idle GPUs are a budget line. Also teams who want to be able to answer the question "how off-policy was this run", because that answer is a configured number here rather than an archaeology project.
The honest caveat
Interruption and resumption are the tightest coupling between trainer and generator of anything on this page, and partial trajectories are a category of bug that does not exist in a synchronous loop. Expect to spend real time on the failure modes of resumed generation before you see the throughput.
ROLL — scheduling as the product
What it is
ROLL is Alibaba's scaling library for RL on large language models, with DeepSpeed and Megatron backends and a trainer-generator-environment data flow. Its agentic story is carried by ROLL Flash, which adds native asynchronous post-training built on two stated principles: fine-grained parallelism and rollout-train decoupling.
The numbers, read carefully
ROLL Flash reports up to 2.24× speedup on RLVR tasks and 2.72× on agentic tasks at the same GPU budget as synchronous baselines, with 2.72× on ALFWorld and 1.81× on a software-engineering task, and near-linear throughput scaling — 7.6× with eight times the GPUs on one 8B configuration. It also reports that several off-policy algorithms were implemented and verified to match synchronous training in quality, which is the claim that makes the speedups usable rather than merely large.
Note the spread inside that one paper: 2.72× on ALFWorld against 1.81× on the software task. Same system, same asynchrony, very different environments. That spread is the whole lesson about cross-library speedup comparisons — the ratio is a property of your environment's latency distribution at least as much as of the library, so a number measured on ALFWorld tells you very little about your repository-scale rollouts.
Who should pick it
Teams with large fixed GPU allocations and a scheduling-shaped problem: many heterogeneous environments, wide variation in episode length, and a need to keep the cluster busy. Queue scheduling and environment-level asynchronous execution are exposed as policy here, which is what you want if your bottleneck is a mix of fast and slow environments rather than one slow one.
The honest caveat
The flexibility is in the scheduling layer, which means the configuration surface is in the scheduling layer too. If your deployment is one environment with uniform latency, you are paying for machinery that solves a problem you do not have.
SkyRL — built from the agent task backwards
What it is
SkyRL comes out of NovaSky at Berkeley's Sky Computing Lab and is deliberately modular: a training component, skyrl-gym for tool-use environments, and a separate serving-oriented layer, usable independently. Its origin is the distinguishing fact — the early pipeline was built to train on long-horizon, real-environment tasks like SWE-Bench, with episodes of roughly twenty to fifty turns, on top of an existing agent harness rather than a bespoke one.
Why the harness question matters more than it looks
Most RL libraries ask you to re-express your agent inside their rollout abstraction. If you already have a harness that works — a coding agent with its own prompt scaffolding, retry logic and tool surface — reimplementing it inside a trainer is both weeks of work and a silent source of train-serve skew: you are now optimising a policy for a harness that is not the one you deploy. SkyRL's integration with an existing open-source coding harness is the most direct answer to that on this page, and skyrl-gym generalises it rather than special-casing it.
The number and what produced it
SkyRL-Agent reports 1.55× training throughput from a fine-grained asynchronous pipeline that overlaps tool execution with model generation. That is a smaller headline than ROLL Flash's and it is measured on harder, longer tasks — which is the second instance of the same caution: compare the environments before you compare the ratios.
The honest caveat
Modularity means you are assembling a stack, and the pieces move independently. For a team that wants one configuration file and a documented happy path, that is friction; for a team that wants to swap the trainer out from under a working environment layer, it is the point.
Staleness is a correctness knob wearing a throughput costume
Every asynchrony mechanism on this page has the same underlying trade: the trainer consumes trajectories that an older policy generated. Comparatively, the four libraries converge on a bounded async queue as the common denominator — SkyRL, verl in fully-async mode, ROLL and AReaL all let multiple batches be in flight with staleness bounded by queue capacity — and then diverge on how visible that bound is. AReaL puts it in a controller and pairs it with an objective built for it; ROLL exposes it as queue scheduling policy; verl and SkyRL leave it as a consequence of how you sized the queue.
That visibility difference is worth more than it sounds, because of how this fails. An unbounded or badly-sized queue does not crash. It produces a run that trains, converges to something, and underperforms the synchronous baseline by an amount nobody can attribute — and the usual response is to blame the reward design. Meanwhile the speedup number on the dashboard looks excellent, which is precisely the problem: throughput improved and the thing you were buying with it got worse, in the same experiment, measured by different people.
Three practices follow regardless of which library you pick:
- Report the staleness bound next to every speedup. A 2× with an unstated bound is not a result, it is a setting. This is the same discipline as eval variance and statistical power applied one layer down.
- Run the synchronous baseline to convergence once, properly. It is the only thing that tells you whether your off-policy correction is working, and it is the run everyone skips because it is slow.
- Pair the bound with an objective that expects it. If your library offers a decoupled or otherwise off-policy-tolerant loss, the staleness setting and the loss choice are one decision, not two.
When to pick which
| Situation | Pick | Because | Watch for |
|---|---|---|---|
| Reproducing a published post-training result | verl | Most published recipes and forks target it. | Configuration surface across backends. |
| Math or code RLVR with a fast verifier | verl, synchronous | Asynchrony buys little when the environment returns in milliseconds. | Reaching for async out of habit. |
| Long-horizon repository or browser tasks | SkyRL | Built for 20–50 turn episodes on an existing harness. | Assembling and version-managing a modular stack. |
| You already have a production agent harness | SkyRL | Training the harness you deploy avoids train-serve skew. | Harness changes becoming training changes. |
| Heavy-tailed tool latency, large cluster | AReaL | Interruptible rollout and an explicit staleness controller. | Partial-trajectory failure modes. |
| Many heterogeneous environments at once | ROLL | Queue scheduling and environment-level asynchrony are exposed as policy. | Paying for scheduling you do not need. |
| You need to answer "how off-policy was this?" | AReaL or ROLL | The bound is a configured value, not an emergent one. | Nothing — this should be a requirement more often. |
And the meta-recommendation, because it saves more time than any of the rows: before choosing, measure your environment's per-turn latency distribution, especially the 95th percentile. If the tail is under a second, pick for ecosystem and stay synchronous. If the tail is in minutes, the environment-scheduling axis is the only thing on this page that matters and you should pick on it alone.
FAQ
Does the licence differentiate these at all?
No. All four are Apache-2.0, which is why the comparison has to be architectural. Licence risk in this category comes from the model weights and the environment images you train against, not from the trainer.
Is ROLL Flash's 2.72× directly comparable to SkyRL-Agent's 1.55×?
No, and treating them as comparable is the most common mistake here. They are measured on different environments with different latency profiles, and ROLL Flash's own results span 2.72× on ALFWorld against 1.81× on a software-engineering task. The ratio is a property of your environment as much as of the library.
Do I need asynchronous rollout at all?
Only if your environment is slow relative to generation. For fast deterministic verifiers, synchronous training is simpler, strictly on-policy, and gives up almost nothing. Asynchrony is a fix for idle accelerators, and if yours are not idle it is pure added complexity.
Can I keep my existing agent framework?
Sometimes, and it is worth insisting on. SkyRL's lineage is explicitly about training on an existing harness; verl is extensible to it through its tool interface and third-party projects. Reimplementing your agent inside a trainer risks optimising a policy for a harness you do not ship.
What is the one measurement to take before choosing?
The 95th-percentile per-turn environment latency for your actual task. It determines whether the environment-scheduling axis is decisive or irrelevant, and it takes an afternoon to collect from an existing agent's traces.
Further reading
On this wiki:
- RLVR & GRPO for Agents — the algorithm layer these libraries implement.
- Environment Engineering for RL — the work that dominates the project once the trainer is chosen.
- RL for Tool Use — why multi-turn rollouts break the single-turn loop.
- Reward Design & Reward Hacking — the failure you will misattribute a staleness bug to.
- Trajectories — the unit of data all four of these libraries move around.
Project sources:
- verl / HybridFlow — the RL post-training framework initiated by ByteDance Seed, now community-maintained; Apache-2.0.
- AReaL: a large-scale asynchronous reinforcement learning system for language reasoning — interruptible rollout, staleness control and the decoupled PPO objective.
- Part II: ROLL Flash — accelerating RLVR and agentic training with asynchrony — fine-grained parallelism, rollout-train decoupling, and the 2.24× / 2.72× figures.
- SkyRL-Agent: efficient RL training for multi-turn LLM agents — the fine-grained asynchronous pipeline and the 1.55× throughput result.
- alibaba/ROLL and NovaSky-AI/SkyRL — the repositories, including
skyrl-gym.