AI Blog

verl vs SkyRL vs AReaL vs ROLL

All four are Apache-2.0 and all four ship PPO and GRPO, so neither the licence nor the algorithm list decides anything. What decides it is whether your environment is a separately scheduled participant in the rollout or a callback inside the generator — because every fix for a slow tool buys throughput by training on stale data. Pick on which staleness knob you get, not on whose speedup number is biggest.

By Agentic AI Wiki 16 min read

All four of these libraries are Apache-2.0, and all four implement PPO and GRPO, so the two things teams actually compare decide nothing. The decision that matters is structural: is your environment a separately scheduled participant in the rollout, or a callback buried inside the generator? That one choice sets your step time the moment a tool call takes ninety seconds — and every fix for it buys throughput by training on data the current policy did not produce. So pick on which staleness control you are handed, not on whose speedup headline is larger, because ROLL Flash's reported 2.72× on an agentic benchmark and SkyRL-Agent's 1.55× are measurements of different environments rather than a ranking.

At a glance

These are post-training libraries for reinforcement learning on language models, all four now carrying explicit support for multi-turn, tool-using rollouts. Read the table for shape, not for a winner.

ProjectOriginRollout shapeLeans hardest on
verlByteDance Seed, now community-maintainedTrainer — Generator, tools inside the loopBreadth: FSDP and Megatron, vLLM and SGLang, the largest ecosystem of forks.
AReaLRL Lab, Ant ResearchTrainer — Generator — EnvironmentAsynchrony as the primitive: interruptible rollout, in-flight weight updates, explicit staleness control.
ROLLAlibabaTrainer — Generator — EnvironmentScale-out scheduling: queue policies and environment-level asynchronous execution.
SkyRLNovaSky / Sky Computing Lab, UC BerkeleyModular, environment layer optionalReal long-horizon agent tasks on an existing harness, with a separate gym and serving layer.
Where each library leans hardest A matrix of four agentic RL libraries against four axes: asynchronous rollout, the environment as a separately scheduled participant, staleness exposed as an explicit control, and reuse of an existing agent harness. AReaL and ROLL lean hardest on asynchrony and staleness control, SkyRL on harness reuse, and verl on breadth of training backends. Four libraries, four axes that actually differ ASYNC ROLLOUT SCHEDULED ENV STALENESS KNOB HARNESS REUSE verl Medium — opt-in mode Weak — tool callback Medium — queue depth Medium — via forks AReaL Strong — async-first Strong Strong — controller Weak ROLL Strong — ROLL Flash Strong — env-level Strong — queue policy Weak SkyRL Strong — overlapped Medium — gym layer Medium — queue depth Strong — OpenHands Weak Medium Strong All four are Apache-2.0 and all four ship PPO and GRPO.
Where each one leans hardest. The licence column would have been identical, which is why it is not here.

One more framing before the detail. Three of the four describe their rollout data flow as trainer, generator and environment; verl's canonical description is trainer and generator, with custom environments reached through tools. That is not a deficiency — it is the right shape for verifiable-reward work, where the "environment" is a grader that returns in milliseconds. It becomes a deficiency precisely when the environment is a repository, a browser or a sandbox.

The axis that decides it: where the environment sits

Where the environment sits in an agentic RL rollout Two architectures side by side. In the trainer-generator shape the environment is a tool callback inside the generator, so a slow tool blocks the generation slot. In the trainer-generator-environment shape the environment is a separately scheduled service, so generation and tool execution overlap and finished trajectories flow into a bounded queue the trainer consumes. Trainer — Generator THE ENVIRONMENT IS A CALLBACK Trainer Waits for a full batch Generator (GPU) Decode, then call the tool, then decode The slot is held for the whole turn tool call — 90s, GPU idle Cost of the shape Tail latency of one tool sets the step time Strictly on-policy, and mostly waiting Trainer — Generator — Environment THE ENVIRONMENT IS SCHEDULED Trainer Consumes from a queue Generator Never blocks on a tool Interruptible Environment Own workers Own concurrency Own timeouts Bounded rollout queue Capacity is the staleness bound you chose Throughput paid for in off-policy data — a correctness knob
A tool call that holds the generation slot is a GPU-hours problem wearing a latency costume.

Single-turn RLVR has a comfortable loop: sample completions, score them with a verifier, update. The verifier is fast and deterministic, and synchronous execution costs you almost nothing. Agentic RL breaks exactly one assumption in that loop, and it breaks it badly. A rollout is now twenty to fifty turns, each turn may execute a test suite, drive a browser or wait on a container, and the per-turn latency distribution has a tail measured in minutes.

If the environment is a callback inside the generator, the generation slot is held for the whole turn. Your step time becomes the slowest trajectory in the batch, and the batch is not done until the straggler is. This is the dominant cost in agentic RL and it is not a tuning problem — a single ninety-second test run can idle a GPU through what should have been thousands of decode steps.

If the environment is separately scheduled, three things become possible at once, and they are the real content of every "async" claim in this space:

  • Overlap. Generation for trajectory B proceeds while trajectory A waits on its tool. SkyRL-Agent reports 1.55× throughput from a fine-grained asynchronous pipeline that does precisely this.
  • Independent concurrency and timeouts. The environment pool scales on its own axis, which matters because tool concurrency and GPU concurrency have nothing to do with each other.
  • Partial progress is addressable. A trajectory can be paused, resumed, or dropped mid-flight without stalling a generation slot — the precondition for everything in the next section.

The architectural question to ask a candidate library is therefore not "does it support tools" — all four do. It is: can the environment block a GPU? If yes, your effective throughput is set by your worst tool, and no amount of batch tuning recovers it.

verl — the default, and the one you will most likely extend

What it is

verl (also published as HybridFlow) was initiated by ByteDance's Seed team and is now maintained as a community project. It is the breadth option: FSDP and Megatron for training, vLLM and SGLang for generation, PPO, GRPO, DPO and supervised fine-tuning in one codebase, Ray for orchestration. If you are reading a paper that post-trained an open-weights model this year, there is a good chance the code is verl or a fork of it.

Where it sits on the axis

verl's rollout data flow is canonically trainer and generator, with custom environments reached through its tool interface, and a fully asynchronous mode available as an opt-in rather than the organising principle. For verifiable-reward work that is the correct trade: you get on-policy data, a simpler failure surface, and no off-policy loss to reason about. For a twenty-turn repository task it means you are either running the opt-in async path or paying the straggler tax.

The real reason to pick it

Gravity, and it is a legitimate reason. The surrounding ecosystem is the largest of the four — VerlTool exists specifically to make agentic tool-use RL first-class on top of verl, SkyRL's early work was built on it, and the fork you need for your exotic requirement probably already exists. If your team's constraint is "we need to reproduce a published result and then change one thing", start here and do not overthink it.

The honest caveat

Breadth has a cost in surface area. Four training backends times two inference engines times several algorithm families is a lot of configuration that can be individually correct and jointly wrong, and the fully-async path is the least travelled of those combinations. Budget time for a shakeout run that measures nothing but throughput and policy lag.

AReaL — asynchrony as the primitive, not a mode

What it is

AReaL comes from the RL Lab at Ant Research and was published as a large-scale asynchronous RL system for language reasoning. Its distinguishing property is that asynchrony is the design centre rather than an option: generation streams, rewards are computed as trajectories complete, and the trainer never waits for a synchronised batch.

The three mechanisms worth knowing

  • Interruptible rollout. When a new policy version lands, rollout workers pause or interrupt generation in flight, load the new parameters, and resume the unfinished trajectories. The alternative — discarding in-flight work on every update — is what makes naive async expensive at long horizons.
  • Explicit staleness control. A rollout controller bounds how far behind the current policy the data in flight may be. This is the single most important configuration value in asynchronous RL and AReaL treats it as a first-class dial rather than an emergent property of queue sizes.
  • A decoupled PPO objective. Off-policy data needs an objective that tolerates it. Pairing a staleness bound with a loss designed for that bound is what keeps the throughput win from quietly becoming a quality loss.

Who should pick it

Teams whose environments have long, heavy-tailed latency and who intend to run at a scale where idle GPUs are a budget line. Also teams who want to be able to answer the question "how off-policy was this run", because that answer is a configured number here rather than an archaeology project.

The honest caveat

Interruption and resumption are the tightest coupling between trainer and generator of anything on this page, and partial trajectories are a category of bug that does not exist in a synchronous loop. Expect to spend real time on the failure modes of resumed generation before you see the throughput.

ROLL — scheduling as the product

What it is

ROLL is Alibaba's scaling library for RL on large language models, with DeepSpeed and Megatron backends and a trainer-generator-environment data flow. Its agentic story is carried by ROLL Flash, which adds native asynchronous post-training built on two stated principles: fine-grained parallelism and rollout-train decoupling.

The numbers, read carefully

ROLL Flash reports up to 2.24× speedup on RLVR tasks and 2.72× on agentic tasks at the same GPU budget as synchronous baselines, with 2.72× on ALFWorld and 1.81× on a software-engineering task, and near-linear throughput scaling — 7.6× with eight times the GPUs on one 8B configuration. It also reports that several off-policy algorithms were implemented and verified to match synchronous training in quality, which is the claim that makes the speedups usable rather than merely large.

Note the spread inside that one paper: 2.72× on ALFWorld against 1.81× on the software task. Same system, same asynchrony, very different environments. That spread is the whole lesson about cross-library speedup comparisons — the ratio is a property of your environment's latency distribution at least as much as of the library, so a number measured on ALFWorld tells you very little about your repository-scale rollouts.

Who should pick it

Teams with large fixed GPU allocations and a scheduling-shaped problem: many heterogeneous environments, wide variation in episode length, and a need to keep the cluster busy. Queue scheduling and environment-level asynchronous execution are exposed as policy here, which is what you want if your bottleneck is a mix of fast and slow environments rather than one slow one.

The honest caveat

The flexibility is in the scheduling layer, which means the configuration surface is in the scheduling layer too. If your deployment is one environment with uniform latency, you are paying for machinery that solves a problem you do not have.

SkyRL — built from the agent task backwards

What it is

SkyRL comes out of NovaSky at Berkeley's Sky Computing Lab and is deliberately modular: a training component, skyrl-gym for tool-use environments, and a separate serving-oriented layer, usable independently. Its origin is the distinguishing fact — the early pipeline was built to train on long-horizon, real-environment tasks like SWE-Bench, with episodes of roughly twenty to fifty turns, on top of an existing agent harness rather than a bespoke one.

Why the harness question matters more than it looks

Most RL libraries ask you to re-express your agent inside their rollout abstraction. If you already have a harness that works — a coding agent with its own prompt scaffolding, retry logic and tool surface — reimplementing it inside a trainer is both weeks of work and a silent source of train-serve skew: you are now optimising a policy for a harness that is not the one you deploy. SkyRL's integration with an existing open-source coding harness is the most direct answer to that on this page, and skyrl-gym generalises it rather than special-casing it.

The number and what produced it

SkyRL-Agent reports 1.55× training throughput from a fine-grained asynchronous pipeline that overlaps tool execution with model generation. That is a smaller headline than ROLL Flash's and it is measured on harder, longer tasks — which is the second instance of the same caution: compare the environments before you compare the ratios.

The honest caveat

Modularity means you are assembling a stack, and the pieces move independently. For a team that wants one configuration file and a documented happy path, that is friction; for a team that wants to swap the trainer out from under a working environment layer, it is the point.

Staleness is a correctness knob wearing a throughput costume

Three staleness regimes and what each costs Synchronous rollout keeps data on-policy but leaves GPUs idle behind the slowest tool call. A bounded async queue trades a known amount of off-policy data for throughput. Interruptible rollout with in-flight weight updates gives the lowest staleness at the highest implementation coupling. PAYS IN IDLE GPUS PAYS IN BOUNDED STALENESS PAYS IN COUPLING Synchronous Strictly on-policy Step time = slowest tool No off-policy correction Right for fast verifiers Wrong for a real repo Bounded async queue Staleness = queue capacity One number to tune and report Needs an off-policy loss The default worth starting from Degrades quietly if unset Interruptible rollout Pause, load weights, resume Lowest achievable staleness Partial trajectories to handle Tightest generator coupling Hardest to debug at 3am Report the staleness bound next to the speedup, or the speedup is unreadable.
Each regime buys throughput with a different currency. Only one of them makes the price explicit by default.

Every asynchrony mechanism on this page has the same underlying trade: the trainer consumes trajectories that an older policy generated. Comparatively, the four libraries converge on a bounded async queue as the common denominator — SkyRL, verl in fully-async mode, ROLL and AReaL all let multiple batches be in flight with staleness bounded by queue capacity — and then diverge on how visible that bound is. AReaL puts it in a controller and pairs it with an objective built for it; ROLL exposes it as queue scheduling policy; verl and SkyRL leave it as a consequence of how you sized the queue.

That visibility difference is worth more than it sounds, because of how this fails. An unbounded or badly-sized queue does not crash. It produces a run that trains, converges to something, and underperforms the synchronous baseline by an amount nobody can attribute — and the usual response is to blame the reward design. Meanwhile the speedup number on the dashboard looks excellent, which is precisely the problem: throughput improved and the thing you were buying with it got worse, in the same experiment, measured by different people.

Three practices follow regardless of which library you pick:

  • Report the staleness bound next to every speedup. A 2× with an unstated bound is not a result, it is a setting. This is the same discipline as eval variance and statistical power applied one layer down.
  • Run the synchronous baseline to convergence once, properly. It is the only thing that tells you whether your off-policy correction is working, and it is the run everyone skips because it is slow.
  • Pair the bound with an objective that expects it. If your library offers a decoupled or otherwise off-policy-tolerant loss, the staleness setting and the loss choice are one decision, not two.

When to pick which

SituationPickBecauseWatch for
Reproducing a published post-training resultverlMost published recipes and forks target it.Configuration surface across backends.
Math or code RLVR with a fast verifierverl, synchronousAsynchrony buys little when the environment returns in milliseconds.Reaching for async out of habit.
Long-horizon repository or browser tasksSkyRLBuilt for 20–50 turn episodes on an existing harness.Assembling and version-managing a modular stack.
You already have a production agent harnessSkyRLTraining the harness you deploy avoids train-serve skew.Harness changes becoming training changes.
Heavy-tailed tool latency, large clusterAReaLInterruptible rollout and an explicit staleness controller.Partial-trajectory failure modes.
Many heterogeneous environments at onceROLLQueue scheduling and environment-level asynchrony are exposed as policy.Paying for scheduling you do not need.
You need to answer "how off-policy was this?"AReaL or ROLLThe bound is a configured value, not an emergent one.Nothing — this should be a requirement more often.

And the meta-recommendation, because it saves more time than any of the rows: before choosing, measure your environment's per-turn latency distribution, especially the 95th percentile. If the tail is under a second, pick for ecosystem and stay synchronous. If the tail is in minutes, the environment-scheduling axis is the only thing on this page that matters and you should pick on it alone.

FAQ

Does the licence differentiate these at all?

No. All four are Apache-2.0, which is why the comparison has to be architectural. Licence risk in this category comes from the model weights and the environment images you train against, not from the trainer.

Is ROLL Flash's 2.72× directly comparable to SkyRL-Agent's 1.55×?

No, and treating them as comparable is the most common mistake here. They are measured on different environments with different latency profiles, and ROLL Flash's own results span 2.72× on ALFWorld against 1.81× on a software-engineering task. The ratio is a property of your environment as much as of the library.

Do I need asynchronous rollout at all?

Only if your environment is slow relative to generation. For fast deterministic verifiers, synchronous training is simpler, strictly on-policy, and gives up almost nothing. Asynchrony is a fix for idle accelerators, and if yours are not idle it is pure added complexity.

Can I keep my existing agent framework?

Sometimes, and it is worth insisting on. SkyRL's lineage is explicitly about training on an existing harness; verl is extensible to it through its tool interface and third-party projects. Reimplementing your agent inside a trainer risks optimising a policy for a harness you do not ship.

What is the one measurement to take before choosing?

The 95th-percentile per-turn environment latency for your actual task. It determines whether the environment-scheduling axis is decisive or irrelevant, and it takes an afternoon to collect from an existing agent's traces.

Further reading

On this wiki:

Project sources: