Arise: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning

Preprint

1 ShanghaiTech University 2 Ant Group

* Equal contribution. Kun Feng conducted this work during an internship at Ant Group.

† Corresponding author: renkan [at] shanghaitech.edu.cn

fengkun2025 [at] shanghaitech.edu.cn · yuchen.fyc [at] antgroup.com

As agents improve, what they need to learn changes.
ARISE adapts what to evaluate, how to explore, and which tasks to learn from.

SkillsBench performance versus model size, and an illustration of how evolving capability gaps and repeated rollout failures motivate adaptive rubrics, skills, and task selection.
Learning needs evolve. Fixed criteria can miss new capability gaps, while failed rollouts can hide partial behavioral progress. ARISE turns rollout evidence into evolving evaluation criteria, targeted guidance, and task priorities.

SkillsBench v1.1

23.4% → 45.6%

Qwen3.5-27B → ARISE
+22.2 percentage points

Terminal-Bench v2.1

41.6% → 50.6%

Qwen3.5-27B → ARISE
+9.0 percentage points

Training efficiency

50 vs. 150 steps

Matches outcome-only RL performance
with the same per-step rollout budget

Abstract

As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training priorities can become misaligned with evolving agent capabilities, while sparse task-level feedback makes such misalignment more difficult to detect. Even when capability gaps are identified, rollouts from the current policy may repeatedly reproduce the same failures rather than explore better alternatives. To address this, we introduce Adaptive Rubric–Skill Co-Evolution (ARISE), a reinforcement learning framework that uses rollout evidence to continually adapt evaluation criteria, exploration guidance, and training priorities. Rubrics evolve to reward partial behavioral progress, while their paired skills are refined and selectively activated to guide exploration toward unresolved weaknesses. Alongside this co-evolution, capability-based adaptive sampling prioritizes tasks that target behaviors needing further improvement. Experiments on two challenging long-horizon agent benchmarks, SkillsBench and Terminal-Bench, demonstrate that ARISE successfully enhances both overall task performance and training efficiency.

Overview

A continuous feedback loop aligns behavioral evaluation, exploration guidance, and training data with the agent's changing capabilities.

ARISE framework: capability-based sampling selects tasks, activated skills guide agent rollouts, task verifiers and rubric judges evaluate trajectories, and the resulting evidence updates both the policy and the rubric-skill pool.
The ARISE training loop. Shared rollout evidence drives rubric–skill co-evolution and capability-based adaptive sampling. Task outcomes and behavioral feedback jointly inform policy optimization.

Methodology

Two coupled components keep training focused on behavioral gaps that still offer opportunities for improvement.

Evaluate & guide

Rubric–Skill Co-Evolution

Each behavioral criterion pairs a rubric for evaluating observable progress with a skill that provides actionable guidance. A reflection model uses recent rollout evidence to discover emerging gaps and add new pairs.

AddActivateRefineHideRetire

Low rubric pass rates activate paired skills; persistent failures prompt refinement. As performance improves, guidance is hidden while evaluation continues. Consistently mastered pairs are retired, making room for new learning needs.

Select & learn

Capability-Based Adaptive Sampling

Rubric judgments estimate capability performance within each task type, revealing behavioral weaknesses that task-level success rates alone can obscure.

DiscoverEstimatePrioritize

Coverage-guided discovery explores less-evaluated tasks. Adaptive selection prioritizes task-type–capability pairs with informative success probabilities, balancing learning potential against repeated failure. As rubrics evolve, unchanged evidence is retained and obsolete contributions are discarded.

Training strategy. The rubric–skill pool starts empty, so initial training uses task rewards alone. As criteria emerge, task rewards and per-rubric feedback are normalized separately and combined for GRPO-style policy optimization with importance-ratio filtering. Updates apply to model-generated tokens, excluding tool outputs and environment observations.

Results

ARISE achieves the highest overall pass rates among comparable-scale models on both benchmarks, with gains that transfer from the Kilo Code harness to Terminus-2.

Task pass rates in percent on SkillsBench v1.1 and Terminal-Bench v2.1.
ModelSkillsBench v1.1TB v2.1
SENSOWIPFEMRCSMCOverallOverall
Proprietary Models
GPT-5.4 Mini27.135.745.221.425.941.733.366.734.559.2
GPT-5.563.477.976.257.537.095.069.060.067.384.3
Claude Opus-4.758.383.354.854.844.450.057.153.358.683.1
Open-Weight Models
GPT-OSS-120B—————————26.2
MiniMax-M2.714.659.528.619.033.329.223.813.328.755.4
GLM-5.141.776.269.045.233.337.547.680.053.661.8
Kimi-K2.639.676.257.150.044.466.747.666.755.265.9
DeepSeek-V3.2—————————46.8
DeepSeek-V4-Pro37.573.852.442.929.666.742.980.051.364.8
DeepSeek-V4-Pro-081362.581.064.347.648.150.042.960.059.078.7
Qwen3.5-122B-A10B18.840.523.816.722.233.319.013.324.147.6
Qwen3.5-397B-A17B16.745.238.131.022.237.523.820.030.351.3
Nemotron-3-Ultra-550B-A55B—————————53.9
Comparable-Scale Models
Qwen3.5-27B (Base Model)14.640.526.223.818.529.214.36.723.441.6
VADE25.054.838.131.033.350.023.873.338.743.8
OnlineRubrics29.264.338.131.037.033.328.673.340.246.1
RuscaRL33.357.142.926.240.729.223.873.339.544.9
ARISE41.759.552.438.140.733.323.880.045.650.6

Pass rates (%); higher is better. Bold red values mark the best results within comparable-scale models. Dashes indicate unreported results.
SE: software engineering; NS: natural science; OW: office and white collar; IP: industrial and physical systems; FE: finance and economics; MR: mathematics and operations research/formal reasoning; CS: cybersecurity; MC: media and content production.

Evaluation protocol. SkillsBench uses Kilo Code with benchmark-provided curated skills. These differ from the rubric-paired skills evolved by ARISE, which are withheld on both evaluation benchmarks. Terminal-Bench uses Terminus-2. Our evaluations average three independent runs. SkillsBench GPT-5.5 results and Terminal-Bench results outside the comparable-scale group come from official leaderboards.

Analysis

Each component contributes to learning

Dynamic rubric evolution, skill-guided exploration, and capability-based sampling each contribute to performance. Under the same per-step rollout budget, ARISE reaches the step-150 performance of outcome-only RL by step 50 and continues improving. This comparison measures training steps and rollout efficiency.

Ablation comparison of ARISE against frozen rubrics, no skill injection, alternative sampling, and outcome-only RL, alongside SkillsBench pass rates over training steps.
Ablations and training efficiency. Variants use matched initialization, data, and training budgets. Even the full lifecycle rubric pool, when kept fixed, does not replace dynamic evolution.

Behavioral feedback evolves with the policy

Starting from an empty pool, ARISE introduces 36 criteria and retires 21, leaving 15 active. Supervision changes through replacement, not just accumulation. On a fixed task set, skill guidance reduces all-failure rollout groups compared with the variant without skill injection, with the largest gap early in training.

Training dynamics showing active, added, and retired rubrics; rubric reward and task pass rates; and lower all-failure rollout group rates with skill injection.
Co-evolution dynamics. The active criteria change as capabilities develop, while paired skills help the policy explore beyond repeated failures.

Training priorities shift toward emerging gaps

Early adaptive sampling emphasizes execution and verification. Later stages increasingly target debugging and efficiency as new criteria enter the pool. Blank cells indicate limited coverage or absent active rubrics; discovery sampling continues to explore these tasks.

Heatmaps of task-type and capability sampling probabilities across early, middle, and late training, showing changing priorities from execution and verification toward debugging and efficiency.
Adaptive data allocation. Probabilities exclude discovery sampling. Early, middle, and late stages correspond to steps 1–50, 51–100, and 101–150.

Citation

If you find ARISE useful in your research, please cite our preprint.

@misc{feng2026arise,
  title={ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning},
  author={Feng, Kun and Fang, Yuchen and Tan, Yiyang and Gu, Shuqi and Zhao, Yongxiang
          and Liu, Yu and Lu, Xingyu and Ma, Lintao and Ren, Kan},
  year={2026},
  eprint={2609.35532},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2609.35532}
}