Skip to content

Commit 4d2021c

Browse files
committed
Add timeslice-verl package and fully-async two-RL-jobs guide
Adds the timeslice-verl integration package and a guide that time-slices two verl fully_async_policy RL jobs' trainers on one shared 1-GPU node (each job keeps a dedicated 1-GPU rollout node; 2-pod Ray cluster per job). The package ships TimesliceFullyAsyncTrainer, a subclass of verl's FullyAsyncTrainer that overrides the trainer's empty on_* lifecycle hooks (same template-method convention as verl's v1 trainers). It registers under the trainer name "timeslice" via verl's fully-async trainer registry — the package's verl.plugins entry point makes `import verl` load the registering module — and is selected with the hydra override async_training.trainer_name=timeslice. No meta_path monkey-patching. Placement-group pinning uses verl's per-pool extra bundle resources (ray_pg_extra_resources hydra override) instead of patching placement_group; the trainer refuses to train unpinned when TIMESLICE_REQUIRE_PG_PINNING=1. Both verl features ship as format-patch files in the guide's patches/ dir until they are available upstream (default install path: the feat/fully-async-lifecycle-hooks fork branch). New: pkg/integrations/verl/ (TimesliceFullyAsyncTrainer = TimesliceHooksMixin over FullyAsyncTrainer, + PhaseLocks; 33 tests, pure-python: no verl/ray/ grpc/GPU needed) New: guides/rl-frameworks/verl/ (README, job manifests for a math-RLVR and a code-RLVR job, GPU monitor, platform values, verl feature patches)
1 parent d1b37de commit 4d2021c

11 files changed

Lines changed: 2849 additions & 0 deletions

File tree

Lines changed: 143 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,143 @@
1+
# Time-Slicing Integration Guide for verl Workloads
2+
3+
This guide covers integrating **verl** (Volcano Engine Reinforcement Learning framework) with the **llm-d-rl-time-slicing** platform using the pre-packaged `timeslice-verl` integration (`pkg/integrations/verl/`).
4+
5+
### Motivation: Maximizing GPU Utilization
6+
7+
In verl's **fully-async** recipe (`fully_async_policy`), rollout generation runs continuously while the trainer sits idle whenever it waits for the next batch of samples — in our reference runs that idle fraction was ~55–70% of wall-clock. Cooperative time-slicing backfills those valleys: multiple jobs' trainers share one physical GPU pool, and the platform checkpoints the idle trainer's GPU state to host RAM (`cuda-checkpoint`) at every handoff. Rollout engines keep dedicated GPUs and never stop generating: while a job's trainer is checkpointed off the shared GPU, its rollout engine keeps generating samples into the staleness queue.
8+
9+
### Supported Modes
10+
11+
| verl mode | Status | Notes |
12+
|---|---|---|
13+
| **Fully-async** (`fully_async_policy`) | **Supported** | Via the `TimesliceFullyAsyncTrainer` subclass + fork lifecycle hooks. See the runnable example below. |
14+
| v1 sync modes (e.g. `main_ppo` PPO/GRPO) | Not shipped in this PR | The v1 trainers expose the same `on_*` hook convention; integration is future work. |
15+
16+
For a runnable example, see:
17+
18+
* **[Fully-Async Example](examples/fully-async/README.md)** — two fully-async RL jobs (math RLVR + code RLVR) sharing one trainer GPU, each with a dedicated rollout GPU
19+
20+
---
21+
22+
## Table of Contents
23+
1. [Cluster Prerequisites](#1-cluster-prerequisites)
24+
2. [Deploying the Time-Slicing Platform](#2-deploying-the-time-slicing-platform)
25+
3. [Integrating with verl](#3-integrating-with-verl)
26+
4. [General Troubleshooting](#4-general-troubleshooting)
27+
28+
---
29+
30+
## 1. Cluster Prerequisites
31+
32+
Before deploying cooperative time-slicing for verl, ensure your environment meets the following requirements:
33+
34+
* A Kubernetes cluster (we validated on GKE, COS nodes) with 1-GPU nodes (H100-class recommended). The GPUs used for time-slicing must NOT be in use by anything else.
35+
* Cluster-admin access; `kubectl`, `helm` (v3), `git`, and `envsubst` (from gettext; `brew install gettext` on macOS) on your workstation.
36+
* Nodes can pull from Docker Hub (`verlai/verl` image, ~20 GB) and reach GitHub + HuggingFace from pods (model + dataset download at startup; the model is fetched by HF id per node, no shared volume).
37+
38+
### Node Labeling and Time-Slice Groups
39+
40+
The orchestrator discovers resource pools (*groups*) from node labels. Jobs in the same group take turns holding the group's accelerator lock. Label ONLY the nodes whose GPUs should be time-sliced into the group — the orchestrator discovers groups from this label:
41+
42+
```bash
43+
kubectl label node <trainer-node> group.timeslice.io/trainers=true --overwrite
44+
```
45+
46+
Nodes hosting rollout engines must NOT carry the group label, and rollout pods must not carry `timeslice.io/*` pod labels: only labeled pods on labeled nodes are snapshot candidates, so unlabeled rollout processes are never touched by the platform.
47+
48+
Verify the orchestrator synced the group after labeling:
49+
50+
```bash
51+
kubectl -n timeslice-system logs deploy/timeslice-timesliceorchestrator --tail=50 \
52+
| grep -i trainers | tail -3
53+
```
54+
55+
### Host-RAM Sizing Rule
56+
57+
A head pod's memory limit must fit its normal RSS **plus its entire GPU allocation** — the cuda-checkpoint dump of a trainer's full GPU state lands in pod memory. Size trainer-node host RAM to hold the GPU memory footprint of every job whose trainer can be checkpointed there.
58+
59+
---
60+
61+
## 2. Deploying the Time-Slicing Platform
62+
63+
We deploy the core platform components — **TimeSlice Orchestrator** (Deployment: the gRPC lock service) and **Snapshot Agent** (DaemonSet, one per GPU node: performs `cuda-checkpoint` snapshot/restore of a job's GPU state when the orchestrator hands the lock over) — using the parent Helm chart.
64+
65+
> If you previously installed the platform, do a clean
66+
> `helm uninstall timeslice -n timeslice-system` first — `helm install` fails
67+
> on an existing release name. (Note: uninstalling also
68+
> removes the bundled NVIDIA DRA driver DaemonSet until you reinstall.)
69+
70+
```bash
71+
git clone https://github.com/llm-d-incubation/llm-d-rl-time-slicing.git
72+
cd llm-d-rl-time-slicing
73+
74+
helm dependency update ./deploy
75+
helm install timeslice ./deploy -n timeslice-system --create-namespace
76+
```
77+
78+
> **Release pin:** the chart defaults pull the official `ghcr.io/llm-d-incubation/llm-d-rl-time-slicing/*` images at `latest`. Once the first tagged release is cut, pin it with `--set timesliceorchestrator.image.tag=<version> --set snapshot-agent.image.tag=<version>`.
79+
80+
Verify (orchestrator Running; one agent pod per GPU node):
81+
82+
```bash
83+
kubectl -n timeslice-system get pods -o wide
84+
```
85+
86+
---
87+
88+
## 3. Integrating with verl
89+
90+
### How It Works
91+
92+
The `timeslice-verl` package ships `TimesliceFullyAsyncTrainer`, a subclass of verl's `fully_async_policy` trainer that overrides the trainer's empty `on_*` lifecycle hooks to acquire/release the orchestrator lock at the trainer's natural wait points. It registers under the trainer name `timeslice` (verl's `register_trainer` registry; the package's `verl.plugins` entry point makes `import verl` load the registering module) and is selected with a single hydra override: `async_training.trainer_name=timeslice`. No monkey-patching and no verl source edits in your job — it requires a verl build with the fully-async lifecycle hooks + trainer registry and per-pool placement-group bundle resources (currently the `feat/fully-async-lifecycle-hooks` fork branch, until the commits land upstream).
93+
94+
Lock protocol:
95+
96+
| Trainer lifecycle point | Lock action |
97+
|---|---|
98+
| `init_workers` (model load) | ACQUIRE → load → YIELD |
99+
| initial weight sync (pre-fit) | ACQUIRE → sync → YIELD |
100+
| a batch of samples is ready | ACQUIRE (this is the resume point) |
101+
| weight update + param sync done | YIELD |
102+
103+
### Placement-Group Pinning
104+
105+
Trainer/rollout placement is pinned with Ray **custom resources** — the head starts ray with `--resources='{"trainer_node": 100}'`, the worker with `--resources='{"rollout_node": 100}'`, and verl's per-pool placement-group bundle resources (hydra override `ray_pg_extra_resources={trainer_pool: {trainer_node: 1}, rollout_pool: {rollout_node: 1}}`) make verl's trainer placement group request `{trainer_node: 1}` and the rollout PG `{rollout_node: 1}`.
106+
107+
### Installing verl (Pinned Fork)
108+
109+
Each job pod clones the [`feat/fully-async-lifecycle-hooks`](https://github.com/aishukamal/verl/tree/feat/fully-async-lifecycle-hooks) branch of `github.com/aishukamal/verl` — upstream `983cb0f24443f87b3d161fad318445130a620b07` + two feature commits (fully-async lifecycle hooks + trainer registry; per-pool PG bundle resources) — and installs it with `pip install -e ".[gpu]"`. This is a temporary fork until the commits land upstream; once they do, point the `VERL_REPO`/`VERL_REF` pod envs at mainline verl instead.
110+
111+
### Installing the `timeslice-verl` Package (ConfigMap)
112+
113+
Publish the integration package (`pkg/integrations/verl/`) into the cluster as ConfigMaps — the job pods install it from these, no GitHub fetch:
114+
115+
```bash
116+
# Package source: the repo you cloned in §2
117+
PKG=llm-d-rl-time-slicing/pkg/integrations/verl
118+
kubectl create configmap timeslice-verl-root --from-file=$PKG/pyproject.toml
119+
kubectl create configmap timeslice-verl-src --from-file=$PKG/timeslice_verl/
120+
```
121+
122+
ConfigMap keys cannot contain `/`, so the package tree is split into two maps (project root + module sources) and reassembled inside the pod; this also works in air-gapped clusters.
123+
124+
### Required NCCL Environment
125+
126+
Set `NCCL_CUMEM_ENABLE=0` and `NCCL_NVLS_ENABLE=0` on every job pod — the NVLS/cuMem NCCL transports don't survive cuda-checkpoint, and a trainer will die right after a restore without them (the example manifests already set both).
127+
128+
---
129+
130+
## 4. General Troubleshooting
131+
132+
Every entry below is a failure we actually hit while building the integration. For failures specific to the runnable demo's topology and manifests, see the [example's troubleshooting section](examples/fully-async/README.md#8-troubleshooting).
133+
134+
| Symptom | Cause | Fix |
135+
|---|---|---|
136+
| Trainer dies right after a restore: `NCCL ... unhandled cuda error` in `transport/nvls.cc` | `NCCL_CUMEM_ENABLE`/`NCCL_NVLS_ENABLE` not set to `0` — NVLS/cuMem NCCL transports don't survive cuda-checkpoint | Set both to `0` on all pods (§3; the example manifests do) — keep them if you edit |
137+
| `ImportError` from `timeslice` client at startup | grpcio/protobuf too old in the base image | Upgrade to `grpcio>=1.81 protobuf>=7.35` in the job pods (the example manifests do) — keep that step |
138+
| Trainer placement group lands on the rollout node (or rollout work on the trainer node) | PG pinning not active — `--resources` missing on a `ray start` line, or the `ray_pg_extra_resources` hydra override is missing/typo'd, or the verl build lacks the feature | Verify both `ray start` lines carry `--resources`, keep the `ray_pg_extra_resources={trainer_pool: ...}` override in the head launch script intact, and use the pinned verl branch (§3); check `ray status` custom resources |
139+
| Rollout throughput collapses when the other job trains; agent log shows cuda-checkpoint activity on a ROLLOUT node | A rollout pod carries `timeslice.io/*` labels, or a rollout node is labeled into the group | Rollout pods/nodes must stay unlabeled — only head pods and the trainer node (§1). A rollout node must never show cuda-checkpoint activity |
140+
| `Found multiple active Ray instances` warning on a shared node | With `hostPID`, Ray's `address='auto'` discovery scans `/proc` and sees the OTHER job's GCS — a driver can join the wrong ray cluster | `export RAY_ADDRESS="$(hostname -i):6379"` right after `ray start --head` (the example manifests do) — keep that line if you edit |
141+
| Hydra error `Key 'use_trainer_do_validate' is not in struct trainer` | Wrong config key path | These flags live under `async_training.*`, not `trainer.*` (the example manifests are correct) |
142+
| `save_freq` / checkpoint-save-skipped warning in log | Checkpoint saving would RPC a possibly-frozen worker | By design: the timeslice trainer vetoes saves when `save_freq > 0`; keep `trainer.save_freq=-1` |
143+
| Steps complete but one job's staleness queue near `staleness_threshold × batch` | Very asymmetric job speeds | Raise `async_training.staleness_threshold` (we use 8) or budget shorter runs; queue overflow drops samples |

0 commit comments

Comments
 (0)