On September 28, 2026, Hugging Face announced RL Environment discovery on the Hub, giving agent builders a shared place to publish and find tasksets for reinforcement-learning training and evaluation. The RL Environment filter is available as of October 5. This guide explains how to package a taskset, declare its framework compatibility and check what its rewards actually measure.
Hugging Face Hub makes tasksets discoverable; framework runtimes execute tasks and verifiers assign rewards. The Hub’s dataset-card documentation is explicit: these remain dataset repositories. Tags add discovery and framework loading commands; they neither execute the environment nor establish that its scoring is correct.
Decide whether the taskset is for evaluation or training
Start with intended use, before choosing a loading command. The announcement’s Terminal-Bench 2.1 example is an evaluation benchmark. Its dataset card asks corpus builders to exclude its data from training. Use it to check evaluator setup, not as the training corpus in this walkthrough.
For a training release, document where tasks came from, their actual license, permitted reuse and how training material is separated from held-out evaluation. Listing an artifact under RL Environments does not resolve those questions. Nor does successful evaluation establish that its recorded trajectories contain everything a trainer needs.
1. Package the files your framework consumes
For Harbor, a task is more than a prompt. The task specification combines instructions, configuration, an environment definition and tests that produce a reward. A reference solution can support validation. A small repository could look like this:
environment-repo/
README.md
registry.json # needed for named -d selection
tasks/
example-task/
instruction.md
task.toml
environment/Dockerfile # one supported environment option
tests/test.sh
solution/solve.sh # reference solution, when providedThis is an illustrative layout, not a complete runnable task. Populate the files with your task’s implementation. Harbor supports environment specifications other than Dockerfiles when the consumer supports them. Its repository loader requires a registry for named dataset selection with -d; directory-based loading is also supported.
In the README, describe the starting state, available actions, termination conditions, reward calculation, known omissions and runtime dependencies. Explain whether a failed build, timeout or missing reward is an infrastructure error or a scored outcome. That description is the reward contract another team needs to review.
2. Add discovery metadata and upload
Add rl-environment and only the framework tags the repository actually supports. The supported tags are harbor, verifiers, openenv and nemo-gym. For a Harbor-only package, an illustrative README header is:
---
pretty_name: Example Harbor tasks
tags:
- rl-environment
- harbor
---Add accurate license and provenance information; do not copy another dataset’s license by default. Multiple framework tags describe existing compatibility—they do not convert task files into another format.
Create the dataset repository with the intended visibility and configure local authentication through the normal Hugging Face tooling. Review the upload directory, then adapt this documented upload pattern:
hf upload YOUR_ORG/YOUR_ENV ./environment-repo . --repo-type datasetUploading files does not execute or validate the taskset. Inspect the uploaded card, files and generated loading command. For a public release, also check discovery in the RL Environment filter. Keep the resulting repository revision so consumers can identify the same taskset.
3. Check one task before running a model
With Harbor and Docker installed and the task’s runtime requirements satisfied, a reference-solution run is a useful first check. The following command is adapted from Hugging Face’s example and Harbor’s repository syntax; it has not been executed for this guide. Check the installed version’s help and record its version before using it.
harbor run \
--repo https://huggingface.co/datasets/harborframework/terminal-bench-2.1@3e235dff6880252a587fa479c09bcb1e16edf2eb \
-d terminal-bench-2.1@2.1.0 \
-i regex-log \
-a oracle --env dockerThe pinned manifest declares the dataset version and includes regex-log. The repository commit fixes the Hub files; -d chooses the named dataset version within the registry. Use the full Hugging Face URL: a bare organization/repository name defaults to GitHub in Harbor.
The oracle executes the supplied reference solution without model inference, but it still needs runtime resources. Inspect the verifier output and logs; do not assume that a completed command means a successful task. A passing reference establishes one working path, not the absence of shortcuts or missing checks.
4. Audit the reward before trusting it
Harbor’s verifier contract requires a numeric reward in /logs/verifier/reward.txt or labeled numeric metrics in /logs/verifier/reward.json. If both exist, Harbor prefers the JSON file. Record the metric names and their meanings rather than treating every reward as a universal pass rate.
A concrete limitation appears in NVIDIA’s structured-output verifier documentation: its checks concern schema adherence, not the factual correctness of generated content. For a hypothetical extraction task, a well-formed response containing the wrong value could satisfy the structural criterion without satisfying the user’s request. Factual extraction therefore needs a separate content check.
Proposed validation controls, not results from this guide:
Reference solution: confirm that the intended solution receives the task’s documented reward.
No-op and plausible wrong output: check whether the verifier rejects doing nothing and completing the wrong objective.
Alternative valid solution: check that the verifier accepts another legitimate route to the requested outcome.
Broken runtime: deliberately interrupt a dependency and verify that the resulting error remains distinguishable from a model’s scored failure.
Hugging Face’s Repo2RLEnv documentation uses reference-solution and no-op controls with expected rewards of 1 and 0 for its binary example. Those values are not universal. A finite set of controls also cannot prove that every incorrect answer will fail. For task-generation workflows, the AutoSynthData guide discusses positive and negative verifier checks in more detail.
Framework compatibility does not guarantee comparable scores
The same task files can run under different execution and recording policies. These documented differences matter when selecting an adapter or comparing results:
Framework | Documented distinction | What to check |
|---|---|---|
The task’s verifier writes one or more numeric rewards. | Preserve verifier logs, metric definitions and infrastructure errors. | |
Authored agent and verifier timeouts are ignored by default. | Set | |
Training-ready capture needs engine token IDs and sampling-policy log probabilities, beyond ordinary evaluation traces. | Inspect | |
Gym produces interactions and rewards; a separate training framework such as NeMo RL performs post-training. | Choose and configure the trainer separately from the environment. |
These differences mean that load compatibility alone cannot establish score parity. For a comparison, hold task revision, harness, model settings, resource budgets, timeout policy and error handling constant—or report the differences explicitly. This is a methodological recommendation, not a measured cross-framework result.
Also check where evaluation records go. Verifiers’ current evaluation documentation describes uploads to Prime Intellect as the default; --no-push keeps the run’s results local. That setting concerns result uploads, not whether model calls or the selected runtime use remote services.
Release evidence that another team can use
A repository commit does not pin dependencies fetched outside that repository. Alongside the taskset, publish a concise validation record with task IDs, container digests, verifier dependencies, framework and harness versions, model settings when applicable, timestamps, raw rewards, errors and limitations. Separate reference-solution checks from model evaluations, and distinguish planned checks from completed ones.
Budget execution separately from hosting. Hub storage allowances and charges do not establish the price of model inference, sandbox builds or verifier calls. For a pilot, track all attempted-run costs and report both valid graded episodes and infrastructure failures; otherwise failed runs disappear from the apparent cost.
Start with one taskset and one framework. Publish its reward contract and versioned files, then add compatibility tags as other loading paths are validated. The useful release is one that another team can load, inspect and evaluate under clearly stated conditions.
Methodology: This AI-assisted guide uses primary announcements, dataset cards and framework documentation checked on October 5, 2026. No environment execution, model inference, reward-control tests or training runs were performed. Commands and proposed checks are documentation-derived, not an end-to-end tested tutorial.
