AI Skill Hub 强烈推荐:Gym AI工作流评估框架 是一款优质的Agent工作流。已获得 1.0k 颗 GitHub Star,AI 综合评分 8.2 分,在同类工具中表现稳健。如果你正在寻找可靠的Agent工作流解决方案,这是一个值得深入了解的选择。
Gym AI工作流评估框架 是一套完整的 AI Agent 自动化工作流方案。通过可视化的节点编排,将复杂的多步骤任务拆解为清晰的自动化流程,实现全程无人值守的智能处理。支持与数百种外部服务和 API 无缝集成,适合构建数据处理管线、业务自动化和 AI 辅助决策系统。
Gym AI工作流评估框架 是一套完整的 AI Agent 自动化工作流方案。通过可视化的节点编排,将复杂的多步骤任务拆解为清晰的自动化流程,实现全程无人值守的智能处理。支持与数百种外部服务和 API 无缝集成,适合构建数据处理管线、业务自动化和 AI 辅助决策系统。
# 方式一:pip 安装(推荐)
pip install gym
# 方式二:虚拟环境安装(推荐生产环境)
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install gym
# 方式三:从源码安装(获取最新功能)
git clone https://github.com/NVIDIA-NeMo/Gym
cd Gym
pip install -e .
# 验证安装
python -c "import gym; print('安装成功')"
# 命令行使用
gym --help
# 基本用法
gym input_file -o output_file
# Python 代码中调用
import gym
# 示例
result = gym.process("input")
print(result)
# gym 配置文件示例(config.yml) app: name: "gym" debug: false log_level: "INFO" # 运行时指定配置文件 gym --config config.yml # 或通过环境变量配置 export GYM_API_KEY="your-key" export GYM_OUTPUT_DIR="./output"
Requirements • Quick Start • Environment Tutorials • Available Environments • Documentation & Resources • Community & Support • Citations
NeMo Gym is a library for evaluating and improving models and agents using environments. NeMo Gym provides infrastructure to develop environments, scalably run evaluation and training, and a collection of popular benchmarks and training environments.
An environment is the complete system an agent interacts with to complete a task. It consists of a dataset (tasks to solve), an agent harness (how the model interacts with the world), a verifier (task completion scoring), and state (per-task execution context).
NeMo Gym is designed to run on standard development machines:
| Hardware Requirements | Software Requirements |
|---|---|
| **GPU**: Not required for NeMo Gym library operation<br>• GPU may be needed for specific resources servers or model inference (see individual server documentation) | **Operating System**:<br>• Linux (Ubuntu 20.04+, or equivalent)<br>• macOS (11.0+ for x86_64, 12.0+ for Apple Silicon)<br>• Windows (via WSL2) |
| **CPU**: Any modern x86_64 or ARM64 processor (e.g., Intel, AMD, Apple Silicon) | **Python**: 3.13.14 or higher |
| **RAM**: Minimum 8 GB (16 GB+ recommended for larger environments) | **Git**: For cloning the repository |
| **Storage**: Minimum 5 GB free disk space for installation and basic usage | **Internet Connection**: Required for downloading dependencies and API access |
Additional Requirements
Requires Python 3.13.14+ on x86_64 or ARM64 (Linux, macOS, Windows via WSL2). No GPU required. See the Getting Started docs for a more comprehensive walkthrough.
Install NeMo Gym:
Requires uv and Python 3.13.14+.
git clone git@github.com:NVIDIA-NeMo/Gym.git
cd Gym
uv venv --python 3.13.14 && source .venv/bin/activate
uv sync
Configure your model:
This quickstart uses OpenAI. NeMo Gym supports local and hosted inference — see Configure Model for vLLM, Fireworks, OpenRouter, and others.
Create env.yaml in the project root:
policy_base_url: https://api.openai.com/v1
policy_api_key: <your-openai-api-key>
policy_model_name: gpt-4.1-2025-04-14
Learn how to build custom environments through hands-on tutorials. Here are popular starting points:
| Name | Demonstrates |
|---|---|
| [Single Step](https://docs.nvidia.com/nemo/gym/main/environment-tutorials/single-step-environment) | Basic single-step tool calling |
| [Multi Step](https://docs.nvidia.com/nemo/gym/main/environment-tutorials/multi-step-environment) | Multi-step tool calling |
| [Session State](https://docs.nvidia.com/nemo/gym/main/environment-tutorials/stateful-environment) | Session state management (in-memory) |
| [Multi Reward](https://docs.nvidia.com/nemo/gym/main/build-verifiers/multi-reward-verification) | Multiple reward components for evaluation and multi-objective RL (e.g. GDPO) |
See all environment tutorials for additional patterns and advanced topics.
Environments for training and evaluation.
Each resources server includes example data, configuration files, and tests. See each server's README for details.
The Dataset column links to publicly available datasets (e.g., on HuggingFace). A - means the train/validation data has not been publicly released yet, or that it is procedurally generated using a provided script. If no data is released yet, new data can be generated, or the environment can be used as a reference. Each server includes 5 example tasks in data/example.jsonl.
| Environment | Domain | Description | Value | Train | Validation | License | Config | Dataset |
|---|---|---|---|---|---|---|---|---|
| Aalcr | other | - | - | - | - | - | <a href='resources_servers/aalcr/configs/aalcr.yaml'>aalcr.yaml</a> | - |
| Abstention | rlhf | Train models to abstain when unsure using three-tier reward on HotPotQA with LLM judge | Improve calibration by rewarding abstention over incorrect answers | ✓ | ✓ | Creative Commons Attribution-ShareAlike 4.0 International | <a href='resources_servers/abstention/configs/abstention.yaml'>abstention.yaml</a> | - |
| Agentif | instruction_following | AgentIF instruction-following benchmark (707 agentic scenarios) scored with CSR/ISR using an LLM judge for llm/llm_conditional_check constraints and code exec for code constraints. | Improve instruction-following in agentic scenarios across unconditional, conditional, and example-driven constraint dimensions. | - | ✓ | - | <a href='resources_servers/agentif/configs/agentif.yaml'>agentif.yaml</a> | - |
| Anyswe Agent | coding | SWE-bench run by a NeMo Fabric-selected harness inside each task environment. | Evaluate software-engineering agents with a harness-neutral Fabric lifecycle. | - | - | - | <a href='responses_api_agents/anyswe_agent/configs/anyswe_nemo_fabric.yaml'>anyswe_nemo_fabric.yaml</a> | - |
| Anyswe Agent | coding | SWE-bench run by Claude Code natively inside the task container. | Eval software engineering capabilities on SWE-bench with any Gym agent. | - | - | - | <a href='responses_api_agents/anyswe_agent/configs/anyswe_claude_code.yaml'>anyswe_claude_code.yaml</a> | - |
| Anyswe Agent | coding | SWE-bench run by Hermes Agent natively inside the task container. | Eval software engineering capabilities on SWE-bench with any Gym agent. | - | - | - | <a href='responses_api_agents/anyswe_agent/configs/anyswe_hermes.yaml'>anyswe_hermes.yaml</a> | - |
| Anyswe Agent | coding | SWE-bench run by OpenClaw natively inside the task container. | Eval software engineering capabilities on SWE-bench with any Gym agent. | - | - | - | <a href='responses_api_agents/anyswe_agent/configs/anyswe_openclaw.yaml'>anyswe_openclaw.yaml</a> | - |
| Anyswe Agent | coding | SWE-bench run by OpenCode inside the task container. | Eval software engineering capabilities on SWE-bench with any Gym agent. | - | - | - | <a href='responses_api_agents/anyswe_agent/configs/anyswe_opencode.yaml'>anyswe_opencode.yaml</a> | - |
| Anyswe Agent | coding | SWE-bench run by Pi inside the task container. | Eval software engineering capabilities on SWE-bench with any Gym agent. | - | - | - | <a href='responses_api_agents/anyswe_agent/configs/anyswe_pi.yaml'>anyswe_pi.yaml</a> | - |
| Anyswe Agent | coding | SWE-bench run by the Cline CLI natively inside the task container. | Eval software engineering capabilities on SWE-bench with any Gym agent. | - | - | - | <a href='responses_api_agents/anyswe_agent/configs/anyswe_cline.yaml'>anyswe_cline.yaml</a> | - |
| Anyterminal Agent | coding | Terminal Bench run by a NeMo Fabric harness inside the task sandbox. | Evaluate terminal-task agents through a harness-neutral Fabric lifecycle. | - | - | - | <a href='responses_api_agents/anyterminal_agent/configs/anyterminal_nemo_fabric.yaml'>anyterminal_nemo_fabric.yaml</a> | - |
| Anyterminal Agent | coding | Terminal Bench run by claude-code natively inside the task container. | Evaluate terminal-task capabilities on Terminal Bench with any Gym agent. | - | - | - | <a href='responses_api_agents/anyterminal_agent/configs/anyterminal_claude_code.yaml'>anyterminal_claude_code.yaml</a> | - |
| Anyterminal Agent | coding | Terminal Bench run by OpenClaw natively inside the task container. | Evaluate terminal-task capabilities on Terminal Bench with any Gym agent. | - | - | - | <a href='responses_api_agents/anyterminal_agent/configs/anyterminal_openclaw.yaml'>anyterminal_openclaw.yaml</a> | - |
| Anyterminal Agent | coding | Terminal Bench run by Terminus-2 natively inside the task container. | Evaluate terminal-task capabilities on Terminal Bench with any Gym agent. | - | - | - | <a href='responses_api_agents/anyterminal_agent/configs/anyterminal_terminus_2.yaml'>anyterminal_terminus_2.yaml</a> | - |
| Anyterminal Agent | coding | Terminal Bench run by the Hermes agent inside the task container. | Evaluate terminal-task capabilities on Terminal Bench with any Gym agent. | - | - | - | <a href='responses_api_agents/anyterminal_agent/configs/anyterminal_hermes.yaml'>anyterminal_hermes.yaml</a> | - |
| Arc Agi | knowledge | Solve puzzles designed to test intelligence. See https://arcprize.org/arc-agi. | Improve puzzle-solving capabilities. | - | ✓ | - | <a href='resources_servers/arc_agi/configs/arc_agi.yaml'>arc_agi.yaml</a> | - |
| Arena | rlhf | LMArena proxy v2 chat evaluation benchmark | Measure general chat quality via win rate against baseline responses | - | ✓ | - | <a href='resources_servers/arena/configs/lmarena_v2.yaml'>lmarena_v2.yaml</a> | - |
| Arena | rlhf | LMArena proxy v3 chat evaluation benchmark | Measure general chat quality via win rate against baseline responses | - | ✓ | - | <a href='resources_servers/arena/configs/lmarena_v3.yaml'>lmarena_v3.yaml</a> | - |
| Arena Judge | - | - | - | - | - | <a href='resources_servers/arena_judge/configs/arena_judge.yaml'>arena_judge.yaml</a> | - | |
| Asr With Pc | other | ASR with WER scoring (standard, case-sensitive, punctuation+capitalization) | Improve transcription quality with structural detail | - | - | - | <a href='resources_servers/asr_with_pc/configs/asr_with_pc.yaml'>asr_with_pc.yaml</a> | - |
| Bbq | safety | BBQ comparative QA with Answer and Explanation Quality checks | Evidence-grounded and fairness-safe comparative reasoning | - | - | - | <a href='resources_servers/bbq/configs/bbq.yaml'>bbq.yaml</a> | - |
| Bigcodebench | coding | Verifies model-generated Python solutions against the BigCodeBench unittest suite. | Improve practical, library-rich Python coding capabilities. | - | - | - | <a href='resources_servers/bigcodebench/configs/bigcodebench.yaml'>bigcodebench.yaml</a> | - |
| Bird Sql | coding | Text-to-SQL with execution-based evaluation on BIRD dev (1534 SQLite tasks). Binary reward from unordered result-set equality. | Improve text-to-SQL capabilities on BIRD's realistic dev split using execution-based binary reward without an LLM judge. | - | - | - | <a href='resources_servers/bird_sql/configs/bird_sql.yaml'>bird_sql.yaml</a> | - |
| Blackjack | games | Blackjack. Model hits or stands. Reward +1 win, 0 draw, -1 loss/bust. | Example gymnasium-style multi-step environment | - | - | - | <a href='resources_servers/blackjack/configs/blackjack.yaml'>blackjack.yaml</a> | - |
| Browsecomp Advanced Harness | agent | Model uses search tools to satisfy a user query. | Measure agentic search capability | - | - | - | <a href='resources_servers/browsecomp_advanced_harness/configs/browsecomp_advanced_harness.yaml'>browsecomp_advanced_harness.yaml</a> | - |
| Bunsenbench Chemistry Mcq | knowledge | Public BunsenBench chemistry multiple-choice benchmark verifier | Measure chemistry MCQ reasoning with source and taxonomy breakdowns | - | - | - | <a href='resources_servers/bunsenbench_chemistry_mcq/configs/bunsenbench_chemistry_mcq.yaml'>bunsenbench_chemistry_mcq.yaml</a> | - |
| Calendar | agent | Multi-turn calendar scheduling dataset. User states events and constraints in natural language; model schedules events to satisfy all constraints. | Improve multi-turn instruction following capabilities | ✓ | ✓ | Apache 2.0 | <a href='resources_servers/calendar/configs/calendar.yaml'>calendar.yaml</a> | <a href='https://huggingface.co/datasets/nvidia/Nemotron-RL-agent-calendar_scheduling'>Nemotron-RL-agent-calendar_scheduling</a> |
| Calendar | agent | Multi-turn calendar scheduling dataset. User states events and constraints in natural language; model schedules events to satisfy all constraints. | Improve multi-turn instruction following capabilities | ✓ | ✓ | Creative Commons Attribution 4.0 International | <a href='resources_servers/calendar/configs/calendar_v2.yaml'>calendar_v2.yaml</a> | <a href='https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Calendar-v2'>Nemotron-RL-Instruction-Following-Calendar-v2</a> |
| Circle Click | other | Click on circles in images | Improve visual grounding and spatial reasoning | - | - | - | <a href='resources_servers/circle_click/configs/circle_click.yaml'>circle_click.yaml</a> | - |
| Circle Count | other | Count circles of a given color in images | Improve visual counting and color recognition | - | - | - | <a href='resources_servers/circle_count/configs/circle_count.yaml'>circle_count.yaml</a> | - |
| Citation If | instruction_following | Citation instruction-following reward checker | Train models to follow citation format instructions in search-grounded synthesis tasks | - | - | - | <a href='resources_servers/citation_if/configs/citation_if.yaml'>citation_if.yaml</a> | - |
| Code Fim | coding | Code Fill-in-the-Middle judged by HumanEval-Infilling test suite (single_line, multi_line, random_span, random_span_light) | Improve Python code-infilling capabilities (prefix + completion + suffix) | - | - | - | <a href='resources_servers/code_fim/configs/code_fim.yaml'>code_fim.yaml</a> | - |
| Code Gen | coding | Model must submit the right code to solve a problem | Improve competitive coding capabilities | ✓ | ✓ | Apache 2.0 | <a href='resources_servers/code_gen/configs/code_gen.yaml'>code_gen.yaml</a> | <a href='https://huggingface.co/datasets/nvidia/nemotron-RL-coding-competitive_coding'>nemotron-RL-coding-competitive_coding</a> |
| Competitive Coding Challenges | coding | Execution of competitive programming competition questions | Improve competitive coding capabilities on contest-style problems | - | - | - | <a href='resources_servers/competitive_coding_challenges/configs/competitive_coding_challenges.yaml'>competitive_coding_challenges.yaml</a> | - |
| Conversational Tool Use Simulation | agent | Conversational tool-use simulation environment. | - | - | - | - | <a href='resources_servers/conversational_tool_use_simulation/configs/conversational_tool_use_simulation.yaml'>conversational_tool_use_simulation.yaml</a> | - |
| Critpt | other | Research-level physics problems scored by the Artificial Analysis API | Evaluate model performance on research-level physics reasoning | - | - | - | <a href='resources_servers/critpt/configs/critpt.yaml'>critpt.yaml</a> | - |
| Equivalence Llm Judge | agent | Short bash command generation questions with LLM-as-a-judge | Improve foundational bash and IF capabilities | ✓ | ✓ | GNU General Public License v3.0 | <a href='resources_servers/equivalence_llm_judge/configs/nl2bash-equivalency.yaml'>nl2bash-equivalency.yaml</a> | - |
| Equivalence Llm Judge | knowledge | Short answer questions with LLM-as-a-judge | Improve knowledge-related benchmarks like GPQA / HLE | - | - | - | <a href='resources_servers/equivalence_llm_judge/configs/equivalence_llm_judge.yaml'>equivalence_llm_judge.yaml</a> | - |
| Equivalence Rule | knowledge | Question - Answering with rule-based reward | Improve retrieval and counting capabilities | - | - | - | <a href='resources_servers/equivalence_rule/configs/lc.yaml'>lc.yaml</a> | - |
| Ether0 | knowledge | ether0 chemistry benchmark verifiers | Evaluate chemistry knowledge and reasoning with ether0 benchmark | - | ✓ | - | <a href='resources_servers/ether0/configs/ether0.yaml'>ether0.yaml</a> | - |
| Evalplus | coding | Function-completion code judged by EvalPlus base + plus tests (HumanEval+, MBPP+) | Improve Python function-completion capabilities | - | - | - | <a href='resources_servers/evalplus/configs/evalplus.yaml'>evalplus.yaml</a> | - |
| Finance Agent V2 |
成熟的AI评估框架,1k星标体现社区认可度。工作流和基准测试功能完整,适合规模化模型评估。代码质量和维护状态良好。
AI Skill Hub 为第三方内容聚合平台,本页面信息基于公开数据整理,不对工具功能和质量作任何法律背书。
建议在沙箱或测试环境中充分验证后,再部署至生产环境,并做好必要的安全评估。
✅ Apache 2.0 — 宽松开源协议,可商用,需保留版权声明和 NOTICE 文件,含专利授权条款。
总体来看,Gym AI工作流评估框架 是一款质量优秀的Agent工作流,在同类工具中具备一定竞争力。AI Skill Hub 将持续追踪其更新动态,建议收藏备用,结合自身场景选择合适时机引入使用。
| 原始名称 | Gym |
| 原始描述 | 开源AI工作流:Evaluate and improve models and agents using environments。⭐1.0k · Python |
| Topics | AI评估工作流智能体基准测试环境模拟 |
| GitHub | https://github.com/NVIDIA-NeMo/Gym |
| License | Apache-2.0 |
| 语言 | Python |
收录时间:2026-06-30 · 更新时间:2026-07-03 · License:Apache-2.0 · AI Skill Hub 不对第三方内容的准确性作法律背书。
选择 Agent 类型,复制安装指令后粘贴到对应客户端