Skip to content

Introducing AgentEnv: An Open-Source Framework for Building RL Environments

Agents learn by practicing in environments that behave like the real world: real apps, realistic data, and tasks that range from a single question to a full simulated workplace. Scale AI built AgentEnv Framework to create these environments, and it’s now open source.

AgentEnv Frameworkintro · 1:00

Why we built AgentEnv Framework

The problem

Creating realistic RL environments requires collaboration between researchers, engineers, and domain experts across many dimensions: artifacts, environment tools, dynamism of the environment, reproducibility, and more. There is no open source framework for building these environments effectively. Until now.

The approach

Most frameworks start from the task. AgentEnv Framework starts from the world, so one Environment can host as many Tasks and as many different agents as you like.

Can my agent… handle refund requests?

AgentEnv FrameworkEngineer adds the apps
Dana wants $98 backservers
Help desk
empty
Email
empty
Customer records
empty
environment inputs
clockMon 09:00
triggerDana changes her mind: store credit instead
task
taskResolve Dana’s refund request · rubric, 4 criteria

Illustrative

Open source

Build your environments with the same tools we use to build ours. The framework is on GitHub under the Apache-2.0 license.

uv tool install agentenv-frameworkagent-env run hello

Designed for interoperability

AgentEnv Framework keeps each piece a separate, versioned primitive: apps, data, the composed environment, tasks, agents and verifiers.

Each user works on their piece independently, so no one is blocked, and every task pins its versions so a rerun reproduces the original run.

Core Concepts →

Servers
Builds the app serversPriya
Composes the environmentPriya, Lena
Data
Writes the seed dataPriya, Marcus
Rules of the World
Sets the clockSam
Adds the triggersSam
Decides who can use which toolsSam
Task
Writes the promptMarcus
Writes the rubricMarcus, Sam
Runs
Deploys it on any sandboxLena
Runs the tasksLena

Agent Agnostic

Environments expose their tools over MCP, REST and a generated CLI, and agents connect through A2A, so the same environment and task run unchanged against any agent.

Sandboxes are pluggable. We have day-one support for four sandbox providers (local Docker, Modal containers, and Modal or E2B virtual machines), and you can register your own by name. This allows us to fairly and consistently test a variety of agents against the same task in the same environment, agnostic of the underlying infrastructure.

Agents →

Infrastructure Agnostic

An environment an engineer deploys on a laptop runs unchanged on a shared platform. AgentEnv Framework ships a registry for environments, tasks, data and agents, and connects to your infrastructure through a single config file. AWS and Google Cloud are supported from day one, and we partnered with Modal, the default sandbox on our own platform.

Ours runs on our internal platform, where we export hundreds of tasks at a time as self-contained bundles that other teams run on their own infrastructure.

Registry →

Your laptop
registers the support desk
Registry · AWSenvironments · tasks · data · agents
documentMongoDBbuilt in
objectAmazon S3built in
imageAmazon ECRbuilt in
secretSecrets Managerbuilt in
Teammate’s laptop
deploys it by id, same bytes
Your platform
runs it on a cloud sandbox

One config file points your install at its stores. Every install that points at the same ones sees the same environments.

Dynamic and Composable

Every app, like Slack or email, is its own Environment. Snap several together and the result is still one Environment, from a support desk to an entire workplace.

Apps are built once and shared: an accounting firm, a tech company and a support desk can all use the same Slack, each loaded with its own data. Build your own with the AgentEnv Framework SDK.

Composing environments →

Accounting firmLedgerPayrollSlackEmailTech companyReposIssuesSlackEmailSupport deskHelp deskCustomersSlackEmailBUILT ONCESlackEmail

Defining the Rules of the World

The real world is ever-changing, occasionally obfuscated and fundamentally non-deterministic. RL environments need to replicate these behavioral primitives while still keeping the mechanical integrity and determinism of RL training.

Virtual Clock

The Virtual Clock is the controller of time in the environment. A task sets it at the start of a run, and it can run faster than real time, up to one virtual day per second.

Controlling time means we can set the same “tomorrow” across unique runs, and an email reply that takes days finishes in seconds.

Virtual Clock →

Mon 09/02/2024
09:00:00
a virtual hour per second
Help desk
reads 09:00
Email
reads 09:00
Slack
reads 09:00
Customer records
reads 09:00

Same date, same time, in every app and for the agent.

Triggers

A trigger is a rule: when something happens in the world, something else follows. It can fire at a set time, on an agent’s action, or when the world reaches a given state.

This means we can set service “rate limits”, dynamically change the agent’s action space, have predicate-based conversations and more!

Triggers →

Fires when the virtual clock reaches a time you set, once or on a repeat, for work that arrives on a schedule, like a morning rush or a deadline.

EXAMPLE

WHEN

The clock reaches 09:30.

THEN Reveal data

A second complaint lands in the help desk.

A TRIGGER CANreveal datahide datagrant accesstake access awayact in plain languagespeak as a person

RBAC

RBAC (role-based access control) gives each role in the environment its own set of tools, the way one teammate has GitHub access and another doesn’t.

On our support desk, a support agent works tickets but can’t issue refunds; a billing lead can. Tools outside a role never appear in its tool list, and any call to them is refused.

RBAC →

Support desk · 6 tools

  • list_tickets
  • read_inbox
  • send_email
  • post_message
  • lookup_account
  • issue_refund
Support agent4 of 6
  • list_tickets
  • read_inbox
  • send_email
  • post_message
Billing lead3 of 6
  • list_tickets
  • read_inbox
  • issue_refund

Blue: only the billing lead has it. The rest is shared with the support agent.

1/5

Build any task, step by step

Tasks are DAGs of primitive steps. The framework ships 49 built-in step types, from deploying an environment to grading a trajectory, so most tasks need no code. Plugins add more.

Our tasks range from a single prompt to an entire simulated workplace. The cookbooks build two from an empty folder.

Tasks →

Give an agent a brief and a sandbox with Blender, and it builds and renders a 3D scene. No environment or grading needed: just infrastructure and an agent.

Step 1 of 4
1/6
outputnothing yet
The render appears as the agent works.

Anything is an Environment with Plugins

Plugins allow you to extend and build on top of AgentEnv and share with the community, turning anything an agent can act on into an environment. Watch agents play Civilization III in the OpenCiv3 RL Env and drive a real iPhone in the iOS Mobile RL Env.

Come build your first environment with AgentEnv Framework today!

No lock-in: any agent, any sandbox, any cloud

$ uv tool install agentenv-framework$ agent-env run hello

Python 3.11+. Without installing: uvx --from agentenv-framework agent-env run hello. With pip: pip install agentenv-framework.