Skip past the demo

RLForge · reinforcement learning in C++20, from the tensor up

We built the tensor engine so we didn't have to trust anyone else's gradients. (We built the tensor engine so we didn't have to trust anyone else's gradients.)

A from-scratch reinforcement-learning library: its own tensor and reverse-mode autograd, NN layers and optimizers, tabular Q-learning, DQN and PPO, sync and threaded vector environments, and optional CBLAS and CUDA backends. No import torch. No Eigen. No borrowed autograd.

  • C++20
  • 148 tests
  • zero dependencies
  • -Werror

Watch an agent learn GridWorld, from −100 to optimal.

Start (0,0), goal (4,4). −1 per step, +10 at the goal, 100-step limit. Hover or tap a cell for its Q-values.

GridWorld 5×5 · tabular Q-learning · 4 lanes
seed 0 · train seed 42
+10
Q(x=0, y=0, ·)greedy → Up
Up
0.00
Down
0.00
Left
0.00
Right
0.00
trainer step
0
/ 20,000
explore ε
1.000
0 transitions
episodes done
0
across 4 lanes
greedy return
−100.0
stuck at a wall
0−50−100optimal 3.00101001k10k20ktrainer step, log scale →−100.0
mean greedy return, 10 eval episodes (seed 1000), after every trainer step

A TypeScript port, not the C++ library running in your browser. Same GridWorld, 4-lane SyncVectorEnvironment, TabularQLearningAgent and Trainer, same config as tests/test_trainer_grid_world.cpp: lr 0.1, γ 0.99, ε 1.0→0.05 over 20,000 transitions, slip 0, ties to Up. Checked checkpoint-for-checkpoint against a native run in web/scripts/parity.

The diamond bug

The bug hiding in most “from scratch” autograd engines.Not a crash. Not an exception. Just a quietly wrong number.

x feeds two branches that meet again at c. The usual mistake: process a's branch, write x's gradient, move on, then overwrite it when b's branch reaches x a moment later. It works on every straight-line example and breaks the instant your graph isn't one.

RLForge's rule: a node isn't processed until every consumer has deposited its contribution. Flip the toggle and watch the wrong gradient appear.

x = [1, 2, 3, 4] · backward passgrad += contribution
×2×3++mean+0.25x[1, 2, 3, 4]grad ·a = 2x[2, 4, 6, 8]grad ·b = 3x[3, 6, 9, 12]grad ·c = a + b[5, 10, 15, 20]grad ·loss = mean(c)12.5grad 1
  1. loss → c0.25
  2. c → a0.25
  3. c → b0.25
  4. a → x0.5
  5. b → x0.5 + 0.75 = 1.25

check_grad, per element of x

analytic
…
numerical (ε=1e-5)
1.250000
rel. error
…

Gradients are still flowing back toward x (0 of 2 deposits in).

the exact graph, from the C++ test suite
auto x = Tensor::from_data({1.0, 2.0, 3.0, 4.0}, {4});
x.requires_grad_(true);

auto a = x.mul(2.0);   // a = 2x
auto b = x.mul(3.0);   // b = 3x
auto c = a.add(b);     // c = 5x  — x used via two paths
auto loss = c.mean();  // scalar

How RLForge guarantees it

  1. 1. Topological order by iterative post-order DFS. No recursion, so deep graphs can't blow the stack.
  2. 2. Every node's incoming_grad is cleared, then contributions are +='d into it by distribute_grad.
  3. 3. A node's backward_fn runs once, after all its consumers, under no_grad().

Two rules we refused to break

Correctness you can't opt out of.Both are cheap to skip. Neither is optional here.

Rule 1

terminated and truncated are never the same flag.

An agent that died and an episode that hit a time limit look identical if all you track is done. The Bellman equation disagrees: one zeroes the bootstrap, the other keeps it. Every environment, buffer and batch conversion in RLForge keeps them apart.

s₀a₀, r = 5s₁ · max Q = 50
y = r + γ · (1 − terminated) · maxₐ Q(s′, a)
y = 5 + 0.9 · 1 · 50 = 50
Q(s₀, a₀) ← 0 + 0.5 · (50 − 0) = 25

The world didn't end, the clock did. s₁ is real, so its value stays in the target. The test expects exactly 25.

Rule 2

Mutate a tensor after forward, and backward throws.

Every storage carries a version counter; every backward closure captures the version it saw at forward time. If they disagree, RLForge throws instead of handing you a gradient computed against data that no longer exists.

  1. 30auto x = Tensor::from_data({1.0, 2.0, 3.0}, {3});
  2. 31x.requires_grad_(true);
  3. 33auto loss = x.mul(x).mean();
  4. 37auto buf = x.data_mutable();
  5. 38buf[0] = 99.0; // mutation
  6. 41loss.backward();
x.storage
·
version ·
mul backward closure
·
saved version ·

Build the graph first. The closure for x·x records the storage version it saw.

“…skipping it means your bugs show up as ‘training is unstable’ three weeks later instead of a stack trace today.”

The proof

−100 → 3.0, checked arithmetically.Not “the loss went down.” A proven-optimal policy.

tests/test_trainer_grid_world.cpp trains a tabular agent on a 4-way vectorized 5×5 GridWorld for 20,000 steps with every seed fixed, and asserts two exact numbers.

The arithmetic

  1. move 1−1
  2. move 2−1
  3. move 3−1
  4. move 4−1
  5. move 5−1
  6. move 6−1
  7. move 7−1
  8. move 8+10

7 × (−1) + 10 = 3

Reward is −1 per step and +10 on the step that reaches the goal (src/envs/grid_world.cpp:90). Deterministic dynamics mean a converged greedy policy hits this exactly on all 10 evaluation episodes, so the mean is exactly 3.0.

Architecture

Two stacks, one loop.Experience on one side, math on the other, algorithms where they meet.

ReplayBuffer's source never names a concrete storage type, so a new layout is a new TransitionStorage, not an edit. Trainer never asks whether an algorithm is step-based or rollout-based; it just asks should_update(). Pick a box to see what it uses and what depends on it.

Experience

rl::core · envs · vector_envs · replay_buffers

The loop

rl::core::Trainer · rl::agents

Math

rl::tensor · nn · optim · data

Trainer

act → step → make_transitions → observe → update when should_update(). Evaluates on a separate env.

uses

VectorEnvironment, Environment, Transition, Agent

used by

the caller

Trainer::train(), one step
actions = agent.act(current_observations, /*explore=*/true);
step_result = train_env.step(actions);
transitions = make_transitions(current_observations, actions, step_result); // Milestone 3's bridge
agent.observe_transitions(transitions);
if (agent.should_update()) {
    metrics = agent.update();
}
current_observations = step_result.observations;

The build log

Ten milestones. Each one had a hard part.Built incrementally, with every deferral written down.

Milestones 1 through 7 (the environment stack, the tensor and autograd engine, and DQN) were built end-to-end by Nikhil Mourya. Milestones 8 through 10 (PPO, the threading layer, and the CUDA/BLAS backends) were built together with aprv10, who led the concurrency and GPU work.

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7

Built together with aprv10, who led concurrency and GPU work.

  1. 8
  2. 9
  3. 10

The test suite

One binary. Tagged by subsystem.“The math looks right” and “the math checks out numerically” are different claims.

Don't take our word for any of this. Clone it, build it, run ctest yourself.

TEST_CASEs in rl_tests, counted from tests/*.cpp

10 groups · 148 TEST_CASEs

The triple every core op has to pass.

Each Milestone 5 op (add, sub, elementwise and scalar mul, matmul, relu, square, mean, gather) has all three. Linear, mse_loss, transpose, broadcasting add and max_last_dim get numerical gradient checks too.

01

Forward

The value is right.

Hand-computed expected outputs: square([−3, 0, 2, 4]) must come out as [9, 0, 4, 16].

02

Analytic backward

The derivative is right.

The closed-form gradient each backward_fn must produce: d(x²)/dx at [1, 2, 3] is [2, 4, 6].

03

Numerical gradient

The math checks out.

Central difference with ε = 1e-5 against backward(), relative tolerance 1e-5.

square(x) at x = 1.50

forward
2.2500
analytic
3.000000
numerical
3.000000
rel. error
6.6e-12

✓ within 1e-5: the analytic gradient checks out numerically.

Computed live in double precision with the suite's central difference: (f(x+ε) − f(x−ε)) / 2ε, ε = 1e-5.

Scars

We're not going to pretend this shipped clean.Honesty is a feature. So is -Werror.

-Wall -Wextra -Wpedantic -Werror has been on since commit one. At one point a stray backslash at the end of a comment in a test file, meant as harmless ASCII art, got read by the compiler as a line continuation and broke the build on a fresh clone. -Werror caught it immediately.

The bug that costs you five minutes today is the one that would've cost someone else an afternoon of “why won't this compile” next month.

// x
// / \← the line ends in a backslash…
auto a = x.mul(2.0);…so this line is now part of the comment
An illustration of the failure mode, not the original file. GCC reports it under -Wcomment (on with -Wall); -Werror turns it into a stop.

What we didn't build (yet)

Documented, not hidden.

Every deferred piece is written down: what it is, why it isn't here, and where it belongs.

If a homemade library's README doesn't have a section like this, it either doesn't have limitations or isn't telling you about them.

Quickstart

The real API. No hidden setup.No config files. Just the headers and a loop.

This is what training an agent actually looks like, straight from the README.

train.cpp
#include "rl/envs/grid_world.hpp"
#include "rl/vector_envs/sync_vector_environment.hpp"
#include "rl/agents/tabular_q_learning_agent.hpp"
#include "rl/core/trainer.hpp"

using namespace rl;

std::vector<vector_envs::EnvFactory> factories = {
    [] { return std::make_unique<envs::GridWorld>(); }
};
vector_envs::SyncVectorEnvironment train_env(factories);
envs::GridWorld eval_env;

agents::TabularQLearningAgent agent(train_env.action_space());
core::Trainer trainer(train_env, eval_env, agent);

auto result = trainer.train(20000);
The exact config behind −100 → 3.0▸

The quickstart runs on defaults: one env, GridWorld's 0.1 slip chance, ε decaying over 10,000 steps, no seeds. The numbers on this page come from the test's pinned config: 4 lanes, no slip, every seed fixed.

GridWorld::Config config;
config.size = 5;
config.max_episode_steps = 100;
config.slip_probability = 0.0f;

SyncVectorEnvironment train_env(make_grid_world_factories(4, config));
GridWorld eval_env(config);

TabularQLearningConfig agent_config;
agent_config.learning_rate = 0.1f;
agent_config.discount_factor = 0.99f;
agent_config.epsilon_start = 1.0f;
agent_config.epsilon_end = 0.05f;
agent_config.epsilon_decay_steps = 20000;
agent_config.seed = 0;
TabularQLearningAgent agent(train_env.action_space(), agent_config);

Trainer trainer(train_env, eval_env, agent);
auto baseline = trainer.evaluate(/*num_episodes=*/10, /*seed=*/1000);  // -100.0
auto train_result = trainer.train(/*num_steps=*/20000, /*seed=*/42);
auto after_training = trainer.evaluate(/*num_episodes=*/10, /*seed=*/1000);  // 3.0
build
git clone https://github.com/TryingtobeingNikhil/RLForge.git
cd RLForge

cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)
run the tests
cd build && ctest --output-on-failure

You'll need

  • A C++20 compiler (Clang 14+ / GCC 12+)
  • CMake 3.20+
  • An internet connection the first time, since Catch2 is fetched via FetchContent
optional backends (CPU stays the default)
# Enable CBLAS (links Accelerate on macOS)
cmake -S . -B build -DRL_ENABLE_BLAS=ON

# Enable CUDA and cuBLAS
cmake -S . -B build -DRL_ENABLE_CUDA=ON