Tutorials and reproductions#

Choose a notebook based on what you want to learn.

Feature guides#

Learn one XDRL capability at a time, in the suggested order.

Interpret one module

Run an unchanged TorchRL policy through one interpreted component.

Interpret one module
Native TDHook workflow

Execute a TDHook workflow through an interpreted TorchRL component.

Native TDHook workflow
Repeated module calls

Select a repeated call with TDHook’s native occurrence support.

Repeated module calls belong to TDHook
Intervention

Apply a focused TDHook intervention through an interpreted component.

Intervene on an activation

Complete workflows#

Follow an end-to-end policy investigation that combines multiple capabilities.

End-to-end investigation

Run matched diagnosis and intervention workflows on one policy.

Compose an end-to-end policy investigation

Paper-inspired examples#

Learn the mechanics behind published interpretability methods on constructed examples.

Functional modules

Detect and prune modules in a synthetic classifier.

Functional-module workflow
Recurrent planning probes

Probe constructed future labels across recurrent calls.

Recurrent planning-probe workflow
Multi-agent concept policies

Intervene on concepts in a supervised policy example.

Multi-agent concept-policy workflow
Spatial goal steering

Patch engineered goal channels in an open-grid policy.

Spatial goal-steering workflow
Additive value decomposition

Inspect unary and pairwise terms on generated tensors.

Additive value-decomposition workflow